NotableotherNVIDIA

Nvidia's SoL-Pi System Reduces Token Usage by Up to 54.3% Compared to Competitors

Published
Sep 26, 2026 — 10:30 UTC

Nvidia's SoL-Pi System Reduces Token Usage by Up to 54.3% Compared to Competitors

Nvidia's SoL-Pi system achieves a 50% reduction in token usage compared to Codex and a 54.3% reduction compared to Claude Code. This system, which costs $894, is a significant improvement over the $1,339 Pi system, offering estimated savings of $8.75 to $13.50 per hour when using Codex and Claude Code, and $4.36 to $5.71 per hour with Pi.

SoL-Pi operates across 535 executable environments, generating 3,000 total runs and facilitating 60,000 agent-environment interactions. In tests on the Terminal-Bench 4, SoL-Pi solved 15 tasks, while Codex and Pi each solved 18 tasks. Performance retention on the Opus 5 model is reported at 94.3% of Pi's performance.

The system's efficiency gains are attributed to recursive self-improvement, although Nvidia researchers caution that shorter context lengths can reduce prompt cache reuse. The average user instructions preserved by compression is 17%, and nearly 70% of token usage comes from cached prompts.

Peter Walker, an analyst, noted that the cost per solved task varied by nearly 3x, despite the same model performing the tasks. Eric Provencher, a developer, emphasized that using more than two sub-agents typically increases token consumption without improving output quality. This follows a trend of 14x growth in agentic token usage since February 2026, highlighting the need for optimization in AI coding agents.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: The Decoder