Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

By Jonathan Kemper
Publication Date: 2026-09-26 10:30:00

A new study from Nvidia researchers tackles these costs not at the model level but at the harness, the control layer between the model and its environment used by systems like Codex, Claude Code, or OpenClaw.

A research AI analyzes agent traces, proposes harness changes, and keeps only those that maintain performance while cutting costs. SoL-Pi saves 50 percent compared to Codex and 54.3 percent compared to Claude Code on EdgeBench. | Image: Nvidia

The harness controls how an agent sees states, runs actions, and processes feedback. Most efficiency methods so far have focused on cutting the cost per token through faster attention kernels and serving infrastructure, model compression like quantization, or swapping in cheaper models.

AI explores 152 directions to find leaner control logic

Optimizing the harness is hard in practice because tool usage, context management, verification, and abort logic are all tightly coupled. A change that saves tokens in one place can trigger errors elsewhere or just push costs into a later phase. Typically, humans sift through long execution traces and translate recurring failure patterns into code.

The system, called SoL-Pi, automates that work. A research agent watches another agent’s traces, proposes changes, and tests them in prepared environments. Capability and efficiency checks determine which candidates survive. The approach draws on recursive self-improvement, according to the authors.

Flowchart of the search pipeline from trajectory rollouts through map-reduce analysis, mechanism proposal, candidate implementation, independent review, and development validation to a separate held-out evaluation.
The held-out evaluation happens only after the…