NVIDIA AVO Pushes Claude Opus 5 To A Perfect ARC-AGI-3 Benchmark Score

NVIDIA AVO Pushes Claude Opus 5 To A Perfect ARC-AGI-3 Benchmark Score

By Jon Markman
Publication Date: 2026-08-24 15:36:00

Nvidia took a frontier AI model that completes about 30 percent of one of the field’s hardest benchmarks and got a perfect score out of it.

The model was Anthropic’s Claude Opus 5, and nothing inside it changed: no retraining, no fine-tuning, not one adjusted weight. Everything that improved sat outside the model, in a software system Nvidia calls AVO that manages the model’s memory, plans its next moves, and watches for mistakes. Better software pulled three times more performance out of a model that already exists.

What AVO Actually Is

The software layer between a model and a task is called a harness, and Nvidia’s researchers describe its job in one clean passage:

A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.

AVO, short for Agentic Variation…