The Harness Did What The Model Couldn't

NVIDIA's AVO (Agentic Variation Operators) system posted a perfect 100.00 score on ARC-AGI-3, completing all 183 levels across the benchmark's 25 public environments [1]. The striking part isn't the score itself, it's the delta: the same underlying model, Anthropic's Claude Opus 5, manages only about 30.2% when tested bare, with no surrounding system around it [2]. AVO is not a new or fine-tuned model, it's a general-purpose software harness wrapped around Opus 5, and NVIDIA's own framing is blunt about what that implies: 'The model is only one component of an agent, and the harness around it - memory, tool use, recovery from failure, sustained context across a long task - is what determines whether that underlying capability translates into completed levels.' [3]Reaction on Reddit converged on the identical read: this is a harness result, not evidence of a smarter model, with one r/accelerate commenter putting it plainly - the models already have more capability than assumed, and the way they're being used has been the actual bottleneck. That reframing matters for anyone tracking AI progress: roughly 70 percentage points of task completion came from scaffolding engineering, not from a model upgrade.


