Particle.news
Download on the App Store

Nvidia’s AVO Agent Scores 100% on ARC-AGI-3 Public Set

The company says its Agentic Variation Operators system turned Anthropic’s Claude Opus 5 into a sustained, memory-backed problem solver that completed all 183 public levels, a result that highlights system design but is limited to public tests.

Overview

  • Nvidia reported on Friday, August 21, 2026 that AVO, its Agentic Variation Operators system paired with Claude Opus 5, achieved a perfect 100.00 RHAE score by clearing all 183 levels across 25 ARC-AGI-3 public environments.
  • AVO combines persistent memory and a programmatic supervisor to preserve state across steps, run iterative plan–act–observe loops, and intervene when searches stall so the agent can pursue long-horizon problems without losing progress.
  • Independent analysis from the ARC Prize and reporting this July found Claude Opus 5 alone scored about 30% on the ARC-AGI-3 public set, underscoring that the large jump to 100% came from the agent architecture wrapped around the model.
  • Nvidia disclosed results only on the benchmark’s public dataset and has not released semi-private or private-set outcomes or plans to commercialize AVO, leaving questions about overfitting and broader generalization unresolved.
  • AVO previously improved GPU-kernel optimization in internal tests and, if it generalizes, could strengthen Nvidia’s software value to customers by enabling models to sustain multi-step workflows; observers should watch for private-test results and any product moves.