Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark.

NVIDIA AVO completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.

Sign in to suggest edits

Key sources

  1. DISCUSSIONdsrtslnd23news.ycombinator.com
  2. SOURCE@nvidiaai“completed all 183 levels across all 25 public environments”x.com
  3. SOURCE@nvidiaai“sustain progress across long-running tasks rather than starting over with each model context”x.com
  4. SUPPORT@wallstengine“finished in 6,624 actions, about 12% fewer than VISTA’s 7,542-action run”x.com
  5. SUPPORT@mark_k“produced kernels that beat FlashAttention-4 by up to 10.5%”x.com
  6. SUPPORT@hesamation“this is the PUBLIC SET, not the actual private benchmark”x.com
  7. SUPPORT@clementdelangue“general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark”x.com
  8. SOURCEhuggingnewshuggingnews.com
Markdown