Nvidia research: the harness, not the model, wins long-horizon tasks
Nvidia researchers argue that the harness — tools, memory, and a supervisor — matters more than the underlying model for long-horizon work. Claude Opus 5 with a custom harness scored 100% on ARC-AGI-3, versus 30% without it.

On August 21, 2026, TechCrunch reported Nvidia research suggesting that for long-horizon tasks the software wrapper around a model — the harness of tools, memory management, and rules — matters more than the model itself. A harness is what turns a raw model into something that can act on its own.
What we know
- With a custom harness tuned for memory and a supervisor component, Claude Opus 5 scored 100% on the interactive benchmark ARC-AGI-3.
- Without the harness, Opus 5 scored 30%, the top result among the models tested in that condition.
- Nvidia researchers built their own harness, Agentic Variation Operators; this is research, not a new Nvidia product.
- Adel El Hallak of Nvidia told TechCrunch that open harnesses, like open models, let users turn more knobs and keep control of the agent stack.
Takeaways
- The 100% ARC-AGI-3 score is Opus 5 plus a custom harness, not the base model alone.
- 30% without a harness was still the best among models tested that way.
- Nvidia’s public argument is that open harnesses now matter as much as open models.
Source: TechCrunch / Nvidia


