Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

MIT CSAIL's VISTA harness lifts Claude Opus 5.0 from 40.68 to 100 on ARC-AGI-3 efficiency

MIT CSAIL's VISTA harness raised Claude Opus 5.0's ARC-AGI-3 action-efficiency score from 40.68 to 100 without retraining or replacing the model. The comparison holds the model fixed, so the lift comes from the interface built around it.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying MIT CSAIL's VISTA harness lifts Claude Opus 5.0 from 40.68 to 100 on ARC-AGI-3 efficiency
Generated illustration

What happened

  • Claude cleared all 25 public games using 57.4% fewer actions than first-time human players, figures the post attributes to the paper's abstract.
  • VISTA swaps the official 64x64 numeric text grid for 512x512 screenshots and archives every frame unaltered so the model can pull any of them back mid-reasoning.
  • Early coverage cited by the post reports that screenshots alone lifted GPT-5.6 Sol from 13.33 to 47.32 with the model and harness unchanged.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Changing what an agent sees and stores is a cheaper first experiment than a model swap, and on these figures it moved scores further than switching models did.
  • precedent ARC-AGI-3 results that name a model without its harness become hard to compare, so expect serious entries to report both.
  • capability Rules kept as three sentences of notes can be read and corrected by a person; a 4,000-line generated program is far harder to review.

Lookups into VISTA's frame archive do not count as steps. Only actions in the game do [7]. The agent can reread old frames, zoom in and check pixel values as often as it likes, and the benchmark charges it only when it moves [6][7]. Claude's run came to a reported 7,302 steps, about 0.43 times the human average [13]. On the game m0r0, humans needed 1,107 steps and Claude needed 219 [14], roughly a fifth of the human count [18]. The post does not report Claude's token spend or wall-clock time under VISTA.

The one token figure in the post comes from the GPT-5.6 Sol comparison [15]. On images, GPT-5.6 Sol used 30.7 million tokens per game. On the text grid it used 71.9 million [16]. That is about 57% fewer tokens for the screenshot version [22].

The post's own numbers allow a rough test of model against harness. One caveat: only Claude's headline row comes from the paper, and the post labels the other rows as early coverage [10]. The harness moved Claude 59.32 points [12] and the 320B open-weight model 65.04 [21]. Under the official interface, Claude led that model by 38.79 points [19]. Under VISTA, it still led by 33.07 [20].

For both models, changing the harness moved the score more than the gap between the two models [12][19][21]. In my view model comparisons still matter. I would run them inside the harness I intended to ship.

The paper, arXiv:2610.02200, was posted October 1 by Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He [1]. The design is clean. VISTA has no trained components, and every game gets the same four-sentence prompt template. The model writes down the rules it infers as natural-language notes [8]. For the checkers game LF52, the post says the Schema approach captured the rules in about 4,000 lines of Python, while VISTA's notes took three sentences [9].

I think the result transfers to an agent when three conditions hold. The state has to be spatial enough that a text encoding hides it. The environment has to let you keep every observation. And the cost you care about has to sit in external actions, since RHAE counts nothing else [7]. ARC-AGI-3's grid-world puzzles and checkers variants, learned by playing with no tutorial, meet all three [11]. An agent billed per token meets the first two at best. For that agent, every free lookup still shows up on the invoice.

What to watch

  • The paper's full tables, to confirm the GPT-5.6 Sol and 320B open-weight figures the post attributes to early coverage.
  • Token and wall-clock totals for Claude under VISTA, to show whether the step savings hold up under a per-token bill.
  • Whether ARC-AGI-3 scoring starts counting frame retrieval toward an agent's action budget.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories