Skip to content

Invest1 publisher2 min readPublished

Meta, Duke and UC Davis researchers lift Gemini 3 Flash to 62% on Olympiad math with a better harness

Meta, Duke and UC Davis researchers lifted Gemini 3 Flash from 46% to 62% on Olympiad math by searching for a better harness around the same model. That puts the wrapper on an operator's budget ahead of a bigger model, at least for math.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Meta, Duke and UC Davis researchers lift Gemini 3 Flash to 62% on Olympiad math with a better harness
Generated illustration

What happened

  • The method splits harness search into specialized branches, each evolving on its own subset of development data with its own policy for proposing changes.
  • The authors report that the search ran only on development-set data and never had access to the test set.
  • The work, posted to arXiv on September 29, 2026, extends the earlier Meta-Harness system and has not necessarily cleared formal peer review.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • decision A team running a small model on reasoning-heavy work can spend engineering time on the wrapper before paying for a larger tier; on math, that spend bought 16 points.
  • cost Adopters have to keep several evolving harnesses and a router running, an ongoing engineering cost that a straight model upgrade would not add.
  • exposure A harness tuned to Gemini 3 Flash is tied to one model, so a team may have to redo the search at each model change; this method overtook its own predecessor within about six months.

The authors lead with 34.8%, which is a relative figure. The underlying move is 16 points [1]. Counted in errors, Gemini 3 Flash went from getting 54% of Olympiad-level problems wrong to getting 38% wrong [2], so the search removed about three in ten of its mistakes [3]. According to the preprint, nobody touched the model's weights [2].

The search works on the wrapper: the prompts the model receives, the tools it can call, the context it sees and how its actions are executed [4]. Each branch keeps the cases where it beats its siblings and refines its strategy from earlier attempts [6]. At deployment a router sends each input to the best branch head [7]. The listed authors include Haoyu Dong (Meta and Duke) and Zihao Lin (Meta and UC Davis), with Lizhu Zhang and Zhuokai Zhao as co-last authors [12].

The spread across benchmarks supports two readings. The plainest is that the method helps most on closed-form math. If the coding figure is relative like the math one, the Olympiad gain is about nine times the SWE-bench Lite gain [4], with Terminal-Bench 2.0 between them [9]. The other reading is that the comparison is looser than it looks. CryptoBriefing's account does not give the search's compute cost or a larger model's score on the same problems. Nor does it say whether the 46.0% starting point is the bare model or Meta-Harness, the March system the new one is reported to beat [10].

I think the wrapper is a real alternative to a bigger model for math-style reasoning on a small model. For coding agents it is a weak one, given the ninefold gap. The counter-case is that harness gains stack on whatever model a team buys. In that case harness spending sits on top of the model budget and replaces none of it.

A bare larger model scoring above 62.0% on the same Olympiad set would break the view, because then the wrapper bought less than an upgrade does, even on math. A double-digit gain on SWE-bench Lite would extend it to coding agents.

What to watch

  • A published compute cost for running the branches and router, the figure needed to set the 16-point gain against the price of a larger model.
  • Results from the same method on models other than Gemini 3 Flash, showing whether a tuned harness carries over.
  • Peer review or an independent replication of the 62.0% Olympiad result.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories