Invest1 publisher2 min readPublished
Meta, Duke and UC Davis researchers lift Gemini 3 Flash to 62% on Olympiad math with a better harness
Meta, Duke and UC Davis researchers lifted Gemini 3 Flash from 46% to 62% on Olympiad math by searching for a better harness around the same model. That puts the wrapper on an operator's budget ahead of a bigger model, at least for math.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The method splits harness search into specialized branches, each evolving on its own subset of development data with its own policy for proposing changes.
- The authors report that the search ran only on development-set data and never had access to the test set.
- The work, posted to arXiv on September 29, 2026, extends the earlier Meta-Harness system and has not necessarily cleared formal peer review.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- decision A team running a small model on reasoning-heavy work can spend engineering time on the wrapper before paying for a larger tier; on math, that spend bought 16 points.
- cost Adopters have to keep several evolving harnesses and a router running, an ongoing engineering cost that a straight model upgrade would not add.
- exposure A harness tuned to Gemini 3 Flash is tied to one model, so a team may have to redo the search at each model change; this method overtook its own predecessor within about six months.
The authors lead with 34.8%, which is a relative figure. The underlying move is 16 points [1]. Counted in errors, Gemini 3 Flash went from getting 54% of Olympiad-level problems wrong to getting 38% wrong [2], so the search removed about three in ten of its mistakes [3]. According to the preprint, nobody touched the model's weights [2].
The search works on the wrapper: the prompts the model receives, the tools it can call, the context it sees and how its actions are executed [4]. Each branch keeps the cases where it beats its siblings and refines its strategy from earlier attempts [6]. At deployment a router sends each input to the best branch head [7]. The listed authors include Haoyu Dong (Meta and Duke) and Zihao Lin (Meta and UC Davis), with Lizhu Zhang and Zhuokai Zhao as co-last authors [12].
The spread across benchmarks supports two readings. The plainest is that the method helps most on closed-form math. If the coding figure is relative like the math one, the Olympiad gain is about nine times the SWE-bench Lite gain [4], with Terminal-Bench 2.0 between them [9]. The other reading is that the comparison is looser than it looks. CryptoBriefing's account does not give the search's compute cost or a larger model's score on the same problems. Nor does it say whether the 46.0% starting point is the bare model or Meta-Harness, the March system the new one is reported to beat [10].
I think the wrapper is a real alternative to a bigger model for math-style reasoning on a small model. For coding agents it is a weak one, given the ninefold gap. The counter-case is that harness gains stack on whatever model a team buys. In that case harness spending sits on top of the model budget and replaces none of it.
A bare larger model scoring above 62.0% on the same Olympiad set would break the view, because then the wrapper bought less than an upgrade does, even on math. A double-digit gain on SWE-bench Lite would extend it to coding agents.
What to watch
- A published compute cost for running the branches and router, the figure needed to set the 16-point gain against the price of a larger model.
- Results from the same method on models other than Gemini 3 Flash, showing whether a tuned harness carries over.
- Peer review or an independent replication of the 62.0% Olympiad result.