Build2 publishers3 min readPublished
Best agent in Databricks' live document-reasoning contest scored 63.3%
Eleven academic teams spent two months tuning on OfficeQA, then met a fresh benchmark on the day. The winner cleared 63.3 percent, and 18.8 percent of questions defeated every entrant.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Databricks hosted the inaugural Grounded Reasoning Cup, a first-of-its-kind live AI competition to evaluate AI agents' ability to reason over complex, enterprise-style document collections.
- The Grounded Reasoning Cup was designed to help answer how well performance improvements on a benchmark generalize to similar, real-world tasks.
- The competition brought together 11 top academic teams from across the U.S. and Canada, paired with resources and mentorship from frontier labs including OpenAI, Anthropic, and Google DeepMind.
- Teams of 2-4 people representing their academic institution had approximately two months to build an agent using any approaches they saw fit, with the one constraint that they must use their partner lab's model family exclusively to power their agent.
- Teams developed and optimized their agents on OfficeQA, Databricks' flagship grounded-reasoning benchmark designed to reflect economically valuable enterprise workflows.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Databricks ran the first Grounded Reasoning Cup, a live competition in which academic teams pointed purpose-built agents at complex, enterprise-style document collections [1]. Stanford won with 63.3 percent accuracy [7], which is another way of saying the best system in the room got roughly one question in three wrong [1].
The structure is what makes the number interesting. Eleven academic teams from the U.S. and Canada, in groups of two to four, were paired with OpenAI, Anthropic, or Google DeepMind for model access and mentorship, and each team was required to build exclusively on its partner lab's model family [3][4]. They had about two months to optimize against OfficeQA, Databricks' public grounded-reasoning benchmark, which the company says reflects economically valuable enterprise workflows [5]. On competition day they ran those same systems in real time against OfficeQA Pro V2, a freshly released benchmark whose whole purpose was to test whether the two months of tuning transferred [6]. Databricks framed the event around exactly that question: how far do benchmark improvements generalize to similar real-world tasks [2].
The answer is partially, and unevenly. The average team scored about 41 percent, the top three exceeded 50 percent, and Stanford's win was about 22 points above the team average and about 35 points above the average frontier-agent offline baseline [8][9]. Work the arithmetic backward and that baseline sits somewhere near 28 percent [2], so purpose-built scaffolding roughly doubled what an off-the-shelf agent managed. That is a real engineering result. It is not a deployment result.
The harder number is 18.8 percent: questions that no team solved at all [10]. Pool every one of the eleven systems, take the best answer available from any of them, and the ceiling is 81.2 percent [3]. That residue is not a ranking artifact or a prompt-tuning gap. It is a class of enterprise document question that current agent design, with frontier models and two months of dedicated effort behind it, does not answer.
What separated the leaders was mostly not the model call. Databricks reports the strong systems combined document preprocessing, targeted retrieval, parallel agents, structured tool use, and answer verification [11], and that performance depended less on a single call than on the surrounding machinery: how documents were parsed, how evidence was retrieved, how intermediate calculations were done, and how answers were checked before submission [12]. Stanford's method, per Databricks, was to trace wrong answers on the public benchmark back to the exact misstep and convert those failure patterns into reusable skills for its Claude Code agent, covering things like table localization and answer formatting [13]. That is debugging discipline, not novelty, and it is the transferable lesson for anyone building the same thing internally.
Databricks' own read is that grounded reasoning over enterprise corpora has improved in the seven months since OfficeQA shipped but is still far from solved [16][14]. Operators should treat 63.3 percent as the ceiling for a well-resourced team with two months and lab support, on a corpus built to look like their own filings and reports.
Watch whether OfficeQA Pro V2 scores move once teams can tune on it directly, and by how much: that delta is the size of the generalization gap. Watch the unsolved 18.8 percent for a taxonomy of what breaks. And watch whether the next iteration reports per-question cost and latency alongside accuracy, since parallel agents and verification passes are not free.