Build2 distinct publishers3 min readUpdated
Eleven academic teams spent two months tuning on OfficeQA, then met a fresh benchmark on the day. The winner cleared 63.3 percent, and 18.8 percent of questions defeated every entrant.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Databricks ran the first Grounded Reasoning Cup, a live competition in which academic teams pointed purpose-built agents at complex, enterprise-style document collections [1]. Stanford won with 63.3 percent accuracy [7], which is another way of saying the best system in the room got roughly one question in three wrong [1].
The structure is what makes the number interesting. Eleven academic teams from the U.S. and Canada, in groups of two to four, were paired with OpenAI, Anthropic, or Google DeepMind for model access and mentorship, and each team was required to build exclusively on its partner lab's model family [3][4]. They had about two months to optimize against OfficeQA, Databricks' public grounded-reasoning benchmark, which the company says reflects economically valuable enterprise workflows [5]. On competition day they ran those same systems in real time against OfficeQA Pro V2, a freshly released benchmark whose whole purpose was to test whether the two months of tuning transferred [6]. Databricks framed the event around exactly that question: how far do benchmark improvements generalize to similar real-world tasks [2].
The answer is partially, and unevenly. The average team scored about 41 percent, the top three exceeded 50 percent, and Stanford's win was about 22 points above the team average and about 35 points above the average frontier-agent offline baseline [8][9]. Work the arithmetic backward and that baseline sits somewhere near 28 percent [2], so purpose-built scaffolding roughly doubled what an off-the-shelf agent managed. That is a real engineering result. It is not a deployment result.
The harder number is 18.8 percent: questions that no team solved at all [10]. Pool every one of the eleven systems, take the best answer available from any of them, and the ceiling is 81.2 percent [3]. That residue is not a ranking artifact or a prompt-tuning gap. It is a class of enterprise document question that current agent design, with frontier models and two months of dedicated effort behind it, does not answer.
What separated the leaders was mostly not the model call. Databricks reports the strong systems combined document preprocessing, targeted retrieval, parallel agents, structured tool use, and answer verification [11], and that performance depended less on a single call than on the surrounding machinery: how documents were parsed, how evidence was retrieved, how intermediate calculations were done, and how answers were checked before submission [12]. Stanford's method, per Databricks, was to trace wrong answers on the public benchmark back to the exact misstep and convert those failure patterns into reusable skills for its Claude Code agent, covering things like table localization and answer formatting [13]. That is debugging discipline, not novelty, and it is the transferable lesson for anyone building the same thing internally.
Databricks' own read is that grounded reasoning over enterprise corpora has improved in the seven months since OfficeQA shipped but is still far from solved [16][14]. Operators should treat 63.3 percent as the ceiling for a well-resourced team with two months and lab support, on a corpus built to look like their own filings and reports.
Watch whether OfficeQA Pro V2 scores move once teams can tune on it directly, and by how much: that delta is the size of the generalization gap. Watch the unsolved 18.8 percent for a taxonomy of what breaks. And watch whether the next iteration reports per-question cost and latency alongside accuracy, since parallel agents and verification passes are not free.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Databricks hosted the inaugural Grounded Reasoning Cup, a first-of-its-kind live AI competition to evaluate AI agents' ability to reason over complex, enterprise-style document collections.
The Grounded Reasoning Cup was designed to help answer how well performance improvements on a benchmark generalize to similar, real-world tasks.
The competition brought together 11 top academic teams from across the U.S. and Canada, paired with resources and mentorship from frontier labs including OpenAI, Anthropic, and Google DeepMind.
Teams of 2-4 people representing their academic institution had approximately two months to build an agent using any approaches they saw fit, with the one constraint that they must use their partner lab's model family exclusively to power their agent.
Teams developed and optimized their agents on OfficeQA, Databricks' flagship grounded-reasoning benchmark designed to reflect economically valuable enterprise workflows.
On competition day, teams were challenged to apply their systems in real time to a newly released grounded-reasoning benchmark, OfficeQA Pro V2, designed to test whether their improvements generalized.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed primary disclosure, single-organizer measurement
Both sources agree on every headline figure and the host published unusually granular method detail: score distribution, unsolved-question share, per-team architectures, Stanford's 57-of-88 attempt record and its mid-event verifier change. A participant institution independently confirmed its placement and the date. Evidence stops short of high confidence because one vendor authored, hosted and scored the evaluation, no per-question data or replication package is reported, and the partner-lab constraint means the scores cannot be read as model comparisons.
Real event participation, no production usage shown
Concrete, dated participation exists: eleven university teams, three frontier labs supplying models and mentorship, a held-out ~120,000-page corpus, a released second benchmark and a published results recap corroborated by one participant. But adoption is confined to a single organizer-run competition. Neither source reports enterprise deployments, customers running these pipelines, or OfficeQA use outside Databricks' own evaluation, so real-world uptake stays low.
Numbers mostly self-limiting, framing slightly ahead
The disclosed figures cut against overstatement: the winner missed about one in three questions, the average team sat at ~41%, and the host volunteered that 18.8% of questions defeated everyone. The modest positive gap comes from framing rather than data - 'first-of-its-kind', the generalization premise resting on one held-out corpus from a single document domain, and an asserted seven-month improvement with no prior-period baseline published. The secondary report narrows the gap by labelling the event vendor-run and cautioning that the same-model-family spread does not isolate model quality.
Organizer owns benchmark, venue and scoring
Databricks created OfficeQA, calls it its flagship benchmark, built the held-out OfficeQA Pro V2, hosted the contest at its own summit and published the results narrative - a stack of aligned promotional interests. The three partner labs supplied models and mentorship and each had exclusive attribution over its teams' results, giving them a stake in how scores are read. Winning universities also gain reputationally. The secondary outlet discloses the vendor-run nature, which is why this is not scored higher.
Numbers solid, generality unproven
Confidence in the reported figures is high - two sources agree, the host published method detail, and a participant confirmed its placement. Confidence in the broader inference is lower: the systems lesson rests on eleven agents in one time-pressured event on one Treasury-document corpus, scores are entangled with speed bonuses and double-value rounds, and no cost, throughput or production data is available to test whether the winning pipelines transfer.
product
Wu says Cognition is not for sale. The more useful fact is who bought Cursor last week.1 distinct publisher
build
OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit3 distinct publishers
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026
1 article · August 18, 2026