Published · 6d agoScience2 min read
4.5 Million Graded Answers Say the Fact Is on the Shelf and the Model Cannot Find It
Google Research profiled 2,150 Wikipedia facts across 13 models and reports frontier systems encode 95-98 percent of them while failing to directly recall a quarter to a third.
Written for builders.See today for builders

What happened
- The paper is titled "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality", posted to arXiv under Computer Science > Computation and Language.
- The arXiv abstract states the findings are drawn from 4 million responses from 13 LLMs.
- The evaluation produced approximately 4.5 million responses.
- 13 LLMs were evaluated, each both with and without thinking.
- For each model, fact and task, eight responses were sampled, and responses were graded automatically by prompted LLM autoraters.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
The figure is a sample size, not a score: roughly 4.5 million graded model responses, produced by Google Research scientists Nitay Calderon and Gal Yona to separate facts a model never learned from facts it learned and cannot retrieve [3][12]. That split matters because the two failures have different fixes and different bills.
The arithmetic is worth seeing, because it bounds what the number supports. WikiProfile holds 2,150 facts extracted from Wikipedia pages, each paired with ten tasks: two probing encoding, four probing knowledge, four multiple-choice items probing recognition [4][14]. Thirteen models were evaluated, each with and without thinking, with eight responses sampled per model, fact and task [c3b][c3c]. That product is about 4.47 million responses [7]. The scale comes from repetition and task variants, not breadth of subject matter: 2,150 facts examined very hard, with the arXiv abstract citing 4 million responses across 13 models [2].
What it turns on: for Gemini-3-Pro and GPT-5, 95 to 98 percent of facts are encoded, and yet those same models fail to directly recall 26 to 34 percent of them [5]. With thinking, they still fail on 11 to 12 percent [6]. Comparing the reported range endpoints, thinking closes roughly three fifths of the direct-recall gap [8]. The authors report the failures are systematic, falling disproportionately on long-tail facts and reverse questions [9].
The decision this changes is which remedy to fund. Encoding failures call for scaling model size or expanding data coverage; recall failures point instead to post-training and inference-time methods that use what is already stored [10]. The paper's conclusion is that future factuality gains may depend less on scaling than on utilisation [13].
One caution on the instrument: the benchmark was built by an automated pipeline driven by a prompted Gemini-2.5-Pro with thinking, filtered against a search engine, then manually validated, and responses were graded by prompted LLM autoraters [11][c3c].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The paper is titled "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality", posted to arXiv under Computer Science > Computation and Language.
ReportedView cited source - [2]
The arXiv abstract states the findings are drawn from 4 million responses from 13 LLMs.
ReportedView cited source - [c3c]
For each model, fact and task, eight responses were sampled, and responses were graded automatically by prompted LLM autoraters.
ReportedView cited source - [4]
WikiProfile contains 2,150 facts after automated filtering and a final manual validation step, and each fact is paired with 10 tasks: two for encoding, four for knowledge evaluation, and four multiple-choice variants for recognition.
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- research.google6d agoEmpty shelves or lost keys? Recall is the bottleneck for parametric factuality
- research.google6d agoDatasets – Google Research


