Invest1 publisher3 min readPublished
Google says frontier models already know the facts they get wrong. That is a budget decision.
Google Research reports Gemini-3-Pro and GPT-5 encode 95-98% of tested facts yet fail to recall 26-34% of them, moving the fix from pretraining scale toward post-training and inference.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A post dated August 12, 2026 on research.google was written by Nitay Calderon and Gal Yona, Research Scientists at Google Research.
- The post summarises a paper titled "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality," which introduces knowledge profiling, a behavioral framework that measures both encoding and recall.
- For Gemini-3-Pro and GPT-5, 95-98% of facts are encoded, yet these models still fail to directly recall 26-34% of facts.
- Even with thinking, Gemini-3-Pro and GPT-5 still fail on 11-12% of facts.
- The authors state that encoding failures call for scaling model size or expanding data coverage, while recall failures might also point to post-training and inference-time methods that help LLMs better utilize what they already encode.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Google Research published a post on August 12, 2026 by research scientists Nitay Calderon and Gal Yona summarising a paper called "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality," which introduces a behavioural framework the authors call knowledge profiling [1][2]. Their headline result: on their own benchmark, Gemini-3-Pro and GPT-5 encode 95-98% of facts but still fail to directly recall 26-34% of them [3]. If that generalises, the marginal dollar for factual reliability belongs to post-training and inference plumbing, not to another pretraining run.
The framing matters because standard accuracy scores collapse two different failures into one number [13]. The paper separates encoding (the fact is represented in the weights), recall (the model produces it with no external cue), and recognition (the model picks it out from alternatives) [6], and then classifies each fact into one of five profiles rather than scoring each question [7]. The authors are explicit about the consequence: encoding failures argue for more parameters or more data coverage, while recall failures also open the door to post-training and inference-time methods that help a model use what it already holds [5].
The measurement scaffolding is substantial. WikiProfile contains 2,150 Wikipedia-derived facts, each paired with ten tasks: two for encoding, four for knowledge evaluation, four multiple-choice variants for recognition [8], which is 21,500 questions [1]. Thirteen models were run with and without thinking, eight samples per model, fact and task, producing roughly 4.5 million responses graded by prompted autoraters [9]. That arithmetic checks out: 21,500 questions at eight samples across thirteen models in two modes is about 4.47 million [5].
Now the ratio that should drive planning. If 95-98% of facts are encoded, the encoding gap is 2-5%, while the direct-recall gap is 26-34%, so recall failure is roughly five to seventeen times the size of the problem that scale is supposed to solve [2]. Thinking helps and does not close it: failures drop to 11-12% [4], which removes about 58-65% of the direct-recall failures [3] but leaves a residual still two to six times the encoding gap [4]. Thinking here means eliciting intermediate computation before the answer, including chain-of-thought prompting and thinking-optimised models [12], so that improvement is bought with tokens and latency on every request, not once at training time.
Two caveats before anyone reallocates a roadmap. This is Google evaluating models including its own: of the four frontier systems named in the post, three are Gemini variants [11][6], and the benchmark itself was generated by a prompted Gemini-2.5-Pro with thinking, filtered through a search engine and then manually validated [10]. And the facts are Wikipedia-derived [8], which bounds the claim to encyclopedic knowledge that pretraining corpora cover well. Nothing here says a model has encoded your contract terms or your inventory table, so retrieval over private data remains a separate expense.
What to watch: whether recall-failure rates move in the next round of post-training releases, since that is now the visible dial; whether vendors start pricing or exposing thinking budgets against measured factual recall rather than reasoning benchmarks; whether anyone applies knowledge profiling to a proprietary corpus, where the encoding-versus-recall split is unknown; and the paper's recognition numbers, because architectures that generate candidates and verify them depend on recognition being much cheaper than recall [6].