Invest1 distinct publisher3 min readUpdated
Google Research reports Gemini-3-Pro and GPT-5 encode 95-98% of tested facts yet fail to recall 26-34% of them, moving the fix from pretraining scale toward post-training and inference.
The Investor · Invest desk
Compiled by The InvestorSomething wrong?How this is made
Google Research published a post on August 12, 2026 by research scientists Nitay Calderon and Gal Yona summarising a paper called "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality," which introduces a behavioural framework the authors call knowledge profiling [1][2]. Their headline result: on their own benchmark, Gemini-3-Pro and GPT-5 encode 95-98% of facts but still fail to directly recall 26-34% of them [3]. If that generalises, the marginal dollar for factual reliability belongs to post-training and inference plumbing, not to another pretraining run.
The framing matters because standard accuracy scores collapse two different failures into one number [13]. The paper separates encoding (the fact is represented in the weights), recall (the model produces it with no external cue), and recognition (the model picks it out from alternatives) [6], and then classifies each fact into one of five profiles rather than scoring each question [7]. The authors are explicit about the consequence: encoding failures argue for more parameters or more data coverage, while recall failures also open the door to post-training and inference-time methods that help a model use what it already holds [5].
The measurement scaffolding is substantial. WikiProfile contains 2,150 Wikipedia-derived facts, each paired with ten tasks: two for encoding, four for knowledge evaluation, four multiple-choice variants for recognition [8], which is 21,500 questions [1]. Thirteen models were run with and without thinking, eight samples per model, fact and task, producing roughly 4.5 million responses graded by prompted autoraters [9]. That arithmetic checks out: 21,500 questions at eight samples across thirteen models in two modes is about 4.47 million [5].
Now the ratio that should drive planning. If 95-98% of facts are encoded, the encoding gap is 2-5%, while the direct-recall gap is 26-34%, so recall failure is roughly five to seventeen times the size of the problem that scale is supposed to solve [2]. Thinking helps and does not close it: failures drop to 11-12% [4], which removes about 58-65% of the direct-recall failures [3] but leaves a residual still two to six times the encoding gap [4]. Thinking here means eliciting intermediate computation before the answer, including chain-of-thought prompting and thinking-optimised models [12], so that improvement is bought with tokens and latency on every request, not once at training time.
Two caveats before anyone reallocates a roadmap. This is Google evaluating models including its own: of the four frontier systems named in the post, three are Gemini variants [11][6], and the benchmark itself was generated by a prompted Gemini-2.5-Pro with thinking, filtered through a search engine and then manually validated [10]. And the facts are Wikipedia-derived [8], which bounds the claim to encyclopedic knowledge that pretraining corpora cover well. Nothing here says a model has encoded your contract terms or your inventory table, so retrieval over private data remains a separate expense.
What to watch: whether recall-failure rates move in the next round of post-training releases, since that is now the visible dial; whether vendors start pricing or exposing thinking budgets against measured factual recall rather than reasoning benchmarks; whether anyone applies knowledge profiling to a proprietary corpus, where the encoding-versus-recall split is unknown; and the paper's recognition numbers, because architectures that generate candidates and verify them depend on recognition being much cheaper than recall [6].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
A post dated August 12, 2026 on research.google was written by Nitay Calderon and Gal Yona, Research Scientists at Google Research.
The post summarises a paper titled "Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality," which introduces knowledge profiling, a behavioral framework that measures both encoding and recall.
For Gemini-3-Pro and GPT-5, 95-98% of facts are encoded, yet these models still fail to directly recall 26-34% of facts.
Even with thinking, Gemini-3-Pro and GPT-5 still fail on 11-12% of facts.
The authors state that encoding failures call for scaling model size or expanding data coverage, while recall failures might also point to post-training and inference-time methods that help LLMs better utilize what they already encode.
The paper uses encoding to denote parametric representation of facts, recall to denote retrieving encoded facts without external cues, and recognition to denote identifying the correct fact when it is presented among alternatives.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party methodology, no external check
The single source is a primary research explainer with an unusually explicit measurement design: named framework, 2,150-fact benchmark, ten tasks per fact, 13 models with and without thinking, eight samples per cell, ~4.5 million graded responses, and internally consistent arithmetic. Against that, everything rests on one self-published post by the lab that built both the benchmark and three of the four frontier systems tested; autorater accuracy, artefact availability and replication are not evidenced in the cluster, and the post text is truncated mid-argument on the reversal-curse section.
No third-party uptake evidence
The only adoption-adjacent event in the cluster is the publisher's own benchmark and internal evaluation run. There is no evidence of external users, downloads, dataset release terms, citations, product changes or deployments applying knowledge profiling, so no adoption level can be measured without inference.
Framing runs ahead of verification
The reported numbers are precise and the paper's own language is hedged ('might also point to' post-training and inference-time methods), but the surrounding framing that the bottleneck has shifted from acquisition to utilisation, and that this is effectively a spend-allocation conclusion, extends beyond what a single vendor benchmark on Wikipedia-derived facts can establish. With no independent replication, no cost data for thinking, and no third-party adoption, the narrative is moderately overstated relative to the evidence rather than fabricated.
Vendor research favouring its own stack
Google Research publishes the framework, builds the benchmark with its own prompted Gemini-2.5-Pro, and evaluates a frontier set in which three of four systems are Gemini variants. The conclusion that models already encode nearly all facts and that remaining errors are utilisation problems is compatible with the publisher's interest in presenting its deployed models as knowledgeable and in promoting thinking-mode inference. Nothing in the cluster indicates fabrication, but the framing and model selection are aligned with the publisher's commercial position.
Single-source, well-specified, unreplicated
Confidence is limited by cluster breadth rather than by internal quality: one publisher, one document, no contradicting or corroborating evidence, and an adoption dimension that cannot be measured at all. The specificity and internal arithmetic consistency of the reported figures support moderate confidence in what was claimed, not in its generality.
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
security
Encrypted injection walks past Grok's filters and out through its own browser1 distinct publisher
build
Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter1 distinct publisher
build
ChatGPT tops Google's paid-click share at 4.75%, and growth teams should reprice the auction1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026