Google Research·· 2026-08-12
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
AI summary
Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.
Selection record
Threshold 60Official, first-handFirst 38Second 38
Not admittedSum of both 76 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:研究LLM事实编码与召回瓶颈
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google