Skip to content
Google Research·· 2026-08-12

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

AI summary

Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.

Selection record

Not admittedSum of both 76 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:研究LLM事实编码与召回瓶颈

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Google Research · research.google