Google Research·· 2026-03-17
Testing LLMs on superconductivity research questions
Testing LLMs on superconductivity research questions
AI summary
Google and Cornell published a PNAS study asking GPT-4o, Perplexity, Claude 3.5, Gemini Advanced Pro 1.5, NotebookLM and a custom RAG system 67 expert superconductivity questions. Twelve international experts blindly assessed balance, comprehensiveness, concision, evidence, image relevance and qualitative feedback.
Selection record
Threshold 60Official, first-handFirst 38Second 38
Not admittedSum of both 76 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:评测LLM回答超导研究问题,实质AI研究
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google