Skip to content
Evaluation sources

GPQA Diamond

Epoch AI / Knowledge · Graduate-level physics, chemistry and biology knowledge

Official evaluation
In NewsRanked
Evidence budget2%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Only Epoch's own re-run mean_score is used, retaining the reasoning tier and standard error. The 198 science knowledge questions require reasoning but do not represent all world knowledge.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6-astra_maxOpenAI95.8%Max reasoning
2claude-sonnet-5-5_maxanthropic95.6%Max reasoning
3gpt-6.1-sol_max—95.4%Max reasoning
3gemini-3.8-flash_highGoogle95.4%High reasoning
5gemini-3.7-flash_highGoogle94.8%High reasoning
6gpt-5.4-pro-2026-03-05_xhighOpenAI94.6%xHigh reasoning
7gemini-3.1-pro-preview_highGoogle94.4%High reasoning
8gpt-6-sol_maxOpenAI94.3%Max reasoning
9gemini-3.6-flash_highGoogle94.1%High reasoning
14claude-opus-5_maxAnthropic93.9%Max reasoning
15gpt-5.6-sol_maxOpenAI93.5%Max reasoning
16grok-4.5_highxAI93.4%High reasoning
17gpt-5.6-terra_maxOpenAI93.3%Max reasoning
18gpt-5.4-2026-03-05_xhighOpenAI93.3%xHigh reasoning
19grok-4.6_xhighxAI93.2%xHigh reasoning
20kimi-k3_maxMoonshot AI93.1%Max reasoning
22gemini-3.5-flash_highGoogle92.8%High reasoning
23qwen3.8-max_xhighAlibaba92.7%xHigh reasoning
24gemini-3-pro-previewGoogle92.6%Source default configuration
25qwen3.8-max-0902_xhighAlibaba92.3%xHigh reasoning
27glm-5.2_maxZ.ai91.9%Max reasoning
28deepseek-v4-pro-0813_maxDeepSeek91.7%Max reasoning
29gpt-5.6-luna_maxOpenAI91.6%Max reasoning
30gpt-5.2-2025-12-11_xhighOpenAI91.4%xHigh reasoning
31claude-opus-4-8_maxAnthropic91.0%Max reasoning
31deepseek-v4-flash-0731_maxDeepSeek91.0%Max reasoning
33glm-5.3_maxZ.ai90.9%Max reasoning
33MiniMax-M3MiniMax90.9%Source default configuration
33qwen3.7-maxAlibaba90.9%Source default configuration
37kimi-k2.6Moonshot AI90.8%Source default configuration
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.