Skip to content
Evaluation sources

SimpleQA Verified

Epoch AI / Knowledge · Factual accuracy on short questions

Official evaluation
In NewsRanked
Evidence budget4%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Only Epoch's own re-run mean_score is used, with a fixed reasoning tier; the ECI and Best score aggregate columns are not read.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6-astra_maxOpenAI75.6%Max reasoning
2gpt-6.1-sol_max—73.9%Max reasoning
3gemini-3.1-pro-preview_highGoogle73.5%High reasoning
4claude-opus-5-5_maxAnthropic72.2%Max reasoning
5claude-fable-5-1_maxAnthropic70.8%Max reasoning
6claude-fable-5_xhighAnthropic70.7%xHigh reasoning
7gemini-3.8-flash_highGoogle69.7%High reasoning
7gpt-5.6-sol_maxOpenAI69.7%Max reasoning
9gemini-3.7-flash_highGoogle69.2%High reasoning
10gemini-3-flash-preview_highGoogle66.8%High reasoning
11gemini-3.6-flash_highGoogle66.2%High reasoning
11gemini-3.5-flash_highGoogle66.2%High reasoning
13gpt-5.5_xhighOpenAI63.0%xHigh reasoning
14gpt-6-sol_maxOpenAI60.7%Max reasoning
15muse-spark-1.2_xhighMeta60.3%xHigh reasoning
16claude-opus-5_maxAnthropic59.9%Max reasoning
17muse-spark-1.1Meta57.8%Source default configuration
18qwen3.7-maxAlibaba55.8%Source default configuration
19claude-opus-4-8_maxAnthropic53.0%Max reasoning
20deepseek-v4-pro-0813_maxDeepSeek52.9%Max reasoning
21qwen3.6-max-previewAlibaba52.0%Source default configuration
22claude-opus-4-7_xhighAnthropic51.7%xHigh reasoning
23kimi-k3_maxMoonshot AI50.6%Max reasoning
24gpt-5-2025-08-07_highOpenAI50.1%High reasoning
25o3-2025-04-16_highOpenAI49.4%High reasoning
27grok-4.6_xhighxAI48.9%xHigh reasoning
28qwen3-max-2025-09-23Alibaba48.7%Source default configuration
29grok-4.5_highxAI48.3%High reasoning
30gpt-5.1-2025-11-13_highOpenAI48.0%High reasoning
31qwen3.8-max-0902_xhighAlibaba47.3%xHigh reasoning
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.