Evaluation sourcesOfficial evaluation
FrontierMath v2 · Tiers 1–3
Epoch AI / Reasoning · Research-level mathematical reasoning
In NewsRanked
Evidence budget3.6%
Upstream data as of09/30 06:00
Last synced10/01 20:05
What it measures, and how
Only Tiers 1–3 of v2 are used; older questions and Tier 4 are not mixed into the same metric.
How this evidence is used
Officially scored; the same question family shares a fixed budget.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gpt-6.1-sol_max— | 93.7% | Max reasoning |
| 1 | gpt-6-astra_maxOpenAI | 93.7% | Max reasoning |
| 3 | claude-opus-5-5_maxAnthropic | 91.2% | Max reasoning |
| 4 | claude-fable-5-1_maxAnthropic | 90.2% | Max reasoning |
| 5 | gpt-6-sol_maxOpenAI | 89.8% | Max reasoning |
| 6 | gpt-5.6-sol_maxOpenAI | 89.1% | Max reasoning |
| 7 | claude-sonnet-5-5_maxanthropic | 88.8% | Max reasoning |
| 8 | gpt-5.5-pro_xhighOpenAI | 87.7% | xHigh reasoning |
| 9 | claude-fable-5_maxAnthropic | 87.0% | Max reasoning |
| 10 | gpt-5.6-terra_maxOpenAI | 86.0% | Max reasoning |
| 11 | claude-opus-5_maxAnthropic | 85.6% | Max reasoning |
| 12 | gpt-5.5_xhighOpenAI | 85.3% | xHigh reasoning |
| 13 | gpt-5.4-pro-2026-03-05_xhighOpenAI | 82.5% | xHigh reasoning |
| 14 | gpt-5.6-luna_maxOpenAI | 82.1% | Max reasoning |
| 15 | claude-opus-4-8_maxAnthropic | 80.0% | Max reasoning |
| 16 | gpt-6-luna_maxOpenAI | 78.9% | Max reasoning |
| 17 | gpt-5.4-2026-03-05_xhighOpenAI | 78.6% | xHigh reasoning |
| 18 | qwen3.8-max_xhighAlibaba | 74.7% | xHigh reasoning |
| 20 | muse-spark-1.3_maxMeta | 74.0% | Max reasoning |
| 21 | gpt-5.2-pro-2025-12-11_xhighOpenAI | 74.0% | xHigh reasoning |
| 22 | kimi-k3_maxMoonshot AI | 72.2% | Max reasoning |
| 23 | gemini-3.7-flash_highGoogle | 71.6% | High reasoning |
| 24 | claude-opus-4-7_maxAnthropic | 70.2% | Max reasoning |
| 25 | glm-5.3_maxZ.ai | 68.8% | Max reasoning |
| 26 | gemini-3.8-flash_highGoogle | 68.4% | High reasoning |
| 27 | gpt-5.2-2025-12-11_xhighOpenAI | 67.4% | xHigh reasoning |
| 28 | claude-opus-4-6_maxAnthropic | 66.0% | Max reasoning |
| 28 | grok-4.6_xhighxAI | 66.0% | xHigh reasoning |
| 30 | claude-sonnet-5_maxAnthropic | 65.6% | Max reasoning |
| 30 | qwen3.8-max-0902_xhighAlibaba | 65.6% | xHigh reasoning |
Limits and data attribution
Officially scored; the same question family shares a fixed budget.
Data licence: CC BY 4.0 · Epoch AI's own evaluation data
Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.