Skip to content
Evaluation sources

FrontierMath v2 · Tiers 1–3

Epoch AI / Reasoning · Research-level mathematical reasoning

Official evaluation
In NewsRanked
Evidence budget3.6%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Only Tiers 1–3 of v2 are used; older questions and Tier 4 are not mixed into the same metric.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6.1-sol_max—93.7%Max reasoning
1gpt-6-astra_maxOpenAI93.7%Max reasoning
3claude-opus-5-5_maxAnthropic91.2%Max reasoning
4claude-fable-5-1_maxAnthropic90.2%Max reasoning
5gpt-6-sol_maxOpenAI89.8%Max reasoning
6gpt-5.6-sol_maxOpenAI89.1%Max reasoning
7claude-sonnet-5-5_maxanthropic88.8%Max reasoning
8gpt-5.5-pro_xhighOpenAI87.7%xHigh reasoning
9claude-fable-5_maxAnthropic87.0%Max reasoning
10gpt-5.6-terra_maxOpenAI86.0%Max reasoning
11claude-opus-5_maxAnthropic85.6%Max reasoning
12gpt-5.5_xhighOpenAI85.3%xHigh reasoning
13gpt-5.4-pro-2026-03-05_xhighOpenAI82.5%xHigh reasoning
14gpt-5.6-luna_maxOpenAI82.1%Max reasoning
15claude-opus-4-8_maxAnthropic80.0%Max reasoning
16gpt-6-luna_maxOpenAI78.9%Max reasoning
17gpt-5.4-2026-03-05_xhighOpenAI78.6%xHigh reasoning
18qwen3.8-max_xhighAlibaba74.7%xHigh reasoning
20muse-spark-1.3_maxMeta74.0%Max reasoning
21gpt-5.2-pro-2025-12-11_xhighOpenAI74.0%xHigh reasoning
22kimi-k3_maxMoonshot AI72.2%Max reasoning
23gemini-3.7-flash_highGoogle71.6%High reasoning
24claude-opus-4-7_maxAnthropic70.2%Max reasoning
25glm-5.3_maxZ.ai68.8%Max reasoning
26gemini-3.8-flash_highGoogle68.4%High reasoning
27gpt-5.2-2025-12-11_xhighOpenAI67.4%xHigh reasoning
28claude-opus-4-6_maxAnthropic66.0%Max reasoning
28grok-4.6_xhighxAI66.0%xHigh reasoning
30claude-sonnet-5_maxAnthropic65.6%Max reasoning
30qwen3.8-max-0902_xhighAlibaba65.6%xHigh reasoning
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.