Skip to content
Evaluation sources

FrontierMath v2 · Tier 4

Epoch AI / Reasoning · Harder mathematical research problems

Official evaluation
In NewsRanked
Evidence budget1.2%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Tier 4 is kept separate and shares the mathematical evidence budget with Tiers 1–3; sampling error is retained.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6.1-sol_max—100.0%Max reasoning
2gpt-6-astra_maxOpenAI97.6%Max reasoning
6claude-opus-5-5_maxAnthropic95.0%Max reasoning
7claude-fable-5_maxAnthropic90.2%Max reasoning
8gpt-6-sol_maxOpenAI90.0%Max reasoning
9claude-fable-5-1_maxAnthropic87.8%Max reasoning
11gpt-5.6-sol_maxOpenAI82.9%Max reasoning
13claude-sonnet-5-5_maxanthropic80.5%Max reasoning
15gpt-5.5-pro_xhighOpenAI78.0%xHigh reasoning
17claude-opus-5_maxAnthropic73.2%Max reasoning
18gpt-5.5_xhighOpenAI72.5%xHigh reasoning
19gpt-5.6-terra_maxOpenAI70.7%Max reasoning
20gpt-5.6-luna_maxOpenAI61.0%Max reasoning
21gpt-5.4-pro-2026-03-05_xhighOpenAI58.5%xHigh reasoning
22gpt-6-luna_maxOpenAI56.1%Max reasoning
22claude-opus-4-8_maxAnthropic56.1%Max reasoning
24gpt-5.4-2026-03-05_xhighOpenAI49.0%xHigh reasoning
25muse-spark-1.3_maxMeta46.3%Max reasoning
25qwen3.8-max_xhighAlibaba46.3%xHigh reasoning
27gpt-5.2-pro-2025-12-11_xhighOpenAI46.0%xHigh reasoning
29kimi-k3_maxMoonshot AI39.0%Max reasoning
30gemini-3.7-flash_highGoogle36.6%High reasoning
31qwen3.8-max-0902_xhighAlibaba34.1%xHigh reasoning
31qwen3.7-maxAlibaba34.1%Source default configuration
33claude-opus-4-7_maxAnthropic31.7%Max reasoning
33grok-4.6_xhighxAI31.7%xHigh reasoning
35gpt-5.2-2025-12-11_xhighOpenAI31.7%xHigh reasoning
36claude-sonnet-5_maxAnthropic29.3%Max reasoning
36glm-5.3_maxZ.ai29.3%Max reasoning
36glm-5.2_maxZ.ai29.3%Max reasoning
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.