Skip to content
Evaluation sources

Chess Puzzles

Epoch AI / Reasoning · Finds the correct move from a written board position

Official evaluation
In NewsRanked
Evidence budget1.8%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Generated chess positions are scored against Stockfish reference answers; the input is chess position text and is not classified as visual understanding.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6-astra_maxOpenAI72.0%Max reasoning
4gemini-3.8-flash_highGoogle61.0%High reasoning
4gpt-6.1-sol_max—61.0%Max reasoning
6gpt-5.4-pro-2026-03-05_xhighOpenAI58.6%xHigh reasoning
7gpt-5.6-sol_maxOpenAI55.0%Max reasoning
9gpt-5.6-terra_maxOpenAI54.0%Max reasoning
11gemini-3.5-flash_highGoogle50.0%High reasoning
12gpt-5.2-2025-12-11_xhighOpenAI49.0%xHigh reasoning
12gemini-3.1-pro-preview_highGoogle49.0%High reasoning
14deepseek-v4-pro-0813_maxDeepSeek47.0%Max reasoning
14gemini-3.7-flash_highGoogle47.0%High reasoning
14claude-fable-5-1_maxAnthropic47.0%Max reasoning
18gpt-5.4-2026-03-05_xhighOpenAI44.0%xHigh reasoning
21claude-opus-5_maxAnthropic42.0%Max reasoning
22claude-fable-5_maxAnthropic41.0%Max reasoning
24gemini-3-flash-preview_highGoogle40.0%High reasoning
24qwen3.8-max-0902_xhighAlibaba40.0%xHigh reasoning
24gemini-3.6-flash_highGoogle40.0%High reasoning
24gpt-5.6-luna_maxOpenAI40.0%Max reasoning
31kimi-k3_maxMoonshot AI39.0%Max reasoning
32muse-spark-1.3_maxMeta38.0%Max reasoning
37gpt-5-2025-08-07_highOpenAI37.0%High reasoning
38grok-4.5_highxAI36.0%High reasoning
42o3-2025-04-16_highOpenAI34.0%High reasoning
42claude-opus-4-8_maxAnthropic34.0%Max reasoning
44deepseek-v4-flash-0731_maxDeepSeek33.0%Max reasoning
46gpt-5.1-2025-11-13_highOpenAI32.0%High reasoning
47gemini-3-pro-previewGoogle31.0%Source default configuration
47grok-4.6_xhighxAI31.0%xHigh reasoning
49gpt-5-mini-2025-08-07_highOpenAI30.0%High reasoning
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.