Skip to content
Evaluation sources

Mystery Game Puzzles

Epoch AI / Reasoning · Deriving correct actions from unfamiliar rules

Official evaluation
In NewsRanked
Evidence budget1.2%
Upstream data as of09/30 06:00
Last synced10/01 20:05

What it measures, and how

Only the mean_score of the Epoch standard protocol is used; the restriction that questions are not public is explicitly retained.

How this evidence is used

Officially scored; the same question family shares a fixed budget.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6-astra_maxOpenAI84.0%Max reasoning
2gpt-6.1-sol_max—80.0%Max reasoning
3claude-opus-5-5_maxAnthropic71.0%Max reasoning
4claude-sonnet-5-5_maxanthropic65.0%Max reasoning
5claude-opus-5_maxAnthropic59.0%Max reasoning
6gpt-5.6-sol_maxOpenAI58.0%Max reasoning
6claude-fable-5-1_maxAnthropic58.0%Max reasoning
8gpt-5.5_xhighOpenAI56.0%xHigh reasoning
8gpt-6-sol_maxOpenAI56.0%Max reasoning
10claude-fable-5_maxAnthropic52.0%Max reasoning
12gemini-3.8-flash_highGoogle47.0%High reasoning
13deepseek-v4-pro-0813_maxDeepSeek43.0%Max reasoning
14qwen3.8-max_xhighAlibaba38.0%xHigh reasoning
15gemini-3.7-flash_highGoogle37.0%High reasoning
15gpt-5.4-2026-03-05_xhighOpenAI37.0%xHigh reasoning
18claude-opus-4-8_maxAnthropic36.0%Max reasoning
19claude-sonnet-5_maxAnthropic35.0%Max reasoning
19gpt-5.6-terra_maxOpenAI35.0%Max reasoning
21deepseek-v4-flash-0731_maxDeepSeek34.0%Max reasoning
21gemini-3.1-pro-preview_highGoogle34.0%High reasoning
21grok-4.6_xhighxAI34.0%xHigh reasoning
24glm-5.3_maxZ.ai33.0%Max reasoning
26qwen3.7-maxAlibaba32.0%Source default configuration
26gemini-3.5-flash_highGoogle32.0%High reasoning
30gemini-3.6-flash_highGoogle30.0%High reasoning
31o3-2025-04-16_highOpenAI29.0%High reasoning
33claude-opus-4-7_maxAnthropic28.0%Max reasoning
37kimi-k3_maxMoonshot AI26.0%Max reasoning
41claude-opus-4-6_maxAnthropic25.0%Max reasoning
41muse-spark-1.3_maxMeta25.0%Max reasoning
Limits and data attribution

Officially scored; the same question family shares a fixed budget.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.