Skip to content
Evaluation sources

MirrorCode

Epoch AI / Coding · Long-horizon software reproduction

Official evaluation
In NewsReference only
Evidence budgetNot scored
Upstream data as of09/23 00:42
Last synced10/01 20:05

What it measures, and how

The published data covers only a few models and run times are very long; the sample and run conditions are insufficient to form a classification consensus.

How this evidence is used

For reference only; not counted towards overall or category scores.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1claude-opus-5-5_maxAnthropic77.4%Max reasoning
2claude-fable-5-1_highAnthropic73.3%High reasoning
3claude-fable-5_highAnthropic63.9%High reasoning
4gpt-6-astra_highOpenAI46.7%High reasoning
5claude-opus-4-7_highAnthropic31.1%High reasoning
6gpt-5.6-sol_highOpenAI20.0%High reasoning
7gpt-5.4-2026-03-05_highOpenAI15.6%High reasoning
8gpt-5.5_highOpenAI10.0%High reasoning
9gemini-3.1-pro-preview_highGoogle8.9%High reasoning
Limits and data attribution

For reference only; not counted towards overall or category scores.

Data licence: CC BY 4.0 · Epoch AI's own evaluation data

Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.