Evaluation sourcesOfficial evaluation
MirrorCode
Epoch AI / Coding · Long-horizon software reproduction
In NewsReference only
Evidence budgetNot scored
Upstream data as of09/23 00:42
Last synced10/01 20:05
What it measures, and how
The published data covers only a few models and run times are very long; the sample and run conditions are insufficient to form a classification consensus.
How this evidence is used
For reference only; not counted towards overall or category scores.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | claude-opus-5-5_maxAnthropic | 77.4% | Max reasoning |
| 2 | claude-fable-5-1_highAnthropic | 73.3% | High reasoning |
| 3 | claude-fable-5_highAnthropic | 63.9% | High reasoning |
| 4 | gpt-6-astra_highOpenAI | 46.7% | High reasoning |
| 5 | claude-opus-4-7_highAnthropic | 31.1% | High reasoning |
| 6 | gpt-5.6-sol_highOpenAI | 20.0% | High reasoning |
| 7 | gpt-5.4-2026-03-05_highOpenAI | 15.6% | High reasoning |
| 8 | gpt-5.5_highOpenAI | 10.0% | High reasoning |
| 9 | gemini-3.1-pro-preview_highGoogle | 8.9% | High reasoning |
Limits and data attribution
For reference only; not counted towards overall or category scores.
Data licence: CC BY 4.0 · Epoch AI's own evaluation data
Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.