Evaluation sourcesOfficial evaluation
EBR-bench
Epoch AI / Professional work · Learn new rules and use memory during long tasks
In NewsReference only
Evidence budgetNot scored
Upstream data as of09/30 02:13
Last synced10/01 20:05
What it measures, and how
The public aggregate mixes memory-compression configurations; it is for reference only until comparable scores under a fixed configuration are obtained.
How this evidence is used
For reference only; not counted towards overall or category scores.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gpt-6-astra_maxOpenAI | 76.2% | Max reasoning |
| 2 | claude-opus-5-5_maxAnthropic | 71.4% | Max reasoning |
| 3 | claude-fable-5-1_maxAnthropic | 57.1% | Max reasoning |
| 4 | gpt-6.1-sol_max— | 54.3% | Max reasoning |
| 5 | gpt-6-sol_maxOpenAI | 53.3% | Max reasoning |
| 6 | claude-opus-5_maxAnthropic | 45.7% | Max reasoning |
| 7 | gpt-5.6-sol_maxOpenAI | 44.8% | Max reasoning |
| 8 | claude-fable-5_maxAnthropic | 39.5% | Max reasoning |
| 9 | gpt-5.5_xhighOpenAI | 34.3% | xHigh reasoning |
| 10 | grok-4.6_xhighxAI | 30.5% | xHigh reasoning |
| 11 | claude-opus-4-8_maxAnthropic | 28.6% | Max reasoning |
| 12 | gpt-5.4-2026-03-05_xhighOpenAI | 25.4% | xHigh reasoning |
| 13 | gpt-5.2-2025-12-11_xhighOpenAI | 23.0% | xHigh reasoning |
| 14 | claude-opus-4-7_maxAnthropic | 19.0% | Max reasoning |
| 15 | gemini-3.1-pro-previewGoogle | 14.3% | Source default configuration |
| 15 | claude-opus-4-5-20251101_128KAnthropic | 14.3% | Source default configuration · 128k |
| 17 | gpt-5-2025-08-07_highOpenAI | 12.7% | High reasoning |
| 17 | claude-opus-4-6_maxAnthropic | 12.7% | Max reasoning |
| 19 | qwen3.7-maxAlibaba | 9.5% | Source default configuration |
| 19 | glm-5.2_maxZ.ai | 9.5% | Max reasoning |
| 21 | claude-opus-4-1-20250805Anthropic | 7.9% | Source default configuration |
| 22 | gemini-3.5-flash_highGoogle | 4.8% | High reasoning |
| 23 | kimi-k2.6Moonshot AI | 2.4% | Source default configuration |
| 23 | claude-sonnet-4-5-20250929Anthropic | 2.4% | Source default configuration |
Limits and data attribution
For reference only; not counted towards overall or category scores.
Data licence: CC BY 4.0 · Epoch AI's own evaluation data
Scores published by Epoch AI. Raw scores and the News consensus score use different scales and cannot be added directly.