LiveBench Global Average
LiveBench · The question bank is continually replaced and answers can be checked automatically. Used to see whether a model still holds up after the questions change.
What it measures, and how
Keeps the official Global Average as an overall reference; official scores use non-overlapping specialist tasks.
How this evidence is used
The original Global Average is kept for reference; official votes switch to non-overlapping specialist tasks.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | claude-fable-5-1-max-effortAnthropic | 83.4% | Max reasoning |
| 2 | claude-opus-5-5-max-effortAnthropic | 83.2% | Max reasoning |
| 3 | claude-fable-5-max-effortAnthropic | 83.0% | Max reasoning |
| 4 | gpt-6-astra-maxOpenAI | 82.2% | Max reasoning |
| 6 | gpt-6.1-sol-max— | 81.6% | Max reasoning |
| 7 | muse-spark-1.3-xhighMeta | 81.6% | xHigh reasoning |
| 9 | deepseek-v4.1-flash-maxDeepSeek | 81.1% | Source default configuration |
| 10 | gpt-5.6-sol-maxOpenAI | 81.1% | Max reasoning |
| 11 | gpt-5.5-xhighOpenAI | 80.2% | xHigh reasoning |
| 12 | claude-opus-5-max-effortAnthropic | 80.1% | Max reasoning |
| 13 | smaug-agenticAbacus.AI | 79.5% | Source default configuration |
| 14 | gpt-6-sol-maxOpenAI | 79.2% | Max reasoning |
| 15 | kimi-k3Moonshot AI | 79.2% | Source default configuration |
| 16 | gemini-3.7-flash-highGoogle | 78.8% | High reasoning |
| 17 | qwen3.8-maxAlibaba | 78.5% | Source default configuration |
| 18 | grok-4.6xAI | 78.0% | Source default configuration |
| 19 | gpt-5.4-xhighOpenAI | 78.0% | xHigh reasoning |
| 20 | muse-spark-1.2-xhighMeta | 78.0% | xHigh reasoning |
| 21 | gpt-5.6-terra-maxOpenAI | 77.9% | Max reasoning |
| 23 | deepseek-v4-pro-0813DeepSeek | 77.4% | Source default configuration |
| 24 | smaug-flashAbacus.AI | 77.4% | Source default configuration |
| 25 | grok-4.7-xhighxAI | 77.4% | xHigh reasoning |
| 26 | gemini-3.1-pro-preview-highGoogle | 77.0% | High reasoning |
| 27 | smaug-miniAbacus.AI | 76.9% | Source default configuration |
| 28 | deepseek-v4-flash-vision-expDeepSeek | 76.8% | Source default configuration |
| 29 | claude-opus-4-7-xhigh-effortAnthropic | 76.5% | xHigh reasoning |
| 30 | claude-opus-4-8-max-effortAnthropic | 76.2% | Max reasoning |
| 31 | qwen3.8-flash-nextAlibaba | 76.2% | Source default configuration |
| 32 | glm-5.3Z.ai | 76.1% | Source default configuration |
| 33 | claude-sonnet-5-xhigh-effortAnthropic | 76.0% | xHigh reasoning |
Limits and data attribution
The overall average dilutes changes in specific abilities; it is for cross-reference only and is not double-counted. Organisational participant Abacus.AI also develops the Smaug model, so its exclusive results need supplementing from other independent sources.
Data licence: Apache 2.0
Scores published by LiveBench. Raw scores and the News consensus score use different scales and cannot be added directly.