LiveBench · reasoning and maths
LiveBench / Reasoning · Average of the Reasoning and Mathematics items, each weighted 50%.
What it measures, and how
First calculate the Reasoning and Mathematics category scores separately, then average them at 50% each and keep one decimal place. The two-item average differs from the official site's score when either category is selected alone.
How this evidence is used
The average of the two items participates in the ranking as one piece of evidence, and the dedicated categories share the LiveBench evidence family; the Global Average is for reference only and is not counted twice.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gpt-6.1-sol-max— | 94.7% | Max reasoning |
| 2 | gpt-6-astra-maxOpenAI | 94.7% | Max reasoning |
| 3 | claude-opus-5-5-max-effortAnthropic | 94.6% | Max reasoning |
| 4 | claude-fable-5-1-max-effortAnthropic | 94.3% | Max reasoning |
| 6 | gpt-5.6-sol-maxOpenAI | 93.9% | Max reasoning |
| 7 | claude-sonnet-5-5-max-effortanthropic | 93.9% | Max reasoning |
| 9 | claude-opus-5-max-effortAnthropic | 93.5% | Max reasoning |
| 10 | claude-fable-5-max-effortAnthropic | 92.8% | Max reasoning |
| 11 | muse-spark-1.3-xhighMeta | 92.8% | xHigh reasoning |
| 12 | gpt-5.6-terra-maxOpenAI | 92.8% | Max reasoning |
| 13 | gpt-5.5-xhighOpenAI | 92.8% | xHigh reasoning |
| 14 | gpt-6-sol-maxOpenAI | 92.5% | Max reasoning |
| 15 | claude-opus-4-8-max-effortAnthropic | 91.8% | Max reasoning |
| 17 | grok-4.6xAI | 91.5% | Source default configuration |
| 18 | gpt-5.4-xhighOpenAI | 91.1% | xHigh reasoning |
| 19 | claude-sonnet-5-xhigh-effortAnthropic | 90.8% | xHigh reasoning |
| 20 | gemini-3.7-flash-highGoogle | 90.6% | High reasoning |
| 21 | muse-spark-1.2-xhighMeta | 90.6% | xHigh reasoning |
| 22 | deepseek-v4-pro-0813DeepSeek | 90.5% | Source default configuration |
| 23 | gemini-3.8-flash-highGoogle | 90.4% | High reasoning |
| 24 | claude-opus-4-7-xhigh-effortAnthropic | 90.0% | xHigh reasoning |
| 25 | deepseek-v4.1-flash-maxDeepSeek | 90.0% | Source default configuration |
| 26 | qwen3.8-maxAlibaba | 89.8% | Source default configuration |
| 27 | grok-4.7-xhighxAI | 89.2% | xHigh reasoning |
| 28 | grok-4.5xAI | 89.0% | Source default configuration |
| 29 | claude-opus-4-6-thinking-auto-high-effortAnthropic | 89.0% | High reasoning |
| 30 | gpt-5.2-2025-12-11-highOpenAI | 88.2% | High reasoning |
| 31 | smaug-flashAbacus.AI | 88.1% | Source default configuration |
| 32 | kimi-k3Moonshot AI | 87.6% | Source default configuration |
| 33 | gemini-3.1-pro-preview-highGoogle | 87.5% | High reasoning |
Limits and data attribution
The average of the two items measures combined performance and must not be treated as a Reasoning single-item score. Scores and rankings change when the official page switches categories or whether fine-tuned models are included. Organisational participant Abacus.AI also develops the Smaug model, and its exclusive scores need to be supplemented by other independent sources.
Data licence: Apache 2.0
Item scores are published by LiveBench, and News aggregates and ranks the above two categories with equal weight. The two-item average and the News consensus score use different scales and cannot be added directly.