Skip to content
Evaluation sources

LiveBench · reasoning and maths

LiveBench / Reasoning · Average of the Reasoning and Mathematics items, each weighted 50%.

Official evaluation
In NewsRanked
Evidence budget4.2%
Upstream data as of09/30 03:01
Last synced10/01 20:05

What it measures, and how

First calculate the Reasoning and Mathematics category scores separately, then average them at 50% each and keep one decimal place. The two-item average differs from the official site's score when either category is selected alone.

How this evidence is used

The average of the two items participates in the ranking as one piece of evidence, and the dedicated categories share the LiveBench evidence family; the Global Average is for reference only and is not counted twice.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gpt-6.1-sol-max—94.7%Max reasoning
2gpt-6-astra-maxOpenAI94.7%Max reasoning
3claude-opus-5-5-max-effortAnthropic94.6%Max reasoning
4claude-fable-5-1-max-effortAnthropic94.3%Max reasoning
6gpt-5.6-sol-maxOpenAI93.9%Max reasoning
7claude-sonnet-5-5-max-effortanthropic93.9%Max reasoning
9claude-opus-5-max-effortAnthropic93.5%Max reasoning
10claude-fable-5-max-effortAnthropic92.8%Max reasoning
11muse-spark-1.3-xhighMeta92.8%xHigh reasoning
12gpt-5.6-terra-maxOpenAI92.8%Max reasoning
13gpt-5.5-xhighOpenAI92.8%xHigh reasoning
14gpt-6-sol-maxOpenAI92.5%Max reasoning
15claude-opus-4-8-max-effortAnthropic91.8%Max reasoning
17grok-4.6xAI91.5%Source default configuration
18gpt-5.4-xhighOpenAI91.1%xHigh reasoning
19claude-sonnet-5-xhigh-effortAnthropic90.8%xHigh reasoning
20gemini-3.7-flash-highGoogle90.6%High reasoning
21muse-spark-1.2-xhighMeta90.6%xHigh reasoning
22deepseek-v4-pro-0813DeepSeek90.5%Source default configuration
23gemini-3.8-flash-highGoogle90.4%High reasoning
24claude-opus-4-7-xhigh-effortAnthropic90.0%xHigh reasoning
25deepseek-v4.1-flash-maxDeepSeek90.0%Source default configuration
26qwen3.8-maxAlibaba89.8%Source default configuration
27grok-4.7-xhighxAI89.2%xHigh reasoning
28grok-4.5xAI89.0%Source default configuration
29claude-opus-4-6-thinking-auto-high-effortAnthropic89.0%High reasoning
30gpt-5.2-2025-12-11-highOpenAI88.2%High reasoning
31smaug-flashAbacus.AI88.1%Source default configuration
32kimi-k3Moonshot AI87.6%Source default configuration
33gemini-3.1-pro-preview-highGoogle87.5%High reasoning
Limits and data attribution

The average of the two items measures combined performance and must not be treated as a Reasoning single-item score. Scores and rankings change when the official page switches categories or whether fine-tuned models are included. Organisational participant Abacus.AI also develops the Smaug model, and its exclusive scores need to be supplemented by other independent sources.

Data licence: Apache 2.0

Item scores are published by LiveBench, and News aggregates and ranks the above two categories with equal weight. The two-item average and the News consensus score use different scales and cannot be added directly.