Skip to content
Evaluation sources

LiveBench · coding overall

LiveBench / Coding · Average of the Coding and Agentic Coding items, 50% each.

Official evaluation
In NewsRanked
Evidence budget2.4%
Upstream data as of09/30 03:01
Last synced10/01 20:05

What it measures, and how

First calculate the Coding and Agentic Coding category scores separately, then average them at 50% each and keep one decimal place. The two-item average differs from the official site's score when either category is selected alone.

How this evidence is used

The average of the two items participates in the ranking as one piece of evidence, and the dedicated categories share the LiveBench evidence family; the Global Average is for reference only and is not counted twice.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1claude-opus-5-5-max-effortAnthropic80.5%Max reasoning
2deepseek-v4.1-flash-maxDeepSeek78.7%Source default configuration
4claude-fable-5-1-max-effortAnthropic76.2%Max reasoning
5claude-fable-5-max-effortAnthropic74.1%Max reasoning
6claude-sonnet-5-5-max-effortanthropic73.8%Max reasoning
7smaug-agenticAbacus.AI73.6%Source default configuration
8claude-opus-5-max-effortAnthropic73.3%Max reasoning
9muse-spark-1.3-xhighMeta72.6%xHigh reasoning
10kimi-k3Moonshot AI71.8%Source default configuration
11gpt-5.6-sol-maxOpenAI70.1%Max reasoning
12claude-sonnet-5-xhigh-effortAnthropic70.0%xHigh reasoning
13glm-5.3Z.ai69.9%Source default configuration
14smaug-flashAbacus.AI69.1%Source default configuration
15gpt-6-astra-maxOpenAI68.8%Max reasoning
16qwen3.8-maxAlibaba68.8%Source default configuration
18gemini-3.7-flash-highGoogle68.6%High reasoning
19qwen3.8-27bAlibaba68.5%Source default configuration
20smaug-miniAbacus.AI68.1%Source default configuration
21gpt-5.5-xhighOpenAI68.1%xHigh reasoning
22glm-5.3-flashZ.ai67.9%Source default configuration
23muse-spark-1.1-xhighMeta67.8%xHigh reasoning
24muse-spark-1.2-xhighMeta67.6%xHigh reasoning
25gpt-6.1-sol-max—67.5%Max reasoning
26gpt-6-sol-maxOpenAI67.3%Max reasoning
27qwen3.8-flash-nextAlibaba67.1%Source default configuration
28grok-4.6xAI66.9%Source default configuration
29deepseek-v4-flash-vision-expDeepSeek66.7%Source default configuration
30gpt-5.6-terra-maxOpenAI66.6%Max reasoning
31gpt-5.2-codexOpenAI66.5%Source default configuration
32claude-opus-4-7-xhigh-effortAnthropic66.4%xHigh reasoning
Limits and data attribution

The average of the two items measures combined performance and must not be treated as a Coding single-item score. Scores and rankings change when the official page switches categories or whether fine-tuned models are included. Organisational participant Abacus.AI also develops the Smaug model, and its exclusive scores need to be supplemented by other independent sources.

Data licence: Apache 2.0

Item scores are published by LiveBench, and News aggregates and ranks the above two categories with equal weight. The two-item average and the News consensus score use different scales and cannot be added directly.