Skip to content
Evaluation sources

LiveBench · language and instruction

LiveBench · Average of the Language and IF items, 50% each.

Official evaluation
In NewsRanked
Evidence budget3%
Upstream data as of09/30 03:01
Last synced10/01 20:05

What it measures, and how

First calculate the Language and IF category scores separately, then average them at 50% each and keep one decimal place. The two-item average differs from the official site's score when either category is selected alone. Language measures language understanding and IF measures instruction and format constraints; this is not a basis for judging fiction prose.

How this evidence is used

Keeps 3% of the overall leaderboard's writing and expression budget; not counted in the literary creative writing category.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gemini-3.8-flash-highGoogle84.6%High reasoning
2claude-fable-5-max-effortAnthropic83.2%Max reasoning
3gemini-3.7-flash-highGoogle82.7%High reasoning
4gpt-6-astra-maxOpenAI82.5%Max reasoning
5gemini-3.1-pro-preview-highGoogle82.2%High reasoning
6gpt-6.1-sol-max—82.1%Max reasoning
7claude-fable-5-1-max-effortAnthropic81.2%Max reasoning
8muse-spark-1.3-xhighMeta80.4%xHigh reasoning
9gemini-3.5-flash-highGoogle80.1%High reasoning
11gpt-5.6-sol-maxOpenAI79.8%Max reasoning
12gemini-3.6-flash-highGoogle79.6%High reasoning
13gpt-5.5-xhighOpenAI79.0%xHigh reasoning
14kimi-k3Moonshot AI78.4%Source default configuration
15grok-4.6xAI77.8%Source default configuration
16grok-4.7-xhighxAI77.7%xHigh reasoning
17smaug-agenticAbacus.AI77.7%Source default configuration
18grok-4.5xAI77.2%Source default configuration
20gpt-6-sol-maxOpenAI76.9%Max reasoning
21qwen3.7-maxAlibaba76.9%Source default configuration
22qwen3.8-maxAlibaba76.9%Source default configuration
23muse-spark-1.2-xhighMeta76.4%xHigh reasoning
24gpt-5.4-xhighOpenAI76.4%xHigh reasoning
26claude-opus-5-max-effortAnthropic76.2%Max reasoning
27claude-opus-5-5-max-effortAnthropic76.0%Max reasoning
28qwen3.8-flash-nextAlibaba75.9%Source default configuration
29claude-opus-4-8-max-effortAnthropic75.8%Max reasoning
30deepseek-v4-flash-vision-expDeepSeek75.7%Source default configuration
31deepseek-v4.1-flash-maxDeepSeek75.6%Source default configuration
32smaug-miniAbacus.AI75.5%Source default configuration
33deepseek-v4-pro-0813DeepSeek74.9%Source default configuration
Limits and data attribution

The average of the two items measures combined performance and must not be treated as a Language single-item score. Scores and rankings change when the official page switches categories or whether fine-tuned models are included. Organisational participant Abacus.AI also develops the Smaug model, and its exclusive scores need to be supplemented by other independent sources.

Data licence: Apache 2.0

Item scores are published by LiveBench, and News aggregates and ranks the above two categories with equal weight. The two-item average and the News consensus score use different scales and cannot be added directly.