LiveBench · language and instruction
LiveBench · Average of the Language and IF items, 50% each.
What it measures, and how
First calculate the Language and IF category scores separately, then average them at 50% each and keep one decimal place. The two-item average differs from the official site's score when either category is selected alone. Language measures language understanding and IF measures instruction and format constraints; this is not a basis for judging fiction prose.
How this evidence is used
Keeps 3% of the overall leaderboard's writing and expression budget; not counted in the literary creative writing category.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gemini-3.8-flash-highGoogle | 84.6% | High reasoning |
| 2 | claude-fable-5-max-effortAnthropic | 83.2% | Max reasoning |
| 3 | gemini-3.7-flash-highGoogle | 82.7% | High reasoning |
| 4 | gpt-6-astra-maxOpenAI | 82.5% | Max reasoning |
| 5 | gemini-3.1-pro-preview-highGoogle | 82.2% | High reasoning |
| 6 | gpt-6.1-sol-max— | 82.1% | Max reasoning |
| 7 | claude-fable-5-1-max-effortAnthropic | 81.2% | Max reasoning |
| 8 | muse-spark-1.3-xhighMeta | 80.4% | xHigh reasoning |
| 9 | gemini-3.5-flash-highGoogle | 80.1% | High reasoning |
| 11 | gpt-5.6-sol-maxOpenAI | 79.8% | Max reasoning |
| 12 | gemini-3.6-flash-highGoogle | 79.6% | High reasoning |
| 13 | gpt-5.5-xhighOpenAI | 79.0% | xHigh reasoning |
| 14 | kimi-k3Moonshot AI | 78.4% | Source default configuration |
| 15 | grok-4.6xAI | 77.8% | Source default configuration |
| 16 | grok-4.7-xhighxAI | 77.7% | xHigh reasoning |
| 17 | smaug-agenticAbacus.AI | 77.7% | Source default configuration |
| 18 | grok-4.5xAI | 77.2% | Source default configuration |
| 20 | gpt-6-sol-maxOpenAI | 76.9% | Max reasoning |
| 21 | qwen3.7-maxAlibaba | 76.9% | Source default configuration |
| 22 | qwen3.8-maxAlibaba | 76.9% | Source default configuration |
| 23 | muse-spark-1.2-xhighMeta | 76.4% | xHigh reasoning |
| 24 | gpt-5.4-xhighOpenAI | 76.4% | xHigh reasoning |
| 26 | claude-opus-5-max-effortAnthropic | 76.2% | Max reasoning |
| 27 | claude-opus-5-5-max-effortAnthropic | 76.0% | Max reasoning |
| 28 | qwen3.8-flash-nextAlibaba | 75.9% | Source default configuration |
| 29 | claude-opus-4-8-max-effortAnthropic | 75.8% | Max reasoning |
| 30 | deepseek-v4-flash-vision-expDeepSeek | 75.7% | Source default configuration |
| 31 | deepseek-v4.1-flash-maxDeepSeek | 75.6% | Source default configuration |
| 32 | smaug-miniAbacus.AI | 75.5% | Source default configuration |
| 33 | deepseek-v4-pro-0813DeepSeek | 74.9% | Source default configuration |
Limits and data attribution
The average of the two items measures combined performance and must not be treated as a Language single-item score. Scores and rankings change when the official page switches categories or whether fine-tuned models are included. Organisational participant Abacus.AI also develops the Smaug model, and its exclusive scores need to be supplemented by other independent sources.
Data licence: Apache 2.0
Item scores are published by LiveBench, and News aggregates and ranks the above two categories with equal weight. The two-item average and the News consensus score use different scales and cannot be added directly.