Arena Text Style-Controlled
LMArena · Humans compare answers anonymously in pairs to see which they would rather read. This adds the overall experience that multiple-choice questions cannot show.
What it measures, and how
Only the Text Overall style-controlled main leaderboard in the official release snapshot is used; the arena.ai live page is not scraped.
How this evidence is used
The human preference budget of 10% is split into 5% overall and 5% creative; it still shares an evidence family with Arena's coding and visual evaluations. Newness is judged by the timestamp of official score data.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gemini-4-argon-highGoogle | 1,524.8 | High reasoning |
| 2 | claude-opus-4-6-highAnthropic | 1,505.5 | High reasoning |
| 3 | claude-fable-5-highAnthropic | 1,504.9 | High reasoning |
| 4 | claude-opus-5.5-highAnthropic | 1,504 | High reasoning |
| 5 | claude-opus-4-7-highAnthropic | 1,501.7 | High reasoning |
| 6 | claude-fable-5.1-maxAnthropic | 1,501.2 | Max reasoning |
| 8 | muse-spark-1.3-maxMeta | 1,494.7 | Max reasoning |
| 10 | muse-spark-1.2 (xHigh)Meta | 1,494.2 | xHigh reasoning |
| 11 | gemini-3.8-flash-highGoogle | 1,494.1 | High reasoning |
| 12 | muse-spark-1.1Meta | 1,491.8 | Source default configuration |
| 14 | claude-opus-5-maxAnthropic | 1,489.1 | Max reasoning |
| 15 | muse-sparkMeta | 1,489.1 | Source default configuration |
| 16 | gemini-3.7-flash-highGoogle | 1,488.3 | High reasoning |
| 17 | kimi-k3-maxMoonshot AI | 1,488 | Source default configuration |
| 18 | gemini-3.1-pro-previewGoogle | 1,487 | Source default configuration |
| 19 | gemini-3-proGoogle | 1,485.5 | Source default configuration |
| 20 | gpt-5.6-sol-xhighOpenAI | 1,483.8 | xHigh reasoning |
| 21 | gemini-3.6-flash-highGoogle | 1,483.1 | High reasoning |
| 22 | gpt-5.5-highOpenAI | 1,481.3 | High reasoning |
| 23 | claude-opus-4-8-highAnthropic | 1,481.3 | High reasoning |
| 24 | qwen3.8-maxAlibaba | 1,480.9 | Source default configuration |
| 25 | mimo-v2.6-proXiaomi | 1,479.6 | Source default configuration |
| 26 | glm-5.3-maxZ.ai | 1,479.4 | Source default configuration |
| 27 | gemini-3.5-flash-highGoogle | 1,477.1 | High reasoning |
| 30 | gpt-6-astra-maxOpenAI | 1,476.2 | Max reasoning |
| 31 | gpt-5.2-chat-latest-20260210OpenAI | 1,475.8 | Source default configuration |
| 32 | glm-5.2-maxZ.ai | 1,475.4 | Source default configuration |
| 33 | gpt-5.4-highOpenAI | 1,475.2 | High reasoning |
| 34 | grok-4.20-beta1xAI | 1,474.6 | Source default configuration |
| 35 | qwen3.7-max-previewAlibaba | 1,474.6 | Source default configuration |
Limits and data attribution
Human preference is shaped by style, user composition and prompt distribution. Overall and creative categories overlap in voting and share the existing human-preference budget, so they are not treated as two independent evaluators.
Data licence: CC BY 4.0
Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.