Arena Vision Style-Controlled
LMArena · Humans blindly select answers with images to observe the actual experience of visual understanding.
What it measures, and how
Only the overall of the official dataset vision_style_control is read; image generation and image editing are not included.
How this evidence is used
Visual budget 6%; this is still human-preference evidence and cannot be equated with strict image accuracy.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | claude-fable-5-highAnthropic | 1,310.1 | High reasoning |
| 2 | qwen3.8-maxAlibaba | 1,300.9 | Source default configuration |
| 3 | claude-opus-4-6-highAnthropic | 1,299.4 | High reasoning |
| 5 | claude-opus-4-7-highAnthropic | 1,298.3 | High reasoning |
| 6 | gemini-3.7-flash-highGoogle | 1,295 | High reasoning |
| 8 | muse-sparkMeta | 1,293.9 | Source default configuration |
| 9 | muse-spark-1.2 (xHigh)Meta | 1,292.4 | xHigh reasoning |
| 10 | muse-spark-1.3-maxMeta | 1,289.6 | Max reasoning |
| 11 | gemini-3-proGoogle | 1,289.1 | Source default configuration |
| 12 | claude-opus-5-highAnthropic | 1,288.8 | High reasoning |
| 13 | gpt-5.6-sol-xhighOpenAI | 1,286.8 | xHigh reasoning |
| 14 | claude-fable-5.1-maxAnthropic | 1,286.6 | Max reasoning |
| 16 | claude-opus-4-8-highAnthropic | 1,286.5 | High reasoning |
| 17 | gemini-3.8-flash-highGoogle | 1,286.3 | High reasoning |
| 18 | gpt-6-astra-maxOpenAI | 1,285.1 | Max reasoning |
| 19 | gemini-3.5-flash-highGoogle | 1,284.5 | High reasoning |
| 20 | gpt-5.4-highOpenAI | 1,284.1 | High reasoning |
| 22 | glm-5.3-flashZ.ai | 1,281.2 | Source default configuration |
| 23 | gpt-5.5-highOpenAI | 1,281 | High reasoning |
| 24 | gemini-3.6-flash-highGoogle | 1,280.1 | High reasoning |
| 25 | gemini-3.1-pro-previewGoogle | 1,279.6 | Source default configuration |
| 26 | muse-spark-1.1Meta | 1,279.5 | Source default configuration |
| 28 | grok-4.5xAI | 1,278.7 | Source default configuration |
| 30 | gpt-5.2-chat-latest-20260210OpenAI | 1,278.3 | Source default configuration |
| 31 | gpt-5.5-instantOpenAI | 1,276.1 | Source default configuration |
| 32 | claude-sonnet-4-6Anthropic | 1,275.4 | Source default configuration |
| 33 | gemini-3-flashGoogle | 1,273 | Source default configuration |
| 34 | gpt-5.6-terra-xhighOpenAI | 1,267.9 | xHigh reasoning |
| 35 | qwen3.7-plusAlibaba | 1,265.8 | Source default configuration |
| 36 | kimi-k2.6Moonshot AI | 1,264.8 | Source default configuration |
Limits and data attribution
Human preference is shaped by style, user composition and prompt distribution. Overall and creative categories overlap in voting and share the existing human-preference budget, so they are not treated as two independent evaluators.
Data licence: CC BY 4.0
Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.