Skip to content
Evaluation sources

Arena Vision Style-Controlled

LMArena · Humans blindly select answers with images to observe the actual experience of visual understanding.

Official evaluation
In NewsRanked
Evidence budget6%
Upstream data as of09/28 08:00
Last synced10/01 20:05

What it measures, and how

Only the overall of the official dataset vision_style_control is read; image generation and image editing are not included.

How this evidence is used

Visual budget 6%; this is still human-preference evidence and cannot be equated with strict image accuracy.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1claude-fable-5-highAnthropic1,310.1High reasoning
2qwen3.8-maxAlibaba1,300.9Source default configuration
3claude-opus-4-6-highAnthropic1,299.4High reasoning
5claude-opus-4-7-highAnthropic1,298.3High reasoning
6gemini-3.7-flash-highGoogle1,295High reasoning
8muse-sparkMeta1,293.9Source default configuration
9muse-spark-1.2 (xHigh)Meta1,292.4xHigh reasoning
10muse-spark-1.3-maxMeta1,289.6Max reasoning
11gemini-3-proGoogle1,289.1Source default configuration
12claude-opus-5-highAnthropic1,288.8High reasoning
13gpt-5.6-sol-xhighOpenAI1,286.8xHigh reasoning
14claude-fable-5.1-maxAnthropic1,286.6Max reasoning
16claude-opus-4-8-highAnthropic1,286.5High reasoning
17gemini-3.8-flash-highGoogle1,286.3High reasoning
18gpt-6-astra-maxOpenAI1,285.1Max reasoning
19gemini-3.5-flash-highGoogle1,284.5High reasoning
20gpt-5.4-highOpenAI1,284.1High reasoning
22glm-5.3-flashZ.ai1,281.2Source default configuration
23gpt-5.5-highOpenAI1,281High reasoning
24gemini-3.6-flash-highGoogle1,280.1High reasoning
25gemini-3.1-pro-previewGoogle1,279.6Source default configuration
26muse-spark-1.1Meta1,279.5Source default configuration
28grok-4.5xAI1,278.7Source default configuration
30gpt-5.2-chat-latest-20260210OpenAI1,278.3Source default configuration
31gpt-5.5-instantOpenAI1,276.1Source default configuration
32claude-sonnet-4-6Anthropic1,275.4Source default configuration
33gemini-3-flashGoogle1,273Source default configuration
34gpt-5.6-terra-xhighOpenAI1,267.9xHigh reasoning
35qwen3.7-plusAlibaba1,265.8Source default configuration
36kimi-k2.6Moonshot AI1,264.8Source default configuration
Limits and data attribution

Human preference is shaped by style, user composition and prompt distribution. Overall and creative categories overlap in voting and share the existing human-preference budget, so they are not treated as two independent evaluators.

Data licence: CC BY 4.0

Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.