Skip to content
Evaluation sources

Arena Text Style-Controlled

LMArena · Humans compare answers anonymously in pairs to see which they would rather read. This adds the overall experience that multiple-choice questions cannot show.

Official evaluation
In NewsRanked
Evidence budget5%
Upstream data as of09/30 08:00
Last synced10/01 20:05

What it measures, and how

Only the Text Overall style-controlled main leaderboard in the official release snapshot is used; the arena.ai live page is not scraped.

How this evidence is used

The human preference budget of 10% is split into 5% overall and 5% creative; it still shares an evidence family with Arena's coding and visual evaluations. Newness is judged by the timestamp of official score data.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gemini-4-argon-highGoogle1,524.8High reasoning
2claude-opus-4-6-highAnthropic1,505.5High reasoning
3claude-fable-5-highAnthropic1,504.9High reasoning
4claude-opus-5.5-highAnthropic1,504High reasoning
5claude-opus-4-7-highAnthropic1,501.7High reasoning
6claude-fable-5.1-maxAnthropic1,501.2Max reasoning
8muse-spark-1.3-maxMeta1,494.7Max reasoning
10muse-spark-1.2 (xHigh)Meta1,494.2xHigh reasoning
11gemini-3.8-flash-highGoogle1,494.1High reasoning
12muse-spark-1.1Meta1,491.8Source default configuration
14claude-opus-5-maxAnthropic1,489.1Max reasoning
15muse-sparkMeta1,489.1Source default configuration
16gemini-3.7-flash-highGoogle1,488.3High reasoning
17kimi-k3-maxMoonshot AI1,488Source default configuration
18gemini-3.1-pro-previewGoogle1,487Source default configuration
19gemini-3-proGoogle1,485.5Source default configuration
20gpt-5.6-sol-xhighOpenAI1,483.8xHigh reasoning
21gemini-3.6-flash-highGoogle1,483.1High reasoning
22gpt-5.5-highOpenAI1,481.3High reasoning
23claude-opus-4-8-highAnthropic1,481.3High reasoning
24qwen3.8-maxAlibaba1,480.9Source default configuration
25mimo-v2.6-proXiaomi1,479.6Source default configuration
26glm-5.3-maxZ.ai1,479.4Source default configuration
27gemini-3.5-flash-highGoogle1,477.1High reasoning
30gpt-6-astra-maxOpenAI1,476.2Max reasoning
31gpt-5.2-chat-latest-20260210OpenAI1,475.8Source default configuration
32glm-5.2-maxZ.ai1,475.4Source default configuration
33gpt-5.4-highOpenAI1,475.2High reasoning
34grok-4.20-beta1xAI1,474.6Source default configuration
35qwen3.7-max-previewAlibaba1,474.6Source default configuration
Limits and data attribution

Human preference is shaped by style, user composition and prompt distribution. Overall and creative categories overlap in voting and share the existing human-preference budget, so they are not treated as two independent evaluators.

Data licence: CC BY 4.0

Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.