Skip to content
Evaluation sources

Arena WebDev

LMArena · Also a human blind choice, it compares whether web pages are usable and deliverable.

Official evaluation
In NewsRanked
Evidence budget2.4%
Upstream data as of09/30 08:00
Last synced10/01 20:05

What it measures, and how

Only the WebDev overall leaderboard in the official release snapshot is used; the arena.ai live page is not scraped. Together with Text it forms one Arena family vote.

How this evidence is used

2.4% of the coding and design budget; measures preference for web work in the aesthetics category, not equivalent to code correctness. Still shares an evidence family with other Arena tasks.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1claude-opus-5.5-maxAnthropic1,817.8Max reasoning
2gpt-6-astra-maxOpenAI1,789.1Max reasoning
3gpt-6.1-sol-max—1,758.7Max reasoning
4claude-fable-5.1-maxAnthropic1,750.7Max reasoning
5claude-sonnet-5.5-highanthropic1,708.7High reasoning
6claude-opus-5-maxAnthropic1,693.7Max reasoning
7gpt-6-sol-maxOpenAI1,689.3Max reasoning
8gemini-4-argon-highGoogle1,679High reasoning
9qwen3.8-maxAlibaba1,671.2Source default configuration
10qwen3.8-max-0902Alibaba1,669.7Source default configuration
12kimi-k3-maxMoonshot AI1,657.8Source default configuration
13muse-spark-1.3-maxMeta1,655.4Max reasoning
14qwen3.8-flash-nextAlibaba1,637.8Source default configuration
15grok-4.7-xhighxAI1,636.3xHigh reasoning
16hy4-previewTencent1,633.2Source default configuration
17claude-fable-5-highAnthropic1,625.6High reasoning
19glm-5.3-maxZ.ai1,622Source default configuration
20grok-4.6-highxAI1,620.1High reasoning
21deepseek-v4.1-flash-maxDeepSeek1,620Source default configuration
22gpt-5.6-sol-xhigh (codex-harness)OpenAI1,619.3xHigh reasoning · codex-harness
23mimo-v2.6-proXiaomi1,618.5Source default configuration
24glm-5.3-flashZ.ai1,614.8Source default configuration
25glm-5.2-maxZ.ai1,604.6Source default configuration
26gemini-3.7-flash-highGoogle1,592.2High reasoning
27qwen3.8-27bAlibaba1,590.4Source default configuration
28gemini-3.8-flash-highGoogle1,582.7High reasoning
29deepseek-v4-pro-high-20260813DeepSeek1,582.3Source default configuration
30gpt-6-luna-maxOpenAI1,581.9Max reasoning
31deepseek-v4-flash-highDeepSeek1,581Source default configuration
32step-5-preview-highStepFun1,569.6Source default configuration
Limits and data attribution

Correlates with the Text preference construct.

Data licence: CC BY 4.0

Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.