Arena WebDev
LMArena · Also a human blind choice, it compares whether web pages are usable and deliverable.
What it measures, and how
Only the WebDev overall leaderboard in the official release snapshot is used; the arena.ai live page is not scraped. Together with Text it forms one Arena family vote.
How this evidence is used
2.4% of the coding and design budget; measures preference for web work in the aesthetics category, not equivalent to code correctness. Still shares an evidence family with other Arena tasks.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | claude-opus-5.5-maxAnthropic | 1,817.8 | Max reasoning |
| 2 | gpt-6-astra-maxOpenAI | 1,789.1 | Max reasoning |
| 3 | gpt-6.1-sol-max— | 1,758.7 | Max reasoning |
| 4 | claude-fable-5.1-maxAnthropic | 1,750.7 | Max reasoning |
| 5 | claude-sonnet-5.5-highanthropic | 1,708.7 | High reasoning |
| 6 | claude-opus-5-maxAnthropic | 1,693.7 | Max reasoning |
| 7 | gpt-6-sol-maxOpenAI | 1,689.3 | Max reasoning |
| 8 | gemini-4-argon-highGoogle | 1,679 | High reasoning |
| 9 | qwen3.8-maxAlibaba | 1,671.2 | Source default configuration |
| 10 | qwen3.8-max-0902Alibaba | 1,669.7 | Source default configuration |
| 12 | kimi-k3-maxMoonshot AI | 1,657.8 | Source default configuration |
| 13 | muse-spark-1.3-maxMeta | 1,655.4 | Max reasoning |
| 14 | qwen3.8-flash-nextAlibaba | 1,637.8 | Source default configuration |
| 15 | grok-4.7-xhighxAI | 1,636.3 | xHigh reasoning |
| 16 | hy4-previewTencent | 1,633.2 | Source default configuration |
| 17 | claude-fable-5-highAnthropic | 1,625.6 | High reasoning |
| 19 | glm-5.3-maxZ.ai | 1,622 | Source default configuration |
| 20 | grok-4.6-highxAI | 1,620.1 | High reasoning |
| 21 | deepseek-v4.1-flash-maxDeepSeek | 1,620 | Source default configuration |
| 22 | gpt-5.6-sol-xhigh (codex-harness)OpenAI | 1,619.3 | xHigh reasoning · codex-harness |
| 23 | mimo-v2.6-proXiaomi | 1,618.5 | Source default configuration |
| 24 | glm-5.3-flashZ.ai | 1,614.8 | Source default configuration |
| 25 | glm-5.2-maxZ.ai | 1,604.6 | Source default configuration |
| 26 | gemini-3.7-flash-highGoogle | 1,592.2 | High reasoning |
| 27 | qwen3.8-27bAlibaba | 1,590.4 | Source default configuration |
| 28 | gemini-3.8-flash-highGoogle | 1,582.7 | High reasoning |
| 29 | deepseek-v4-pro-high-20260813DeepSeek | 1,582.3 | Source default configuration |
| 30 | gpt-6-luna-maxOpenAI | 1,581.9 | Max reasoning |
| 31 | deepseek-v4-flash-highDeepSeek | 1,581 | Source default configuration |
| 32 | step-5-preview-highStepFun | 1,569.6 | Source default configuration |
Limits and data attribution
Correlates with the Text preference construct.
Data licence: CC BY 4.0
Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.