Arena Creative Writing
LMArena · Humans blindly select stories, poems and other creative expressions to observe whether the works engage readers.
What it measures, and how
Read the creative_writing category of the official text_style_control dataset and use the style-controlled Bradley–Terry scores and confidence intervals; scores from the overall text or professional writing categories are not borrowed.
How this evidence is used
Officially scored; 5% is split from the original 10% human-preference share, overall text keeps 5%, and the total budget does not increase. Shares an evidence family with other Arena tasks.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | gemini-4-argon-highGoogle | 1,521.8 | High reasoning |
| 2 | claude-opus-5.5-highAnthropic | 1,515 | High reasoning |
| 3 | claude-fable-5-highAnthropic | 1,503 | High reasoning |
| 4 | claude-opus-4-6-highAnthropic | 1,501.2 | High reasoning |
| 5 | gemini-3.7-flash-highGoogle | 1,495.6 | High reasoning |
| 6 | claude-opus-4-7-highAnthropic | 1,489.3 | High reasoning |
| 7 | gemini-3-proGoogle | 1,484.4 | Source default configuration |
| 8 | gemini-3.8-flash-highGoogle | 1,484.4 | High reasoning |
| 10 | claude-fable-5.1-maxAnthropic | 1,481.6 | Max reasoning |
| 11 | gemini-3.1-pro-previewGoogle | 1,480.3 | Source default configuration |
| 14 | gemini-3.6-flash-highGoogle | 1,471.8 | High reasoning |
| 15 | claude-opus-5-maxAnthropic | 1,469.5 | Max reasoning |
| 16 | claude-opus-4-8-highAnthropic | 1,469.4 | High reasoning |
| 18 | claude-opus-4-5-20251101-high-32kAnthropic | 1,469.1 | High reasoning · 32k |
| 19 | gpt-5.6-sol-xhighOpenAI | 1,468.9 | xHigh reasoning |
| 20 | qwen3.8-maxAlibaba | 1,467.8 | Source default configuration |
| 21 | muse-sparkMeta | 1,464.9 | Source default configuration |
| 22 | gemini-3.5-flash-highGoogle | 1,463.8 | High reasoning |
| 24 | grok-4.20-beta1xAI | 1,462.5 | Source default configuration |
| 26 | kimi-k3-maxMoonshot AI | 1,459.7 | Source default configuration |
| 27 | muse-spark-1.3-maxMeta | 1,458.6 | Max reasoning |
| 28 | gemini-3-flashGoogle | 1,456.5 | Source default configuration |
| 29 | gpt-5.5-instantOpenAI | 1,456.3 | Source default configuration |
| 30 | glm-5.2-maxZ.ai | 1,455.3 | Source default configuration |
| 31 | glm-5.3-maxZ.ai | 1,455.3 | Source default configuration |
| 33 | muse-spark-1.2 (xHigh)Meta | 1,452.7 | xHigh reasoning |
| 34 | gpt-5.5-highOpenAI | 1,451.4 | High reasoning |
| 35 | grok-4.5xAI | 1,450.9 | Source default configuration |
| 36 | claude-sonnet-4-5-20250929-high-32kAnthropic | 1,450.1 | High reasoning · 32k |
| 37 | claude-sonnet-4-6Anthropic | 1,449.4 | Source default configuration |
Limits and data attribution
Open user prompts are classified by a classifier and are not all long-form fiction or Chinese creative writing; votes include personal preference. There is overlap with Arena's overall text and professional writing categories, so it cannot serve as an additional independent institution.
Data licence: CC BY 4.0 · Arena official Hugging Face leaderboard-dataset; retain attribution and source links and explain aggregation changes.
Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.