Skip to content
Evaluation sources

Arena Creative Writing

LMArena · Humans blindly select stories, poems and other creative expressions to observe whether the works engage readers.

Official evaluation
In NewsRanked
Evidence budget5%
Upstream data as of09/30 08:00
Last synced10/01 20:05

What it measures, and how

Read the creative_writing category of the official text_style_control dataset and use the style-controlled Bradley–Terry scores and confidence intervals; scores from the overall text or professional writing categories are not borrowed.

How this evidence is used

Officially scored; 5% is split from the original 10% human-preference share, overall text keeps 5%, and the total budget does not increase. Shares an evidence family with other Arena tasks.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1gemini-4-argon-highGoogle1,521.8High reasoning
2claude-opus-5.5-highAnthropic1,515High reasoning
3claude-fable-5-highAnthropic1,503High reasoning
4claude-opus-4-6-highAnthropic1,501.2High reasoning
5gemini-3.7-flash-highGoogle1,495.6High reasoning
6claude-opus-4-7-highAnthropic1,489.3High reasoning
7gemini-3-proGoogle1,484.4Source default configuration
8gemini-3.8-flash-highGoogle1,484.4High reasoning
10claude-fable-5.1-maxAnthropic1,481.6Max reasoning
11gemini-3.1-pro-previewGoogle1,480.3Source default configuration
14gemini-3.6-flash-highGoogle1,471.8High reasoning
15claude-opus-5-maxAnthropic1,469.5Max reasoning
16claude-opus-4-8-highAnthropic1,469.4High reasoning
18claude-opus-4-5-20251101-high-32kAnthropic1,469.1High reasoning · 32k
19gpt-5.6-sol-xhighOpenAI1,468.9xHigh reasoning
20qwen3.8-maxAlibaba1,467.8Source default configuration
21muse-sparkMeta1,464.9Source default configuration
22gemini-3.5-flash-highGoogle1,463.8High reasoning
24grok-4.20-beta1xAI1,462.5Source default configuration
26kimi-k3-maxMoonshot AI1,459.7Source default configuration
27muse-spark-1.3-maxMeta1,458.6Max reasoning
28gemini-3-flashGoogle1,456.5Source default configuration
29gpt-5.5-instantOpenAI1,456.3Source default configuration
30glm-5.2-maxZ.ai1,455.3Source default configuration
31glm-5.3-maxZ.ai1,455.3Source default configuration
33muse-spark-1.2 (xHigh)Meta1,452.7xHigh reasoning
34gpt-5.5-highOpenAI1,451.4High reasoning
35grok-4.5xAI1,450.9Source default configuration
36claude-sonnet-4-5-20250929-high-32kAnthropic1,450.1High reasoning · 32k
37claude-sonnet-4-6Anthropic1,449.4Source default configuration
Limits and data attribution

Open user prompts are classified by a classifier and are not all long-form fiction or Chinese creative writing; votes include personal preference. There is overlap with Arena's overall text and professional writing categories, so it cannot serve as an additional independent institution.

Data licence: CC BY 4.0 · Arena official Hugging Face leaderboard-dataset; retain attribution and source links and explain aggregation changes.

Scores published by LMArena. Raw scores and the News consensus score use different scales and cannot be added directly.