Skip to content
Evaluation sources

Short-Story · Short fiction

Lech Mazur · Weave fixed characters, plot and style requirements naturally into a complete short story.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

600–800 word short pieces, 10 creative requirements; same prompt, three judges, two-way comparison. Thurstone main scores are used, BT is diagnostic only, and votes are not counted twice; multiple tiers are selected by fixed rules.

How this evidence is used

Under observation: formal data reuse boundaries must be completed and the mismatch between the README and machine scores resolved before integration.

Limits and data attribution

Model judges still have literary preferences; some variants did not complete all stories. Short-story ability is not the same as long-form fiction or Chinese prose.

Data licence: No clear permission to reuse the leaderboard data was found; public on GitHub does not mean openly licensed.

Scores published by Lech Mazur. Raw scores and the News consensus score use different scales and cannot be added directly.