ToneBench · Style and script
Towards AI · Writes video scripts in a specified author's style, testing the opening, structure, expression and spoken pacing.
What it measures, and how
10 real scripted tasks, multiple generations, three judges; only identical tasks, rubrics, judges and run fingerprints are compared. Mixed-model fallback rows must be excluded, and configurations are not chosen by the best score in each column.
How this evidence is used
Under observation: the data protocol and current-period coverage are good; the score re-display boundary needs confirmation and the pure-model configuration needs verification.
Limits and data attribution
Leans towards the single-channel style of Towards AI, with exact questions and reference scripts undisclosed; the current first place includes a Fable 5 + Opus 4.8 fallback and cannot be treated as a single-model score.
Data licence: The public JSON is readable, but no clear data-aggregation and re-display terms have been found.
Scores published by Towards AI. Raw scores and the News consensus score use different scales and cannot be added directly.