Skip to content
Evaluation sources

ToneBench · Style and script

Towards AI · Writes video scripts in a specified author's style, testing the opening, structure, expression and spoken pacing.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

10 real scripted tasks, multiple generations, three judges; only identical tasks, rubrics, judges and run fingerprints are compared. Mixed-model fallback rows must be excluded, and configurations are not chosen by the best score in each column.

How this evidence is used

Under observation: the data protocol and current-period coverage are good; the score re-display boundary needs confirmation and the pure-model configuration needs verification.

Limits and data attribution

Leans towards the single-channel style of Towards AI, with exact questions and reference scripts undisclosed; the current first place includes a Fable 5 + Opus 4.8 fallback and cannot be treated as a single-model score.

Data licence: The public JSON is readable, but no clear data-aggregation and re-display terms have been found.

Scores published by Towards AI. Raw scores and the News consensus score use different scales and cannot be added directly.