Skip to content
Evaluation sources

SimpleBench

SimpleBench / Reasoning · Uses everyday common sense and logic questions to test whether a model can identify the actual constraints in a problem.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Distinguishes multiple-choice from open-ended answers, fixing AVG@5, temperature and the prompt protocol; the date of joining the leaderboard is not treated as the per-question evaluation time.

How this evidence is used

Candidate observation: new-product coverage is timely, but the reuse boundary of official-site scores and a freezeable data-version statement are awaited; code licences do not apply, and collection and scoring are not yet automated.

Limits and data attribution

Everyday common-sense questions add value, but question representativeness and multiple-choice versus open-answer conditions both affect interpretation.

Data licence: The sample-question and evaluation-code repository is MIT; the boundary for re-displaying scripts of the latest scores published independently on the official site is unconfirmed.

Scores published by SimpleBench. Raw scores and the News consensus score use different scales and cannot be added directly.