Skip to content
Evaluation sources

PRBench · Legal

Scale AI / Professional work · Answering complex questions in real legal work, testing professional judgement and the quality of argument.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Either the full Legal 500 questions or the Hard 250 questions is fixed; search tools, reasoning tier and judge version are kept consistent. The rubric aggregate score of that question set is used, the full set and the hard subset are not counted twice, and professional Q&A is not described as full file delivery.

How this evidence is used

Under observation: can add legal professional judgement; the problem set, tool conditions and score batches must first be verified, so only the source is recorded for now, with no automatic collection or scoring.

Limits and data attribution

Judged by model referees, with referee bias; legal and financial specialities from the same institution should share an impact boundary and cannot serve as two independent pieces of evidence. It does not directly measure Word, Excel or slide creation ability.

Data licence: The official question bank and evaluation code are public; their licence does not automatically cover live leaderboard scores, and the boundaries for stable export and re-display of the latest scores are unconfirmed.

Scores published by Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.