PRBench · Legal
Scale AI / Professional work · Answering complex questions in real legal work, testing professional judgement and the quality of argument.
What it measures, and how
Either the full Legal 500 questions or the Hard 250 questions is fixed; search tools, reasoning tier and judge version are kept consistent. The rubric aggregate score of that question set is used, the full set and the hard subset are not counted twice, and professional Q&A is not described as full file delivery.
How this evidence is used
Under observation: can add legal professional judgement; the problem set, tool conditions and score batches must first be verified, so only the source is recorded for now, with no automatic collection or scoring.
Limits and data attribution
Judged by model referees, with referee bias; legal and financial specialities from the same institution should share an impact boundary and cannot serve as two independent pieces of evidence. It does not directly measure Word, Excel or slide creation ability.
Data licence: The official question bank and evaluation code are public; their licence does not automatically cover live leaderboard scores, and the boundaries for stable export and re-display of the latest scores are unconfirmed.
Scores published by Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.