PRBench · Finance
Scale AI / Professional work · Answering real financial decision questions, testing professional judgement, reasoning and risk explanation.
What it measures, and how
Either the full Finance 600 questions or the Hard 300 questions is fixed; search tools, reasoning tier and judge version are kept consistent. The rubric aggregate score of that question set is used, the full set and the hard subset are not counted twice, and professional Q&A is not treated as an end-to-end office task.
How this evidence is used
Under observation: can add financial professional judgement; the problem set, tool conditions, stable exports and re-display boundaries need verification, so only the source is recorded for now, with no automatic collection or scoring.
Limits and data attribution
Judged by model referees and does not mean financial modelling or spreadsheet creation has been completed; it belongs to the same Scale as PRBench Legal and MCP Atlas, and institutional weight cannot be raised by adding more leaderboards.
Data licence: The official question bank and evaluation code are public; their licence does not automatically cover live leaderboard scores, and the boundaries for stable export and re-display of the latest scores are unconfirmed.
Scores published by Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.