FinanceBenchmark
FinanceBenchmark / Professional work · Complete financial pricing, capital calculation and risk control tasks, verified with numerical and code execution results.
What it measures, and how
Question sets v2 and v3 are kept separate from the runtime environment version; the condition of three attempts per question is retained. Tasks passed on the official site means at least one success, yet it is also called Pass@1 and cannot be mixed with the average single-attempt accuracy; Finance Index is also a different aggregate metric.
How this evidence is used
Under observation: can add finance evidence with definite answers and code verification; the metric definition still needs checking, and the official Sol results JSON download returned a 404 this time, so only the source is recorded for now, with no automatic collection or scoring.
Limits and data attribution
Succeeding at least once in three attempts is not the same as a single-attempt success rate; intervals must correspond to the correct metric, and rolling model aliases also need version verification. A public question bank and download button do not mean the original scores are obtainable.
Data licence: The official repository code is MIT, and the website requires retention of the FinanceBenchmark citation; the usage boundaries for a complete export of live scores are unconfirmed, and the code licence is not treated as a substitute for the leaderboard data licence.
Scores published by FinanceBenchmark. Raw scores and the News consensus score use different scales and cannot be added directly.