Evaluation sourcesOfficial evaluation
FrontierSWE v2
Proximal Labs / Coding · Long-running engineering implementation and research tasks
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet
What it measures, and how
Only the unified Proximus harness of v2 is compared; it is not mixed with older versions or best-of under different conditions.
How this evidence is used
Under observation: coverage of new models on long-horizon tasks is still limited; a stable raw-data entry point and run configuration need to be added.
Limits and data attribution
Under observation: coverage of new models on long-horizon tasks is still limited; a stable raw-data entry point and run configuration need to be added.
Data licence: Leaderboard data usage boundaries to be confirmed
Scores published by Proximal Labs. Raw scores and the News consensus score use different scales and cannot be added directly.