SWE-rebench
Nebius / Coding · Continuously collects real repository issues to compare models' software repair abilities across multiple programming languages.
What it measures, and how
The task window, ReAct environment and text/tools interaction protocol are fixed, using the average solve rate and standard error over multiple runs; a full Agent product and a base model are not conflated as the same object.
How this evidence is used
Candidate observation: real repository tasks are worth adding, but a verifiable run date, a stable scoring protocol and re-display boundaries are awaited; automated collection and scoring are not yet enabled.
Limits and data attribution
Rolling questions, run environments and task windows affect comparability; recent site maintenance does not mean all models were retested.
Data licence: The official Hugging Face CC BY 4.0 dataset is the task question bank; the boundaries for stable export and re-display of final web page scores are not yet confirmed.
Scores published by Nebius. Raw scores and the News consensus score use different scales and cannot be added directly.