Skip to content
Evaluation sources

SWE-rebench

Nebius / Coding · Continuously collects real repository issues to compare models' software repair abilities across multiple programming languages.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

The task window, ReAct environment and text/tools interaction protocol are fixed, using the average solve rate and standard error over multiple runs; a full Agent product and a base model are not conflated as the same object.

How this evidence is used

Candidate observation: real repository tasks are worth adding, but a verifiable run date, a stable scoring protocol and re-display boundaries are awaited; automated collection and scoring are not yet enabled.

Limits and data attribution

Rolling questions, run environments and task windows affect comparability; recent site maintenance does not mean all models were retested.

Data licence: The official Hugging Face CC BY 4.0 dataset is the task question bank; the boundaries for stable export and re-display of final web page scores are not yet confirmed.

Scores published by Nebius. Raw scores and the News consensus score use different scales and cannot be added directly.