Skip to content
Evaluation sources

MCP Atlas

Scale AI / Professional work · Retrieve material, operate documents and databases through real MCP services to complete cross-software tasks.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

The operating protocol adjusted in April 2026 and the 100 tool-call budget are fixed; the full 1,000 questions and the public 500 questions are kept separate. Only the task pass rate is used, and answer-condition coverage is not counted twice; judges, retries and tool versions must be locked with the score batch.

How this evidence is used

Under observation: can add tool-execution evidence across office software; stable score exports, the actual evaluation batch and re-display boundaries need confirmation, so only the source is recorded for now, with no automatic collection or scoring.

Limits and data attribution

Old method descriptions for the web and live scores may be out of sync; public problem-set code is not a live score export. It belongs to the same organisation as other Scale leaderboards and does not add to the count of independent organisations.

Data licence: The official environment code uses MIT; that licence does not automatically cover live web scores, and the boundaries for re-displaying leaderboard data are unconfirmed.

Scores published by Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.