Skip to content
Evaluation sources

Terminal-Bench Science

Harbor / Professional work · Completing scientific research tasks in the terminal

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Distinguishes systems such as Claude Code, Codex and mini-SWE; scores from different systems cannot be treated directly as model capability.

How this evidence is used

Under observation: 12 models with mixed run systems; new frontier models are still incomplete.

Limits and data attribution

Under observation: 12 models with mixed run systems; new frontier models are still incomplete.

Data licence: Leaderboard data usage boundaries to be confirmed

Scores published by Harbor. Raw scores and the News consensus score use different scales and cannot be added directly.