Skip to content
Evaluation sources

Terminal-Bench 4 · System reference

Harbor / Terminal-Bench / Coding · Keeps the original scores of models paired with different agents on terminal tasks for cross-checking.

Official evaluation
In NewsReference only
Evidence budgetNot scored
Upstream data as of09/03 10:43
Last synced10/01 20:05

What it measures, and how

Terminal-Bench 4.0 is fixed, retaining the model, Agent version, reasoning tier and number of runs separately; different tool systems are not interpreted as unified model conditions.

How this evidence is used

Original score reference: automatically collected after locking the official submission list and data version, retaining each different system configuration and not selecting a representative model score. For cross-checking only; not counted towards overall or coding rankings.

Results

The following are system scores for models paired with different agents; run conditions differ, so they are for reference only. Each item shows at most 30 configurations, and anonymous test models are not shown.

Rank thereModel thereRaw scoreRun configuration
1claude-fable-5-1Anthropic57.9%Claude Code 2.1.257 · max reasoning
2claude-opus-5Anthropic51.8%Claude Code 2.1.231 · max reasoning
3claude-fable-5Anthropic44.5%Claude Code 2.1.231 · max reasoning
4glm-5.3Z.ai41.8%Claude Code 2.1.207 · max reasoning
5gpt-5.6-solOpenAI37.3%Codex 0.149.1 · max reasoning
6claude-opus-4-8Anthropic23.6%Claude Code 2.1.231 · max reasoning
7gpt-5.6-terraOpenAI21.5%Codex 0.149.1 · max reasoning
8grok-4.6xAI20.3%Grok Build 1.0.5 · source did not report a reasoning tier
9gemini-3.8-flashGoogle19.1%mini-SWE-agent 2.4.6 · high reasoning
10gpt-5.6-lunaOpenAI17.3%Codex 0.149.1 · max reasoning
11grok-4.5xAI12.4%Grok Build 1.0.5 · source did not report a reasoning tier
11claude-sonnet-5Anthropic12.4%Claude Code 2.1.231 · max reasoning
13gemini-3.7-flashGoogle11.2%mini-SWE-agent 2.4.6 · high reasoning
Limits and data attribution

Different models use different tools such as Claude Code, Codex, GrokBuild or mini-SWE; AA v4.3 already includes Terminal-Bench v4 and must not be counted again in the overall vote.

Data licence: Apache 2.0 · result submissions in the official harbor-framework/terminal-bench repository; retain attribution, licence and modification notes.

Scores published by Harbor / Terminal-Bench. Raw scores and the News consensus score use different scales and cannot be added directly.