Terminal-Bench 4 · System reference
Harbor / Terminal-Bench / Coding · Keeps the original scores of models paired with different agents on terminal tasks for cross-checking.
What it measures, and how
Terminal-Bench 4.0 is fixed, retaining the model, Agent version, reasoning tier and number of runs separately; different tool systems are not interpreted as unified model conditions.
How this evidence is used
Original score reference: automatically collected after locking the official submission list and data version, retaining each different system configuration and not selecting a representative model score. For cross-checking only; not counted towards overall or coding rankings.
Results
The following are system scores for models paired with different agents; run conditions differ, so they are for reference only. Each item shows at most 30 configurations, and anonymous test models are not shown.
| Rank there | Model there | Raw score | Run configuration |
|---|---|---|---|
| 1 | claude-fable-5-1Anthropic | 57.9% | Claude Code 2.1.257 · max reasoning |
| 2 | claude-opus-5Anthropic | 51.8% | Claude Code 2.1.231 · max reasoning |
| 3 | claude-fable-5Anthropic | 44.5% | Claude Code 2.1.231 · max reasoning |
| 4 | glm-5.3Z.ai | 41.8% | Claude Code 2.1.207 · max reasoning |
| 5 | gpt-5.6-solOpenAI | 37.3% | Codex 0.149.1 · max reasoning |
| 6 | claude-opus-4-8Anthropic | 23.6% | Claude Code 2.1.231 · max reasoning |
| 7 | gpt-5.6-terraOpenAI | 21.5% | Codex 0.149.1 · max reasoning |
| 8 | grok-4.6xAI | 20.3% | Grok Build 1.0.5 · source did not report a reasoning tier |
| 9 | gemini-3.8-flashGoogle | 19.1% | mini-SWE-agent 2.4.6 · high reasoning |
| 10 | gpt-5.6-lunaOpenAI | 17.3% | Codex 0.149.1 · max reasoning |
| 11 | grok-4.5xAI | 12.4% | Grok Build 1.0.5 · source did not report a reasoning tier |
| 11 | claude-sonnet-5Anthropic | 12.4% | Claude Code 2.1.231 · max reasoning |
| 13 | gemini-3.7-flashGoogle | 11.2% | mini-SWE-agent 2.4.6 · high reasoning |
Limits and data attribution
Different models use different tools such as Claude Code, Codex, GrokBuild or mini-SWE; AA v4.3 already includes Terminal-Bench v4 and must not be counted again in the overall vote.
Data licence: Apache 2.0 · result submissions in the official harbor-framework/terminal-bench repository; retain attribution, licence and modification notes.
Scores published by Harbor / Terminal-Bench. Raw scores and the News consensus score use different scales and cannot be added directly.