Skip to content
Evaluation sources

Long-Horizon Terminal-Bench

Tencent HY / Lehigh and others / Coding · Use the terminal continuously over a long period to complete complex engineering and tool tasks.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Budgets of 90 minutes, 2 hours and 3 hours are compared separately; judging protocols before and after fixes must not be mixed, nor should the best result of different agents be treated as a single model's score.

How this evidence is used

Candidate observation: awaiting a new batch with an explicitly stated fix protocol; existing old scores are not automatically collected into the public leaderboard or scored, and this is not grounds to claim every old result exploited a vulnerability.

Limits and data attribution

Officially disclosed judging information leakage; it currently cannot be proven that existing scores belong to the same protocol after isolated judging. Source overlap between some tasks and other professional evaluations remains to be checked.

Data licence: Apache 2.0 · official IntelligenceLab/LHTB-leaderboard results dataset; retain attribution and change notes.

Scores published by Tencent HY / Lehigh and others. Raw scores and the News consensus score use different scales and cannot be added directly.