Long-Horizon Terminal-Bench
Tencent HY / Lehigh and others / Coding · Use the terminal continuously over a long period to complete complex engineering and tool tasks.
What it measures, and how
Budgets of 90 minutes, 2 hours and 3 hours are compared separately; judging protocols before and after fixes must not be mixed, nor should the best result of different agents be treated as a single model's score.
How this evidence is used
Candidate observation: awaiting a new batch with an explicitly stated fix protocol; existing old scores are not automatically collected into the public leaderboard or scored, and this is not grounds to claim every old result exploited a vulnerability.
Limits and data attribution
Officially disclosed judging information leakage; it currently cannot be proven that existing scores belong to the same protocol after isolated judging. Source overlap between some tasks and other professional evaluations remains to be checked.
Data licence: Apache 2.0 · official IntelligenceLab/LHTB-leaderboard results dataset; retain attribution and change notes.
Scores published by Tencent HY / Lehigh and others. Raw scores and the News consensus score use different scales and cannot be added directly.