Humanity’s Last Exam
CAIS / Scale AI / Knowledge · High-difficulty academic knowledge and reasoning across disciplines.
What it measures, and how
The question set, tool conditions and judge are fixed; HLE, HLE-Rolling and tool-assisted results are compared separately. AA v4.3 already includes HLE, and if a dedicated score becomes available in future it will be used only for classification, not counted twice in the overall leaderboard.
How this evidence is used
Under observation: the official side has added some new products, but the raw scores still lack verifiable evaluation dates; the Epoch self-evaluation data licence is not applied, and existing AA overall evidence is not credited again.
Limits and data attribution
Under observation: the official side has added some new products, but the raw scores still lack verifiable evaluation dates; the Epoch self-evaluation data licence is not applied, and existing AA overall evidence is not credited again.
Data licence: Evaluation code is MIT; the boundaries for re-displaying official leaderboard scores are still unconfirmed.
Scores published by CAIS / Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.