Skip to content
Evaluation sources

Humanity’s Last Exam

CAIS / Scale AI / Knowledge · High-difficulty academic knowledge and reasoning across disciplines.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

The question set, tool conditions and judge are fixed; HLE, HLE-Rolling and tool-assisted results are compared separately. AA v4.3 already includes HLE, and if a dedicated score becomes available in future it will be used only for classification, not counted twice in the overall leaderboard.

How this evidence is used

Under observation: the official side has added some new products, but the raw scores still lack verifiable evaluation dates; the Epoch self-evaluation data licence is not applied, and existing AA overall evidence is not credited again.

Limits and data attribution

Under observation: the official side has added some new products, but the raw scores still lack verifiable evaluation dates; the Epoch self-evaluation data licence is not applied, and existing AA overall evidence is not credited again.

Data licence: Evaluation code is MIT; the boundaries for re-displaying official leaderboard scores are still unconfirmed.

Scores published by CAIS / Scale AI. Raw scores and the News consensus score use different scales and cannot be added directly.