Skip to content
Back to the leaderboard

Evaluation sources

What each source measures, how it updates and whether it counts toward the ranking.

Ranked this round
20
Sources and specialised evaluations reviewed
69
How the ranking works

Overall evaluation and experience

17

Cross-capability comprehensive tests and human blind selection.

Coding

13

From writing code to modifying repositories, testing whether a model can actually build software.

Reasoning

11

Maths, logic and unfamiliar rules show whether a model can think through new problems.

Knowledge

5

Factual Q&A and graduate-level scientific knowledge, testing knowledge mastery and answer accuracy.

Professional work

19

Financial analysis, legal advice and banking, to see whether professional tasks can be completed.

Visual understanding

2

Evidence of understanding images remains in the overall ranking; reading images and design aesthetics are different abilities.

Chinese and multilingual

2

Supplementary evidence on language fundamentals; English writing performance is not treated as Chinese ability.

“Observing” means its data, running conditions or limits are still being checked; it is not part of the overall or category boards. Repeated collection or display slices of one evaluation add no weight; overlap between different evaluations' question sets is still being checked.