How the ranking works
Many public evaluations, combined: the evidence and method behind the ranking.
0–100
Shared evidence, a clear order.
It combines several public evaluations and keeps the complete ranking in as little conflict with known results as possible. The overall board and the four category boards each show at most the top 30.
The consensus index turns the evidence gaps behind the ranking into a 0–100 scale for easy comparison; it is not an accuracy rate or a percentage gap in ability. Close support can produce equal scores, and the rank is still set by the complete evidence.
- 01
First, confirm it is the same model
Names and versions are made consistent across boards. When one model has several reasoning levels, one representative configuration is chosen by a rule fixed in advance, never by the highest score. Anonymous test code names, and results that used other models to finish the task, are excluded; a Preview version that has been publicly released can take part.
- 02
Compare only models that were actually tested
Only real results within the same evaluation are compared. When both sides publish an error margin, a small difference inside it counts as close to a tie, not a clear win. No measurement, no comparison.
- 03
Set the weights first, then combine the shared view
Each evaluation has a budget fixed in advance. The overall board needs at least three evaluation organisations, three evidence families and three specialised evaluations; repeated collection from one source, or display slices of it, add no weight.
- 04
Find the complete order with the fewest conflicts
Evaluations can disagree. We choose the order that goes against the least net weight of shared evidence, then compute the display index separately. Missing results are not filled with zero, and unpublished results are not guessed.
Broad tests, blind human votes and specialised evaluations, read together.
Broad evaluations take 30%, blind human votes 10%, and specialised evaluations 60% in total. A new evaluation is first checked for overlap against its published notes, then given part of the matching share; a missing share is not passed on to other evaluations.
- Overall evaluation30%AA Index
- Human blind selection10%Arena Text · Arena creative blind selection
- Coding and design12%Arena WebDev · DeepSWE v1.1 · TapTap Maker · LiveBench · coding overall
- Writing and expression9%LiveBench · language and instruction · Creative Writing v3 · Longform Writing
- Maths and reasoning12%LiveBench · reasoning and maths · FrontierMath v2 · Tiers 1–3 · FrontierMath v2 · Tier 4 · Chess Puzzles · Mystery Game Puzzles
- Knowledge and facts6%SimpleQA Verified · GPQA Diamond
- Visual understanding6%Arena Vision
- Tools and office6%APEX-Agents 1.1 · τ³-Banking
- Chinese and multilingual6%AA Chinese / multilingual
- Industry professional tasks3%Vals Finance Agent
Coding, reasoning, knowledge and professional work each draw on their own real evaluations, and each category score is computed separately. Visual understanding and multilingual evidence stay in the overall board; renaming a category gives the same result no extra vote. The overall score is not the average of the category scores. A category usually appears only with at least two valid evaluations and five comparable models. Knowledge currently rests on two Epoch evaluations, with the single-organisation limit stated. Evidence such as creative preference and web development still counts toward the overall board; this first version has no separate aesthetics or writing board.
You might also ask
Why can a model with little evidence make the board?
What does “evidence-sensitive” mean?
Does a higher model always win its head-to-head comparisons?
Why might the ranking differ from my experience?
When does a new model appear?
What happens when a board is temporarily unreachable?
Does the overall index double-count specialised evaluations?
Do price or speed affect the ranking?
Calculation details and current version
Method version: 2026.09-public-consensus-v15. It uses a weighted incomplete Kemeny ranking. Each pair of models with shared evaluations sums its net support M; the goal is to minimise the total net support that the order reverses. The integer optimisation returns its optimality status and the bounds of the objective, and only a result that passes full validation is published.
When both sides state a standard error, net support is 2Φ(score gap / combined standard error) − 1, assuming zero covariance; other comparisons take only the direction of the raw lead. Unknown error is not zero error, and treating small gaps as ordinal remains a limit. Weight applies to potential model pairs, so a source that covers more models uses more comparison slots; the nominal budget cannot be read as a precise contribution to the final rank.
The display index keeps the original order. For each pair of adjacent models it reverses their order, lets the other models reorder, and computes the least added opposing net support. These non-negative gaps are summed along the original ranking, then mapped to 0–100 with a sigmoid relative to a fixed reference group. When an alternative order is equally optimal the gap stays at zero; scores show one decimal, with no artificial minimum gap. The index is never used to reorder the ranking; changes in which models are on the board, the reference models and the evidence still move it, and a score gap is not a true distance in ability. Ties in cost are broken by a fixed model ID order. The integer optimisation uses the HiGHS solver (highs 1.15.3), and the normal distribution function for error uses the same algorithm as SciPy norm.cdf. Sources, terms, eligibility and each round's inputs are all versioned. Where models are not connected by shared evidence, no falsely precise order across components is published.
Fixed reference models: claude-fable-5, gpt-5-6-sol, kimi-k-3, qwen-3-8-max, gpt-5-4, claude-opus-4-8, claude-sonnet-5, grok-4-5, gemini-3-5-flash, glm-5-2, claude-sonnet-4-6, qwen-3-7-max, kimi-k-2-6, deepseek-v-4-pro, qwen-3-6-plus, minimax-m-3, gpt-5-4-mini, grok-4-3. The reference group sets the scale of the index; it does not decide where any provider should rank. Indexes from different categories are not directly comparable.