Skip to content
Back to the leaderboard

How the ranking works

Many public evaluations, combined: the evidence and method behind the ranking.

0–100

Shared evidence, a clear order.

It combines several public evaluations and keeps the complete ranking in as little conflict with known results as possible. The overall board and the four category boards each show at most the top 30.

The consensus index turns the evidence gaps behind the ranking into a 0–100 scale for easy comparison; it is not an accuracy rate or a percentage gap in ability. Close support can produce equal scores, and the rank is still set by the complete evidence.

  1. 01

    First, confirm it is the same model

    Names and versions are made consistent across boards. When one model has several reasoning levels, one representative configuration is chosen by a rule fixed in advance, never by the highest score. Anonymous test code names, and results that used other models to finish the task, are excluded; a Preview version that has been publicly released can take part.

  2. 02

    Compare only models that were actually tested

    Only real results within the same evaluation are compared. When both sides publish an error margin, a small difference inside it counts as close to a tie, not a clear win. No measurement, no comparison.

  3. 03

    Set the weights first, then combine the shared view

    Each evaluation has a budget fixed in advance. The overall board needs at least three evaluation organisations, three evidence families and three specialised evaluations; repeated collection from one source, or display slices of it, add no weight.

  4. 04

    Find the complete order with the fewest conflicts

    Evaluations can disagree. We choose the order that goes against the least net weight of shared evidence, then compute the display index separately. Missing results are not filled with zero, and unpublished results are not guessed.

Broad tests, blind human votes and specialised evaluations, read together.

Broad evaluations take 30%, blind human votes 10%, and specialised evaluations 60% in total. A new evaluation is first checked for overlap against its published notes, then given part of the matching share; a missing share is not passed on to other evaluations.

  • Overall evaluation30%AA Index
  • Human blind selection10%Arena Text · Arena creative blind selection
  • Coding and design12%Arena WebDev · DeepSWE v1.1 · TapTap Maker · LiveBench · coding overall
  • Writing and expression9%LiveBench · language and instruction · Creative Writing v3 · Longform Writing
  • Maths and reasoning12%LiveBench · reasoning and maths · FrontierMath v2 · Tiers 1–3 · FrontierMath v2 · Tier 4 · Chess Puzzles · Mystery Game Puzzles
  • Knowledge and facts6%SimpleQA Verified · GPQA Diamond
  • Visual understanding6%Arena Vision
  • Tools and office6%APEX-Agents 1.1 · τ³-Banking
  • Chinese and multilingual6%AA Chinese / multilingual
  • Industry professional tasks3%Vals Finance Agent

Coding, reasoning, knowledge and professional work each draw on their own real evaluations, and each category score is computed separately. Visual understanding and multilingual evidence stay in the overall board; renaming a category gives the same result no extra vote. The overall score is not the average of the category scores. A category usually appears only with at least two valid evaluations and five comparable models. Knowledge currently rests on two Epoch evaluations, with the single-organisation limit stated. Evidence such as creative preference and web development still counts toward the overall board; this first version has no separate aesthetics or writing board.

You might also ask

Why can a model with little evidence make the board?
The number of evaluations and ability are different things. The overall board needs at least three organisations and evidence across capabilities. Most categories need two organisations; knowledge uses two evaluations from one organisation, and says so. A missing result does not count as zero, but it can still bias the outcome.
What does “evidence-sensitive” mean?
The flag appears when removing one evaluation or one organisation, moving a single weight up or down by 20%, or changing how error is handled widens the rank range to three places or more, makes the model ineligible in some scenarios, or leaves a comparison incomplete. The scenario range on a model's page is not a 95% confidence interval and does not include unknown results.
Does a higher model always win its head-to-head comparisons?
No. A can beat B, B beat C, and C beat A. The full board weighs these conflicts, so ranks that are not neighbours can differ from a head-to-head comparison. Mathematically optimal means the fewest total conflicts under the current rules, not a proven order of ability across all real work.
Why might the ranking differ from my experience?
The board brings public evaluations together and reflects the overall ability that evidence supports. Real use also depends on the product version, reasoning level, tools and stability over long tasks. We test old rankings against new evaluations; when scores are close, a gap of one or two places is not a clear difference in strength.
When does a new model appear?
News checks upstream results four times a day. A new model can be ranked only after an evaluator actually publishes its results; a page refresh, a price change or our own collection time does not count as a new evaluation.
What happens when a board is temporarily unreachable?
A source snapshot that is still valid can stay in use. Within the same evaluation version, a row that goes missing temporarily can carry over a verified record from the last seven days at most; a valid result that stays on the public board is not removed just because its value has not changed in a long time. An official withdrawal, correction or new version invalidates or replaces it, and no historical best is kept. If a whole round fails, the last valid board and its time stay in place.
Does the overall index double-count specialised evaluations?
We check overlap against the published question sets and method notes, and cap the total share of related sources; correlation we cannot confirm remains a limit. The AA index currently takes 30%. Arena's overall text and creative writing boards share the original 10% human-preference share, 5% each; their votes overlap, so they still count as one evidence family, not as two independent evaluations.
Do price or speed affect the ranking?
Neither does. Price is there only to show what the API costs to use, always per million tokens, with the provider's official source. It does not cover subscription fees.
Calculation details and current version

Method version: 2026.09-public-consensus-v15. It uses a weighted incomplete Kemeny ranking. Each pair of models with shared evaluations sums its net support M; the goal is to minimise the total net support that the order reverses. The integer optimisation returns its optimality status and the bounds of the objective, and only a result that passes full validation is published.

When both sides state a standard error, net support is 2Φ(score gap / combined standard error) − 1, assuming zero covariance; other comparisons take only the direction of the raw lead. Unknown error is not zero error, and treating small gaps as ordinal remains a limit. Weight applies to potential model pairs, so a source that covers more models uses more comparison slots; the nominal budget cannot be read as a precise contribution to the final rank.

The display index keeps the original order. For each pair of adjacent models it reverses their order, lets the other models reorder, and computes the least added opposing net support. These non-negative gaps are summed along the original ranking, then mapped to 0–100 with a sigmoid relative to a fixed reference group. When an alternative order is equally optimal the gap stays at zero; scores show one decimal, with no artificial minimum gap. The index is never used to reorder the ranking; changes in which models are on the board, the reference models and the evidence still move it, and a score gap is not a true distance in ability. Ties in cost are broken by a fixed model ID order. The integer optimisation uses the HiGHS solver (highs 1.15.3), and the normal distribution function for error uses the same algorithm as SciPy norm.cdf. Sources, terms, eligibility and each round's inputs are all versioned. Where models are not connected by shared evidence, no falsely precise order across components is published.

Fixed reference models: claude-fable-5, gpt-5-6-sol, kimi-k-3, qwen-3-8-max, gpt-5-4, claude-opus-4-8, claude-sonnet-5, grok-4-5, gemini-3-5-flash, glm-5-2, claude-sonnet-4-6, qwen-3-7-max, kimi-k-2-6, deepseek-v-4-pro, qwen-3-6-plus, minimax-m-3, gpt-5-4-mini, grok-4-3. The reference group sets the scale of the index; it does not decide where any provider should rank. Indexes from different categories are not directly comparable.

See every piece of evaluation evidence