Evaluation sources
What each source measures, how it updates and whether it counts toward the ranking.
- Ranked this round
- 20
- Sources and specialised evaluations reviewed
- 69
Overall evaluation and experience
17Cross-capability comprehensive tests and human blind selection.
- AA IndexRe-run several current evaluations in the same environment to produce one overall capability score. Price is for reference only and does not count towards the total.Artificial AnalysisRankedEvidence budget 30%
- LiveBenchThe question bank is continually replaced and answers can be checked automatically. Used to see whether a model still holds up after the questions change.LiveBenchCross-reference
- Arena TextHumans compare answers anonymously in pairs to see which they would rather read. This adds the overall experience that multiple-choice questions cannot show.LMArenaRankedEvidence budget 5%
- Arena WebDevAlso a human blind choice, it compares whether web pages are usable and deliverable.LMArenaRankedEvidence budget 2.4%
- LiveBench · language and instructionAverage of the Language and IF items, 50% each.LiveBenchRankedEvidence budget 3%
- Arena creative blind selectionHumans blindly select stories, poems and other creative expressions to observe whether the works engage readers.LMArenaRankedEvidence budget 5%
- Design Arena · visual designHumans blindly select design work for web pages, interfaces, charts and slides.Design Arena / IntelligenceObserving
- EQ-Bench 4 · emotional and interpersonal understandingUnderstanding characters' emotions, implicit needs and different communication preferences in multi-turn exchanges.EQ-BenchObserving
- Short-Story · Short fictionWeave fixed characters, plot and style requirements naturally into a complete short story.Lech MazurObserving
- ToneBench · Style and scriptWrites video scripts in a specified author's style, testing the opening, structure, expression and spoken pacing.Towards AIObserving
- TubeLab · Video scriptWrite scripts for different video channels and compare story organisation, style, pacing and clarity.TubeLabObserving
- NC Bench · Writer workflowHelp writers rewrite, continue, summarise, translate and organise creative ideas.NovelcrafterObserving
- OpenVibeEval · Web worksCompare web work, accessibility and blind preference to observe design performance across different tools.OpenVibeEvalObserving
- SVGBench · Vector graphic blind selectionBlind-selects SVGs under the same prompt, testing visual clarity and graphic expression.SVGBenchObserving
- Rapidata SVGHumans compare vector graphics generated by models and directly rate visual preference.RapidataObserving
- Creative Writing v3Short-form creation and expressionEQ-BenchRankedEvidence budget 3.6%
- Longform WritingCoherence and completeness of long-form storiesEQ-BenchRankedEvidence budget 2.4%
Coding
13From writing code to modifying repositories, testing whether a model can actually build software.
- DeepSWE v1.1Modifying repositories and fixing defects in a unified coding assistant environment; what is compared is the model that writes the code itself.DataCurveRankedEvidence budget 3.6%
- TapTap MakerWriting working features in Lua inside a real game engine. The score is judged by engine execution, not by another model.TapTap MakerRankedEvidence budget 3.6%
- LiveBench · coding overallAverage of the Coding and Agentic Coding items, 50% each.LiveBenchRankedEvidence budget 2.4%
- Hyper-τ-benchBuilds and tests a working customer service agent from business materials.SierraObserving
- SWE-rebenchContinuously collects real repository issues to compare models' software repair abilities across multiple programming languages.NebiusObserving
- LHTBUse the terminal continuously over a long period to complete complex engineering and tool tasks.Tencent HY / Lehigh and othersObserving
- Terminal-Bench 4Keeps the original scores of models paired with different agents on terminal tasks for cross-checking.Harbor / Terminal-BenchReference only
- MirrorCodeLong-horizon software reproductionEpoch AIReference only
- Code MigrationSoftware migration and interface reworkVals AIObserving
- ProgramBenchImplementation and delivery of complete programsVals AIObserving
- FrontierCodeQuality, testing and modification scope on real repository tasksCognitionObserving
- FrontierSWE v2Long-running engineering implementation and research tasksProximal LabsObserving
- WeirdMLImplement machine learning solutions on unfamiliar dataWeirdMLObserving
Reasoning
11Maths, logic and unfamiliar rules show whether a model can think through new problems.
- LiveBench · reasoning and mathsAverage of the Reasoning and Mathematics items, each weighted 50%.LiveBenchRankedEvidence budget 4.2%
- SimpleBenchUses everyday common sense and logic questions to test whether a model can identify the actual constraints in a problem.SimpleBenchObserving
- FrontierMath v2 · Tiers 1–3Research-level mathematical reasoningEpoch AIRankedEvidence budget 3.6%
- FrontierMath v2 · Tier 4Harder mathematical research problemsEpoch AIRankedEvidence budget 1.2%
- Chess PuzzlesFinds the correct move from a written board positionEpoch AIRankedEvidence budget 1.8%
- Mystery Game PuzzlesDeriving correct actions from unfamiliar rulesEpoch AIRankedEvidence budget 1.2%
- MathArena · ArXivLeanVerifying mathematical proofs with LeanMathArena / ETH ZurichObserving
- MathArena · BrokenArXivFinding and fixing errors in mathematical proofsMathArena / ETH ZurichObserving
- MathArena · ArXivMathNew questions distilled from recent mathematics papersMathArena / ETH ZurichObserving
- ARC-AGI-2Learning and generalising from unfamiliar rulesARC PrizeObserving
- ARC-AGI-3Learning and generalising from unfamiliar rulesARC PrizeObserving
Knowledge
5Factual Q&A and graduate-level scientific knowledge, testing knowledge mastery and answer accuracy.
- FACTS ParametricAnswers world-fact questions without search tools, testing the accuracy of the model's own knowledge.Google DeepMind / KaggleObserving
- Humanity’s Last ExamHigh-difficulty academic knowledge and reasoning across disciplines.CAIS / Scale AIObserving
- BullshitBench · false premisesIdentify false premises in a question and observe whether the model keeps fabricating along the error.Peter Gostev / BullshitBenchObserving
- SimpleQA VerifiedFactual accuracy on short questionsEpoch AIRankedEvidence budget 4%
- GPQA DiamondGraduate-level physics, chemistry and biology knowledgeEpoch AIRankedEvidence budget 2%
Professional work
19Financial analysis, legal advice and banking, to see whether professional tasks can be completed.
- APEX-Agents 1.1Long tasks from investment banking, consulting and law, testing whether a model can finish a complete piece of work as an assistant.MercorRankedEvidence budget 4%
- Vals Finance AgentEveryday office tasks for financial analysts, to see how work such as research reports and modelling is done.Vals AIRankedEvidence budget 3%
- OfficeQA ProRetrieving, calculating and answering questions from large volumes of real financial documents.DatabricksObserving
- Harvey Legal Agent BenchmarkProcess legal materials and deliver reviewable files such as Word, Excel and slides.Vals AI / HarveyObserving
- MCP AtlasRetrieve material, operate documents and databases through real MCP services to complete cross-software tasks.Scale AIObserving
- FinanceBenchmarkComplete financial pricing, capital calculation and risk control tasks, verified with numerical and code execution results.FinanceBenchmarkObserving
- PRBench · LegalAnswering complex questions in real legal work, testing professional judgement and the quality of argument.Scale AIObserving
- PRBench · FinanceAnswering real financial decision questions, testing professional judgement, reasoning and risk explanation.Scale AIObserving
- SpreadsheetBench 2Complete a full spreadsheet workload covering financial modelling, workbook fixes and data visualisation.SpreadsheetBench / Renmin University / AfterQueryObserving
- Vending-Bench 2Run a simulated shop for a year, handling procurement, negotiation, inventory and pricing.Andon LabsObserving
- τ³-BankingConducting banking business in a unified tool environment, testing rule understanding, customer communication and task completion.SierraAwaiting resultsEvidence budget 2%
- EBR-benchLearn new rules and use memory during long tasksEpoch AIReference only
- Excel · EMBSpreadsheet editing and Excel office workVals AIObserving
- Legal ResearchLegal retrieval and argumentationVals AIObserving
- TaxAgentBenchTax speciality questionsVals AIObserving
- MedScribeClinical record organisation and writingVals AIObserving
- BioMysteryBenchBiology research reasoningVals AIObserving
- AutomationBenchAutomating work across office softwareZapierObserving
- Terminal-Bench ScienceCompleting scientific research tasks in the terminalHarborObserving
Visual understanding
2Evidence of understanding images remains in the overall ranking; reading images and design aesthetics are different abilities.
Chinese and multilingual
2Supplementary evidence on language fundamentals; English writing performance is not treated as Chinese ability.
“Observing” means its data, running conditions or limits are still being checked; it is not part of the overall or category boards. Repeated collection or display slices of one evaluation add no weight; overlap between different evaluations' question sets is still being checked.