Skip to content
Back to Overall

Muse Spark 1.1

Meta · released 2026-07-09 · updated 10/01 14:05

Overall consensus index61.0No. 22 overall
Category results2/ 4 categories
Results on record13evaluations
Context window—Token
API input / output · per million tokens¥8.38 / ¥28.49list $1.25 / $4.25Provider's official price

Strengths, seen one at a time.

Each capability is computed on its own. Without enough measurements, it stays empty.

Each category's score reflects ranking support within its own reference group; the scores cannot be added up or used to compare how strong different capabilities are.

How stable is the overall rank?

Remove one evaluation or one organisation at a time, change single weights and the handling of error, and see how the rank moves.

Rank after eligibility is checked againNo. 21–32Evidence-sensitive

It stayed eligible in every completed comparison. With the candidates held fixed it ranks 22–32. This range is not a confidence interval and does not include results that were never published.

Every result has a source.

The model's published aggregate results. Open one to see its run configuration and how it was used.

Overall evaluation and experience6

Coding2

Reasoning1

Knowledge1

Professional work2

Visual understanding1

Scored evaluations without a result

These evaluations have published no result for this model; a missing result does not count as zero.

What is still unknown

Missing evaluations do not count as zero. The rank moves with new evidence; when scores are close, do not read much into small gaps.

How it is calculated