Skip to content
Back to Overall

DeepSeek V4 Pro Preview

DeepSeek · released 2026-04-24 · updated 10/01 14:05

Overall consensus index35.4Beyond the overall top 30
Category results3/ 4 categories
Results on record14evaluations
Context window—Token
API input / output · per million tokens¥2.92 / ¥5.83cached ¥0.0243list $0.435 / $0.87Provider's official price

Strengths, seen one at a time.

Each capability is computed on its own. Without enough measurements, it stays empty.

Each category's score reflects ranking support within its own reference group; the scores cannot be added up or used to compare how strong different capabilities are.

How stable is the overall rank?

Remove one evaluation or one organisation at a time, change single weights and the handling of error, and see how the rank moves.

Rank after eligibility is checked againNo. 40–49Evidence-sensitive

It stayed eligible in every completed comparison. With the candidates held fixed it ranks 40–49. This range is not a confidence interval and does not include results that were never published.

Every result has a source.

The model's published aggregate results. Open one to see its run configuration and how it was used.

Overall evaluation and experience4

Coding2

Reasoning5

Knowledge2

Professional work1

Scored evaluations without a result

These evaluations have published no result for this model; a missing result does not count as zero.

What is still unknown

Missing evaluations do not count as zero. The rank moves with new evidence; when scores are close, do not read much into small gaps.

How it is calculated