DeepSeek V4 Flash Preview
DeepSeek · released 2026-04-24 · updated 10/01 14:05
Strengths, seen one at a time.
Each capability is computed on its own. Without enough measurements, it stays empty.
Each category's score reflects ranking support within its own reference group; the scores cannot be added up or used to compare how strong different capabilities are.
How stable is the overall rank?
Remove one evaluation or one organisation at a time, change single weights and the handling of error, and see how the rank moves.
In 6 scenarios the evidence was not enough to rank it. With the candidates held fixed it ranks 59–68. This range is not a confidence interval and does not include results that were never published.
Every result has a source.
The model's published aggregate results. Open one to see its run configuration and how it was used.
Overall evaluation and experience4
Coding2
Reasoning1
Scored evaluations without a result
These evaluations have published no result for this model; a missing result does not count as zero.
What is still unknown
Missing evaluations do not count as zero. The rank moves with new evidence; when scores are close, do not read much into small gaps.
How it is calculated