TapTap Maker Benchmark
TapTap Maker / Coding · Writing working features in Lua inside a real game engine. The score is judged by engine execution, not by another model.
What it measures, and how
Only the official main_board L2 pass rate is used; cost, speed, Held-out and category scores are not voted on again. For the same base model, the representative row is chosen by a fixed reasoning-tier priority, and the choice does not look at scores.
How this evidence is used
3.6% of the coding and design budget; engine and L2 metrics fixed.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | claude-fable-5-1Anthropic | 90.5% | High reasoning |
| 2 | claude-fable-5Anthropic | 89.0% | High reasoning |
| 3 | gpt-6-astraOpenAI | 88.9% | xHigh reasoning |
| 4 | claude-opus-5Anthropic | 88.9% | xHigh reasoning |
| 6 | grok-4.6@xhighxAI | 87.3% | xHigh reasoning |
| 7 | gemini-3.8-flashGoogle | 86.8% | High reasoning |
| 8 | gpt-5.6-solOpenAI | 86.5% | xHigh reasoning |
| 11 | glm-5.3Z.ai | 83.6% | Max reasoning |
| 12 | qwen3.8-max-0902Alibaba | 83.3% | xHigh reasoning |
| 13 | claude-opus-4-8Anthropic | 82.8% | xHigh reasoning |
| 14 | qwen3.8-maxAlibaba | 82.5% | xHigh reasoning |
| 15 | deepseek-v4.1-flashDeepSeek | 82.5% | Max reasoning |
| 16 | qwen3.8-flashAlibaba | 82.5% | xHigh reasoning |
| 17 | gpt-5.6-terraOpenAI | 81.8% | xHigh reasoning |
| 18 | glm-5.3-flashZ.ai | 81.5% | Max reasoning |
| 19 | hy4-previewTencent | 81.0% | High reasoning |
| 20 | kimi-k3Moonshot AI | 79.7% | Source default configuration |
| 21 | grok-4.5xAI | 78.3% | High reasoning |
| 22 | doubao-seed-evolvingByteDance | 78.3% | High reasoning |
| 23 | gpt-5.6-lunaOpenAI | 77.8% | xHigh reasoning |
| 24 | deepseek-v4-proDeepSeek | 77.8% | Source default configuration |
| 25 | claude-sonnet-5Anthropic | 76.2% | xHigh reasoning |
| 27 | deepseek-v4-flashDeepSeek | 74.6% | Max reasoning |
| 28 | qwen3.8-27bAlibaba | 68.3% | xHigh reasoning |
| 29 | gemini-3.1-pro-previewGoogle | 64.7% | Source default configuration |
| 30 | MiniMax-M3MiniMax | 60.1% | Source default configuration |
| 31 | claude-haiku-4-5Anthropic | 43.6% | High reasoning |
Limits and data attribution
Single engine, single language, and cannot be generalised to other tech stacks; different models may use different agent drivers, and a closed-source harness cannot be recomputed question by question.
Data licence: Official public results; public re-display retains official attribution, and full redistribution authorisation remains subject to Yìwán/TapTap terms
Scores published by TapTap Maker. Raw scores and the News consensus score use different scales and cannot be added directly.