Skip to content
Evaluation sources

TapTap Maker Benchmark

TapTap Maker / Coding · Writing working features in Lua inside a real game engine. The score is judged by engine execution, not by another model.

Official evaluation
In NewsRanked
Evidence budget3.6%
Upstream data as of09/10 08:00
Last synced10/01 20:05

What it measures, and how

Only the official main_board L2 pass rate is used; cost, speed, Held-out and category scores are not voted on again. For the same base model, the representative row is chosen by a fixed reasoning-tier priority, and the choice does not look at scores.

How this evidence is used

3.6% of the coding and design budget; engine and L2 metrics fixed.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1claude-fable-5-1Anthropic90.5%High reasoning
2claude-fable-5Anthropic89.0%High reasoning
3gpt-6-astraOpenAI88.9%xHigh reasoning
4claude-opus-5Anthropic88.9%xHigh reasoning
6grok-4.6@xhighxAI87.3%xHigh reasoning
7gemini-3.8-flashGoogle86.8%High reasoning
8gpt-5.6-solOpenAI86.5%xHigh reasoning
11glm-5.3Z.ai83.6%Max reasoning
12qwen3.8-max-0902Alibaba83.3%xHigh reasoning
13claude-opus-4-8Anthropic82.8%xHigh reasoning
14qwen3.8-maxAlibaba82.5%xHigh reasoning
15deepseek-v4.1-flashDeepSeek82.5%Max reasoning
16qwen3.8-flashAlibaba82.5%xHigh reasoning
17gpt-5.6-terraOpenAI81.8%xHigh reasoning
18glm-5.3-flashZ.ai81.5%Max reasoning
19hy4-previewTencent81.0%High reasoning
20kimi-k3Moonshot AI79.7%Source default configuration
21grok-4.5xAI78.3%High reasoning
22doubao-seed-evolvingByteDance78.3%High reasoning
23gpt-5.6-lunaOpenAI77.8%xHigh reasoning
24deepseek-v4-proDeepSeek77.8%Source default configuration
25claude-sonnet-5Anthropic76.2%xHigh reasoning
27deepseek-v4-flashDeepSeek74.6%Max reasoning
28qwen3.8-27bAlibaba68.3%xHigh reasoning
29gemini-3.1-pro-previewGoogle64.7%Source default configuration
30MiniMax-M3MiniMax60.1%Source default configuration
31claude-haiku-4-5Anthropic43.6%High reasoning
Limits and data attribution

Single engine, single language, and cannot be generalised to other tech stacks; different models may use different agent drivers, and a closed-source harness cannot be recomputed question by question.

Data licence: Official public results; public re-display retains official attribution, and full redistribution authorisation remains subject to Yìwán/TapTap terms

Scores published by TapTap Maker. Raw scores and the News consensus score use different scales and cannot be added directly.