Google Research·· 2026-06-27
Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction
Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction
AI summary
Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.
Selection record
Threshold 60Official, first-handFirst 58Second 62
AdmittedSum of both 120 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:介绍Gemini Nano模型MTP加速技术
- Why it was chosen
- Adding MTP heads to frozen Gemini Nano v3 shows how on-device speculative decoding avoids the memory overhead of a separate draft model.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google