Skip to content
Google Research·· 2026-06-27

Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction

Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction

AI summary

Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.

Selection record

AdmittedSum of both 120 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:介绍Gemini Nano模型MTP加速技术
Why it was chosen
Adding MTP heads to frozen Gemini Nano v3 shows how on-device speculative decoding avoids the memory overhead of a separate draft model.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Google Research · research.google