Google DeepMind·· 2026-06-09
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
AI summary
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Selection record
Threshold 60Official, first-handFirst 82Second 78
AdmittedSum of both 160 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:发布Gemma 4多模态模型,含架构与开发工具
- Why it was chosen
- Gemma 4 12B fits multimodal capabilities into 16GB of VRAM through an encoder-free architecture, enabling performance comparisons with 26B MoE.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google DeepMind · deepmind.google