Skip to content
Google DeepMind·· 2026-06-09

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

AI summary

Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.

Selection record

AdmittedSum of both 160 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:发布Gemma 4多模态模型,含架构与开发工具
Why it was chosen
Gemma 4 12B fits multimodal capabilities into 16GB of VRAM through an encoder-free architecture, enabling performance comparisons with 26B MoE.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Google DeepMind · deepmind.google