Google DeepMind Unveils Gemma 4: Advanced, Efficient, and Versatile AI Models for All Devices
September 24, 2026
Gemma 4 is Google DeepMind's open model family available in five sizes from mobile to workstation, designed around intelligence-per-parameter rather than merely largest size.
The 12B Unified model uniquely omits encoders for images and audio by embedding pixel and waveform directly, enabling a single model to handle text, images, and audio.
Gemma 4 supports multi-token prediction to speed up inference by up to about three times without sacrificing quality, and all models are Apache 2.0 licensed with no usage caps.
Android readiness and developer previews place Gemma 4 on mobile, with E4B and E2B demos promising roughly fourfold speed and around 60% battery savings versus the prior generation.
Google benchmarks show clear improvements over Gemma 3, illustrating how model size correlates with tasks like MMLU, AIME, and code benchmarks (LiveCodeBench v6).
Community and ecosystem impact is strong, with Gemma 4 open variants surpassing 1 billion downloads by August 2026 and more than 400 million downloads since the first release, signaling active engagement and rapid downstream fine-tuning.
Gemma 4 aims for broader accessibility of open models with practical hardware guidance, including a baseline 16 GB RAM and 24–32 GB for larger models, recommending starting with the 26B A4B for real workloads.
Among the five options, the 26B A4B Mixture-of-Experts model operates with roughly 4B effective parameters per token, delivering high intelligence with lower compute than typical 26B models.
Strategically, Google positions Gemma 4 as a cost-conscious, efficient open-model family that emphasizes total cost of ownership, price competitiveness, API access, and on-device viability.
The five sizes are E2B (mobile, 1.1 GB), E4B (mobile, 2.5 GB), 12B Unified (multimodal, ~6.7 GB), 26B A4B (workstation-grade, ~14.4 GB), and 31B Dense (highest quality, ~17.5 GB).
Summary based on 1 source