Google DeepMind launched Gemma 4 on April 2, 2026, a family of open models released under the Apache 2.0 license in five parameter sizes from E2B to 31B. Per Google's documentation, the models run on-device rather than only in the cloud, enabling agents that plan and act without a network connection.
What does Gemma 4 actually do per its documentation?
Gemma 4 is a set of open-weight language models designed for on-device AI development. Google's announcement states the family supports multi-step planning, autonomous action, offline code generation and audio-visual processing without specialized fine-tuning, with support for more than 140 languages. The smallest variants, E2B and E4B, target ultra-mobile and browser deployment on hardware like Pixel phones and Chrome; the dense 31B model bridges toward server use; and a 26B A4B variant mixes the two regimes with sparse activation.
The documentation also specifies the engineering behind the speed claims. Every Gemma 4 model includes a dedicated draft model for speculative decoding, a technique where a small model proposes tokens that the large model verifies, cutting inference latency with no quality loss according to Google. Google publishes approximate memory requirements per size and precision, from 0.84 GB for a text-only mobile E2B build up to 69.9 GB for a 16-bit 31B deployment, which lets developers match a variant to a device before writing any code.
Why does the Apache 2.0 switch matter?
The license is the quiet headline. Previous Gemma generations shipped under Google's custom terms, which developers found restrictive, and Ars Technica's report explains that the switch to Apache 2.0 removes the overbearing terms of use and commercial restrictions that had made many developers apprehensive about building on Gemma. Apache 2.0 is a permissive standard license that Google cannot unilaterally reinterpret later, which matters for companies betting products on open weights they do not control at the source.
The release also lands in a specific competitive context. Open-weight families compete on what developers may legally build, and a permissive license widens the pool of commercial users who can ship Gemma-derived products without counsel reviewing custom terms. That is a distribution decision as much as a technical one.
What can these models not do yet?
Open models trade capability for portability. Gemma 4's largest documented size, 31B parameters, is far below frontier-scale closed models, and its headline capabilities have not been evaluated in the announcement against named third-party benchmarks. On-device agents also inherit the memory ceilings of phones and laptops, so the audio-visual and planning features depend on which of the five variants a device can host.
What the release demonstrably provides, per the documentation, is a licensed, multi-size model family built for edge deployment, with the Agent Skills application in Google's AI Edge Gallery showing multi-step workflows, such as querying Wikipedia and building study flashcards, running entirely on a phone. The gap between that demonstration and unsupervised daily-use agents remains the open question, and it is a question of memory, not of licensing.

