Google Open-Sources Gemma 4: Four Sizes, 256K Context, Runs Locally
Google releases the Gemma 4 open model family (E2B/E4B/26B MoE/31B Dense) under Apache 2.0, with 256K context, native multimodality, and local deployment channels.


Google Open-Sources Gemma 4: Four Sizes, 256K Context, Runs Locally
Google releases the Gemma 4 open model family (E2B/E4B/26B MoE/31B Dense) under Apache 2.0, with 256K context, native multimodality, and local deployment channels.
Google has released Gemma 4 — its strongest open model family to date. It shares lineage with the closed-source Gemini 3 but ships under Apache 2.0: you can download the weights and run them on your own machines, in your own cloud accounts, with data never leaving your premises. For developers who need data sovereignty, offline inference, or just lower API bills, this generation finally brings "frontier capability" and "runs out of the box locally" together.
Four Sizes — Pick by Hardware
Gemma 4 isn't one model but four versions optimized for different hardware:
| Model | Type | Best-suited hardware |
|---|---|---|
| E2B | Edge, effective 2B parameters | Phones, Raspberry Pi, NVIDIA Jetson Orin Nano |
| E4B | Edge, effective 4B parameters | Same as above, more capable |
| 26B MoE | Mixture-of-experts, only 3.8B active at inference | Consumer GPUs, built for low latency |
| 31B Dense | Dense model, the largest version | A single 80GB H100; runs on consumer GPUs once quantized |
The 26B MoE takes the "low latency" route — 26B total parameters with only 3.8B active at inference, so tokens-per-second is fast. The 31B Dense takes the "highest quality + fine-tuning base" route. Google says the 31B currently ranks 3rd among open models worldwide on the Arena AI text leaderboard, with the 26B at 6th.
Core Capabilities
- Long context: 128K on the edge models, 256K on the 26B/31B — enough to fit an entire code repository or a long document into a single prompt
- Native multimodality: all four versions natively process video and images (variable resolution, OCR, chart understanding); the E2B/E4B also support native audio input
- Agent workflows: native function calling, structured JSON output, and system instructions — well suited to building automation agents
- Code generation: works as a local coding assistant, writing code offline
- 140+ languages trained natively
- Reasoning: multi-step planning and deep logic, with marked gains on math and instruction-following tasks
How to Get It Running
Google has spread availability wide, covering nearly every mainstream local inference stack:
Try in the Cloud (Zero Setup)
- Google AI Studio: run the 31B and 26B MoE directly
- Google AI Edge Gallery: run the E4B and E2B
Local Deployment
Weights are published on Hugging Face, Kaggle, and Ollama, with day-one support across:
- Hugging Face (Transformers, TRL, Transformers.js, Candle)
- Ollama (
ollama run gemma4:31b-it-q4_K_M) - LM Studio
- vLLM, llama.cpp, MLX
- NVIDIA NIM / NeMo
- Unsloth, SGLang, Cactus, Baseten, Docker, MaxText, Tunix, Keras
Android Development
The edge models are callable via the AICore Developer Preview, forward-compatible with Gemini Nano 4. You can also prototype agent flows with Android Studio's Agent Mode.
💡 Tip: for a quick try, the fastest path is a single Ollama command:
ollama run gemma4:31b-it-q4_K_M. For an in-IDE coding assistant, LM Studio plus the Continue plugin is the smoothest combination.
Fine-Tuning
Gemma 4 is designed to be fine-tunable on consumer hardware (a gaming GPU, say). Available platforms include:
- Google Colab (free/paid)
- Vertex AI
- Your own GPUs
The official examples cite two fine-tuning cases: INSAIT trained the Bulgarian-first BgGPT with it, and Yale University used it for Cell2Sentence-Scale to discover new cancer treatment pathways.
Ecosystem and Hardware Support
- NVIDIA: optimized across the line, from Jetson Orin Nano to Blackwell GPUs
- AMD: integrated via the open-source ROCm stack
- TPU: deployable at scale on Trillium and Ironwood TPUs
- Google Cloud: Vertex AI, Cloud Run, GKE, Sovereign Cloud, and TPU-accelerated inference
License
Apache 2.0 — commercially friendly, with no usage gates. You can deploy the model locally or in any cloud environment, keeping full control over your data. That's a notably bigger opening from Google than in previous generations.
Who It's For
- Teams that need data sovereignty: finance, healthcare, government — anywhere data can't be sent to a third-party API
- Developers looking to cut API costs: deploy once, call without limit, zero marginal cost
- Offline scenarios: mobile, IoT, and environments with unreliable networks
- Researchers: Apache 2.0 plus multiple sizes makes experimentation and fine-tuning easy
References
Related articles

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Qwen3.8-Max Preview: A Hands-On Early Access Guide
Qwen's 2.4T-parameter flagship preview is live on Token Plan, Qoder, and the Qwen website; officially rated second only to Fable 5 overall.

Making a Full-Scene Infographic with SenseNova U1 Pro
A hands-on tutorial for SenseTime's flagship multimodal model U1 Pro: turn raw data into a deliverable 8K infographic and full-match panoramic visual, automatically.

Claude for Teachers: A Free AI Teaching Assistant for Every K-12 Teacher in the US
Anthropic launches a free AI teaching assistant for certified US K-12 teachers, wired into 50 states' standards, with lesson-plan alignment, differentiated tiering, and student data analysis.

A Step-by-Step Codex Desktop Pet Tutorial: Build an Animated Desk Companion with Hatch Pet
Build a custom desktop pet for Codex with Hatch Pet from the OpenAI Skills repo — 5 steps from reference image to install and wake-up, roughly 1 hour and 60% of weekly usage in our test.

Turning an n8n Workflow into a Codex Skill: A Meta-Skill's Five-Step Pipeline
With the n8n-to-codex-skill meta-skill, convert any of n8n's 10,000+ public templates into a native Codex Skill — includes two real case studies: an SEO audit skill and a GEO content skill.