Google Open-Sources Gemma 4: Four Sizes, 256K Context, Runs Locally

·Toolin Editorial Team

Google releases the Gemma 4 open model family (E2B/E4B/26B MoE/31B Dense) under Apache 2.0, with 256K context, native multimodality, and local deployment channels.

Google Open-Sources Gemma 4: Four Sizes, 256K Context, Runs Locally

Google has released Gemma 4 — its strongest open model family to date. It shares lineage with the closed-source Gemini 3 but ships under Apache 2.0: you can download the weights and run them on your own machines, in your own cloud accounts, with data never leaving your premises. For developers who need data sovereignty, offline inference, or just lower API bills, this generation finally brings "frontier capability" and "runs out of the box locally" together.

Four Sizes — Pick by Hardware

Gemma 4 isn't one model but four versions optimized for different hardware:

ModelTypeBest-suited hardware
E2BEdge, effective 2B parametersPhones, Raspberry Pi, NVIDIA Jetson Orin Nano
E4BEdge, effective 4B parametersSame as above, more capable
26B MoEMixture-of-experts, only 3.8B active at inferenceConsumer GPUs, built for low latency
31B DenseDense model, the largest versionA single 80GB H100; runs on consumer GPUs once quantized

The 26B MoE takes the "low latency" route — 26B total parameters with only 3.8B active at inference, so tokens-per-second is fast. The 31B Dense takes the "highest quality + fine-tuning base" route. Google says the 31B currently ranks 3rd among open models worldwide on the Arena AI text leaderboard, with the 26B at 6th.

Core Capabilities

  • Long context: 128K on the edge models, 256K on the 26B/31B — enough to fit an entire code repository or a long document into a single prompt
  • Native multimodality: all four versions natively process video and images (variable resolution, OCR, chart understanding); the E2B/E4B also support native audio input
  • Agent workflows: native function calling, structured JSON output, and system instructions — well suited to building automation agents
  • Code generation: works as a local coding assistant, writing code offline
  • 140+ languages trained natively
  • Reasoning: multi-step planning and deep logic, with marked gains on math and instruction-following tasks

How to Get It Running

Google has spread availability wide, covering nearly every mainstream local inference stack:

Try in the Cloud (Zero Setup)

  • Google AI Studio: run the 31B and 26B MoE directly
  • Google AI Edge Gallery: run the E4B and E2B

Local Deployment

Weights are published on Hugging Face, Kaggle, and Ollama, with day-one support across:

  • Hugging Face (Transformers, TRL, Transformers.js, Candle)
  • Ollama (ollama run gemma4:31b-it-q4_K_M)
  • LM Studio
  • vLLM, llama.cpp, MLX
  • NVIDIA NIM / NeMo
  • Unsloth, SGLang, Cactus, Baseten, Docker, MaxText, Tunix, Keras

Android Development

The edge models are callable via the AICore Developer Preview, forward-compatible with Gemini Nano 4. You can also prototype agent flows with Android Studio's Agent Mode.

💡 Tip: for a quick try, the fastest path is a single Ollama command: ollama run gemma4:31b-it-q4_K_M. For an in-IDE coding assistant, LM Studio plus the Continue plugin is the smoothest combination.

Fine-Tuning

Gemma 4 is designed to be fine-tunable on consumer hardware (a gaming GPU, say). Available platforms include:

  • Google Colab (free/paid)
  • Vertex AI
  • Your own GPUs

The official examples cite two fine-tuning cases: INSAIT trained the Bulgarian-first BgGPT with it, and Yale University used it for Cell2Sentence-Scale to discover new cancer treatment pathways.

Ecosystem and Hardware Support

  • NVIDIA: optimized across the line, from Jetson Orin Nano to Blackwell GPUs
  • AMD: integrated via the open-source ROCm stack
  • TPU: deployable at scale on Trillium and Ironwood TPUs
  • Google Cloud: Vertex AI, Cloud Run, GKE, Sovereign Cloud, and TPU-accelerated inference

License

Apache 2.0 — commercially friendly, with no usage gates. You can deploy the model locally or in any cloud environment, keeping full control over your data. That's a notably bigger opening from Google than in previous generations.

Who It's For

  • Teams that need data sovereignty: finance, healthcare, government — anywhere data can't be sent to a third-party API
  • Developers looking to cut API costs: deploy once, call without limit, zero marginal cost
  • Offline scenarios: mobile, IoT, and environments with unreliable networks
  • Researchers: Apache 2.0 plus multiple sizes makes experimentation and fine-tuning easy

References

Related articles

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
AI Products

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion

Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Toolin Editorial Team
Qwen3.8-Max Preview: A Hands-On Early Access Guide
AI Products

Qwen3.8-Max Preview: A Hands-On Early Access Guide

Qwen's 2.4T-parameter flagship preview is live on Token Plan, Qoder, and the Qwen website; officially rated second only to Fable 5 overall.

Toolin Editorial Team
Making a Full-Scene Infographic with SenseNova U1 Pro
AI Tutorials

Making a Full-Scene Infographic with SenseNova U1 Pro

A hands-on tutorial for SenseTime's flagship multimodal model U1 Pro: turn raw data into a deliverable 8K infographic and full-match panoramic visual, automatically.

Toolin Editorial Team
Claude for Teachers: A Free AI Teaching Assistant for Every K-12 Teacher in the US
AI Products

Claude for Teachers: A Free AI Teaching Assistant for Every K-12 Teacher in the US

Anthropic launches a free AI teaching assistant for certified US K-12 teachers, wired into 50 states' standards, with lesson-plan alignment, differentiated tiering, and student data analysis.

Toolin Editorial Team
A Step-by-Step Codex Desktop Pet Tutorial: Build an Animated Desk Companion with Hatch Pet
AI Tutorials

A Step-by-Step Codex Desktop Pet Tutorial: Build an Animated Desk Companion with Hatch Pet

Build a custom desktop pet for Codex with Hatch Pet from the OpenAI Skills repo — 5 steps from reference image to install and wake-up, roughly 1 hour and 60% of weekly usage in our test.

Toolin Editorial Team
Turning an n8n Workflow into a Codex Skill: A Meta-Skill's Five-Step Pipeline
AI Tutorials

Turning an n8n Workflow into a Codex Skill: A Meta-Skill's Five-Step Pipeline

With the n8n-to-codex-skill meta-skill, convert any of n8n's 10,000+ public templates into a native Codex Skill — includes two real case studies: an SEO audit skill and a GEO content skill.

Toolin Editorial Team