Google Gemma
Open SourceGoogle DeepMind's open-weight LLM family, from phones to workstations.
Published 29 May 2026 · Last updated 27 September 2026
Scores
Popularity4/5
The google-deepmind/gemma GitHub repository has ~5.1K stars; Gemma models on Hugging Face have millions of downloads. Widely used in on-device AI, edge deployment, and fine-tuning research. Strong adoption in the open-source ML community, though Meta Llama commands a larger absolute user base and mindshare.
Learning Curve2/5
Running Gemma locally is accessible to any developer: `ollama run gemma3` requires no account, no API key, and no GPU on smaller models. The Hugging Face Transformers integration is industry-standard. Fine-tuning via LoRA or full fine-tuning adds moderate complexity but is well-documented.
Flexibility5/5
Multiple sizes (1B to 31B), multiple architectures (dense and MoE), multimodal variants (PaliGemma), code-specialised variants (CodeGemma), domain fine-tunes (MedGemma), and embedding models (EmbeddingGemma). Deployable on any hardware — phones, laptops, consumer GPUs, cloud VMs. Runs with Ollama, vLLM, llama.cpp, JAX, PyTorch, Keras, and Transformers. Apache 2.0 (Gemma 4) allows full fine-tuning and redistribution.
Performance4/5
Gemma models consistently top their weight class on standard benchmarks. Gemma 2 27B rivalled 70B-class models at release; Gemma 3 27B set new small-model records on MMLU and reasoning. Gemma 4 31B reaches 85.2% MMLU Pro and 89.2% AIME 2026. However, top-of-market reasoning (GPT-5, Claude Opus, Gemini 3.1 Pro) still exceeds the 31B ceiling.
Portability5/5
Maximum portability — Gemma models run on CPUs, consumer GPUs, server GPUs, TPUs, mobile phones (E2B/E4B variants), and Raspberry Pi-class hardware. GGUF quantisation via llama.cpp enables deployment anywhere. No cloud lock-in; weights are self-contained.
About Google Gemma
Gemma is Google DeepMind's family of open-weight models, built from the same research as Gemini and published on Hugging Face and Kaggle. The current generation, Gemma 4, comes in five core configurations: E2B and E4B for phones and edge devices, a 12B Unified model that handles text, image, audio, and video natively without separate encoders, a 31B dense flagship, and a 26B mixture-of-experts model with 3.8B active parameters that runs on a single consumer GPU. Context windows range from 128K to 256K tokens, and the models support configurable thinking modes, native function calling, and multi-token prediction.
Gemma 4 is released under Apache 2.0, replacing the custom Gemma Terms of Use that governed earlier generations, so it can be fine-tuned, redistributed, and built into products without extra restrictions. Specialised variants cover specific jobs: PaliGemma 2 for vision-language, MedGemma for medical text and images, TranslateGemma for translation, EmbeddingGemma for retrieval, ShieldGemma 2 for safety classification, CodeGemma for code, and DiffusionGemma, which generates text through diffusion rather than token by token for much higher throughput.
Gemma runs locally through Ollama, llama.cpp and Gemma.cpp, vLLM, LM Studio, and Unsloth, or in PyTorch, JAX, and Keras, and Android Studio bundles it for offline coding assistance. Google offers free, rate-limited access in AI Studio and scalable serving on Google Cloud, but sells no separately priced Gemma API; per-token prices on aggregators such as OpenRouter come from third-party inference providers.
Key Features
- Gemma 4 in five configs: E2B, E4B, 12B Unified, 31B dense, and 26B MoE (3.8B active)
- 12B Unified handles text, image, audio, and video natively and runs on a 16GB laptop
- 128K to 256K context depending on configuration
- Configurable thinking modes, native function calling, and multi-token prediction
- Apache 2.0 license for Gemma 4
- Specialised variants: PaliGemma 2, MedGemma, TranslateGemma, EmbeddingGemma, ShieldGemma 2, CodeGemma, DiffusionGemma
- Self-host via Ollama, llama.cpp, vLLM, LM Studio, PyTorch, JAX, or Keras
- Free access in Google AI Studio; scalable serving on Google Cloud
Pros
- Strong performance for its parameter count compared with similarly sized open models
- The 26B MoE gives near-31B quality with only 3.8B active parameters, cutting inference cost
- Runs on phones, laptops, consumer GPUs, and the cloud
- Apache 2.0 removes commercial restrictions: fine-tune, redistribute, and ship freely
- Broad support across Ollama, vLLM, llama.cpp, Hugging Face, Keras, JAX, and PyTorch
- Specialised variants give task-tuned starting points for fine-tuning
Cons
- Gemma 1 to 3 remain under the custom Gemma Terms of Use, with prohibited-use clauses that flow down to derivatives
- The largest variant is 31B, so teams needing 70B-class reasoning must look to Llama, Qwen, or DeepSeek
- On-device deployment needs quantisation and memory tuning that cloud APIs avoid
- CodeGemma and PaliGemma trail newer specialist models on task benchmarks
- Fast release cadence means older Gemma versions go stale quickly
- Fine-tuning the larger variants still needs substantial GPU memory even with LoRA
Google Gemma Pricing
Open SourceTech Stacks with Google Gemma
Self-Hosted AI with Ollama and Open WebUI
InfrastructureA private, ChatGPT-style assistant on hardware you own. Ollama runs open-weight models such as Google Gemma or Qwen, and Open WebUI gives them a chat interface in the browser with accounts, document search, and model management. Both run in Docker, and nothing leaves your machine.
Local LLM Dev Stack
DeveloperA coding agent that works against a model on your own hardware. A local runtime such as Ollama or LM Studio serves an open-weight coding model such as Qwen or Mistral's Devstral, and an agent such as OpenCode or Cline does the editing. No code leaves the machine and there is no per-token bill.
Tools Related to Google Gemma
Works well with Google Gemma(2)
Unsloth maintains quantized builds and fine-tuning support for Gemma models.
Gemma's small open-weight models fit on consumer laptops, and LM Studio downloads and runs them locally with a chat interface and a local API server.
Alternatives to Google Gemma(9)
Google Gemma and Meta Llama are open-weight alternatives from major AI labs. Gemma excels at small model efficiency (e.g. the 26B MoE fits on a consumer GPU) while Llama covers sizes from 1B to 405B+ with a larger fine-tune ecosystem.
Gemma is Google's compact open family for on-device and self-hosted use; Qwen spans sub-1B models to a 2.4T flagship across reasoning, coding, and vision, mostly under Apache 2.0. Gemma for small models, Qwen for range and scale.
GLM is Z.ai's bilingual Chinese and English family with an OpenAI-compatible API and free Flash models; Gemma is Google DeepMind's family built for self-hosting and on-device use up to 31B. GLM for hosted bilingual use, Gemma for compact self-hosted models.
Gemma targets self-hosting and on-device use with models up to 31B; MiniMax M3 is a large model with a 1M-token context and native multimodal input. Gemma for small, efficient deployments, MiniMax for long-context agentic work.
Gemma focuses on compact models for self-hosting and edge devices; Xiaomi MiMo's 1T MoE flagship targets agentic tasks with omnimodal input under MIT. Gemma for lightweight deployment, MiMo for frontier-scale open weights.
Google Gemma (open-weight) and Google Gemini (proprietary API) are alternatives from Google. Gemma provides self-hostable weights (Apache 2.0 for Gemma 4) while Gemini offers a richer managed API with multimodal capabilities, grounding, and code execution.