Model Family
What is Gemma 4?
Google DeepMind's open-weight model family released May 2026 under the Apache 2.0 license, built from Gemini 3 research. Gemma 4 includes four variants: E2B (2.3B active, ~155 tok/s, runs on iPhone), E4B (4B effective, ~125 tok/s, multimodal+audio), 26B-A4B (MoE with 128 experts, 3.8B active, Arena AI #6), and 31B Dense (Arena AI #3, strongest open model under 100B). All variants support 256K context, text+image input, and native function calling. The E2B/E4B models use Per-Layer Embeddings (PLE) and also accept audio input. According to the LLMCheck index, Gemma 4 26B-A4B scores 67/100 on the leaderboard, the highest of any model.
Where Gemma 4 comes up on LLMCheck
- Local AI Troubleshooting Hub — Fix Common LLM Issues on Mac
- How to Fine-Tune a Local LLM on Mac with MLX (LoRA) — 2026 Guide
- Best Local LLMs for MacBook Air (2026) — M1 to M5, 8–24 GB
- Function Calling & Tool Use with Local LLMs on Mac
- Local LLM Not Using GPU on Mac? How to Enable Metal Acceleration
- Why Is My Local LLM Slow on Mac? 7 Fixes for Faster Inference
Browse all 81 models in the index →