Architecture
What is Per-Layer Embeddings (PLE)?
A technique introduced in Google's Gemma 4 E2B and E4B models that feeds a secondary embedding signal into every decoder layer. PLE allows a 2.3B-active model to carry the representational depth of 5.1B parameters while fitting in under 1.5 GB of memory with quantization. This is why Gemma 4 E2B and E4B punch well above their weight class, achieving quality closer to models 2–3x their active parameter count. PLE is distinct from MoE — it adds depth without adding width.
Where Per-Layer Embeddings (PLE) comes up on LLMCheck
- How to Run Google Gemma 4 on Mac: Complete Setup Guide & Benchmarks
- Gemma 4 E2B & E4B: Run Google's AI on iPhone, iPad & Mac Mini
- Gemma 4 vs Qwen 3.5: Which Is the Best Local LLM for Mac in 2026?
- Qwen 3.6 vs Gemma 4: Deep Technical Comparison for Mac (2026)
- Gemma 4 E4B on M6: 55 tokens/sec — Benchmark & Setup (2026)
- Gemma 4 E4B on M4 Pro: 84 tokens/sec — Benchmark & Setup (2026)
Browse all 81 models in the index →