Home/Glossary/Per-Layer Embeddings (PLE)

Architecture

What is Per-Layer Embeddings (PLE)?

A technique introduced in Google's Gemma 4 E2B and E4B models that feeds a secondary embedding signal into every decoder layer. PLE allows a 2.3B-active model to carry the representational depth of 5.1B parameters while fitting in under 1.5 GB of memory with quantization. This is why Gemma 4 E2B and E4B punch well above their weight class, achieving quality closer to models 2–3x their active parameter count. PLE is distinct from MoE — it adds depth without adding width.

Where Per-Layer Embeddings (PLE) comes up on LLMCheck

Browse all 81 models in the index →

Related terms

All 37 terms in the LLMCheck glossary →