Modular Encoders and Matryoshka Truncation: Squeezing Four Modalities Into Half a Gigabyte

EmbeddingGemma 2's trick isn't one giant network - it's a 270M-parameter text/code core that can load an optional 170M-parameter vision encoder or a 300M-parameter audio encoder on demand [2]. That modularity is why a phone can run the text-only configuration in roughly 191MB of active RAM, while the full multimodal stack with all encoders loaded costs about 567MB on a Pixel 11 Pro [3]. On top of that, every embedding it produces starts at 768 dimensions but can be truncated down to 512, 256, or 128 dimensions using Matryoshka Representation Learning, cutting local vector-database storage by up to 6x [2]. In practice, that works because the technique trains the vector so its first N dimensions carry the most meaning on their own, letting you cut the tail off without starting over. The context window also grew 4x to 8,000 tokens, enough to process about 5.5 minutes of audio, 29 images, or 58 video frames in a single pass [7]- in practice, that's what makes whole-meeting or whole-video search on a phone feasible instead of theoretical.




