Google releases EmbeddingGemma 2, an open multimodal embedding model
Google DeepMind released EmbeddingGemma 2 on October 6, 2026 — an open 740M-parameter model that maps text, code, images, video and audio into a single 768-dimensional vector space and runs on-device. It is modular, scaling from 270M parameters for text to 740M for all modalities.
- Single 768-dimensional embedding space across five modalities
- Modular encoders: text 270M, +vision 440M, +audio 570M, full 740M
- 8192-token context: up to 29 images, 58 video frames or ~5.5 min audio
- MTEB Code rose from 68.76 to 78.68; Apache 2.0 license
Read next
AI