Signum
Feed
Useful signal6 Oct 2026high confidence

Google DeepMind releases EmbeddingGemma 2, an open 740M-parameter multimodal embedding model under Apache 2.0

Google DeepMind released EmbeddingGemma 2 (740M parameters, built on Gemma 4 architecture, Apache 2.0) with weights on Hugging Face and Kaggle. It maps text, code, images, audio and video into one embedding space. It is modular: 270M text-only, with optional 170M vision and 300M audio encoders. It has an 8K context window (4x EmbeddingGemma 1) and MRL truncation from 768 to 512, 256 or 128 dimensions. Quantized on a Pixel 11 Pro it needs about 191MB RAM (text) to 567MB (multimodal). MTEB Code rose from 68.76 to 78.68. Vertex-style Model Garden availability (Gemini Enterprise Agent Platform) is coming later. Integrations include transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LMStudio, Qdrant, Unsloth, LiteRT and MediaPipe.

CapabilityAccessAdoptionInfrastructure

Entities: Google DeepMind, EmbeddingGemma 2, EmbeddingGemma, Gemma 4, Gemini Embedding, Hugging Face

74Useful signal
2 sources
1 primary
Was this useful?
01

What happened

Google DeepMind released EmbeddingGemma 2, an open-weights embedding model (Apache 2.0) with weights on Hugging Face and Kaggle. It has 740M parameters in total and maps text, code, images, audio and video into one shared space. It is modular: a 270M text-only core, with optional 170M vision and 300M audio encoders. The context window is 8K tokens (4x the first version), and vectors can be shortened from 768 to 512, 256 or 128 dimensions. Google reports MTEB Code rising from 68.76 to 78.68, and quantised memory use on a Pixel 11 Pro of about 191MB (text) to 567MB (multimodal). Hosted availability through Google's enterprise platform is promised later, with no date given.

02

Why it matters

Developers building search or retrieval features that must run offline or keep data on the device now have a permissively licensed multimodal option, with day-one support in common tooling such as sentence-transformers, llama.cpp, Ollama, vLLM and Qdrant. The modular design lets teams pay the memory cost only for the modalities they need, which matters on phones and small servers. The Apache 2.0 licence removes the usage restrictions that complicate some other open models. The real impact depends on independent testing, because the headline numbers come from Google and cover a narrow set of benchmarks.

03

What is noise

"Most capable on-device multimodal embedder" and "best-in-class under 1B" are Google's own claims, and the comparisons were chosen by Google. "Sometimes outperforms models twice its size" is a hedge that says nothing about how often. The only concrete benchmark gain quoted is for code retrieval, not for image, audio or video, which are the new selling point. The memory figures are for one flagship phone with quantisation, so they will not transfer to typical hardware. This is an incremental successor in a crowded field, not a shift in who leads embeddings.

04

Watch next

  1. 01Independent results on MTEB and multimodal retrieval benchmarks (image-text, audio, video), compared with Gemini Embedding and other open sub-1B models, within the next few weeks.
  2. 02Real-world reports on mid-range Android and iOS devices: latency, memory and battery use for the multimodal configuration, versus Google's Pixel 11 Pro figures.
  3. 03Whether the promised Model Garden or enterprise platform release gets a date and pricing, and whether Hugging Face downloads and adoption in vector database and RAG projects show uptake over the next one to three months.

Coverage

2 stories

More capability signals

Full feed →