One space for text, photos, video & sound
What an embedding is, what makes it good, and how Google’s EmbeddingGemma 2 takes CLIP’s 2021 idea of a shared vector space and extends it to every sense. All numbers on this page come from a real run on a laptop CPU.
Meaning, written as numbers
An embedding is a list of numbers that captures what something means. A model reads the input and places it at a point in a space with 768 directions. Things that mean similar things end up close together.
These are the real first values the model produced. Each square is one number: orange = positive, blue = negative, darker = larger. No single number means “dog”. The meaning lives in the whole pattern.
Similar meaning → closer vectors
A good model puts a sentence and its paraphrase close together, even when they share almost no words, and puts unrelated sentences further away. “Closeness” is measured by cosine similarity, the angle between two vectors: the smaller the angle, the closer to 1.
A and B share no key words (dog / puppy, beach / sand by the sea), yet score . A and C score . Raw scores aren’t percentages: with this model even unrelated text sits fairly high. What matters is the ranking, and the paraphrase clearly wins.
One open model, one space for every kind of input
Google DeepMind’s EmbeddingGemma 2 (6 Oct 2026) puts text, images, video and audio into the same 768-number space, in a model small enough to run on a laptop or phone.
Same paradigm, wider reach
CLIP (OpenAI, Feb 2021) showed that two encoders trained on 400 million image–caption pairs can share one space, so “a photo of a dog” lands near dog photos, with zero-shot classification as a result. EmbeddingGemma 2 keeps that joint-space idea and widens it.
The sentence was never paired with any of these photos. It still picks the right one, ahead of the next. That’s the joint space at work.
Same core idea, more senses, and it runs where your data lives
CLIP showed that photos and text can share one vector space. EmbeddingGemma 2 keeps that idea and adds audio and video: every input becomes the same kind of 768-number vector. The model is small and open, so the whole pipeline, from embedding your files to searching them, runs on your own laptop, phone or server. Your files never have to leave it.
🔊 sounds · 🎙 voice notes
💬 chat messages
each file → 768 numbers
optional: keep only the first 256 to save space
→ 768 numbers → nearest items
💬 Bruno chat
🔊 the bark
A small “phone” with 20–30 everyday photos, sounds, voice notes and chat messages, all embedded with EmbeddingGemma 2. Type any sentence and get back the closest items of every kind, ranked, with no keywords, tags or captions.