Embeddings & Vector Similarity
After this lesson, you will be able to:
- Understand how embedding models turn text into lists of numbers (vectors) where similar meanings end up close together in high-dimensional space
- Calculate cosine similarity between two vectors and interpret the score as 'how related are these two pieces of text?', and know why cosine beats Euclidean distance for text
- Compare popular embedding models (OpenAI text-embedding-3, BGE, E5) and understand the trade-offs (speed vs. accuracy vs. cost vs. dimensionality) when picking one for your project
- Explain how contrastive learning trains embedding models using positive and negative pairs, and why asymmetric query/document embeddings improve RAG retrieval
Before You Start
Embeddings are one of those ideas that sounds intimidating but is actually beautifully simple once it clicks. You are going to have a genuine "aha" moment in this lesson — most people do.
Active Recall
Before we step into RAG: recall from the vectors-and-spaces lesson, write the formula for cosine similarity between two vectors a and b in one line. Then state in one sentence why we normalize by the vector magnitudes. Don't scroll back — commit your answer first. This formula is the entire mathematical engine of embedding retrieval; if it's not on the tip of your tongue, this lesson is the place to lock it in.
Write your answer in your own words — don't look back at the lesson. This is the most effective way to remember what you just learned.
Type your explanation above
#Meaning as a Point in Space
#From Text to Vectors
An embedding model takes a string of text and outputs a vector — a list of floating-point numbers, typically between 256 and 3072 dimensions. Each dimension captures some aspect of meaning, though individual dimensions are not human-interpretable.
#How Embedding Models Learn
Modern embedding models are trained using contrastive learning. The training process goes like this:
-
Positive pairs: Take pairs of text that should be similar (a question and its correct answer, a sentence and its paraphrase, a query and a relevant document).
-
Negative pairs: Take pairs that should not be similar (a question paired with an irrelevant document).
-
Train the model to produce vectors that are close for positive pairs and far apart for negative pairs.
After training on millions of such pairs, the model learns to map semantically similar text to nearby points in vector space, regardless of the specific words used.
Try it! Go to projector.tensorflow.org and type words like "king," "queen," "man," and "woman" into the search. Watch how related words cluster together in 3D space. Rotate the visualization and notice the geometric patterns — this is exactly what embedding models learn to produce.
#The Geometry of Meaning
If 'king' - 'man' + 'woman' = ?, what vector would you expect to be closest to the result?
This works for many relationships:
- Paris - France + Japan = Tokyo (capital-of relationship)
- walked - walk + swim = swam (past-tense relationship)
- bigger - big + small = smaller (comparative relationship)
These are not programmed rules. They emerge naturally from the geometry of the learned embedding space.
#How Embeddings Work: End to End
#Step 1: Input Text
A piece of text enters the embedding model. It can be a single word like "king," a sentence, or an entire paragraph. The model needs to convert this raw text into a numerical representation that captures its meaning.
#Step 2: Lookup in Embedding Matrix
#Step 3: Dense Vector Output
[0.2, -0.5, 0.8, 0.1, ...]. For modern models, this is 768 to 3072 numbers. Each dimension encodes some abstract aspect of meaning, though no single dimension is human-interpretable on its own.#Step 4: Similar Words Are Nearby in Space
Words with similar meanings end up at nearby coordinates in this high-dimensional space. "King" is close to "monarch," "ruler," and "sovereign." It is far from "bicycle" or "tomato." The distance between points reflects semantic relatedness.
#Step 5: Arithmetic on Meaning
#Step 6: Cosine Similarity Measures Closeness
#Cosine Similarity: Measuring Meaning
#Why Cosine, Not Euclidean Distance?
#Interpreting Similarity Scores
In practice, cosine similarity scores for text embeddings cluster in specific ranges:
| Score Range | Interpretation | Example |
|---|---|---|
| 0.90 - 1.00 | Near-identical or paraphrase | "The cat is on the mat" vs "A cat sits on the mat" |
| 0.75 - 0.90 | Highly related, same topic | "Machine learning models" vs "AI training algorithms" |
| 0.50 - 0.75 | Somewhat related | "Python programming" vs "Data science tools" |
| 0.20 - 0.50 | Vaguely related | "Weather forecast" vs "Climate change research" |
| < 0.20 | Unrelated | "Chocolate cake recipe" vs "Quantum entanglement" |
#Other Distance Metrics
While cosine similarity dominates RAG, other metrics exist:
#Dot Product
#Euclidean Distance (L2)
#Embedding Models: The Landscape
The choice of embedding model dramatically affects RAG quality. Here are the key categories:
#Proprietary Models
- OpenAI text-embedding-3-large (3072 dims): Strong general performance. Supports Matryoshka embeddings (you can truncate dimensions).
- Cohere embed-v3 (1024 dims): Excellent multilingual support. Separate query/document embeddings.
- Google Gemini embeddings (768 dims): Integrated with Google Cloud ecosystem.
#Open-Source Models
- BGE-large-en-v1.5 (1024 dims): Strong English embedding model from BAAI.
- E5-mistral-7b (4096 dims): LLM-based embeddings, very high quality but slow.
- nomic-embed-text (768 dims): Good performance with a permissive license.
- GTE-large (1024 dims): Alibaba's general text embedding model.
#Key Trade-offs
| Factor | Small Model (384d) | Medium Model (768-1024d) | Large Model (1536-4096d) |
|---|---|---|---|
| Embedding speed | Fast | Moderate | Slow |
| Storage per vector | 1.5 KB | 3-4 KB | 6-16 KB |
| Retrieval quality | Good | Very good | Excellent |
| Cost | Low | Medium | High |
#The Evolution of Embeddings
Understanding the history helps appreciate why modern embeddings work so well.
#Word2Vec (2013): The Beginning
Mikolov et al. at Google trained shallow neural networks to predict a word from its context (or vice versa). The hidden layer weights became the embeddings. For the first time, word relationships became geometric: king - man + woman = queen.
#Sentence-BERT (2019): Sentence-Level Meaning
Reimers and Gurevych fine-tuned BERT with a Siamese network structure to produce meaningful sentence embeddings. For the first time, you could embed entire sentences and compare them meaningfully with cosine similarity.
#Modern Embedding Models (2023+): Instruction-Tuned
Models like E5 and GTE introduced instruction-tuned embeddings. You prefix the text with a task instruction: "Represent this query for retrieving relevant documents:" vs "Represent this document for retrieval:". The model produces different embeddings for the same text depending on the task, dramatically improving RAG retrieval quality.
#Multimodal Embeddings
Embeddings are not limited to text. Modern models embed images, audio, and video into the same vector space as text:
- CLIP (OpenAI): Embeds images and text in a shared space. "A photo of a cat" and an actual photo of a cat have similar vectors.
- ImageBind (Meta): Extends to 6 modalities: text, image, audio, depth, thermal, and IMU data.
- Multimodal RAG: Index images alongside text. A query "diagram showing neural network architecture" retrieves relevant diagrams, not just text descriptions.
#Try It Yourself
Tests · Verify cosine similarity between 'cat on mat' and 'feline on rug' is > 0.95. Verify similarity between 'cat on mat' and 'stock market' is < 0.5.
#The Math of Embedding Geometry
Word-vector analogies are real, but they have subtleties. The playground below lets you build a tiny embedding space from scratch, run the king − man + woman experiment numerically, and explore Matryoshka truncation (the 2024 trick that lets you store one 1536-d vector and search at 64, 256, 512, or 1536 dims depending on latency budget).
#Key Takeaways
- Embeddings convert text into dense vectors that capture semantic meaning. Similar concepts map to nearby points in vector space, enabling "semantic search" that finds relevant content even when exact keywords do not match
- Cosine similarity measures semantic relatedness. It compares the angle between vectors regardless of magnitude, giving a score from -1 (opposite meaning) to +1 (same meaning) that powers retrieval ranking
- Embedding model choice significantly impacts RAG quality. Larger models produce better embeddings but are slower and more expensive; the right model depends on your domain, latency requirements, and accuracy needs
- Embeddings are the foundation of modern information retrieval. Every vector database, semantic search engine, and RAG system depends on high-quality embeddings to bridge the gap between natural language queries and stored knowledge
#Quick Check
What do text embeddings fundamentally represent?
A dense vector (typically 256-3072 dimensions) produced by a neural network that encodes the semantic meaning of a piece of text, image, or other input.
A vector where most components are non-zero and each dimension carries continuous semantic information. Contrasts with sparse representations like TF-IDF or bag-of-words.
Training regime used for modern embedding models. Positive pairs (query and relevant doc) are pulled together in vector space while negative pairs are pushed apart, typically via the InfoNCE loss.
Angle-based similarity metric: a . b / (||a|| * ||b||). Returns a value in [-1, 1] and is the default ranking function for text embeddings because it ignores magnitude.
Architecture that encodes query and document independently into vectors, then compares via dot product or cosine. Fast and scalable — the standard for first-stage retrieval.
Architecture that takes query and document as a joint input and outputs a single relevance score. More accurate than a bi-encoder but ~100x slower; used for re-ranking the top candidates.
Embedding where the first N dimensions are themselves a valid lower-dimensional embedding. Lets you run fast search on truncated vectors and re-rank with the full vector — huge storage/speed wins at scale.
Embedding models (Cohere embed-v3, E5) that use different prompts/modes for queries vs. documents, producing vector spaces optimized for retrieval rather than symmetric similarity.
Semantic Search at Notion AI
Notion AI embeds every block across a workspace so users can ask questions in natural language and retrieve relevant notes by meaning, even when no keywords overlap.
Turns millions of docs per workspace into a searchable knowledge base
CLIP Multimodal Embeddings
CLIP embeds images and text into a shared vector space, enabling image-to-text and text-to-image search. 'A photo of a cat' and an actual cat photo land at nearly the same coordinates.
Foundational model behind Stable Diffusion, DALL-E, and most image search systems
RAG over Private Docs at Enterprises
Enterprises index Slack, Confluence, Drive, Jira, and email as embeddings so employees can ask 'what did we decide about X?' and retrieve the actual source passages as context for an LLM.
Cuts time-to-answer for internal questions from hours to seconds