Vector Databases & Indexing
After this lesson, you will be able to:
- Understand why comparing every single vector one by one is too slow for real applications, and how approximate nearest neighbor (ANN) algorithms trade tiny accuracy losses for massive speed gains
- Compare three indexing strategies — HNSW (graph-based for fast queries), IVF (cluster-based for large scale), and product quantization (compression-based for memory efficiency) — and know when to pick each one
- Evaluate popular vector databases — Pinecone (managed), Weaviate (open-source with hybrid search), ChromaDB (prototyping), Qdrant (high-performance), pgvector (PostgreSQL extension) — and choose the right one
- Design pre-filtering strategies that combine metadata constraints with vector search in a single pass to avoid wasting retrieval budget on out-of-scope documents
Before You Start
This lesson covers the infrastructure that powers every AI search feature you have ever used. It might feel technical, but the core ideas are surprisingly intuitive once you see the analogies. You have got this.
#The Needle in a Billion-Dimensional Haystack
#How Vector Search Works: End to End
#Step 1: Documents to Embeddings
During ingestion, every document chunk is converted to a dense vector using an embedding model. These vectors — each a list of hundreds or thousands of floating-point numbers — are stored in the vector database alongside the original text and metadata.
#Step 2: Query Arrives
A user asks a question: "How do I reset my password?" This raw text needs to be compared against millions of stored document vectors to find the most relevant ones.
#Step 3: Embed the Query
#Step 4: ANN Search (HNSW Graph Traversal)
#Step 5: Return Top-K Nearest Vectors
The search returns the K vectors closest to the query (typically K = 5 to 20). Each result includes a similarity score indicating how close the match is. These are the document chunks whose meaning is most similar to the query.
#Step 6: Retrieve Original Documents
The vector database looks up the original text and metadata associated with each returned vector. The chunk text, source document name, page number, and any other metadata are sent back to the RAG pipeline, which feeds them to the LLM as context for generating an answer.
#The Brute-Force Baseline
At small scale (under 100K vectors), brute-force search is perfectly fine. Many RAG prototypes work this way. But production systems with millions to billions of vectors need something smarter.
Try it! Install ChromaDB (pip install chromadb) and add 10 short text snippets. Search with a question and see the results ranked by similarity. You just used a vector database. Now imagine doing this with 10 million documents — that is why the indexing algorithms below exist.
#HNSW: The Highway System for Vectors
If you have a billion vectors and need to find the 10 most similar, how many vectors does HNSW typically compare against?
HNSW typically examines only a few thousand candidates to find near-optimal results among billions. Here is how it works:
#Layer 0: The Full Graph
Every vector is a node. Each node is connected to its M nearest neighbors (typically M=16-64). This is the bottom layer with maximum detail.
#Higher Layers: The Highway System
Random subsets of nodes are promoted to higher layers, forming a hierarchy. Layer 1 might have 10% of nodes. Layer 2 might have 1%. The top layer has very few nodes but provides long-range connections — like highways connecting distant cities.
#Search: Top-Down Navigation
To find nearest neighbors, start at the top layer. Greedily move to the closest node. Drop to the next layer. Greedily move again. Each layer provides finer granularity. By the time you reach Layer 0, you are in the right neighborhood and only need local refinement.
#Why It Works
The hierarchy creates "skip connections" across the vector space. Without them, you would need to traverse many local edges to get from one side to the other. With the highway layers, you can jump across the space in a few hops, then refine locally. This is what makes search logarithmic rather than linear.
#HNSW Parameters
| Parameter | What It Controls | Trade-off |
|---|---|---|
| M (connections per node) | Graph density | Higher M = better recall, more memory |
| ef_construction | Build-time search depth | Higher = better graph quality, slower build |
| ef_search | Query-time search depth | Higher = better recall, slower query |
#IVF: The Clustering Approach
- Build: Run K-means clustering to partition vectors into
nlistclusters (typically 100-10,000). - Search: Compute distance from the query to each cluster centroid. Search only the
nprobeclosest clusters.
#IVF vs HNSW
| Feature | HNSW | IVF |
|---|---|---|
| Build time | Slow (graph construction) | Moderate (K-means + assignment) |
| Query speed | Very fast | Fast (tunable via nprobe) |
| Memory | High (stores graph edges) | Lower (only centroids + lists) |
| Update (add new vectors) | Easy (insert into graph) | Hard (may need re-clustering) |
| Best for | Read-heavy, moderate size | Very large scale, batch workloads |
#Product Quantization: Compressing Vectors
#The Vector Database Landscape
Vector databases wrap these indexing algorithms in a full database system with APIs, persistence, filtering, and scaling.
#Pinecone (Managed)
- Fully managed, serverless option available
- Excellent for teams that want zero infrastructure management
- Supports metadata filtering, namespaces, sparse-dense hybrid
- Pricing by storage + read/write units
#Weaviate (Open Source + Cloud)
- GraphQL and REST APIs
- Built-in vectorization (can embed text automatically)
- Supports hybrid search (BM25 + vector) natively
- Multi-tenancy support for SaaS applications
#ChromaDB (Open Source)
- Python-native, minimal setup
- Perfect for prototyping and small-to-medium workloads
- Runs in-memory or with SQLite persistence
- Simple API: add, query, delete
#Qdrant (Open Source + Cloud)
- Rust-based, very fast
- Rich filtering with payload indexes
- Supports quantization (scalar and product)
- Good documentation and growing community
#pgvector (PostgreSQL Extension)
- Vector search inside your existing PostgreSQL database
- No separate infrastructure needed
- Supports HNSW and IVF indexes
- Best when you already use PostgreSQL and want simplicity
#FAISS (Library, not a database)
- Meta's library for similarity search
- Not a database — no persistence, no API, no filtering out of the box
- The gold standard for raw search performance
- Best as a building block inside your own system
#Try It Yourself
Tests · Verify IVF uses fewer comparisons than brute force. Verify increasing nprobe improves recall. Verify nprobe=NUM_CLUSTERS gives perfect recall.
#Key Takeaways
- Brute-force search is impractical at scale. Comparing a query vector against millions of stored vectors is too slow for production; approximate nearest neighbor (ANN) algorithms trade tiny accuracy losses for massive speed gains
- HNSW is the most popular ANN index. Hierarchical Navigable Small World graphs enable sub-millisecond search over millions of vectors by building a multi-layer graph structure for efficient traversal
- Product quantization compresses vectors for memory efficiency. By splitting vectors into subgroups and quantizing each separately, PQ dramatically reduces memory usage while maintaining good search quality
- Choose your vector database based on scale and requirements. ChromaDB for prototyping, pgvector for PostgreSQL-based systems, Pinecone/Weaviate for managed scale, and FAISS/Qdrant for high-performance self-hosted deployments
#Picking a Vector DB in 2026: The Decision Matrix
The 2026 landscape stabilized: pgvector for <10M and Postgres-native shops, Qdrant or Weaviate for self-hosted with filtering, Pinecone serverless for pay-as-you-go, LanceDB for embedded/edge, and Vespa for late-interaction / ColPali workloads.
| DB | Best for | Index | Filtering | Hybrid | Late-interaction (ColBERT/ColPali) | Pricing model |
|---|---|---|---|---|---|---|
| pgvector 0.7+ | Postgres shops, <10M vec | HNSW + IVFFlat | SQL where | via tsvector | no | self-host |
| Qdrant 1.10+ | self-hosted, rich filters | HNSW + scalar/int8 | yes (payload index) | yes | yes (multi-vector) | self-host or cloud |
| Weaviate 1.25+ | hybrid + ColBERT | HNSW | yes | native (alpha param) | yes (since 2024) | self-host or cloud |
| Pinecone Serverless | pay-as-you-go, no ops | proprietary | yes | sparse-dense built-in | partial | $0.03 / M reads |
| Vespa | Yahoo/Spotify-scale, late-interaction | HNSW + tensor | yes (YQL) | first-class | mature | self-host or cloud |
| LanceDB | embedded, edge, local | IVF-PQ + diskANN | yes | yes | partial | self-host (zero infra) |
| Milvus 2.4+ | distributed, billion-scale | HNSW + IVF + DiskANN | yes | yes | partial | self-host or Zilliz Cloud |
| Turbopuffer | object-storage-backed, cheap cold | proprietary on S3 | yes | partial | no | $0.04 / M reads (~70% cheaper) |
| ChromaDB | prototyping, single-process | HNSW | yes | partial | no | self-host, in-process |
- < 1M vectors and Postgres in stack: just use pgvector. The latency and ops simplicity win.
- Need ColPali / multi-vector retrieval: Vespa, Weaviate, or Qdrant 1.10+.
- Going from zero queries to 10M queries this month: Pinecone Serverless (pay-as-you-go).
- Cold-tier vectors (rarely queried, must be cheap to keep): Turbopuffer or LanceDB-on-S3.
- Anything offline / on-device: LanceDB.
#Quick Check
Why can't we use brute-force cosine similarity search at scale?