Dense retrieval gets you 80% of the way. The last 20% — the difference between "usually right" and "reliably right" — is reranking. Cross-encoders cost 100× more per query and are worth every cent. Cohere Rerank, BGE-Reranker, ColBERTv2. These are the models that turned weekend RAG demos into Perplexity-grade products. Every production RAG above hobby-scale ships with a reranker, and here's why.
Learning Objectives
After this lesson, you will be able to:
Explain why a two-stage retrieve-then-rerank pipeline beats single-stage retrieval — first stage casts a wide net, second stage applies a smarter (slower) model
Use cross-encoder rerankers (Cohere Rerank, BGE-Reranker, MS MARCO MonoT5) to lift retrieval quality 10-30% with a 100ms cost per query
Apply ColBERT/ColBERTv2 late-interaction for the best of both worlds — reranker quality at retrieval speed
Pick the right reranker tier (bi-encoder retrieval, cross-encoder reranking, late-interaction) based on latency and corpus size
Try it: Watch a cross-encoder rerank top-100 → top-10Interactive
Open the /play/reranking-stage playground to step through how a cross-encoder re-scores a candidate set retrieved by bi-encoder. Notice how documents at rank 47 jump to rank 2 — that's the NDCG@10 lift you're paying 100× per query for.
ColBERT (Khattab & Zaharia 2020 SIGIR) sits between bi-encoders and cross-encoders. Like bi-encoders, it precomputes per-token embeddings for documents. Unlike bi-encoders, it computes a MaxSim score at query time:
score(q,d)=i∈q∑j∈dmaxqi⋅dj
ColBERTv2 (Santhanam 2022): adds quantization for storage (per-token embeddings would otherwise blow up memory). Production-ready via libraries like RAGatouille and PyLate.
Tradeoff: ColBERT keeps per-token embeddings (10-100x more storage than single-vector retrieval) but achieves cross-encoder-level quality at retrieval speed. Used by Vespa, Weaviate as a "ColBERT" mode.
A quick check before we get to the production rubric — make sure the cost intuition lands.
Quick check
A bi-encoder serves 1M docs in under 5ms. A cross-encoder reranker takes ~100ms for 100 candidates. Why is the two-stage pipeline still faster than a 'cross-encoder over the whole corpus'?
Your single-stage dense retriever gets 70% recall@10 on a customer-support QA system. You add a cross-encoder reranker. What's the most likely outcome?
The right answer depends on what you mean by recall@10. If you fix the retrieved candidate set to top-10 from stage-1, reranking just reorders. The standard pattern: stage-1 retrieves top-100, reranker selects top-10 from those 100. The reranker promotes good candidates that ranked 30-80 in stage-1 into the final top-10. Recall@10 over the full corpus jumps because more truly-relevant docs make it into the final 10.
Scenario
Recommended pipeline
Hobby project, low traffic
Bi-encoder retrieval only
Production Q&A
Hybrid (BM25 + dense) → cross-encoder rerank top-100 to top-10
Latency-critical (<50ms p95)
ColBERT/ColBERTv2 single-stage
Multilingual customer support
Cohere Rerank 3 (multilingual)
Self-hosted, cost-sensitive
BGE-Reranker-v2-m3 + bi-encoder
#The Math of NDCG, MRR, and the Lift You Actually Get
A reranker shifts gold-relevant docs up the ranking. The way you measure that shift matters — recall@10 only checks "did the doc make the cut," but NDCG and MRR reward where in the top-k it landed. Use the playground below to build intuition for how a reranker's reordering changes each metric.
Tests · Verify stage-2 reranker reorders the top-5 differently from stage-1; verify scores are calibrated probabilities (cross-encoder output is sigmoid).
Why does a cross-encoder achieve higher quality than a bi-encoder for the same query-document pair?
Stage 1 finds candidates fast; stage 2 picks the best. Next: query-side techniques (HyDE, Step-Back, RAG-Fusion) that improve retrieval before any model touches the documents.
Your Reflection
Saves automatically
What’s one thing you learned? What’s still confusing?