Building a RAG App
After this lesson, you will be able to:
- Architect a production RAG app end-to-end — ingestion, embedding, hybrid retrieval, reranking, generation, citations, evaluation, monitoring
- Stitch together LangChain or LlamaIndex + Pinecone/Weaviate + Cohere Rerank + Claude/GPT-4 + RAGAS into a working Perplexity-style application
- Apply production patterns: streaming responses, query caching, fallback handling, cost tracking, A/B testing different retrievers
- Identify the most common production failure modes and the production hardening that prevents them
Before You Start
Open the /play/rag-pipeline playground to walk a query end-to-end: preprocessing → hybrid retrieval → rerank → prompt assembly → streamed generation with citations. Flip the "cache hit" toggle to see how Anthropic prompt caching cuts cost 90% on repeated context.
#The Production RAG Architecture
USER QUERY
│
▼
┌────────────────────────────────────┐
│ 1. Query Preprocessing │
│ - Normalization │
│ - PII redaction │
│ - Query rewriting (optional) │
└────────────────────────────────────┘
│
▼
┌────────────────────────────────────┐
│ 2. Hybrid Retrieval │
│ - BM25 search (sparse) │
│ - Dense vector search │
│ - RRF fusion │
│ - Top-50 candidates │
└────────────────────────────────────┘
│
▼
┌────────────────────────────────────┐
│ 3. Reranker │
│ - Cohere Rerank or BGE │
│ - Top-50 → top-10 │
└────────────────────────────────────┘
│
▼
┌────────────────────────────────────┐
│ 4. LLM Generation │
│ - System prompt with rules │
│ - Top-10 chunks as context │
│ - Streaming response │
│ - Inline citations │
└────────────────────────────────────┘
│
▼
┌────────────────────────────────────┐
│ 5. Post-processing │
│ - Citation validation │
│ - Refusal detection │
│ - Cache write │
└────────────────────────────────────┘
│
▼
USER STREAMING RESPONSE WITH FOOTNOTES
│
└──► Logged to eval pipeline (RAGAS)
#Step-by-Step Implementation
Tests · Verify the FastAPI endpoint streams a response with [1] [2] citations. Check that hybrid_retrieve returns docs from both BM25 and dense search.
A quick check before the production hardening table — make sure you can spot the right place to spend optimization effort.
Your end-to-end Perplexity-clone has 3 second p95 latency. Breakdown: embed 30ms, vector search 25ms, rerank 120ms, LLM TTFT 800ms, LLM full response 2000ms. Where is the highest-leverage latency improvement?
#Production Hardening Checklist
| Concern | Solution |
|---|---|
| Streaming latency | Use SSE (Server-Sent Events); first-token < 1s |
| Cost overruns | Per-user rate limits, semantic caching, model tiering (mini for cheap queries, premium for hard ones) |
| Hallucination | Citation requirement in system prompt; post-process to verify cited claims actually appear in retrieved chunks |
| Refusals | Detect "I don't have information" responses; route to fallback (web search, escalate to human) |
| Stale corpus | Incremental re-embedding pipeline; track doc updated_at; re-embed on change |
| PII leakage | Redact PII before indexing; check responses for accidental PII regurgitation |
| Prompt injection | Sanitize retrieved chunks; treat them as untrusted; never let chunks override system prompt |
| Cost monitoring | Track tokens/query and $/query in real-time; alert on anomalies |
| A/B testing | Run two retrievers in parallel for 5% of traffic; compare RAGAS metrics |
| Eval CI | RAGAS golden-set runs on every PR; fail builds on > 3% regression |
#The Cost Math: Unit Economics of a RAG App
Before you scale, model your unit economics. The playground below computes per-query cost across the realistic 2026 component prices (embed, retrieval, rerank, LLM input, LLM output), shows the cache-hit-rate lever, and finds the volume where self-hosted infra beats hosted APIs.
#Common Production Failure Modes
#Hands-On Walkthrough: Building a Perplexity-Clone
The CodePlayground above gives the bones. To productionize:
-
Replace mocks with real services:
- Vector DB: Pinecone / Weaviate / Qdrant
- Reranker: Cohere Rerank API or BGE-Reranker-v2 self-hosted
- LLM: Claude 3.5 Sonnet (citations API) / GPT-4o-mini / Gemini 2.0 Flash
- Observability: Langfuse / Arize Phoenix / LangSmith
-
Add safety layers:
- Input sanitization
- Output PII redaction
- Refusal detection
- Citation validation (do the cited chunk ids actually contain the claim?)
-
Add streaming: replace mock
stream_llm_responsewith real OpenAI / Anthropic streaming SSE. -
Add caching:
- Exact-match cache (hash query)
- Semantic cache (embedding similarity)
- Per-user invalidation on subscription changes
-
Add evaluation:
- RAGAS pipeline runs on every deploy with a golden set
- Online metrics: thumbs-up/down, regeneration rate, time-on-page
- Track which retrievers / rerankers / LLMs win in A/B tests
-
Add monitoring:
- Cost per query (LLM + reranker + retrieval)
- Latency breakdown (retrieval / rerank / LLM gen)
- Error rates and refusal rates
- User satisfaction signals
#Key Takeaways
- Production RAG is integration, not invention. The components are well-understood; the hard part is wiring them together with caching, streaming, monitoring, and eval
- Stream first-token in < 1s. Users expect immediate feedback; spinners > 5s feel broken
- Citations are required for trust. Every claim must trace to a retrieved chunk; bake this into your system prompt
- A/B test retrievers and rerankers. Production data will reveal which configurations win on YOUR queries; don't trust generic benchmarks
- Eval pipeline is non-negotiable. RAGAS on a golden set in CI; alert on regressions
#Quick Check
Which production hardening provides the largest cost savings for a high-traffic RAG app?