Agentic RAG: When Retrieval Gets Smart
After this lesson, you will be able to:
- Understand how query routing looks at a user's question and automatically picks the best way to find the answer — vector search, web search, database query, or no search at all
- Describe Self-RAG: the AI first decides whether it even needs to search, and after searching, checks if what it found is actually useful before answering
- Compare fallback strategies (Corrective RAG) for what happens when the first search comes up empty or irrelevant
- Design a multi-step retrieval pipeline for complex questions that have several parts requiring different sources
Before You Start
#From Static to Agentic Retrieval
#Query Analysis: Classify Before You Retrieve
Not all queries are created equal. The first step in agentic RAG is analyzing the query to determine the best retrieval strategy:
#Query Types
| Query Type | Example | Best Strategy |
|---|---|---|
| Factual lookup | "What year was Python created?" | Vector search on knowledge base |
| Real-time | "What's the weather in Tokyo right now?" | Live API call (weather service) |
| Analytical | "Compare the GDP growth of India vs China over 10 years" | Multi-step retrieval + data tables |
| Conversational | "Thanks, that makes sense!" | No retrieval needed — respond directly |
| Code-related | "How do I use asyncio.gather?" | Code search + documentation |
| Personal context | "What did I ask you about yesterday?" | Conversation memory / user history |
User asks: 'What's the weather in Tokyo right now?' Should agentic RAG search a vector database or call a weather API?
#Router-Based Retrieval
A query router is the decision-making component that directs queries to the right retrieval source. There are several implementation approaches:
#LLM-Based Router
The simplest approach: ask the LLM itself to classify the query.
System: You are a query router. Classify the user's question into one of:
- VECTOR_SEARCH: factual questions answerable from the knowledge base
- WEB_SEARCH: questions requiring recent/real-time information
- SQL_QUERY: questions about structured data (numbers, comparisons, dates)
- CODE_SEARCH: questions about code, APIs, or documentation
- NO_RETRIEVAL: greetings, clarifications, or questions answerable from context
- CALCULATOR: mathematical computations
Respond with only the category name.
Try it! Copy that system prompt into ChatGPT or Claude and send it a few different questions: "What is the refund policy?", "What is the weather in Tokyo right now?", "Thanks, that helps!", "SELECT * FROM users". See how the LLM classifies each one. You just built a query router in 30 seconds.
#Embedding-Based Router
For lower latency, compute the query embedding and use a classifier trained on labeled query-category pairs. This avoids the cost and latency of an LLM call for every routing decision.
#Multi-Source Retrieval
- "What's the latest Python 3.13 feature?" -> Web search (recency) + docs search (depth)
- "How does our API handle rate limiting?" -> Code search (implementation) + docs search (policy)
#Self-RAG: The Model Decides Everything
Self-RAG (Self-Reflective Retrieval-Augmented Generation) takes agentic RAG further by training the model to make three critical decisions through special tokens:
#Decision 1: Should I Retrieve?
#Decision 2: Is the Retrieved Content Relevant?
#Decision 3: Is My Response Supported?
#Step 1: Query Arrives
The user asks: "What programming language is TensorFlow written in?" The model receives the query and must decide its first action.
#Step 2: Analyze Intent
#Step 3: Route to Source
The query router sends this to the technical documentation vector database. The retriever returns 3 passages about TensorFlow's architecture, codebase, and history.
#Step 4: Retrieve and Filter
The model evaluates each retrieved passage:
- Passage 1 (TensorFlow architecture overview): [Relevant] — mentions core language
- Passage 2 (TensorFlow tutorial for beginners): [Irrelevant] — does not mention implementation language
- Passage 3 (TensorFlow GitHub README): [Relevant] — states languages used
Passage 2 is discarded.
#Step 5: Generate with Relevant Context
Using only the relevant passages, the model generates: "TensorFlow is primarily written in C++ for its core runtime, with Python as the main user-facing API language."
#Step 6: Verify Support
#Corrective RAG (CRAG)
CRAG addresses the question: what do you do when retrieval fails? Instead of generating a potentially hallucinated answer from bad context, CRAG detects failure and takes corrective action.
#The CRAG Pipeline
- Retrieve documents as usual
- Evaluate retrieval quality with a lightweight classifier:
- Correct: At least one document is highly relevant -> proceed to generation
- Ambiguous: Documents are partially relevant -> refine the query and re-retrieve
- Incorrect: No relevant documents found -> trigger fallback strategies
#Fallback Strategies When Retrieval Fails
| Strategy | When to Use | Example |
|---|---|---|
| Query rewriting | Original query was too vague or used different terminology | "JS async patterns" -> "JavaScript Promise async/await best practices" |
| Web search fallback | Knowledge base does not contain the answer | Internal docs search fails -> fall back to web search |
| Query decomposition | Complex question needs to be broken into parts | "Compare X and Y" -> "What is X?" + "What is Y?" + "Differences?" |
| Knowledge graph lookup | Question involves relationships or entities | "Who founded the company that made TensorFlow?" -> entity lookup |
| Honest abstention | No reliable source can answer the question | "I don't have enough information to answer this accurately." |
#Multi-Step Retrieval: Decomposing Complex Questions
Some questions cannot be answered with a single retrieval step. Multi-step retrieval decomposes complex queries into a chain of simpler retrievals:
#Example: "How did the company that created React change its name, and when?"
- Step 1: Retrieve "Who created React?" -> "React was created by Facebook"
- Step 2: Retrieve "When did Facebook change its name?" -> "Facebook rebranded to Meta in October 2021"
- Step 3: Synthesize: "Facebook, the company that created React, rebranded to Meta in October 2021."
Each step builds on the previous answer, and each retrieval is targeted and specific. The agent decides when it has enough information to synthesize a final answer.
#Adaptive Retrieval Depth
Not every query needs the same number of retrieval steps:
- Simple factual: 1 step ("What is the capital of France?")
- Multi-hop reasoning: 2-3 steps ("Who is the CEO of the company that acquired Instagram?")
- Research synthesis: 5+ steps ("Summarize the key arguments for and against remote work, citing recent studies")
The agent monitors its confidence at each step and stops retrieving when it has sufficient evidence.
#Try It Yourself
Tests · Verify 'What is the capital of France?' routes to VECTOR_SEARCH. Verify 'latest news' routes to WEB_SEARCH. Verify 'Hello' routes to NO_RETRIEVAL.
#Key Takeaways
- Agentic RAG adds decision points to the retrieval pipeline — instead of always retrieving from the same source, the system classifies queries, routes to appropriate sources, evaluates results, and retries when retrieval fails
- Self-RAG lets the model decide whether to retrieve, evaluate relevance, and verify support — three critical decision points that reduce unnecessary retrieval, filter noise, and catch hallucinations
- Corrective RAG provides fallback strategies for failed retrieval — query rewriting, source switching, and decomposition ensure the system can recover rather than generating from bad context
- Multi-step retrieval handles complex questions by decomposing them — each step builds on previous results, with the agent adaptively deciding when it has gathered enough evidence to synthesize a final answer
#Quick Check
What is the primary advantage of query routing in agentic RAG?