Contents
Remember the naive RAG pipeline we've been building across this whole series? Retrieve chunks, stuff them in a prompt, generate an answer. It's honest work — but it has the reasoning capacity of a vending machine. Ask it something that needs two lookups, a comparison, or a judgment call about whether to retrieve at all, and it just shrugs and hallucinates. Agentic RAG fixes this by putting an LLM in charge of the retrieval process itself, and honestly? It's the most fun I've had building pipelines in years.
Fair warning upfront: my first agentic RAG system was a disaster. I gave an agent five tools and zero guardrails, and it cheerfully called the retriever eleven times for one question, burned through tokens like a college student with a fresh credit card, and still got the answer wrong. :) So this tutorial comes with both the shiny parts and the scar tissue. Let's build it right.
What Is Agentic RAG, Exactly?
Naive RAG follows a fixed script: query → retrieve → generate. Every question, same path, no matter what. It doesn't matter if the question is trivial, impossible, or secretly three questions in a trench coat — the pipeline marches forward blindly.
Agentic RAG hands the steering wheel to an LLM agent. The agent looks at the question and decides:
- Do I need to retrieve anything at all? (You'd be surprised how often the answer is no)
- What should I search for, and in which index?
- Are these results good enough, or should I reformulate and try again?
- Do I need another lookup to fill a gap?
- Should I stop and answer — or admit defeat?
Same building blocks as everything we've covered — vector stores, retrievers, generators. The difference is the agent orchestrates retrieval dynamically instead of following a pipeline hardcoded by you. You're not writing the control flow anymore; you're writing the decision space.
Why Bother? The Cases Naive RAG Flubs
Let's be concrete, because "agentic" is becoming a buzzword people slap on everything. Real problems where the fixed script breaks:
- Multi-hop questions — Compare the refund policy for annual vs. monthly subscribers requires two separate retrievals. Naive RAG does one and hopes.
- Ambiguous queries — the bug fix could mean five things; an agent can retrieve options and ask which one.
- Self-correction — the first retrieval returns junk; an agent reformulates the query and retries. Naive RAG just feeds junk to the generator and lets it improvise.
- Filtering nonsense — what's the meaning of life? doesn't need retrieval; an agent skips it and saves you a vector search plus a confused answer.
Sound familiar? Some of these are the exact problems we attacked in the query expansion article with multi-query and decomposition. The difference: those were pre-programmed strategies. The agent picks strategies on the fly based on the actual question. IMO, that adaptability is the whole value proposition.
Building It with LangGraph
You could hand-roll an agent loop with raw LLM function calling, and honestly that's a great learning exercise. But LangGraph gives you the structure for free, and structure is what keeps agents from going feral. Install it:
pip install langgraph langchain-openai
Step 1: Define the Agent State
LangGraph thinks in state — a shared dictionary that flows between nodes:
from typing import Annotated
from typing_extensions import TypedDict
from langgraph.graph import StateGraph, END
class AgentState(TypedDict):
question: str
documents: list
retrievals: int
answer: str
That retrievals counter matters more than it looks. It's your circuit breaker, and I'll explain why shortly.
Step 2: Build the Tool and the Agent Node
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
llm = ChatOpenAI(model="gpt-4o", temperature=0)
@tool
def retrieve(query: str):
"""Search the knowledge base for relevant documents."""
docs = vectorstore.similarity_search(query, k=4)
return "\n\n".join(d.page_content for d in docs)
The agent node uses a create_react_agent style setup — or, for a classic corrector pattern, a grading step:
def grade_documents(state: AgentState) -> dict:
"""Judge whether retrieved docs answer the question."""
prompt = (
"Question: {q}\n\nDocuments:\n{d}\n\n"
"Do these documents contain enough information "
"to answer the question? Answer yes or no."
)
verdict = llm.invoke(
prompt.format(q=state["question"], d=state["documents"])
).content.strip().lower()
return {"retrievals": state["retrievals"] + 1,
"answer": "" if "yes" in verdict else "NEEDS_RETRY"}
Step 3: Wire the Graph
workflow = StateGraph(AgentState)
workflow.add_node("agent", agent_node) # decide what to search
workflow.add_node("retrieve", retrieve_node) # call the retriever
workflow.add_node("grade", grade_documents) # judge the results
workflow.add_node("generate", generate_node) # write the answer
workflow.set_entry_point("agent")
workflow.add_edge("agent", "retrieve")
workflow.add_edge("retrieve", "grade")
workflow.add_conditional_edges(
"grade",
decide_to_generate, # routes to "generate" or back to "agent"
{"generate": "generate", "agent": "agent"},
)
workflow.add_edge("generate", END)
app = workflow.compile()
The beauty hides in that conditional edge: when the grader says the documents stink, the flow loops back and the agent reformulates the query. Retrieval stops being a single dice roll and becomes a retry loop. My retrieval metrics on messy user queries improved more from this one loop than from any embedder upgrade — and I say that as someone who has opinions about embedders. :/
Figure 1: Agentic RAG — an LLM agent orchestrates retrieve, grade, and retry loops
Image Alt Text: "Agentic RAG combining AI agents with retrieval using LangGraph and corrective loops"
The Corrective RAG Pattern (CRAG)
Let me name the star pattern, because you'll see it in research papers: Corrective RAG (CRAG). The loop runs like this:
- Retrieve — pull candidate documents
- Grade — an LLM judges each document: relevant, ambiguous, or junk
- Correct — refine the query, drop junk chunks, or trigger a web search fallback when the local corpus fails
- Generate — answer only from graded-relevant material
The grading step reuses a trick you might recognize from the contextual compression article — a cheap LLM making document-level judgments before the expensive stuff happens. These patterns stack, people. That's the real secret of this whole series.
The LlamaIndex Route
Prefer LlamaIndex? It ships agent machinery too, and it slots neatly next to the query engines you already know:
from llama_index.core.agent import ReActAgent
from llama_index.llms.openai import OpenAI
from llama_index.core.tools import QueryEngineTool
query_tool = QueryEngineTool.from_defaults(
query_engine=index.as_query_engine(similarity_top_k=4),
name="knowledge_base",
description="Answers questions about company policies and docs",
)
agent = ReActAgent.from_tools(
[query_tool],
llm=OpenAI(model="gpt-4o"),
verbose=True,
)
response = agent.chat(
"Compare refunds for annual and monthly subscribers."
)
Run it with verbose=True once and watch the agent's thought process scroll by — it'll decompose your comparison question into two retrievals by itself. No decomposition prompt you wrote, no hardcoded logic. The agent just... figured it out. That's the first whoa moment, and I won't pretend otherwise. FYI, I still keep verbose on in dev, because watching an agent reason never stops being interesting — and it's the best debugging tool you have.
The Costs (Or: Why I Now Respect Guardrails)
Deep breath. Here's what the conference talks skip:
- Latency — a naive RAG query is one retrieval plus one generation. An agentic query might run five LLM calls before it answers. Your users will notice.
- Cost — multiply the token bill accordingly, especially with a strong model as the reasoning brain.
- Runaway loops — an agent that grades results as insufficient forever will spin until your API budget files for divorce.
That last one is why the retrievals counter in our state exists. Cap the retry loop — I use three retrievals max — and force the agent to answer or gracefully fail. Also route the cheap judgment calls (grading, reformulation) to a mini model and reserve the big model for the final answer. That hybrid setup cut my agentic costs roughly in half with zero quality loss.
And please, measure. Run your RAGAS baseline, then compare: agentic RAG should lift context recall on multi-hop and ambiguous queries specifically. If it doesn't beat naive RAG on your query distribution, you've added expensive machinery for nothing. Bucket your eval set by query type — same lesson as the HyDE article, and it applies double here.
| Setup | Latency | Cost | Best for |
|---|---|---|---|
| Naive RAG | Lowest | Lowest | Simple FAQ single-shot queries |
| Corrective RAG (CRAG) | Medium | Medium | Self-correction on messy corpora |
| Full agentic | Highest | Highest | Multi-hop, ambiguous, mixed indexes |
| Hybrid routing | Mixed | Controlled | Production default — easy path fast, hard path agentic |
When to Go Agentic (and When Not To)
My honest decision framework:
- Go agentic for multi-hop questions, ambiguous queries, mixed corpora, and self-correction needs
- Stay naive for simple FAQ-style retrieval where single-shot works — an agent there is a Ferrari for a grocery run
- Go hybrid — route easy queries straight through a fast path and send only hard ones to the agent. IMO, hybrid routing is what actually works in production; nobody runs full-agentic on every query with a straight face.
Recommended Books
- Retrieval-Augmented Generation (RAG) — pipeline patterns that set the stage for agent-driven orchestration layers.
- Designing AI Agents — agent architectures, tool use, and guardrails that map directly onto LangGraph and ReActAgent loop design.
- Designing Machine Learning Systems by Chip Huyen — cost, latency, and evaluation trade-offs that justify hybrid routing over always-on agentic calls.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is agentic RAG?
Agentic RAG puts an LLM agent in charge of the retrieval process — deciding whether to retrieve, what to search for, whether results are good enough, and when to reformulate or stop. Unlike fixed query-retrieve-generate pipelines, control flow is chosen dynamically per question.
How does agentic RAG differ from naive RAG?
Naive RAG always runs one retrieval then one generation. Agentic RAG can skip retrieval entirely, run multiple lookups for multi-hop questions, grade retrieved documents, and retry with reformulated queries when results are insufficient.
What is Corrective RAG (CRAG)?
Corrective RAG is a pattern where the pipeline retrieves candidate documents, grades each as relevant/ambiguous/junk with an LLM, corrects by refining the query or dropping bad chunks (optionally web-search fallback), then generates only from graded-relevant material.
How do I build agentic RAG with LangGraph?
Define a TypedDict state (question, documents, retrievals counter, answer), add agent/retrieve/grade/generate nodes to a StateGraph, and wire a conditional edge from grade back to agent on insufficient results — creating a capped retry loop instead of a single dice roll.
How do I prevent runaway agent loops in RAG?
Cap retrievals in state (e.g., three max) and force an answer or graceful fail past the limit. Route cheap judgment calls like grading to mini models and reserve the big model for final generation. Hybrid routing sends only hard queries to the full agent.
When should I use agentic RAG vs naive RAG?
Use agentic for multi-hop questions, ambiguous queries, mixed corpora, and self-correction needs. Keep naive for simple FAQ-style retrieval where single-shot works. In production, hybrid-route easy queries on a fast path and hard ones to the agent.
Wrapping Up
Recap time: agentic RAG puts an LLM in control of retrieval — deciding when to search, judging results, reformulating on failure, and looping until it has what it needs. We built the LangGraph version with a grade-and-retry loop, met the CRAG pattern, ran the LlamaIndex ReActAgent route, and talked honestly about latency, cost, and loop caps — the stuff I wish someone had told me before my agent called the retriever eleven times.
The big picture from this whole series? Naive RAG, compressed context, parent-child chunks, query transformations, and now agents aren't competing techniques — they're stacking layers, each covering the previous layer's blind spots. Agentic RAG is the orchestration layer that ties them together.
So go build one. Start with the grade-and-retry loop, cap it at three, watch the verbose logs, and enjoy the moment the agent reformulates a bad query on its own. Just don't hand it five tools with no limits on day one — learn from my scar tissue, my friend. The token bill from that mistake still haunts my dreams. ;)