Contents
Pop quiz, hotshot. You ask your RAG bot: "Who are the key suppliers connected to our Berlin factory, and which of them also serve our competitors?" Your vector retriever does its thing, pulls up a few chunks mentioning Berlin, a few mentioning suppliers, and your LLM assembles an answer that sounds plausible and is mostly fiction. Why? Because the answer lives in the relationships between facts, not in the facts themselves — and chunks don't do relationships. That's where GraphRAG enters the chat.
I resisted this one for a long time, I'll be honest. Knowledge graphs sounded like something enterprise architects discuss in meetings with catered lunch. But after Microsoft published their GraphRAG approach and I actually tried it, I changed my tune fast. Because when your questions involve connections — who, what connects to what, how does X influence Y — chunk-based retrieval is structurally the wrong tool. Let me show you what I mean. :)
The Problem: Chunks Can't See the Forest
Here's the core weakness we keep circling back to in this series. Vector RAG retrieves chunks of text based on semantic similarity. It's brilliant at find me the paragraph about refunds. It's terrible at questions requiring synthesis across the entire corpus.
Classic failure cases:
- Global questions — What are the main themes across these 500 support tickets? No single chunk contains the answer. The answer is distributed everywhere.
- Relationship questions — Which investors are connected to both startups? The connection exists across documents that share no overlapping text.
- Multi-hop reasoning at corpus scale — If Company A acquires Company B, what happens to B's partnership with Company C? Good luck finding that in one chunk.
Ever asked your RAG bot a big picture question and gotten a confidently narrow answer built from three random chunks? That's not a bug you'll fix with better chunking or fancier embeddings. That's a structural ceiling. And no, another embedding model upgrade won't save you — I checked. :/
What Is GraphRAG?
GraphRAG combines knowledge graphs with retrieval-augmented generation. Instead of only indexing text chunks, you extract entities (people, places, organizations, concepts) and their relationships from your documents, then store them as a graph: nodes for entities, edges for how they connect.
At query time, you can retrieve by walking the graph — follow the edges from Berlin factory to its suppliers, then check which of those nodes connect to competitor entities — instead of fuzzy-matching text blobs.
The payoff comes in two flavors:
- Local search — answer specific questions by pulling the graph neighborhood around the relevant entities, enriched with the original text
- Global search — answer corpus-wide what are the themes? questions by summarizing from community-level descriptions, not raw chunks
That second one is the killer feature IMO. Nobody else in our series could even attempt a global question without dumping the whole corpus into the context window.
| Capability | Vector RAG | GraphRAG |
|---|---|---|
| Factual lookup ("what's the refund policy?") | Excellent | Overkill |
| Relationship queries ("who connects to whom?") | Structurally weak | Graph walk shines |
| Global themes ("what are the main patterns?") | Needs whole-corpus dump | Community summaries |
| Indexing cost | Low | High (LLM entity extraction) |
| Infra | Vector store only | Vector store + graph DB |
The Microsoft GraphRAG Approach
Microsoft open-sourced a GraphRAG implementation that kicked off the whole wave, and it's worth understanding its pipeline even if you don't use their tooling directly:
- Chunk your documents — same as usual, nothing exotic here
- Extract entities and relationships — an LLM reads each chunk and pulls out entities, their attributes, and the relationships between them
- Build the graph — you merge everything into a network, resolving Acme Corp and Acme Inc. into the same node (usually)
- Detect communities — the Leiden algorithm clusters the graph into groups of densely connected entities
- Summarize each community — an LLM writes a summary of every community and its relationships
- Query globally or locally — global questions match against community summaries; local questions walk the graph from specific entities
Sound expensive? It is. More on that shortly — I promise I won't hide the bill from you.
Install and run it:
pip install graphrag
python -m graphrag.index --init --root ./my-project
# Add your docs to ./my-project/input/, set your API key in settings.yaml
python -m graphrag.index --root ./my-project
The indexing step is where the LLM does the heavy lifting — extracting all those entities and relationships costs real tokens. FYI, on a large corpus, the indexing bill can make you sit down slowly. Budget for it, index once, and amortize across thousands of queries.
The Lighter Route: Neo4j + LlamaIndex
Microsoft's tooling is opinionated. If you want more control — and honestly, I usually do — building on Neo4j with LlamaIndex's knowledge graph support gives you a flexible, production-friendly stack:
pip install llama-index llama-index-graph-stores-neo4j
from llama_index.core import KnowledgeGraphIndex, SimpleDirectoryReader
from llama_index.graph_stores.neo4j import Neo4jGraphStore
from llama_index.core import StorageContext
graph_store = Neo4jGraphStore(
username="neo4j", password="your-password", url="bolt://localhost:7687"
)
documents = SimpleDirectoryReader("./data").load_data()
index = KnowledgeGraphIndex.from_documents(
documents,
graph_store=graph_store,
max_triplets_per_chunk=10,
)
query_engine = index.as_query_engine(
include_text=True, # graph relationships PLUS source chunks
response_mode="tree_summarize",
)
Two things I love about this route. First, include_text=True — you keep the graph and the original chunks, so retrieval gets both relationship structure and full document context. Second, you can query Neo4j with plain Cypher, which means your graph serves double duty for analytics teams:
MATCH (f:Facility {name: "Berlin Factory"})-[:SUPPLIED_BY]->(s:Supplier)
-[:SUPPLIES]->(c:Company)
WHERE c.name <> "Our Company"
RETURN s.name, collect(c.name) AS also_serves
That query answers our opening question exactly — no fuzzy similarity, no hoping the chunks overlap. The graph walks the actual relationships. When I demoed this pattern to a stakeholder after weeks of chunk-RAG flailing, I felt like a wizard. A well-rested wizard, for once.
Figure 1: GraphRAG — entities, edges, and communities instead of fuzzy chunk matching
Image Alt Text: "GraphRAG knowledge graphs meeting retrieval augmented generation with Neo4j and Microsoft GraphRAG"
GraphRAG Meets the Rest of Our Stack
Here's where the series pays off, because GraphRAG doesn't replace what we've built — it stacks on top, like everything else:
- Agentic RAG + graphs — an agent can choose between vector retrieval and graph queries per question. Simple factual lookups hit the vector store; relationship or global questions hit the graph. This hybrid routing is, IMO, the actual future of serious RAG systems.
- Query decomposition + graphs — remember our multi-hop decomposition? A graph traversal is a multi-hop retrieval. Your decomposed sub-queries map beautifully onto graph queries.
- Compression + graphs — retrieved graph neighborhoods still benefit from culling before they hit your generator. Old habits still apply.
The Costs (The Part Conference Talks Skip)
Let's be honest about what GraphRAG demands:
- Indexing cost — LLM-driven entity extraction over a big corpus costs serious money. Index a 500-page corpus once and it's fine; re-index it nightly and finance will notice.
- Build time — the first index run on a large corpus takes hours, not seconds.
- Maintenance — documents change. Your graph goes stale. You need a refresh strategy, or your graph confidently asserts last year's org chart.
- Complexity creep — you're now running a vector store and a graph database. Two things to monitor, secure, and pay for.
My honest production advice: don't GraphRAG everything. Use it where relationships and global themes actually matter — legal contracts, organizational data, research corpora, fraud detection. Keep the vector store for your FAQ-style content. Adding a knowledge graph to a chatbot over recipe blogs is enterprise architecture cosplay. :/
Measuring Whether It's Worth It
You know the drill by now — RAGAS baseline before, RAGAS after. But here's the GraphRAG-specific twist: standard RAGAS metrics under-serve global questions, because context recall assumes the answer sits in retrievable chunks. For global queries, the relevant context is the community summaries themselves.
Practical approach: split your eval set into local questions (GraphRAG and vector RAG should both score decently — compare them honestly) and global questions (where GraphRAG should dominate). If it doesn't win on the global bucket, your graph extraction quality is the suspect — check whether your entity extraction prompt is pulling the right relationship types.
Recommended Books
- Designing Data-Intensive Applications by Martin Kleppmann — graph and relational modeling foundations that make knowledge-graph design decisions click.
- Graph Databases by Robinson, Webber, and Eifrem — Neo4j and property-graph fundamentals for the LlamaIndex + Neo4j route in this tutorial.
- Retrieval-Augmented Generation (RAG) — where graph-based retrieval fits alongside chunk-based pipelines in modern LLM systems.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is GraphRAG?
GraphRAG combines knowledge graphs with retrieval-augmented generation. Documents are parsed for entities and relationships, stored as graph nodes and edges, then queried by walking the graph instead of only fuzzy-matching text chunks — enabling relationship and corpus-wide questions.
Why is vector RAG bad at relationship questions?
Vector RAG retrieves semantically similar chunks, but answers to multi-hop or global questions live in relationships across documents that share no overlapping text. No single chunk contains the answer, so better chunking or embeddings cannot fix the structural gap.
How does Microsoft GraphRAG work?
Microsoft GraphRAG chunks documents, uses an LLM to extract entities and relationships, merges them into a graph, runs Leiden community detection, summarizes each community with an LLM, then answers global questions from community summaries or local questions by walking the graph.
What is local vs global search in GraphRAG?
Local search answers specific questions by pulling the graph neighborhood around relevant entities plus original text. Global search answers corpus-wide theme questions by summarizing from community-level descriptions rather than raw chunks.
Can I build GraphRAG with Neo4j and LlamaIndex?
Yes. LlamaIndex KnowledgeGraphIndex with a Neo4jGraphStore extracts triplets into Neo4j. Set include_text=True to keep graph structure plus source chunks, and query Neo4j directly with Cypher for precise relationship traversal.
When should I use GraphRAG instead of vector RAG?
Use GraphRAG when questions involve connections or global themes: legal contracts, organizational data, research corpora, fraud detection. Keep vector RAG for FAQ-style content. Hybrid agentic routing sends factual lookups to vectors and relationship queries to the graph.
Wrapping Up
Let's recap the big idea: vector RAG retrieves facts; GraphRAG retrieves relationships. You extract entities and their connections into a knowledge graph, cluster them into communities with Leiden, summarize those communities, and suddenly your system answers global what are the themes? questions and multi-hop who connects to whom? questions that chunk retrieval structurally cannot. Microsoft's GraphRAG gives you the batteries-included version; Neo4j plus LlamaIndex gives you the flexible version; agentic routing lets you use the right tool per question.
My parting take? GraphRAG is the first technique in this whole series that made me rethink the architecture rather than tune a component. It's also the most expensive and complicated thing we've covered — so earn it. Start with a small corpus, build the graph, ask the questions your vector pipeline kept fumbling, and measure.
So — got a question your chunk-based bot keeps failing? Trace whether the answer lives in a fact or a relationship. If it's a relationship, my friend, you might just need a graph. And if the indexing bill makes you wince? No judgment. My first GraphRAG invoice made me close my laptop and go outside. ;)