Contents
There's a specific kind of relief in finally being able to ask "what's our PTO policy?" in plain English instead of Ctrl+F-ing through a 40-page PDF hoping you spelled it the same way the document did. That's the entire promise of a knowledge base chatbot, and LlamaIndex happens to be the framework built most specifically around making that promise easy to deliver.
I gravitate toward LlamaIndex specifically for this kind of project because its whole design center is "get data into a queryable form," not general-purpose agent orchestration. When the job is genuinely "answer questions over documents," it usually gets you there with less code than the alternatives.
By the end of this tutorial, you'll have a working chatbot that answers questions grounded in your own documents, with conversation memory intact. IMO, watching it correctly answer a follow-up question that depends on the previous one is when this framework really earns its keep :)
Figure 1: LlamaIndex makes building document Q&A chatbots fast and simple
What Makes LlamaIndex Different
Before diving in, it's worth understanding why you'd reach for LlamaIndex specifically instead of, say, LangChain. Both can build RAG systems, but LlamaIndex's core abstractions are built around one job: getting your data into a retrieval-ready form efficiently.
- 160+ built-in data connectors for pulling from PDFs, Notion, Slack, databases, and more without writing custom parsers.
- Multiple index types — vector, keyword, tree, knowledge graph — letting you match retrieval strategy to your actual data shape.
- A chat engine abstraction specifically designed for multi-turn conversation, not just single-shot Q&A.
Ever wondered why so many "chat with your PDF" tools feel like they were built quickly? A lot of them genuinely were, on top of LlamaIndex's chat engine doing the heavy lifting underneath.
Setting Up Your Environment
LlamaIndex has also gone through package restructuring, so make sure you're pulling from the current llama_index.core namespace rather than older tutorials' flat imports.
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install llama-index llama-index-embeddings-openai llama-index-llms-openai
FYI, if you spot a tutorial importing directly with from llama_index import VectorStoreIndex (no .core), that's outdated syntax from an earlier package structure. Current LlamaIndex wants from llama_index.core import VectorStoreIndex instead.
API Key Setup
import os
os.environ["OPENAI_API_KEY"] = "your-key-here"
Standard warning applies here: don't commit this to version control. I've watched this exact mistake get caught by GitHub's secret scanning within minutes — mildly embarrassing, genuinely useful safety net.
Step 1: Loading Your Documents
LlamaIndex's SimpleDirectoryReader handles the tedious part of pulling text out of a whole folder of mixed file types automatically.
from llama_index.core import SimpleDirectoryReader
documents = SimpleDirectoryReader("data").load_data()
print(f"Loaded {len(documents)} documents")
Drop PDFs, markdown files, or plain text into a data/ folder, and this single line handles parsing all of them. You're not writing separate loaders for each file type — that's exactly the friction LlamaIndex exists to remove.
Step 2: Building Your Index
With documents loaded, converting them into a searchable vector index takes one more line.
from llama_index.core import VectorStoreIndex
index = VectorStoreIndex.from_documents(documents)
Under the hood, this chunks your documents, generates embeddings, and stores everything in an in-memory vector index — all the plumbing from the "build RAG from scratch" version of this process, compressed into a single function call. That compression is genuinely the whole value proposition here.
Step 3: Creating a Basic Query Engine
Let's confirm everything works with a simple one-shot query interface before building the full chatbot.
query_engine = index.as_query_engine()
response = query_engine.query("What's our refund policy?")
print(response)
This works, but it's stateless — each query stands alone with no memory of previous questions. For a real chatbot, that's a real limitation. Ever asked a support bot a follow-up question only to have it completely forget what you were just talking about? This is exactly why.
Step 4: Upgrading to a Chat Engine
This is where LlamaIndex's chatbot-specific design actually shows up. Swapping to a chat engine gives you conversation memory almost for free.
chat_engine = index.as_chat_engine(
chat_mode="condense_question",
verbose=True
)
response = chat_engine.chat("What's our refund policy?")
print(response)
response = chat_engine.chat("What about for enterprise customers?")
print(response)
That second question — "what about for enterprise customers?" — only makes sense in context of the first. The condense_question chat mode rewrites follow-up questions into standalone queries using the conversation history, so retrieval still works correctly even when the user's phrasing depends entirely on what came before.
Understanding Chat Modes
LlamaIndex offers a few different chat modes, and picking the right one matters more than it initially seems.
- condense_question — rewrites each new message into a self-contained query using chat history before retrieving. Solid default for most Q&A bots.
- context — retrieves relevant context for every message and includes it directly in the system prompt, without rewriting the question itself.
- react — gives the model reasoning and tool-use capability, useful when your bot needs to do more than just answer from documents.
My honest take: start with condense_question unless you have a specific reason not to. It handles the most common failure mode — follow-up questions losing context — without adding real complexity.
Building an Interactive Chat Loop
Let's wire this into something you can actually talk to from the terminal.
def run_chat():
print("Ask me anything about your documents (type 'exit' to quit)")
while True:
user_input = input("You: ")
if user_input.lower() == "exit":
break
response = chat_engine.chat(user_input)
print(f"Bot: {response}\n")
run_chat()
Run this, and you've got a genuinely conversational interface sitting on top of your document collection. Ask a question, ask a follow-up, and watch it stay coherent across turns — that continuity is the entire difference between a chatbot and a search bar with extra steps.
Adding Source Attribution
A knowledge base chatbot that can't tell you where an answer came from is only half-trustworthy. Let's fix that.
response = chat_engine.chat("What's our vacation policy?")
for node in response.source_nodes:
print(f"Source: {node.metadata.get('file_name', 'unknown')}")
print(f"Excerpt: {node.text[:150]}...")
That source_nodes attribute gives you back the actual retrieved chunks that informed the answer, along with their metadata. Surfacing this to users transforms "trust me" into "here's exactly where I got that." That's a meaningfully bigger deal than it sounds for anything used in a real workplace setting.
Customizing the LLM and System Prompt
The default LLM configuration works, but real deployments usually want more control over tone and model choice.
from llama_index.llms.openai import OpenAI
from llama_index.core import Settings
Settings.llm = OpenAI(model="gpt-4o-mini", temperature=0.2)
chat_engine = index.as_chat_engine(
chat_mode="condense_question",
system_prompt="You are a helpful assistant that answers questions strictly based on the provided documents. If the answer isn't in the documents, say you don't know."
)
That system_prompt instruction, telling the model to admit uncertainty, does more to prevent hallucination than almost any other single change you can make. Skip it, and your bot will happily invent plausible-sounding answers when your documents genuinely don't cover the question.
Persisting Your Index
Rebuilding your index from scratch every time your script restarts wastes both time and, if you're paying per API call, actual money.
index.storage_context.persist(persist_dir="./storage")
Loading it back later:
from llama_index.core import StorageContext, load_index_from_storage
storage_context = StorageContext.from_defaults(persist_dir="./storage")
index = load_index_from_storage(storage_context)
Don't skip this in anything beyond a five-minute experiment. I've re-embedded the same documents unnecessarily more than once out of pure forgetfulness, and it's a genuinely avoidable waste.
Common Mistakes Beginners Make
I've hit a few of these myself, so treat this as a shortcut past my own trial and error.
- Using the stateless query engine for a genuinely conversational use case. If follow-up questions matter, you need
as_chat_engine, notas_query_engine. - Skipping persistence entirely. Rebuilding the index on every run burns API costs for zero benefit once your documents are stable.
- Forgetting the "say you don't know" instruction in your system prompt. Without it, your bot confidently answers questions your documents don't actually cover.
- Using outdated flat imports from older tutorials. Modern LlamaIndex expects
llama_index.core, not barellama_indeximports for these core classes.
When Should You Reach for Something More Advanced?
This setup handles single-document-collection chatbots well, but larger projects sometimes need more: routing between multiple indexes, agent-based tool use, or sub-question decomposition for complex, multi-part queries.
LlamaIndex supports all of that too — FunctionAgent and query-routing tools exist specifically for cases where a single vector index isn't enough, like answering questions that span multiple years of financial filings. That's genuinely worth exploring once your simple chatbot starts hitting its limits, but it's not where you should start.
Frequently Asked Questions
What is LlamaIndex used for?
LlamaIndex is a framework for building RAG applications. It connects LLMs to external data sources, enabling chatbots that answer questions grounded in your own documents with source citations.
How is LlamaIndex different from LangChain?
LlamaIndex focuses specifically on data retrieval and indexing with built-in chat engines. LangChain is more general-purpose for agent orchestration. LlamaIndex typically requires less code for document Q&A use cases.
What chat modes does LlamaIndex support?
LlamaIndex offers condense_question (rewrites follow-ups), context (retrieves context for every message), and react (reasoning + tool use). Start with condense_question for most Q&A bots.
How do I persist a LlamaIndex index?
Use index.storage_context.persist(persist_dir='./storage') to save. Load it back with load_index_from_storage. This avoids re-embedding documents on every script restart.
What is the condense_question chat mode?
condense_question rewrites follow-up questions into standalone queries using conversation history. This allows retrieval to work correctly even when user questions depend on previous context.
How do I add source citations to LlamaIndex answers?
Access response.source_nodes after querying. Each node contains the retrieved text chunk and metadata including file name and page number. Surface this to users for trust and verification.
Wrapping This Up
Building a knowledge base chatbot with LlamaIndex comes down to a tight loop: load your documents, build an index, wrap it in a chat engine for conversational memory, and always surface your sources. The framework handles far more of the plumbing automatically than you might expect coming from a from-scratch approach.
Remember to persist your index so you're not re-embedding on every run, and don't skip the "admit uncertainty" instruction in your system prompt. FYI, once you've built this basic version, the leap to routing across multiple document collections or adding agent tools feels like an extension rather than a rewrite — that's genuinely the framework's biggest strength :)
Now go point this at your own document folder instead of a hypothetical refund policy and vacation days example. That's where a knowledge base chatbot actually starts proving useful.