Sam Austin AI

LangChain RAG Tutorial: Build a Document Q&A Bot

September 1, 2026 13 min read Updated September 2, 2026 Sam Austin
Contents

Every company has that one folder of PDFs nobody wants to actually read — the onboarding docs, the policy manual, the 40-page spec that everyone skims and forgets. RAG exists specifically to make that folder queryable in plain English, and LangChain remains the most common way to wire that whole pipeline together.

I built my first version of this exact bot to answer questions against a pile of technical documentation, and the thing that surprised me most wasn't the AI part — it was how much of the work is just getting your documents loaded and split correctly before the "smart" part even happens. That's the part we're not skipping today.

By the end of this tutorial, you'll have a working document Q&A bot that answers questions using your own PDFs, complete with source citations. IMO, seeing your bot correctly cite which page an answer came from is one of the more satisfying moments in this whole field :)

LangChain RAG Tutorial Document Q&A Bot
LangChain RAG Tutorial Document Q&A Bot

Figure 1: Building a document Q&A bot with LangChain RAG pipeline

What We're Actually Building

Before touching code, let's nail the mental model. A document Q&A bot needs five pieces working together in sequence:

  1. Document loaders — pull text out of your PDFs, markdown, or text files.
  2. A text splitter — break long documents into retrievable chunks.
  3. An embedding model — convert those chunks into vectors.
  4. A vector store — hold those vectors and find relevant matches.
  5. An LLM chain — generate an answer using retrieved context.

LangChain's whole value proposition is gluing these five pieces together with consistent interfaces, so swapping any one piece later doesn't require rewriting everything else.

Setting Up Your Environment

LangChain split into multiple packages a while back, which trips up people following outdated tutorials. Make sure you're installing the current modular packages, not the old monolithic one.

python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

pip install langchain langchain-openai langchain-chroma langchain-community pypdf

FYI, if you find a tutorial importing from langchain.vectorstores or langchain.embeddings directly, that's outdated syntax. Modern LangChain wants you importing from the specific provider package instead — langchain_openai, langchain_chroma, and so on.

API Key Setup

You'll need an OpenAI API key for this version of the tutorial.

import os
os.environ["OPENAI_API_KEY"] = "your-key-here"

Don't hardcode this into anything you commit to version control. I've made this exact mistake once, and GitHub's secret scanning caught it embarrassingly fast — a useful safety net, but an avoidable embarrassment.

Step 1: Loading Your Documents

Let's load some PDFs into memory using LangChain's built-in document loader.

from langchain_community.document_loaders import PyPDFLoader

def load_documents(pdf_paths):
    documents = []
    for path in pdf_paths:
        loader = PyPDFLoader(path)
        documents.extend(loader.load())
    return documents

docs = load_documents(["handbook.pdf", "policy.pdf"])
print(f"Loaded {len(docs)} pages")

Each page comes back as a separate Document object with its own metadata — including which page it came from. That page metadata is what eventually powers your source citations, so don't discard it along the way.

Step 2: Splitting Into Chunks

Raw documents are too long to embed usefully or fit into an LLM's context window wholesale. We need to split them into manageable, retrievable pieces.

from langchain_text_splitters import RecursiveCharacterTextSplitter

text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
    separators=["\n\n", "\n", " ", ""]
)

chunks = text_splitter.split_documents(docs)
print(f"Split into {len(chunks)} chunks")

That separators list matters more than it looks — RecursiveCharacterTextSplitter tries paragraph breaks first, then line breaks, then spaces, only falling back to raw character splitting as a last resort. This respects your document's natural structure instead of slicing arbitrarily through sentences.

Why 1000 Tokens with 200 Overlap?

This is a reasonable starting point, not a universal law. Factoid-heavy content does fine with smaller chunks; explanatory or analytical content usually wants more. Start here, test against your actual documents, and adjust based on what your retrieval quality actually shows you — not based on what a tutorial (this one included) claims is optimal.

Step 3: Embedding and Storing in Chroma

Now let's convert those chunks into vectors and store them in a local Chroma database.

from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

vector_store = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    persist_directory="./chroma_db"
)

That persist_directory argument means your embeddings survive between script runs — you won't be paying to re-embed the same documents every single time you restart. I skipped this step once during early experimentation and burned through API calls needlessly. Don't repeat my mistake.

Step 4: Building the Retriever

With your vector store built, converting it into a retriever LangChain can use downstream takes one line.

retriever = vector_store.as_retriever(search_kwargs={"k": 4})

That k=4 means "retrieve the four most relevant chunks per query." Ever wondered why some RAG bots feel shallow, missing context that's clearly relevant? Often it's just this number set too low for the complexity of the questions being asked.

Step 5: Wiring in the LLM

Now let's connect an LLM that generates answers grounded in whatever the retriever hands it.

from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.2)

prompt = ChatPromptTemplate.from_template("""
Answer the question using only the context below.
If the answer isn't in the context, say you don't know.

Context:
{context}

Question: {question}
""")

def format_docs(docs):
    return "\n\n".join(doc.page_content for doc in docs)

rag_chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | prompt
    | llm
    | StrOutputParser()
)

That instruction telling the model to admit uncertainty does more to reduce hallucination than nearly anything else in this pipeline. Skip it, and the model happily guesses instead of saying "I don't know" — exactly the failure mode RAG exists to prevent.

What's Happening With That Chain Syntax?

This pipe-operator style is LangChain Expression Language (LCEL) — it chains components together declaratively, where each piece's output feeds into the next. It looks unusual at first, but it's genuinely elegant once it clicks: retrieve, format, prompt, generate, parse — one readable pipeline.

Adding Source Citations

A raw answer is useful; an answer with citations is trustworthy. Let's modify our approach to return sources alongside the generated text.

def ask_with_sources(question):
    retrieved_docs = retriever.invoke(question)
    context = format_docs(retrieved_docs)

    answer = llm.invoke(
        prompt.format(context=context, question=question)
    ).content

    sources = list({
        doc.metadata.get("source", "unknown") + f" (page {doc.metadata.get('page', '?')})"
        for doc in retrieved_docs
    })

    return {"answer": answer, "sources": sources}

result = ask_with_sources("What's our refund policy?")
print(result["answer"])
print("Sources:", result["sources"])

Now users can actually verify your bot's answer instead of just trusting it blindly. That verifiability is arguably the whole point of RAG over a plain chatbot — you're not just getting an answer, you're getting a trail back to where it came from.

Wrapping It in an API

For a real deployment, you'll likely want this behind an API rather than a script. Here's a minimal FastAPI wrapper.

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI(title="Document QnA API")

class Question(BaseModel):
    query: str

@app.post("/ask")
def ask(q: Question):
    return ask_with_sources(q.query)

Run it with uvicorn main:app --reload, and you've got a working document Q&A endpoint. Genuinely not much more code than that stands between "prototype" and "something you could demo to a team."

Common Mistakes When Building This

I've made a few of these myself, so consider this a shortcut past my own trial and error.

  • Forgetting persist_directory on your vector store. Without it, you're re-embedding documents on every single script run — wasted time and money.
  • Setting retrieval k too low. Complex questions often need more than 2-3 chunks of context; test against real queries before locking this in.
  • Skipping the "say you don't know" instruction. Without it, your bot will confidently answer questions your documents don't actually cover.
  • Using outdated import paths from old tutorials. LangChain's package split means langchain.vectorstores.Chroma style imports are stale — use langchain_chroma instead.

When Should You Move Beyond This Basic Setup?

This tutorial gets you a genuinely working bot, but production systems typically add a few more layers on top: hybrid search combining keyword and vector retrieval, reranking for precision, and conversation memory for multi-turn dialogue.

LangChain supports all of these as incremental additions rather than rewrites — that's honestly the framework's biggest strength. You're not locked into today's simple version; you're building the foundation the more sophisticated version sits on top of.

Frequently Asked Questions

What is LangChain RAG?

LangChain RAG combines document loading, text splitting, embedding, vector storage, and LLM generation into one pipeline. It lets you query your own documents using natural language with source citations.

How do I load PDFs into LangChain?

Use PyPDFLoader from langchain_community.document_loaders. Each page becomes a Document object with metadata including page number, which powers source citations later.

What chunk size should I use for RAG?

Start with 1000 tokens and 200 overlap as a baseline. Factoid queries may need smaller chunks. Analytical content may need larger chunks. Test against your actual documents and adjust.

Which vector store should I use with LangChain?

Chroma is the easiest for local development with persist_directory support. For production, consider Qdrant, Pinecone, or pgvector depending on your scale and infrastructure needs.

How do I add source citations to RAG answers?

Retrieve the documents with metadata, then extract source and page information from each document's metadata. Include these in your response alongside the generated answer.

How do I reduce hallucination in RAG?

Add an instruction telling the model to say "I don't know" if the answer isn't in the context. This simple prompt engineering reduces hallucination more than most technical fixes.

Wrapping This Up

Building a document Q&A bot with LangChain boils down to five connected pieces: load your documents, split them into chunks, embed and store them, retrieve relevant context per query, and generate a grounded answer with citations. Everything more advanced builds incrementally on this same foundation.

Remember to persist your vector store so you're not re-embedding on every run, and don't skip the instruction telling your model to admit when it doesn't know something. FYI, this exact pattern — load, split, embed, retrieve, generate — is standardized enough now that once you've built it once, every future RAG project feels like assembling familiar pieces rather than starting from scratch :)

Now go point this at your own PDFs instead of my hypothetical handbook and policy documents. That's where this stuff actually starts proving its worth.

Share this article X Facebook LinkedIn Reddit WhatsApp