Sam Austin AI

Building a Local RAG Chatbot with Ollama and LangChain

September 7, 2026 9 min read Sam Austin
Contents
Local RAG chatbot running entirely on personal machine with Ollama and LangChain
Local RAG chatbot running entirely on personal machine with Ollama and LangChain

Somewhere back at the start of this whole journey, the RAG series covered building a document Q&A bot with LangChain — using OpenAI's API, an API key, and a per-token bill. This is that same project, rebuilt with everything not leaving your machine. No API key, no network call, no data ever touching a server you don't control. Just Ollama serving the model and LangChain wiring the pipeline together.

This is genuinely the moment two entire arcs of this series converge — the RAG fundamentals from the first act, and the local LLM tooling from everything since Ollama, llama.cpp, and GGUF. If you've followed both threads, this tutorial should feel less like new material and more like watching two things you already understand snap together.

By the end of this guide, you'll have a fully local RAG chatbot answering questions about your own documents, with zero cloud dependency anywhere in the pipeline. IMO, this is genuinely the most practically useful single project in this entire series — it's the one most people actually want to build for real :)

Why Local RAG Is Worth the Setup

Recall the original LangChain RAG tutorial's five-piece pipeline: load documents, split into chunks, embed and store, retrieve, generate. Every piece of that pipeline can now run locally, and there are genuinely good reasons to do so rather than defaulting to a cloud API.

  • Privacy: if your documents are legal contracts, medical records, or internal company data, they never leave your machine — not even as embeddings sent to a third-party API.
  • Cost: no per-token billing regardless of how many documents you index or how many questions you ask.
  • Offline operation: works genuinely without internet, useful for air-gapped environments or just avoiding a network dependency for something that should work locally.

The tradeoff, stated honestly upfront: local models are generally less capable than frontier cloud models, and local embedding models likewise trail the best hosted options. For most document Q&A use cases, that gap is genuinely tolerable — but it's worth knowing you're trading some capability for these benefits, not getting them for free.

The Stack: What Replaces What

Mapping directly onto the original LangChain RAG tutorial, here's what changes and what stays identical.

Original PieceLocal Replacement
PyPDFLoaderUnchanged — document loading has no cloud dependency either way
RecursiveCharacterTextSplitterUnchanged — chunking logic doesn't care where the model runs
OpenAIEmbeddingsOllamaEmbeddings — local embedding generation
ChatOpenAIChatOllama — local generation
Chroma (persist_directory)Unchanged — Chroma was already running locally

Genuinely, only two lines fundamentally change. Everything else from the original tutorial's architecture — loaders, splitters, the vector store, the LCEL chain syntax — carries over completely unmodified.

Setting Up the Environment

python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

pip install langchain langchain-ollama langchain-chroma langchain-community pypdf

Notice langchain-ollama replacing langchain-openai — this is the actual package swap driving everything else in this tutorial. No API key setup section this time, since there's genuinely nothing to authenticate against.

Before writing any Python, confirm Ollama itself is running with the models you'll need:

ollama pull llama3.2
ollama pull nomic-embed-text

That second model matters specifically — nomic-embed-text is a dedicated embedding model, genuinely more suited to this task than asking a general chat model to produce embeddings. Recall from the embedding models comparison earlier in this series: a model trained specifically for retrieval tasks generally outperforms a generalist model repurposed for the same job.

Step 1-2: Loading and Splitting (Unchanged)

from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter

def load_documents(pdf_paths):
    documents = []
    for path in pdf_paths:
        loader = PyPDFLoader(path)
        documents.extend(loader.load())
    return documents

docs = load_documents(["handbook.pdf", "policy.pdf"])

text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
    separators=["\n\n", "\n", " ", ""]
)
chunks = text_splitter.split_documents(docs)

This is copied directly from the original tutorial, unchanged — recall the chunking strategies article from earlier in this series: recursive character splitting at roughly this chunk size remains a solid, benchmark-validated default regardless of whether your embedding model runs locally or in the cloud.

Step 3: Embedding with Ollama Instead of OpenAI

from langchain_ollama import OllamaEmbeddings
from langchain_chroma import Chroma

embeddings = OllamaEmbeddings(model="nomic-embed-text")

vector_store = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    persist_directory="./chroma_db"
)

This is genuinely the entire embedding-side change — swap OpenAIEmbeddings(model="text-embedding-3-small") for OllamaEmbeddings(model="nomic-embed-text"), and everything downstream (Chroma's storage, the persist directory, the retriever interface) works identically to the original tutorial.

One thing worth setting expectations on: embedding an entire document collection locally will genuinely be slower than an API call, since you're using your own CPU or GPU rather than a data center's. For a handful of PDFs, this difference is trivial; for a genuinely large document collection, budget real time for the initial indexing pass.

Step 4: The Retriever (Unchanged)

retriever = vector_store.as_retriever(search_kwargs={"k": 4})

Identical to the original tutorial — the retriever interface doesn't know or care whether the underlying embeddings came from OpenAI or Ollama.

Step 5: Generation with ChatOllama

from langchain_ollama import ChatOllama
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

llm = ChatOllama(model="llama3.2", temperature=0.2)

prompt = ChatPromptTemplate.from_template("""
Answer the question using only the context below.
If the answer isn't in the context, say you don't know.

Context:
{context}

Question: {question}
""")

def format_docs(docs):
    return "\n\n".join(doc.page_content for doc in docs)

rag_chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | prompt
    | llm
    | StrOutputParser()
)

Notice this is genuinely the exact same LCEL chain from the original tutorial — retrieve, format, prompt, generate, parse — with ChatOllama swapped in for ChatOpenAI. That "say you don't know" instruction still does the same critical hallucination-prevention work it always did, regardless of which model is actually generating the response.

response = rag_chain.invoke("What's our refund policy?")
print(response)

Watch this respond with zero network activity happening anywhere, and you've genuinely built the same document Q&A capability from the very first RAG tutorial in this series, running completely offline.

Choosing the Right Local Model for This Task

Not every Ollama model is equally suited to RAG specifically — recall the embedding models and LLM comparison articles from earlier in this series for the general principles, applied here concretely.

  • llama3.2 (3B) — a genuinely solid default for RAG on modest hardware; fast, and generally reliable at the "answer from context, admit uncertainty" instruction-following this task depends on.
  • Larger models (7B-8B class) — worth trying if your questions require more nuanced synthesis across multiple retrieved chunks, at the cost of slower generation and higher memory requirements.
  • deepseek-r1 or similar reasoning-focused models — genuinely useful if your document Q&A involves multi-step reasoning rather than straightforward lookup, though typically slower per response.

Start with the smallest model that gives acceptable answers, exactly the same principle from the Ollama setup guide earlier in this series — scale up only once you've confirmed you actually need the extra capability.

If you're running larger models for RAG on consumer hardware, a capable GPU speeds up both embedding and generation significantly — the RTX 5070 and RTX 5080 both handle 7B-8B RAG pipelines comfortably, and the faster iteration cycles genuinely matter when tuning retrieval parameters.

Adding Source Citations (Unchanged Logic)

def ask_with_sources(question):
    retrieved_docs = retriever.invoke(question)
    context = format_docs(retrieved_docs)

    answer = llm.invoke(
        prompt.format(context=context, question=question)
    ).content

    sources = list({
        doc.metadata.get("source", "unknown") + f" (page {doc.metadata.get('page', '?')})"
        for doc in retrieved_docs
    })

    return {"answer": answer, "sources": sources}

Copied directly from the original tutorial, unmodified — source citation logic operates on document metadata, which has nothing to do with which model generated the final answer.

Wrapping It in an API (Unchanged)

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI(title="Local Document QnA API")

class Question(BaseModel):
    query: str

@app.post("/ask")
def ask(q: Question):
    return ask_with_sources(q.query)

Genuinely nothing changes here either — the FastAPI wrapper doesn't care what's running underneath the ask_with_sources function it's calling.

Improving This Pipeline With Techniques From Earlier in This Series

Since this project sits at the intersection of two full arcs of this series, several earlier articles apply directly as upgrades.

  • Hybrid search: recall the RRF-based hybrid search article — combining BM25 keyword matching with Ollama's local embeddings would catch exact-term queries (product codes, specific names) that pure vector search alone might miss.
  • Reranking: a local cross-encoder like bge-reranker-v2-m3, self-hosted rather than Cohere's hosted API, keeps the entire pipeline offline while still capturing reranking's meaningful accuracy gains.
  • Quantization: if you're running on modest hardware, recall the GGUF article's guidance — Ollama already handles this automatically by pulling appropriately quantized model tags, but explicitly choosing a Q4_K_M-tagged model over a larger default is worth doing on constrained hardware.

None of these require abandoning the local-first approach — they're genuinely additive improvements, each one already covered in depth earlier in this series.

Want to Go Deeper?

If the RAG pipeline or LangChain LCEL chain concepts clicked and you want to dig into the theory behind retrieval-augmented generation and vector search architectures, Educative's Machine Learning path covers RAG systems in detail alongside the broader LLM application landscape — worth exploring if you're building production RAG pipelines beyond this tutorial.

Common Mistakes People Make

  • Using a chat model for embeddings instead of a dedicated embedding model. nomic-embed-text or similar purpose-built models genuinely outperform repurposing a general chat model for this specific task.
  • Forgetting persist_directory on the Chroma vector store. Exactly the same mistake flagged in the original tutorial — without it, you're re-embedding your entire document collection on every script run.
  • Expecting cloud-model-level quality from a small local model without adjusting expectations. A 3B local model genuinely won't match GPT-4-class synthesis quality — this is the real tradeoff for privacy and cost, not a bug in your setup.
  • Not confirming Ollama is actually running before debugging your Python code. A surprising number of "my RAG chain isn't working" issues trace back to Ollama's background service not being active — check ollama ps before assuming your code is broken.
  • Skipping the "say you don't know" prompt instruction. This matters just as much for local models as it did for GPT-4 in the original tutorial — local models hallucinate too, and this single line remains one of the cheapest mitigations available.

Wrapping This Up

Building this pipeline locally required changing genuinely two lines from the original LangChain RAG tutorial — swap OpenAIEmbeddings for OllamaEmbeddings, swap ChatOpenAI for ChatOllama — and everything else from document loading through chunking, retrieval, and source citation carries over completely unchanged. That's honestly the best evidence that RAG's architecture was never actually coupled to any specific cloud provider in the first place.

Remember to pull a dedicated embedding model rather than repurposing a chat model for that role, and set realistic expectations that a local 3B model trades some synthesis quality for the genuine privacy, cost, and offline benefits this whole approach provides. FYI, if you've read this entire series from the first RAG tutorial through Ollama, GGUF, and quantization, you now have every piece needed to build this without a single new concept — this tutorial was genuinely just the assembly step :)

Now go point this at your own documents, disconnect your network entirely, and confirm it still answers correctly. That's the moment this stops being an abstract "local AI" idea from earlier articles and becomes something concretely working on your own machine.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles