Sam Austin AI

Haystack Tutorial: Build a RAG Pipeline with Deepset's Framework

September 23, 2026 12 min read Sam Austin
Contents

Ever asked an LLM a question about your own company docs and gotten back a confident, beautifully-written answer that's completely wrong? Yeah, me too. That's what happens when you hand a language model a question and hope for the best — it just makes things up and smiles while doing it. RAG (Retrieval-Augmented Generation) fixes this, and Haystack by deepset makes building RAG pipelines genuinely pleasant.

I'll confess something upfront: I used to build RAG systems by hand. I glued together retrievers, embedders, and prompt templates with spaghetti code that I now look back on with a mixture of shame and awe. Then I found Haystack 2.x, and suddenly my glue code disappeared. Let me show you how to build a full RAG pipeline today — no PhD required. :)

What Is Haystack, and Why Should You Care?

Haystack is deepset's open-source framework for building applications with LLMs — think RAG, semantic search, and agents. The keyword here is framework: it gives you pre-built, Lego-like components you snap together into pipelines.

Here's why I'm a fan:

  • Components do one thing well — embedders, retrievers, generators, prompt builders all come pre-packaged
  • Pipelines connect everything — you wire components together and Haystack handles the data flow
  • Provider flexibility — swap OpenAI for a local model without rewriting your whole app
  • Battle-tested — deepset runs production search systems on this thing

Compare that to a DIY approach, where you write your own chunking, embedding, retrieval, and prompt plumbing. IMO, life's too short for plumbing. And if you've followed this series' RAG beginner's guide, Haystack is exactly the framework that packages those concepts into code you don't have to wire yourself.

Installing Haystack

Fair warning: old tutorials online reference Haystack 1.x, which uses a completely different API. If you copy code from a 2023 blog post, it won't work. Search for Haystack 2.x docs, or just trust this tutorial.

Install it with pip:

pip install haystack-ai

Note: the package name is haystack-ai, not haystack. The old name installs the legacy version, and it will silently ruin your afternoon. Ask me how I know. :/

Understanding the Architecture

Before we write code, let's talk about how Haystack thinks. Everything in Haystack 2.x revolves around two ideas:

Components

A component is a self-contained unit with typed inputs and outputs. SentenceTransformersTextEmbedder takes text and produces embeddings. OpenAIGenerator takes a prompt and produces an answer. Each one handles exactly one job.

Pipelines

A pipeline orchestrates components. You add components to a pipeline, then connect one component's output to another component's input. Haystack validates all the connections before running anything — so if you wire things up wrong, it tells you immediately instead of failing mysteriously at runtime. Small thing? Maybe. But after years of debugging silent failures, I appreciate it deeply.

The two pipelines in this tutorial map onto the RAG stages directly:

Pipeline Components Job
Indexing DocumentSplitter → DocumentEmbedder → DocumentWriter Ingest, chunk, and embed documents into the store
Query TextEmbedder → Retriever → PromptBuilder → Generator Embed the question, retrieve chunks, ground the prompt, generate the answer

Building the Indexing Pipeline

RAG needs data before it can retrieve anything, right? So first we'll build a pipeline that ingests documents and creates embeddings.

from haystack import Pipeline
from haystack.components.writers import DocumentWriter
from haystack.components.preprocessors import DocumentSplitter
from haystack.components.embedders import (
    SentenceTransformersDocumentEmbedder,
)
from haystack.datastore.types import DuplicatePolicy
from haystack.document_stores.in_memory import InMemoryDocumentStore

document_store = InMemoryDocumentStore()

indexing = Pipeline()
indexing.add_component("splitter", DocumentSplitter(
    split_by="word", split_length=200, split_overlap=30
))
indexing.add_component("embedder", SentenceTransformersDocumentEmbedder(
    model="sentence-transformers/all-MiniLM-L6-v2"
))
indexing.add_component("writer", DocumentWriter(
    document_store, policy=DuplicatePolicy.SKIP
))

indexing.connect("splitter", "embedder")
indexing.connect("embedder", "writer")

What's Happening Here?

Let's break down each piece:

  • DocumentSplitter — chops long documents into overlapping chunks of 200 words
  • SentenceTransformersDocumentEmbedder — converts each chunk into a vector using a free, local model
  • DocumentWriter — stores everything in the document store
  • InMemoryDocumentStore — Haystack's built-in store, perfect for prototyping

Notice the split_overlap=30? That overlap prevents you from cutting sentences in half and losing context at chunk boundaries. Retrieval quality lives and dies on chunking, so tune this number for your data — same 200–500 token guidance from the building a RAG pipeline tutorial earlier in this series.

Now run it with some documents:

from haystack import Document

docs = [
    Document(content="Haystack is an open-source framework by deepset."),
    Document(content="RAG combines retrieval with generation for accurate answers."),
]

indexing.run({"splitter": {"documents": docs}})

The embedder downloads the model on first run — go grab a coffee, it takes a minute. After that, everything runs locally and fast.

Building the Query Pipeline

Here's where the magic happens. This pipeline takes a user's question, retrieves relevant chunks, and generates an answer grounded in your data.

from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryEmbeddingRetriever
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator

template = """
Given these documents, answer the question.
Documents:
{% for doc in documents %}
    {{ doc.content }}
{% endfor %}
Question: {{ question }}
Answer:
"""

query = Pipeline()
query.add_component("text_embedder", SentenceTransformersTextEmbedder(
    model="sentence-transformers/all-MiniLM-L6-v2"
))
query.add_component("retriever", InMemoryEmbeddingRetriever(document_store))
query.add_component("prompt", PromptBuilder(template=template))
query.add_component("generator", OpenAIGenerator())

query.connect("text_embedder", "embeddings", "retriever.query_embedding")
query.connect("retriever", "documents", "prompt.documents")
query.connect("prompt", "prompt", "generator.prompt")

Why the Same Embedding Model?

I hammered this point in my last tutorial, and I'll hammer it again: use the same embedding model for indexing and querying. Your query vector and your document vectors must live in the same mathematical space. Mix models, and your retriever returns nonsense — confidently. LLMs have taught me that everything fails confidently these days — same warning the embedding models compared article made about picking one model and sticking with it.

Running Your RAG Pipeline

Time for the moment of truth:

question = "What is Haystack?"

results = query.run({
    "text_embedder": {"text": question},
    "prompt": {"question": question},
})

print(results["generator"]["replies"][0])

The flow looks like this:

  1. The question becomes a vector via the text embedder
  2. The retriever finds the most similar document chunks
  3. The prompt builder stuffs those chunks into your template
  4. OpenAIGenerator produces an answer grounded in your actual data

FYI, you'll need your OpenAI API key set as an environment variable (OPENAI_API_KEY) for the generator to work. Don't hardcode keys in your scripts �� future you will thank present you.

Haystack RAG Pipeline Deepset Framework Tutorial

Figure 1: Haystack by deepset — components and pipelines for RAG

Image Alt Text: "Haystack tutorial building a RAG pipeline with deepset components and prompt builders"

Leveling Up: Production Tips

The setup above works great for a prototype. When you get serious, consider these upgrades:

Swap the Document Store

InMemoryDocumentStore lives in RAM and vanishes when your process ends. For real deployments, pick a persistent store. Haystack supports several, including Qdrant, Weaviate, Elasticsearch, and — pleasant callback to last time — Postgres with pgvector. I'd start with Qdrant; it's fast, Docker-friendly, and painless to run — see the Qdrant tutorial and the Supabase pgvector tutorial from this series for both options.

Use Better Models

all-MiniLM-L6-v2 runs fast on a laptop, but it's the entry-level option. If you want sharper retrieval, try BAAI/bge-base-en-v1.5 or an OpenAI embedding model. Better embeddings mean better answers — often more than a fancier generator does.

Add Metadata Filtering

Real pipelines filter by metadata — document source, date, category. Haystack handles this with retrieval filters, so you can restrict search to, say, only HR policies from 2025. Your users will love you for this feature.

Save Your Pipelines

Haystack pipelines serialize to YAML with pipeline.dump(path). You can then load and version them in Git like any other config. This trick has saved my bacon when reproducing a setup on a different machine — no wait, which model did we use? Slack messages required.

Where RAG Shines (and Where It Doesn't)

Let's be honest about use cases, because blind enthusiasm helps nobody. RAG works brilliantly for:

  • Chatbots over your documentation — support bots, internal knowledge assistants
  • Search that understands meaning — not just keyword matching
  • Any LLM app needing current or private data — models can't know your company's Q3 numbers

Where does it struggle? Highly structured analytical queries (compute year-over-year growth from this spreadsheet) need actual computation, not retrieval. And if your documents are messy, your answers will be messy too. RAG amplifies your data quality — good and bad. Garbage in, confident garbage out. :/

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is Haystack?

Haystack is deepset's open-source framework for building LLM applications such as RAG, semantic search, and agents. In Haystack 2.x, you snap pre-built components (embedders, retrievers, generators, prompt builders) into pipelines that validate connections before running.

What is the difference between Haystack 1.x and 2.x?

Haystack 2.x uses a completely different API from 1.x — components with typed inputs/outputs connected in Pipeline objects. Old 2023 tutorial code written for 1.x will not work; search for Haystack 2.x docs specifically.

What package do I install for Haystack?

Install haystack-ai with pip install haystack-ai. The old package name haystack installs the legacy 1.x version, which can silently break code written for the modern API.

How does a Haystack RAG pipeline work?

An indexing pipeline splits documents, embeds them with a model like all-MiniLM-L6-v2, and writes them to a document store. A query pipeline embeds the question, retrieves similar chunks, builds a grounded prompt with PromptBuilder, and passes it to a generator like OpenAIGenerator for the final answer.

Which document store should I use with Haystack?

InMemoryDocumentStore is fine for prototyping but vanishes when the process ends. For production, Haystack supports persistent stores including Qdrant, Weaviate, Elasticsearch, and Postgres with pgvector — Qdrant is a solid Docker-friendly default.

Why must indexing and querying use the same embedding model?

Query vectors and document vectors must live in the same mathematical space. Mixing models means cosine similarity compares unrelated vector spaces, so the retriever returns nonsense results — often with full confidence.

Wrapping Up

Look at what you just built: a complete RAG pipeline with chunking, embeddings, retrieval, prompt templating, and generation — in maybe 60 lines of clean Python. No spaghetti glue code, no silent failures, and every component swappable when your needs change.

Here's my parting advice: prototype with the in-memory store, then swap in Qdrant when you're ready for real traffic. Upgrade your embedding model before you upgrade your generator. And dump your pipelines to YAML so future-you can reproduce everything without archaeology.

Now go build something. Index your own docs, ask your bot ridiculous questions, and watch it answer with citations instead of hallucinations. And if you accidentally install the legacy haystack package and spend an hour confused? No judgment. Welcome to the club — membership is free and the meetings are just error messages. ;)

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles