Contents
Ever asked an LLM a question about your own company docs and gotten back a confident, beautifully-written answer that's completely wrong? Yeah, me too. That's what happens when you hand a language model a question and hope for the best — it just makes things up and smiles while doing it. RAG (Retrieval-Augmented Generation) fixes this, and Haystack by deepset makes building RAG pipelines genuinely pleasant.
I'll confess something upfront: I used to build RAG systems by hand. I glued together retrievers, embedders, and prompt templates with spaghetti code that I now look back on with a mixture of shame and awe. Then I found Haystack 2.x, and suddenly my glue code disappeared. Let me show you how to build a full RAG pipeline today — no PhD required. :)
What Is Haystack, and Why Should You Care?
Haystack is deepset's open-source framework for building applications with LLMs — think RAG, semantic search, and agents. The keyword here is framework: it gives you pre-built, Lego-like components you snap together into pipelines.
Here's why I'm a fan:
- Components do one thing well — embedders, retrievers, generators, prompt builders all come pre-packaged
- Pipelines connect everything — you wire components together and Haystack handles the data flow
- Provider flexibility — swap OpenAI for a local model without rewriting your whole app
- Battle-tested — deepset runs production search systems on this thing
Compare that to a DIY approach, where you write your own chunking, embedding, retrieval, and prompt plumbing. IMO, life's too short for plumbing. And if you've followed this series' RAG beginner's guide, Haystack is exactly the framework that packages those concepts into code you don't have to wire yourself.
Installing Haystack
Fair warning: old tutorials online reference Haystack 1.x, which uses a completely different API. If you copy code from a 2023 blog post, it won't work. Search for Haystack 2.x docs, or just trust this tutorial.
Install it with pip:
pip install haystack-ai
Note: the package name is haystack-ai, not haystack. The old name installs the legacy version, and it will silently ruin your afternoon. Ask me how I know. :/
Understanding the Architecture
Before we write code, let's talk about how Haystack thinks. Everything in Haystack 2.x revolves around two ideas:
Components
A component is a self-contained unit with typed inputs and outputs. SentenceTransformersTextEmbedder takes text and produces embeddings. OpenAIGenerator takes a prompt and produces an answer. Each one handles exactly one job.
Pipelines
A pipeline orchestrates components. You add components to a pipeline, then connect one component's output to another component's input. Haystack validates all the connections before running anything — so if you wire things up wrong, it tells you immediately instead of failing mysteriously at runtime. Small thing? Maybe. But after years of debugging silent failures, I appreciate it deeply.
The two pipelines in this tutorial map onto the RAG stages directly:
| Pipeline | Components | Job |
|---|---|---|
| Indexing | DocumentSplitter → DocumentEmbedder → DocumentWriter | Ingest, chunk, and embed documents into the store |
| Query | TextEmbedder → Retriever → PromptBuilder → Generator | Embed the question, retrieve chunks, ground the prompt, generate the answer |
Building the Indexing Pipeline
RAG needs data before it can retrieve anything, right? So first we'll build a pipeline that ingests documents and creates embeddings.
from haystack import Pipeline
from haystack.components.writers import DocumentWriter
from haystack.components.preprocessors import DocumentSplitter
from haystack.components.embedders import (
SentenceTransformersDocumentEmbedder,
)
from haystack.datastore.types import DuplicatePolicy
from haystack.document_stores.in_memory import InMemoryDocumentStore
document_store = InMemoryDocumentStore()
indexing = Pipeline()
indexing.add_component("splitter", DocumentSplitter(
split_by="word", split_length=200, split_overlap=30
))
indexing.add_component("embedder", SentenceTransformersDocumentEmbedder(
model="sentence-transformers/all-MiniLM-L6-v2"
))
indexing.add_component("writer", DocumentWriter(
document_store, policy=DuplicatePolicy.SKIP
))
indexing.connect("splitter", "embedder")
indexing.connect("embedder", "writer")
What's Happening Here?
Let's break down each piece:
DocumentSplitter— chops long documents into overlapping chunks of 200 wordsSentenceTransformersDocumentEmbedder— converts each chunk into a vector using a free, local modelDocumentWriter— stores everything in the document storeInMemoryDocumentStore— Haystack's built-in store, perfect for prototyping
Notice the split_overlap=30? That overlap prevents you from cutting sentences in half and losing context at chunk boundaries. Retrieval quality lives and dies on chunking, so tune this number for your data — same 200–500 token guidance from the building a RAG pipeline tutorial earlier in this series.
Now run it with some documents:
from haystack import Document
docs = [
Document(content="Haystack is an open-source framework by deepset."),
Document(content="RAG combines retrieval with generation for accurate answers."),
]
indexing.run({"splitter": {"documents": docs}})
The embedder downloads the model on first run — go grab a coffee, it takes a minute. After that, everything runs locally and fast.
Building the Query Pipeline
Here's where the magic happens. This pipeline takes a user's question, retrieves relevant chunks, and generates an answer grounded in your data.
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryEmbeddingRetriever
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator
template = """
Given these documents, answer the question.
Documents:
{% for doc in documents %}
{{ doc.content }}
{% endfor %}
Question: {{ question }}
Answer:
"""
query = Pipeline()
query.add_component("text_embedder", SentenceTransformersTextEmbedder(
model="sentence-transformers/all-MiniLM-L6-v2"
))
query.add_component("retriever", InMemoryEmbeddingRetriever(document_store))
query.add_component("prompt", PromptBuilder(template=template))
query.add_component("generator", OpenAIGenerator())
query.connect("text_embedder", "embeddings", "retriever.query_embedding")
query.connect("retriever", "documents", "prompt.documents")
query.connect("prompt", "prompt", "generator.prompt")
Why the Same Embedding Model?
I hammered this point in my last tutorial, and I'll hammer it again: use the same embedding model for indexing and querying. Your query vector and your document vectors must live in the same mathematical space. Mix models, and your retriever returns nonsense — confidently. LLMs have taught me that everything fails confidently these days — same warning the embedding models compared article made about picking one model and sticking with it.
Running Your RAG Pipeline
Time for the moment of truth:
question = "What is Haystack?"
results = query.run({
"text_embedder": {"text": question},
"prompt": {"question": question},
})
print(results["generator"]["replies"][0])
The flow looks like this:
- The question becomes a vector via the text embedder
- The retriever finds the most similar document chunks
- The prompt builder stuffs those chunks into your template
OpenAIGeneratorproduces an answer grounded in your actual data
FYI, you'll need your OpenAI API key set as an environment variable (OPENAI_API_KEY) for the generator to work. Don't hardcode keys in your scripts �� future you will thank present you.
Figure 1: Haystack by deepset — components and pipelines for RAG
Image Alt Text: "Haystack tutorial building a RAG pipeline with deepset components and prompt builders"
Leveling Up: Production Tips
The setup above works great for a prototype. When you get serious, consider these upgrades:
Swap the Document Store
InMemoryDocumentStore lives in RAM and vanishes when your process ends. For real deployments, pick a persistent store. Haystack supports several, including Qdrant, Weaviate, Elasticsearch, and — pleasant callback to last time — Postgres with pgvector. I'd start with Qdrant; it's fast, Docker-friendly, and painless to run — see the Qdrant tutorial and the Supabase pgvector tutorial from this series for both options.
Use Better Models
all-MiniLM-L6-v2 runs fast on a laptop, but it's the entry-level option. If you want sharper retrieval, try BAAI/bge-base-en-v1.5 or an OpenAI embedding model. Better embeddings mean better answers — often more than a fancier generator does.
Add Metadata Filtering
Real pipelines filter by metadata — document source, date, category. Haystack handles this with retrieval filters, so you can restrict search to, say, only HR policies from 2025. Your users will love you for this feature.
Save Your Pipelines
Haystack pipelines serialize to YAML with pipeline.dump(path). You can then load and version them in Git like any other config. This trick has saved my bacon when reproducing a setup on a different machine — no wait, which model did we use? Slack messages required.
Where RAG Shines (and Where It Doesn't)
Let's be honest about use cases, because blind enthusiasm helps nobody. RAG works brilliantly for:
- Chatbots over your documentation — support bots, internal knowledge assistants
- Search that understands meaning — not just keyword matching
- Any LLM app needing current or private data — models can't know your company's Q3 numbers
Where does it struggle? Highly structured analytical queries (compute year-over-year growth from this spreadsheet) need actual computation, not retrieval. And if your documents are messy, your answers will be messy too. RAG amplifies your data quality — good and bad. Garbage in, confident garbage out. :/
Recommended Books
- Retrieval-Augmented Generation (RAG) — practical RAG architecture patterns that map directly onto Haystack's retrieve-prompt-generate pipeline shape.
- Natural Language Processing with Transformers by Tunstall et al. — deeper grounding on the embedding and generation models Haystack components wrap.
- Building LLM Powered Applications — end-to-end LLM app construction, useful companion context for when to reach for a framework like Haystack versus hand-rolled glue code.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is Haystack?
Haystack is deepset's open-source framework for building LLM applications such as RAG, semantic search, and agents. In Haystack 2.x, you snap pre-built components (embedders, retrievers, generators, prompt builders) into pipelines that validate connections before running.
What is the difference between Haystack 1.x and 2.x?
Haystack 2.x uses a completely different API from 1.x — components with typed inputs/outputs connected in Pipeline objects. Old 2023 tutorial code written for 1.x will not work; search for Haystack 2.x docs specifically.
What package do I install for Haystack?
Install haystack-ai with pip install haystack-ai. The old package name haystack installs the legacy 1.x version, which can silently break code written for the modern API.
How does a Haystack RAG pipeline work?
An indexing pipeline splits documents, embeds them with a model like all-MiniLM-L6-v2, and writes them to a document store. A query pipeline embeds the question, retrieves similar chunks, builds a grounded prompt with PromptBuilder, and passes it to a generator like OpenAIGenerator for the final answer.
Which document store should I use with Haystack?
InMemoryDocumentStore is fine for prototyping but vanishes when the process ends. For production, Haystack supports persistent stores including Qdrant, Weaviate, Elasticsearch, and Postgres with pgvector — Qdrant is a solid Docker-friendly default.
Why must indexing and querying use the same embedding model?
Query vectors and document vectors must live in the same mathematical space. Mixing models means cosine similarity compares unrelated vector spaces, so the retriever returns nonsense results — often with full confidence.
Wrapping Up
Look at what you just built: a complete RAG pipeline with chunking, embeddings, retrieval, prompt templating, and generation — in maybe 60 lines of clean Python. No spaghetti glue code, no silent failures, and every component swappable when your needs change.
Here's my parting advice: prototype with the in-memory store, then swap in Qdrant when you're ready for real traffic. Upgrade your embedding model before you upgrade your generator. And dump your pipelines to YAML so future-you can reproduce everything without archaeology.
Now go build something. Index your own docs, ask your bot ridiculous questions, and watch it answer with citations instead of hallucinations. And if you accidentally install the legacy haystack package and spend an hour confused? No judgment. Welcome to the club — membership is free and the meetings are just error messages. ;)