Sam Austin AI

Query Expansion and Rewriting for Better RAG Retrieval

September 24, 2026 13 min read Sam Austin
Contents

Here's a truth bomb that took me embarrassingly long to accept: most RAG failures aren't retrieval's fault. They're the query's fault. Your user types "can I get my money back," your documents say "refund eligibility criteria," and the embedding similarity between those two phrases is... let's call it polite nodding from across the street. Query expansion and rewriting fix this mismatch before retrieval ever runs, and today I'll show you how.

My awakening moment came from a support bot I built ages ago. It kept failing on the simplest questions, and I kept blaming the embedder, the chunking, the vector store — everything except the obvious. Then I printed the raw user queries next to the document language. Users ask like users; documents read like documents. The retrieval gap is usually a vocabulary gap, and no fancy embedder fully closes it. :/

Why Queries Fail in the First Place

Let's diagnose the problem properly, because you can't fix what you don't understand. Raw user queries break retrieval in a few classic ways:

  • Vocabulary mismatch — users say "money back," docs say "refund" (my favorite example, obviously)
  • Vagueness — "how does that work?" retrieves... what exactly?
  • Pronouns and context loss — a follow-up like "what about the second one?" contains zero searchable content
  • Single-vector weakness — one embedding has to represent a multi-part question, so it represents all of it badly

Ever noticed how your pipeline nails the questions you test with and whiffs on real users' questions? That's not a coincidence. You write test queries in document-speak. Users don't. Rude of them, honestly.

The fix in all these cases? Transform the query before you retrieve. Instead of forcing retrieval to work with a broken input, repair or enrich the input first. It sounds obvious now, right? It didn't sound obvious to me for an embarrassingly long time.

Query Rewriting: Fix the Query First

The simplest technique: pass the raw query through an LLM and have it rewrite the query into something retrieval-friendly. This handles vagueness, pronouns, and conversational debris.

Here's a minimal version with LangChain:

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_openai import ChatOpenAI

rewrite_prompt = ChatPromptTemplate.from_template(
    "Rewrite this user question into a clear, standalone search "
    "query for a document retrieval system. Resolve pronouns and "
    "add likely relevant keywords.\n\nQuestion: {question}\nQuery:"
)

rewriter = (
    rewrite_prompt
    | ChatOpenAI(model="gpt-4o-mini", temperature=0)
    | StrOutputParser()
)

clean_query = rewriter.invoke({"question": "what about the second one?"})
# → "What is the process for the second payment option?"

Notice the prompt resolves pronouns using the conversation history. This matters enormously for chat-style RAG, where follow-up questions are pronoun soup. If you feed "what about the second one?" directly to a retriever, you deserve the garbage you get back — and I say that with love, because I did exactly that for months. :D

Multi-Query Expansion: Retrieve From Many Angles

One query gives you one shot at the right chunk. Why settle for one? Multi-query expansion asks an LLM to generate several variations of the query, retrieves for each, and merges the results.

LangChain ships this as the MultiQueryRetriever:

import logging
from langchain.retrievers.multi_query import MultiQueryRetriever

logging.basicConfig(level=logging.INFO)

retriever = MultiQueryRetriever.from_llm(
    retriever=vectorstore.as_retriever(search_kwargs={"k": 4}),
    llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),
)

docs = retriever.invoke("how do I get my money back?")

Set the logging to INFO once and watch what it generates — you'll see something like:

  • What is the refund process?
  • How do customers request refunds?
  • What is the return and reimbursement policy?

Each variation pulls slightly different chunks, and the union of all three beats any single phrasing. IMO, this is the highest value-per-effort technique in the whole article: five lines of code, one extra LLM call, and retrieval recall improves on messy real-world queries immediately.

HyDE: Hypothetical Document Embeddings

Now for the strangest trick in the bag — and my personal favorite. HyDE stands for Hypothetical Document Embeddings, and the logic goes like this:

Your queries and your documents live in different linguistic worlds. But what if they didn't? What if you made the query look like a document?

Here's the flow:

  1. Ask an LLM to hallucinate a hypothetical answer to the query
  2. Embed that fake answer instead of the question
  3. Retrieve documents similar to the fake answer
  4. Feed the real retrieved documents to your generator

Sounds insane, right? It works because embeddings of fake answers land closer to real answers than embeddings of questions do. Both are answer-shaped text. The hallucination gets discarded immediately — you never show it to anyone — it just serves as a retrieval key.

from langchain_experimental.hyde import HypotheticalDocumentEmbedder

hyde_retriever = HypotheticalDocumentEmbedder.from_llm(
    llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),
    base_embeddings=OpenAIEmbeddings(),
    prompt_key="web_search",
)

docs = vectorstore.similarity_search_by_vector(
    hyde_retriever.embed_query("How does contextual compression work?")
)

Fair warning: HyDE shines on knowledge-heavy, factual queries, but it can misfire on queries where the hypothetical answer steers retrieval in the wrong direction. If the LLM's guess is confidently wrong, you'll retrieve documents about the wrong thing. Test it on your own query distribution before committing — RAGAS will tell you if it helped, more on that below.

Query Expansion Rewriting HyDE Multi-Query RAG Retrieval

Figure 1: Query expansion and rewriting — repair the input before retrieval runs

Image Alt Text: "Query expansion and rewriting techniques for better RAG retrieval with HyDE and multi-query"

Query Decomposition: Split Multi-Part Questions

Some questions are secretly several questions wearing a trench coat. "Compare our Pro and Enterprise plans and tell me which supports SSO" contains two retrieval targets: plan details and SSO support. One embedding for both queries means both halves get retrieved poorly.

The fix: decompose the query into sub-questions, retrieve for each, and combine:

decompose_prompt = ChatPromptTemplate.from_template(
    "Break this question into 2-4 simple standalone sub-questions.\n"
    "Question: {question}\nSub-questions:"
)

sub_questions = [
    q.strip()
    for q in (
        (decompose_prompt | llm | StrOutputParser())
        .invoke({"question": user_question})
        .split("\n")
    )
    if q.strip()
]

all_docs = []
for sq in sub_questions:
    all_docs.extend(base_retriever.invoke(sq))

Then deduplicate the combined results (by chunk ID, please — duplicates waste your context window) and hand the union to your generator. This pattern is the backbone of most agentic RAG systems you see demos of, even if nobody tells you the secret is just... more retrieval calls.

Step-Back Prompting: Zoom Out First

One more gem, especially for technical or reasoning-heavy queries. Step-back prompting generates a broader, more general question alongside the specific one, and retrieves for both.

Ask "why did my model's accuracy drop after I switched to AdamW?" and a step-back query might be "how do optimizer choices affect model training?" The specific question might miss; the general question pulls in the background context your generator needs to reason about the specific case.

You retrieve for both, concatenate, and generate. Simple, and surprisingly effective on the kinds of questions where the literal answer isn't actually in your documents — but the reasoning material is.

Technique What it fixes Extra cost Best for
Query rewriting Pronouns, vagueness, conversational debris 1 LLM call Chat / follow-up queries
Multi-query expansion Single-phrasing recall misses 1 LLM call + N retrievals Messy real-world phrasing
HyDE Query-document language gap 1 LLM call + embed Knowledge-heavy factual Qs
Decomposition Multi-part questions 1 LLM call + N retrievals Comparison / compound asks
Step-back Literal misses needing background 1 LLM call + 1 retrieval Technical reasoning queries

The Costs Nobody Mentions

Let's be adults about this, because every technique above has a price tag:

  • Latency — each expansion adds LLM calls before retrieval even starts; multi-query with 4 variations plus a rewrite means your simple search now runs 5 LLM calls upfront
  • Cost — those calls aren't free at scale, even with mini models
  • Over-retrieval — more query variations pull in more borderline chunks, which can dilute your context quality if you don't dedupe and cap

My honest production advice: layer these techniques, don't stack them all at once. Start with query rewriting for conversational bots (non-negotiable for chat). Add multi-query expansion when recall metrics look sad. Reach for HyDE and decomposition only when you have specific query types failing. A pipeline with rewriting and multi-query and HyDE and decomposition on every single query is a pipeline that bills like a rocket and responds like a glacier. :/

FYI, a cheap trick for taming latency: cache expanded queries. Users ask the same things repeatedly, and a simple LRU cache on rewritten queries skips the LLM call entirely on repeat hits.

Measuring the Wins

You already know where this is going, right? Same scoreboard, same rules: run your RAGAS baseline before adding query transformations, then re-run after. Here's what to watch:

  • Context recall — this is the headline metric for query expansion; more query angles should find documents the raw query missed
  • Context precision — watch for dips; sloppy expansion pulls in junk alongside the gold
  • Answer relevancy — better retrieval should produce better answers

That context recall number is your entire justification. If multi-query doesn't move recall on your data, you're paying extra LLM calls for nothing. If it jumps 15 points? Buy yourself a coffee and enjoy the win. Numbers beat vibes — you know this by now.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is query expansion in RAG?

Query expansion transforms the user's raw query before retrieval — rewriting vague phrasing, generating multiple paraphrases, or decomposing multi-part questions — so the retriever sees search-friendly input instead of conversational or mismatched language.

What does query rewriting fix?

LLM-based query rewriting resolves pronouns, fills in conversational context, and normalizes vocabulary so follow-ups like "what about the second one?" become standalone retrieval queries with actual searchable content.

How does HyDE improve retrieval?

HyDE (Hypothetical Document Embeddings) asks an LLM for a hypothetical answer, embeds that fake answer instead of the question, and retrieves real documents similar to it. Answer-shaped embeddings land closer to real answer text than question embeddings do.

When should I use multi-query expansion?

Use MultiQueryRetriever when context recall looks weak on messy real-world queries. It generates several query variations, retrieves for each, and merges results — usually one extra LLM call for a meaningful recall lift. Dedupe merged chunks before generation.

What is query decomposition for multi-part questions?

Decomposition splits compound questions into 2–4 standalone sub-questions, retrieves for each, then deduplicates and combines the results. It prevents one muddy embedding from representing two unrelated retrieval targets poorly.

Do query transformations improve RAGAS metrics?

Context recall is the headline metric — more query angles should find documents the raw query missed. Watch context precision for dips from sloppy expansion pulling in junk, and expect answer relevancy to climb when retrieval improves. Baseline before and after.

Wrapping Up

Let's recap the toolkit: query rewriting cleans up vague, pronoun-heavy queries. Multi-query expansion retrieves from several angles and merges the results. HyDE embeds a hallucinated answer to bridge the query-document language gap. Decomposition handles multi-part questions. Step-back prompting pulls in general context for reasoning-heavy queries. Each one attacks the same enemy: the gap between how users ask and how documents speak.

My parting take after way too many RAG pipelines: query transformation is the cheapest big win in retrieval. People obsess over embedders, vector stores, and exotic rerankers — and ignore the broken input at the front door. Fix the query first. It's often one prompt template and one extra LLM call standing between you and a dramatically better pipeline.

So go print out your top 20 real user queries this week. Compare them against your document language. If they don't match — and they won't — you now have five weapons to fix it. And if you catch yourself blaming the embedder for a vocabulary mismatch, well... no judgment. I wrote four tutorials of bad embedder-blaming before I figured it out. ;)

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles