Contents
Let's say you're searching a legal database for "can my boss read my Slack messages?" The relevant document probably says Employer monitoring rights extend to workplace communication platforms. Same meaning, completely different language — and your embedding model shrugs politely at the mismatch. HyDE (Hypothetical Document Embeddings) fixes this with one of the weirdest tricks in RAG: it asks an LLM to hallucinate a fake answer, then uses that hallucination as the search key. Yes, really. And yes, it works.
I'll admit my first reaction to HyDE was skepticism with a capital S. You're telling me the fix for hallucinating LLMs is... more hallucination? I grumbled at my screen. Then I tested it on a research-heavy corpus where my queries kept missing, and recall jumped noticeably. I've been a convert ever since, and today you'll see exactly why this counterintuitive trick punches so far above its weight.
The Problem HyDE Solves: The Query-Document Gap
To appreciate HyDE, you need to understand a core weakness of embedding-based retrieval. Queries and documents live in different linguistic worlds.
Think about it:
- Queries are short, interrogative, and telegraphic — why do prices rise?
- Documents are long, declarative, and information-dense — Inflation occurs when demand for goods outpaces supply, driving average prices upward across the economy.
Your embedding model maps both into the same vector space, but the two text styles cluster differently. A short question about a topic often sits closer to other short questions than to the actual answer. It's like asking a librarian your question and having them file you in the questions section instead of walking you to the bookshelf. :/
Why Similarity Alone Fails Here
Traditional retrieval embeds the question and hunts for documents near that question's vector. But if the question space and the answer space don't overlap cleanly, the nearest documents might be topically adjacent junk — documents about the topic that never actually answer the question.
The 2022 paper that introduced HyDE (Gao et al., Precise Zero-Shot Dense Retrieval without Relevance Labels) observed exactly this. Their insight: instead of closing the gap with better training, make the query look like a document. Then similarity compares apples to apples.
How HyDE Works, Step by Step
The pipeline runs in four moves:
- Generate a hypothetical answer — an LLM writes a fake, plausible answer to the query (no retrieval involved)
- Embed the fake answer — you embed that hallucinated text instead of the question
- Retrieve — you search the vector store for documents similar to the fake answer
- Generate the real answer — the actual retrieved documents feed your generator
The critical detail: you throw the fake answer away immediately. It never reaches your user. It's a retrieval key, nothing more — a disposable costume the query wears to sneak past the embedding model's bias.
Sound familiar? It should, if you read my query expansion article. I gave HyDE a quick cameo there; today it gets the full feature treatment it deserves.
Implementing HyDE with LangChain
LangChain ships HyDE in the experimental package. Install it first:
pip install langchain-experimental
Then build the hypothetical embedder:
from langchain_experimental.hyde import HypotheticalDocumentEmbedder
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
hyde = HypotheticalDocumentEmbedder.from_llm(
llm=llm,
base_embeddings=OpenAIEmbeddings(),
prompt_key="web_search",
)
fake_doc_vector = hyde.embed_query(
"How does contextual compression improve RAG?"
)
docs = vectorstore.similarity_search_by_vector(
fake_doc_vector, k=4
)
FYI, that prompt_key picks the persona for the hallucination. web_search writes the fake answer like a search snippet. Other keys generate answer styles for different domains — and honestly, for legal or medical corpora, a domain-specific prompt style matters a lot.
Writing a Custom HyDE Prompt
The default prompts work fine for general text, but custom prompts often beat them. Here's the pattern:
from langchain.prompts import PromptTemplate
custom_prompt = PromptTemplate(
input_variables=["question"],
template=(
"Write a short paragraph answering this question "
"in the style of a technical documentation page. "
"Use precise terminology.\n\n{question}\n\nAnswer:"
),
)
hyde_custom = HypotheticalDocumentEmbedder.from_llm(
llm=llm,
base_embeddings=OpenAIEmbeddings(),
custom_prompt=custom_prompt,
)
Match the fake answer's style to your real documents. If your corpus reads like API docs, hallucinate API-doc-style answers. If it reads like academic papers, hallucinate citations and formal prose. The closer the fake lands to the real documents' linguistic neighborhood, the better retrieval performs. I learned this after a HyDE setup with a chatty, blog-style prompt flailed against a corpus of dry compliance documents — the hallucinations were great reading, useless for retrieval.
Figure 1: HyDE — hallucinate an answer-shaped key, retrieve real documents, discard the fake
Image Alt Text: "HyDE Hypothetical Document Embeddings for RAG closing the query-document gap"
The LlamaIndex Route
Prefer LlamaIndex? It offers HyDE as a query transform, which slots neatly into any query engine:
from llama_index.core.indices.query.query_transform import (
HyDEQueryTransform,
)
from llama_index.llms.openai import OpenAI
from llama_index.core.query_engine import TransformQueryEngine
hyde_transform = HyDEQueryTransform(
llm=OpenAI(model="gpt-4o-mini"),
include_original=True,
)
hyde_query_engine = TransformQueryEngine(
query_engine=base_query_engine,
query_transform=hyde_transform,
)
response = hyde_query_engine.query(
"What is the refund policy?"
)
That include_original=True flag deserves a shoutout: it embeds both the original query and the hypothetical answer, then combines the vectors. I usually enable it — you hedge against the hallucination steering retrieval somewhere silly, while keeping the hallucination's vocabulary-matching superpowers.
When HyDE Shines (and When It Faceplants)
HyDE works brilliantly for:
- Knowledge-heavy, factual queries — how does X work, what causes Y — where answers follow predictable patterns
- Vocabulary-heavy domains — legal, medical, technical — where users speak layman and documents speak jargon
- Zero-shot setups — no training data, no fine-tuning, just a prompt and an embedding model
And it faceplants on:
- Queries where the LLM knows nothing about the topic — a confident hallucination in the wrong direction retrieves wrong-topic documents enthusiastically
- Simple keyword-style queries — error code E-4012 needs literal matching, not creative writing; the fake answer might bury the code in a paragraph
- Ultra-latency-sensitive apps — you've added an LLM generation step before retrieval even starts
IMO, the biggest misconception about HyDE is that it's magic. It's not. It's a conditional technique: brilliant on the queries above, wasteful or harmful on the rest. That's why I never run HyDE unconditionally in production — I gate it by query type, and I'll show you how to check whether yours deserves it.
Measuring HyDE's Real Impact
You already know my rule by now: numbers beat vibes. Before you ship HyDE, run your RAGAS baseline, then re-run with HyDE enabled and compare:
- Context recall — the headline number; HyDE should find documents the raw query embedding missed
- Context precision — watch for dips; a badly-prompted hallucination pulls in topically-wrong neighbors
- Answer relevancy — better retrieval should lift the final answers
One practical tip from my eval logs: bucket your eval set by query type before comparing. HyDE might tank recall 10 points on error code queries while lifting it 20 points on how does it work queries. A single blended average hides both truths and sends you the wrong signal. Split the buckets, see the real story, then gate HyDE accordingly.
HyDE vs. Multi-Query: Which One?
People ask me this constantly, so here's my honest comparison:
| Approach | What it generates | Mismatch it fixes |
|---|---|---|
| Multi-query expansion | Several question phrasings | How the question could have been asked |
| HyDE | One fake answer | Question-space to answer-space gap |
Rule of thumb: multi-query fixes how the question could have been asked; HyDE fixes what the answer probably looks like. They're solving different halves of the mismatch, and you can even combine them. But start with one, measure, and add the second only if your recall still has gaps. Stacking both on every query means paying two LLM round-trips before retrieval, and your latency dashboard will file a formal complaint. :/
Recommended Books
- Retrieval-Augmented Generation (RAG) — pipeline context for where a query transform like HyDE sits before vector search.
- AI-Powered Search by Trey Grainger et al. — query understanding foundations that frame the query-document language gap HyDE addresses.
- Natural Language Processing with Transformers by Tunstall et al. — how LLMs generate the hypothetical answers and how embeddings place them in vector space.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is HyDE in RAG?
HyDE (Hypothetical Document Embeddings) asks an LLM to generate a fake answer to the query, embeds that hypothetical document instead of the question, retrieves real documents similar to it, then discards the fake answer. It bridges the query-document language gap before retrieval.
Why does HyDE work?
Questions and answers sit in different parts of embedding space — short interrogative queries cluster near other queries, not near declarative answer text. Fake answer embeddings land closer to real answer embeddings, so similarity compares answer-shaped text to answer-shaped text.
How do I implement HyDE with LangChain?
Install langchain-experimental, then use HypotheticalDocumentEmbedder.from_llm with an LLM and base embeddings. Call embed_query on your question and pass the resulting vector to similarity_search_by_vector — the hypothetical text never reaches the user.
When does HyDE fail?
HyDE underperforms when the LLM knows little about the topic (confidently wrong hallucinations retrieve wrong documents), on literal keyword-style queries like error codes, and in ultra-latency-sensitive apps where an extra LLM generation step hurts.
HyDE vs multi-query expansion — what's the difference?
Multi-query generates several question phrasings and retrieves for each, covering how the question could have been asked. HyDE generates a fake answer, fixing the question-space to answer-space mismatch. They solve different halves and can be combined if recall still lags.
How do I measure whether HyDE helps?
Run a RAGAS baseline before enabling HyDE, then re-run and compare. Context recall is the headline metric. Bucket eval samples by query type — HyDE may hurt error-code lookups while lifting "how does it work" questions — so blended averages can hide the real story.
Wrapping Up
Let's recap: HyDE hallucinates a hypothetical answer, embeds it instead of the question, retrieves with it, and then discards it. The trick works because fake answers live in the same linguistic neighborhood as real answers, closing the query-document gap that plain embedding similarity can't. Use custom prompts that match your corpus style, keep include_original=True as a safety net, gate HyDE to knowledge-heavy query types, and let RAGAS bucket-by-bucket analysis decide whether it stays.
My parting thought? I love HyDE because it captures something profound about RAG: your system's weakness isn't always retrieval quality — sometimes it's the shape of what you feed retrieval. HyDE reshapes the input, and the whole pipeline improves without touching your vector store, your embedder, or your chunking. One prompt, one LLM call, done.
So go hallucinate something. Your retriever will thank you — and the beauty of it is, nobody ever has to read what the LLM dreamed up. It's the one hallucination in your pipeline that's actually working for you. ;)