Sam Austin AI

Contextual Compression in RAG: Retrieve Less, Answer Better

September 24, 2026 12 min read Sam Austin
Contents

Here's a dirty secret about RAG pipelines: your retriever is probably stuffing your LLM's context window with garbage. You retrieve the top 5 chunks, three of them contain one useful sentence buried in filler, and your poor generator has to dig through all of it like a raccoon searching a dumpster. Contextual compression fixes this — it strips the junk before it reaches your LLM. The result? Better answers, lower costs, and fewer tokens burned on nonsense.

I discovered this the painful way, of course. My first production RAG bot retrieved chunks beautifully but answered poorly, and I blamed everything — the embedder, the prompt, the moon phases. Then I printed out the actual context my LLM received and winced. Half of it was boilerplate and irrelevant paragraphs. The answer was sitting right there; my generator just couldn't see the forest for the trees. :/

What Is Contextual Compression?

Let's define the term, because compression sounds like we're zipping files or something. We're not.

Contextual compression means taking your retrieved documents and extracting only the parts relevant to the query — then passing just those parts to your LLM. The retrieval step stays the same; you simply add a cleanup pass between retrieval and generation.

How It Works

Think of it as a two-layer system:

  • Base retriever — finds candidate documents quickly (your usual vector search)
  • Document compressor — re-reads those documents against the query and keeps only what matters

The base retriever does the broad sweep; the compressor does the fine cut. Together they hand your generator a distilled, focused context instead of a content dump.

Why bother? Because retrieval and relevance aren't the same thing. Vector similarity finds documents that are topically close, but a 500-word chunk might contain exactly one sentence that matters. Feeding all 500 words to your LLM wastes tokens, dilutes attention, and invites the model to latch onto the wrong detail. Ever had a bot answer a question the documents didn't really answer — confidently? That's diluted context at work.

Why Naive RAG Struggles

Let's name the enemies so we know what we're fighting:

  • Noise pollution — chunks contain relevant and irrelevant content, and your LLM can't always tell which is which
  • Context window pressure — more retrieved chunks means more tokens, means higher cost and slower responses
  • The lost-in-the-middle problem — LLMs pay less attention to information buried mid-context
  • Top-k whack-a-mole — increase top_k for better recall, and precision tanks; decrease it, and you miss things. Fun game, right?

IMO, most people treat top-k like a magic dial and hope. Compression lets you keep recall high and precision high, because the compressor does the precision work that a single similarity score can't.

Getting Set Up

Let's build this with LangChain, which ships first-class contextual compression support. If you've followed my earlier tutorials, you already know I like tools that make my life easier rather than harder.

pip install langchain langchain-openai chromadb

FYI, this pattern exists in other frameworks too — LlamaIndex has a very similar ContextualRetrieval style approach — but I'll stick with LangChain since it exposes the machinery most clearly. And if you built the pipeline in the Haystack RAG tutorial, the same compressor-between-retrieval-and-generation idea maps over directly.

The Standard Retrieval Baseline

First, let's appreciate the before picture:

from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import Chroma

vectorstore = Chroma(
    embedding_function=OpenAIEmbeddings(),
    persist_directory="./my_store"
)

retriever = vectorstore.as_retriever(search_kwargs={"k": 5})
docs = retriever.invoke("What is our refund policy?")

This grabs the 5 most similar chunks, full stop. If each chunk runs 400 words, you just fed your LLM 2,000 words to answer a question whose answer occupies 30 of them. That's like photocopying the entire encyclopedia because someone asked about penguins.

Adding the LLM Chain Extractor

Now for the upgrade. LangChain's LLMChainExtractor uses an LLM to rewrite each retrieved document, keeping only the query-relevant parts:

from langchain.retrievers import ContextualCompressionRetriever
from langchain.retrievers.document_compressors import LLMChainExtractor
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
compressor = LLMChainExtractor.from_llm(llm)

compression_retriever = ContextualCompressionRetriever(
    base_retriever=retriever,
    base_compressor=compressor
)

compressed_docs = compression_retriever.invoke(
    "What is our refund policy?"
)

Now your generator receives a handful of tight, focused snippets instead of five rambling chunks. The query stays in the loop — that's the contextual part. The extractor reads each document in light of your specific question, so the same chunk compresses differently for different queries. Clever, right? :)

Contextual Compression RAG LangChain Retrieve Less Answer Better

Figure 1: Contextual compression — strip noise between retrieval and generation

Image Alt Text: "Contextual compression in RAG with LangChain LLMChainExtractor and embeddings filter"

The Trade-Off (Because There's Always One)

Honesty time: LLMChainExtractor runs an LLM call per retrieved document. That adds latency and cost to every query. For a hobby project, whatever. For a high-traffic production bot, you'll notice.

My rule of thumb: start with LLM-based extraction when quality matters most, then switch to cheaper compressors once you're scaling. Speaking of which…

Cheaper Compression Options

Not every compressor needs an LLM. Pick your fighter:

  • LLMChainFilter — an LLM decides yes or no per document, keeping relevant ones untouched. One cheap binary call instead of full rewrites.
  • EmbeddingsFilter — computes embedding similarity between the query and each chunk, drops chunks under a threshold. No LLM calls, nearly instant, costs pennies.
  • DocumentCompressorPipeline — chains several compressors together, so you can filter by embeddings first, then extract with an LLM on the survivors.

Here's that pipeline pattern, which IMO is the sweet spot for production:

from langchain.retrievers.document_compressors import (
    DocumentCompressorPipeline,
    EmbeddingsFilter,
)
from langchain_openai import OpenAIEmbeddings

embeddings_filter = EmbeddingsFilter(
    embeddings=OpenAIEmbeddings(),
    similarity_threshold=0.8
)

pipeline_compressor = DocumentCompressorPipeline(
    compressors=[embeddings_filter, compressor]
)

smart_retriever = ContextualCompressionRetriever(
    base_retriever=retriever,
    base_compressor=pipeline_compressor
)

The flow reads like a bouncer system: the embeddings filter throws out the obvious misfits instantly and for free, then the LLM extractor gives the remaining documents a careful read. Cheap filter first, expensive brain second. This is how you get quality and a bill that doesn't require a support ticket to accounting.

Compressor Cost Latency What it does
EmbeddingsFilter Pennies Near-instant Drops whole chunks under a similarity threshold
LLMChainFilter Low One binary LLM call/doc Keeps or discards each document
LLMChainExtractor Higher Full rewrite per doc Extracts only query-relevant sentences
CrossEncoderReranker Medium Pairwise scoring Joint query-doc relevance, smart ordering + culling

Adding a Reranker to the Mix

Want the full premium setup? Drop a cross-encoder reranker into your compressor pipeline. A reranker scores each query-document pair jointly, which catches relevance that embedding similarity misses:

from langchain_community.cross_encoders import HuggingFaceCrossEncoder
from langchain.retrievers.document_compressors import (
    CrossEncoderReranker,
    DocumentCompressorPipeline,
)

reranker = CrossEncoderReranker(
    model=HuggingFaceCrossEncoder(
        model_name="BAAI/bge-reranker-base"
    ),
    top_n=3
)

final_compressor = DocumentCompressorPipeline(
    compressors=[embeddings_filter, reranker]
)

Rerankers cost more than embedding similarity but far less than LLM extraction, and they boost precision like crazy. I've seen reranking alone turn a mediocre retriever into a good one — same embedder, same chunks, just smarter ordering and culling.

Measuring Whether It Actually Helped

Remember our RAGAS conversation? This is exactly where that scoreboard earns its keep. Compression changes your retrieved context, so it directly moves those retrieval metrics:

  • Context precision should jump — compression literally exists to raise this number
  • Faithfulness often improves — less noise means fewer chances for the generator to grab the wrong detail
  • Answer relevancy typically climbs too — focused context makes focused answers

Run your RAGAS baseline before adding compression, then re-run after. If context precision doesn't move, your compressor isn't earning its keep — tune that threshold or swap strategies. Numbers beat vibes, every time. :D

When Should You Skip Compression?

Let's stay honest, because compression isn't free lunch:

  • Tiny, clean chunks — if your chunking is already tight, compression adds latency for little gain
  • Ultra-low-latency apps — every compressor adds a hop; the embeddings filter is the only nearly-free option
  • Already-broken retrieval — if your base retriever never finds the right document, compression can't rescue what wasn't retrieved. Fix recall first (RAGAS context recall will tell you), then compress.

Compression polishes good retrieval; it doesn't resurrect bad retrieval. Diagnose before you optimize — I learned that one the slow way, and you shouldn't have to.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is contextual compression in RAG?

Contextual compression extracts only the query-relevant parts of retrieved documents before they reach the LLM. A base retriever finds candidates as usual; a document compressor re-reads those chunks against the query and passes only the distilled snippets to generation.

How does LangChain implement contextual compression?

LangChain's ContextualCompressionRetriever wraps a base retriever with a base_compressor such as LLMChainExtractor (LLM rewrites each doc keeping query-relevant parts), EmbeddingsFilter (drops low-similarity chunks), or a DocumentCompressorPipeline combining filters and rerankers.

What is the difference between EmbeddingsFilter and LLMChainExtractor?

EmbeddingsFilter drops entire chunks below an embedding similarity threshold — no LLM calls, nearly instant, cheap. LLMChainExtractor rewrites each document with an LLM to keep only query-relevant sentences — higher quality but one model call per retrieved document, adding latency and cost.

Does contextual compression improve RAGAS metrics?

It targets retrieval quality: context precision should rise because noise is stripped, and faithfulness and answer relevancy often improve as the generator sees less irrelevant material. Run a RAGAS baseline before and after; if context precision doesn't move, tune the compressor.

When should you skip contextual compression?

Skip it when chunks are already tight and clean, when you need ultra-low latency and only the embeddings filter would fit, or when base retrieval itself is broken — compression cannot rescue documents the retriever never found. Fix context recall first.

What does a cross-encoder reranker add to compression?

A reranker like BAAI/bge-reranker-base scores each query-document pair jointly, catching relevance embedding similarity misses. It costs more than embedding filtering but far less than LLM extraction, and often boosts precision dramatically with the same embedder and chunks.

Wrapping Up

Here's the recap: contextual compression sits between retrieval and generation, stripping irrelevant content so your LLM receives only the good stuff. Start with LLMChainExtractor for maximum quality, graduate to a DocumentCompressorPipeline with EmbeddingsFilter plus a reranker for production economics, and verify every change with RAGAS metrics — especially context precision.

The core insight is worth tattooing somewhere: retrieval finds documents, but compression finds the parts that matter. Your generator doesn't need more context — it needs better context. And as a bonus, your token bill shrinks, your latency drops, and your bot stops quoting boilerplate at confused users.

So go print out what your pipeline currently feeds the LLM. I promise it'll be an educational experience — mine certainly was. ;) Then add a compressor, rerun your evals, and enjoy the scoreboard going up. Retrieve less, answer better — it really is that simple once you stop feeding your LLM the whole dumpster.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles