Sam Austin AI

Hybrid Search: Combining Keyword and Vector Search

September 1, 2026 12 min read Updated September 2, 2026 Sam Austin
Contents

Pure vector search has a genuinely embarrassing blind spot: ask it for an exact product code or error string, and it'll confidently hand you something "semantically close" instead of the thing you actually typed. That's not a minor edge case — that's the moment your users stop trusting the search bar entirely.

I ran into this exact failure mode while working with structured technical content, where exact identifiers matter just as much as conceptual meaning. Ever wondered why your RAG chatbot occasionally returns a confidently wrong product? This is usually the culprit, and hybrid search is the fix nobody mentions until it's too late.

By the end of this guide, you'll understand exactly why keyword and vector search need each other, how they actually get combined, and why the fusion method matters more than most tutorials let on. IMO, this is the single highest-impact upgrade you can make to a pure-vector RAG system :)

Hybrid Search Combining Keyword and Vector Search
Hybrid Search Combining Keyword and Vector Search

Figure 1: Hybrid search combines BM25 keyword matching with vector semantic search for better retrieval

Why Neither Approach Wins Alone

Here's the honest truth: neither keyword nor vector search dominates in every case. As Qdrant's engineering team has put it, keyword-based search wins sometimes, vector search wins other times, and pretending one universally beats the other misreads the actual problem.

  • BM25 (keyword search) nails exact terms — product codes, error strings, identifiers, names. It fails on paraphrase, synonymy, and natural-language questions where the user's words don't match the corpus.
  • Dense vector search captures meaning and handles paraphrasing beautifully. It stumbles badly on rare, specific tokens — an embedding model has no special treatment for something like a specific protocol identifier, and will happily return generic related content instead of the one document that mentions it exactly.

A customer types a specific SKU, and the bot describes a different, semantically-related product instead. That's the exact failure hybrid search exists to prevent.

What Hybrid Search Actually Does

Hybrid search runs both retrieval methods side by side — BM25 for lexical matching, dense vectors for semantic matching — and merges the two ranked lists into one final result set.

  • BM25 returns documents ranked by term-frequency scoring.
  • The vector index returns documents ranked by embedding similarity.
  • These produce two independent ranked lists that may overlap heavily or barely at all, depending on the query.

The genuinely clever part isn't running both searches — that's easy. The hard part is combining two fundamentally incompatible scoring systems into one meaningful ranking, and that's where most naive implementations quietly break.

Why You Can't Just Average the Scores

Here's the gotcha that trips up almost everyone building this for the first time. BM25 produces unbounded scores driven by term statistics — they can be any positive number. Dense vector search produces cosine similarities bounded between -1 and 1.

If you combine these with a simple weighted formula, BM25 will always dominate, because its scores exist on a completely different numeric scale. Without explicit normalization, that "weighted blend" you thought you built isn't actually blending anything meaningfully — it's just BM25 wearing a costume.

Reciprocal Rank Fusion (RRF): The Fix

The solution that's become the industry standard is Reciprocal Rank Fusion, and it sidesteps the scaling problem entirely by ignoring raw scores altogether.

score(doc) = sum over each retriever of: 1 / (k + rank_of_doc_in_that_list)

The constant k (commonly 60) prevents top-ranked documents from dominating too aggressively. A document ranked #1 in both BM25 and vector search gets a substantially higher combined score than one ranked #1 in just a single method — and a document that shows up in only one list still gets counted, just with less weight.

RRF works because it only cares about rank position, not the incompatible raw scores. That single design choice is what makes hybrid search actually reliable in production instead of a fragile normalization hack.

Seeing It Work: A Concrete Example

Let's ground this in an actual scenario. Say your documentation contains a rare technical identifier like a specific protocol string.

  • BM25 matches it exactly, since it's a distinctive term appearing in just one or two documents.
  • Vector search may miss it entirely — the embedding model has no special understanding of that identifier and produces a vector semantically closer to generic networking content instead.

With hybrid search and RRF fusion, both documents surface correctly. The conceptual match gets boosted by its strong vector-search rank; the exact-match document gets boosted by its strong BM25 rank. Neither method's individual failure mode dominates the final result. In one evaluated documentation system, this combination achieved a Precision@5 of 0.84, compared to just 0.71 for pure vector search alone — a meaningful, measurable gap.

The Numbers Behind the Hype

I'm generally skeptical of vendor benchmarks, but a few independently reported figures here are worth taking seriously.

  • On an e-commerce benchmark, a tuned hybrid setup reached 0.7497 NDCG — roughly a 7.4% lift over either BM25 or pure vector search running alone.
  • On financial documents mixing text and tables, BM25 alone actually outperformed dense retrieval on every metric except Recall@20 — a genuinely surprising result that challenges the assumption semantic search always wins.
  • Adding reranking on top of hybrid fusion pushed Recall@5 to 0.816, a +17.4% relative improvement over RRF fusion alone, and nearly 39% over dense-only retrieval.

That last stat matters most: fusion alone gets you most of the way, but pairing it with reranking is where the biggest additional gains live.

Setting Up Hybrid Search: The Pipeline

Here's roughly what a production hybrid retrieval pipeline looks like end to end.

  1. Query comes in (optionally expanded or rewritten).
  2. BM25 search returns its top 20 candidates.
  3. Vector search returns its top 20 candidates, independently.
  4. RRF merges both lists into a single ranked set, typically the top 30 unique documents.
  5. A reranker (often a cross-encoder) refines that shortlist down to the final top 5–10.
  6. Those results get passed to the LLM as context.

The mental model worth internalizing: BM25 and vector search are complementary first-stage retrievers, RRF is a fusion algorithm operating purely on ranks, and reranking is a second-stage precision layer applied to a shortlist — not the full index. Conflating these stages is the architectural mistake that trips up most naive implementations.

Tuning the Balance

If you're using a weighted alpha approach instead of pure RRF (some platforms support this), the weighting genuinely depends on your query mix:

  • Weight toward keyword/BM25 (0.8+) for exact lookups — SKUs, error codes, specific identifiers.
  • Weight toward vector search (0.8+) for conceptual, exploratory, or paraphrased queries.
  • True 50/50 works reasonably well for genuinely mixed query intent, which describes most real-world search traffic.

My honest take: don't guess this number — test it against your actual query logs. A common starting point is 0.7 vector / 0.3 keyword, but that's a starting point, not a rule.

Where to Implement This

The good news: you don't need to build RRF from scratch. Several platforms now support it natively.

  • Qdrant, Elasticsearch, OpenSearch, and Weaviate all offer built-in hybrid search support.
  • PostgreSQL users can combine pg_search for BM25 with pgvector for dense retrieval, staying entirely within Postgres.
  • Weaviate is often cited as the fastest path to hybrid search if you're starting fresh, since it's a first-class feature rather than something bolted on.

If you're already running Postgres for other parts of your stack, staying there instead of adding a separate search engine is a genuinely reasonable choice — one less system to operate, monitor, and pay for.

Nothing's free, so let's be upfront about the tradeoffs. Production hybrid search adds roughly 6ms of latency and about 1.4x the storage footprint compared to vector-only search, since you're maintaining two indexes instead of one.

That's a genuinely small price for the accuracy gains involved. In most production RAG systems, a few extra milliseconds is invisible to users, while returning the wrong product because of a missed SKU match absolutely is not.

Common Mistakes People Make

I've seen these repeated across enough implementations to call them patterns.

  • Averaging BM25 and cosine similarity scores directly. This silently lets BM25 dominate every ranking due to scale mismatch — always normalize or use rank-based fusion instead.
  • Skipping reranking entirely. Fusion alone gets you most of the way, but the largest accuracy gains often come from pairing RRF with a reranking step.
  • Assuming vector search always wins on "smart" queries. Some benchmarks show BM25 outperforming dense retrieval even on modern embedding models, especially on documents with mixed text and tabular data.
  • Retrieving too few candidates before fusion. RRF needs a reasonably sized candidate pool from each retriever (commonly top 20 each) to actually produce a meaningful merged ranking.

Frequently Asked Questions

Hybrid search combines BM25 keyword matching with dense vector semantic search. It runs both methods simultaneously and merges results using fusion algorithms like Reciprocal Rank Fusion (RRF) for better accuracy.

Why is hybrid search better than vector search alone?

Vector search fails on exact terms like product codes, error strings, and identifiers. Hybrid search adds BM25 keyword matching to catch these exact matches while retaining semantic understanding for conceptual queries.

What is Reciprocal Rank Fusion (RRF)?

RRF combines ranked results from multiple search methods by using rank positions instead of raw scores. It solves the problem of incompatible score scales between BM25 and vector search. The formula is: score = sum(1/(k+rank)) across retrievers.

How much does hybrid search improve retrieval?

Hybrid search typically improves Precision@5 from 0.71 to 0.84 over pure vector search. Adding reranking on top pushes Recall@5 to 0.816, a 39% improvement over dense-only retrieval.

Qdrant, Elasticsearch, OpenSearch, Weaviate, and PostgreSQL (with pg_search + pgvector) all support hybrid search natively or through extensions.

RRF is generally preferred because it avoids score normalization issues. Weighted scoring requires careful tuning and can let BM25 dominate due to scale differences. RRF works reliably without manual weight tuning.

Wrapping This Up

Hybrid search exists because keyword and vector retrieval fail in genuinely different, complementary ways: BM25 nails exact terms and stumbles on meaning; vector search nails meaning and stumbles on exact terms. Reciprocal Rank Fusion combines them without falling into the score-normalization trap that breaks naive implementations.

Remember that fusion alone gets you most of the accuracy gain, but pairing it with a reranking step is where the largest remaining improvements live. FYI, if your current RAG system runs pure vector search, adding BM25 through RRF is genuinely one of the highest-impact, lowest-effort upgrades available to you right now :)

Now go test this against your own query logs instead of trusting a single alpha-weighting default — your actual mix of exact-match versus conceptual queries is the only benchmark that really matters here.

Share this article X Facebook LinkedIn Reddit WhatsApp