Sam Austin AI

Elasticsearch for Vector Search: kNN and Dense Retrieval

September 23, 2026 14 min read Sam Austin
Contents

Recall the hybrid search article from much earlier in this series — its central argument was that BM25 and vector search fail in opposite, complementary ways, and that Reciprocal Rank Fusion is how you combine them without falling into the score-normalization trap. Every pure-play vector database covered in this series' last four tutorials — Pinecone, ChromaDB, Qdrant, Milvus — added hybrid search as a feature bolted onto a vector-first foundation. Elasticsearch inverts that history entirely: it's been the dominant full-text search engine for over a decade, and vector search is what got added to it. That inversion matters practically, not just historically.

The genuinely distinctive capability worth understanding upfront: Elasticsearch can run BM25, dense vector kNN, and its own learned sparse-vector model together in a single request, fused server-side via RRF, with no client-side merging required at all. Recall the Qdrant tutorial's Fusion.RRF query type from earlier in this series — Elasticsearch does the same thing, but with a third retriever most pure vector databases don't have any equivalent for at all: ELSER, a genuinely different approach to semantic matching that's neither classic BM25 nor a dense embedding.

By the end of this guide, you'll understand dense vector search, ELSER's sparse-vector alternative, and the retriever framework that fuses all three approaches in one request. IMO, "no client-side merging" is genuinely the single sentence that separates Elasticsearch's hybrid search story from the earlier tutorials' patchwork implementations :)

Elasticsearch Vector Search kNN Dense Retrieval ELSER Hybrid

Figure 1: Elasticsearch vector search — dense kNN, ELSER sparse retrieval, and server-side RRF fusion

Image Alt Text: "Elasticsearch kNN vector search and ELSER dense retrieval tutorial for hybrid search"

Dense Vector Search: kNN on dense_vector Fields

Elasticsearch has supported dense vectors since version 7.3, but native, approximate kNN search arrived in 8.0 — worth knowing directly, since anything built before that version needs reindexing with "index": true in its mapping to actually use approximate kNN.

PUT /documents
{
  "mappings": {
    "properties": {
      "content_vector": {
        "type": "dense_vector",
        "dims": 384,
        "index": true,
        "similarity": "cosine"
      },
      "text": { "type": "text" },
      "category": { "type": "keyword" }
    }
  }
}

Notice text staying a standard text field alongside the vector field — recall this being the entire architectural point; the same document carries both its BM25-searchable text and its dense vector representation, in one index, rather than needing a separate keyword-search system synced against a separate vector database.

from sentence_transformers import SentenceTransformer
from elasticsearch import Elasticsearch

model = SentenceTransformer("all-MiniLM-L6-v2")
es = Elasticsearch("http://localhost:9200")

def search(query_text, k=10):
    query_vector = model.encode(query_text).tolist()
    response = es.search(
        index="documents",
        knn={
            "field": "content_vector",
            "query_vector": query_vector,
            "k": k,
            "num_candidates": 100
        }
    )
    return response["hits"]["hits"]

Recall all-MiniLM-L6-v2 directly from the Sentence Transformers article earlier in this series — the same embedding model, plugged into yet another vector store's ingestion pipeline. That num_candidates parameter matters more than it looks: it controls how many candidates each shard's HNSW graph considers before returning the top-k — a genuinely direct analog to the oversampling concept from the hybrid search article's "retrieve top-20-per-method before fusing" guidance, just applied within a single retriever rather than across two.

Pre-Filtering: Applied Before the kNN Search Runs

Recall the Qdrant and Pinecone tutorials' payload/metadata filtering from earlier in this series — Elasticsearch's filter clause works the same way, with one genuinely important mechanical detail worth understanding.

GET /documents/_search
{
  "knn": {
    "field": "content_vector",
    "query_vector": [0.1, 0.2, 0.3],
    "k": 10,
    "num_candidates": 100,
    "filter": {
      "bool": {
        "must": [
          { "term": { "category": "technical" } },
          { "range": { "created_at": { "gte": "2025-01-01" } } }
        ]
      }
    }
  }
}

Filters are applied before the kNN search itself, genuinely reducing the candidate pool the HNSW graph has to search through — this is meaningfully different from a naive "search everything, then filter results" approach, and it's exactly why filtered queries against a large, mostly-irrelevant-to-your-filter dataset perform well rather than searching the full index and discarding most results afterward.

ELSER: The Genuinely Distinct Third Option

This is the capability without a real equivalent in the Pinecone, ChromaDB, Qdrant, or Milvus tutorials from earlier in this series. ELSER (Elastic Learned Sparse EncodeR) performs semantic search using sparse vector representations — arrays where most elements are zero, and the small number of non-zero dimensions each correspond to a specific term or concept, rather than the dense, uninterpretable 384-or-768-dimensional vectors every other tutorial in this series has used.

  • Built directly into Elasticsearch, requiring no external model deployment — genuinely different from every dense embedding model covered in this series, which required you to run or call out to a separate model.
  • Out-of-domain and requiring no fine-tuning — a real, practical advantage for teams without the data science resources the embedding models comparison flagged as a genuine barrier to custom model training.
  • Results are more explainable than dense vectors, since you can genuinely see which specific terms contributed to a match — recall this being a real, distinct advantage over dense embeddings' black-box similarity scores, worth knowing about specifically for use cases where explainability matters to end users or auditors.
  • Currently English-only, and — worth being direct about this real cost — ELSER requires a Platinum or Enterprise license, or Elastic Cloud — the retriever framework itself is available across all tiers, but this specific model sits behind a paywall, a genuinely important detail to check before building an architecture around it.

The Retriever Framework: Fusing Three Approaches in One Request

This is genuinely the payoff of everything above — Elastic's own recommended high-recall configuration combines BM25, ELSER, and dense kNN together as three child retrievers under one rrf retriever, all resolved server-side in a single request.

GET /documents/_search
{
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": { "match": { "text": "optimizing endpoint latency" } }
          }
        },
        {
          "knn": {
            "field": "content_vector",
            "query_vector": [0.1, 0.2, 0.3],
            "k": 10,
            "num_candidates": 100
          }
        },
        {
          "standard": {
            "query": { "sparse_vector": { "field": "ml_tokens", "inference_id": "elser-endpoint", "query": "optimizing endpoint latency" } }
          }
        }
      ],
      "rank_window_size": 50
    }
  }
}

Recall the exact RRF formula from the hybrid search article directly — score(doc) = sum over each retriever of 1/(k + rank) — this is that identical algorithm, now fusing three ranked lists instead of two, entirely inside Elasticsearch's own query planner. No client-side merging, no separate score-normalization logic you'd have to hand-write — recall this being precisely the trap the hybrid search article warned about (BM25 and cosine similarity living on incompatible numeric scales); RRF's rank-based approach sidesteps that problem here exactly as it did in that article's original explanation.

semantic_text: The Simplified Path

Worth knowing about directly for anyone who wants the hybrid search benefit without hand-building retriever queries: Elasticsearch's semantic_text field type lets you query with a simple match query for the simplest approach, or drop down to the explicit knn query when you need finer control.

PUT /documents
{
  "mappings": {
    "properties": {
      "text": { "type": "semantic_text" }
    }
  }
}

This genuinely abstracts away the embedding-model plumbing entirely — Elasticsearch handles vectorization and indexing automatically behind this field type, closer in spirit to Pinecone's integrated embedding feature or Qdrant's FastEmbed convenience layer from the earlier tutorials in this series, just wrapped inside a mapping type rather than a client method call.

GPU-Accelerated Indexing: Building HNSW Graphs Faster

Recall the model compression and GPU-buying articles from earlier in this series' local LLM arc — the same "GPU acceleration matters at real scale" principle applies here directly. Building HNSW graphs is compute-intensive, and Elasticsearch specifically supports GPU-accelerated vector indexing on nodes with compatible NVIDIA GPUs to speed up that construction step — worth knowing about specifically once your indexing time (not just query time) becomes a genuine bottleneck at scale, echoing exactly the same NVIDIA-hardware discussion from this series' GPU and Jetson articles.

Where Elasticsearch Fits Against the Four Prior Vector Database Tutorials

ChromaDB Pinecone Qdrant Milvus Elasticsearch
Primary identity Vector-first, lightweight Vector-first, managed Vector-first, self-hosted Vector-first, billion-scale Full-text search, vectors added
Native BM25 No No No No Yes — its original core
Native sparse semantic model No No No No Yes (ELSER)
Single-request 3-way fusion No No Partial (2-way RRF) Partial Yes (BM25 + ELSER + dense, one retriever)
Existing deployment reuse N/A N/A N/A N/A Genuine — if you already run Elastic Stack for logs/observability

Recall the Drupal AI module case study directly from current sources: teams already running Elasticsearch for logging, observability, or existing full-text search get RAG and semantic search "for free" against infrastructure they already operate — no separate vector database to stand up, patch, or monitor, genuinely the strongest practical case for choosing Elasticsearch specifically when that condition already holds.

A Practical Decision Framework

  1. Do you already run an Elastic Stack for logs, observability, or existing search? This is genuinely the strongest signal to add vector search here rather than standing up a dedicated vector database alongside it.
  2. Do you need explainable semantic matching, not just black-box dense similarity? Recall ELSER's sparse-vector explainability directly — a genuine, distinct advantage worth checking against your actual requirements.
  3. Is single-request, server-side three-way fusion (BM25 + sparse + dense) genuinely valuable for your use case? Recall this being Elastic's own recommended high-recall configuration — if your queries mix exact-term lookups with paraphrased natural language, this is a real, evidenced improvement over two-way fusion alone.
  4. Can you accept ELSER's licensing requirement (Platinum/Enterprise or Elastic Cloud)? Confirm this cost directly before building an architecture that assumes ELSER's availability.
  5. Is your dataset genuinely vector-only, with no meaningful keyword-search need? If yes, recall the earlier four tutorials directly — a purpose-built vector database (Qdrant for filtering, Milvus for billion-scale) may be the simpler, more focused choice.

Common Mistakes People Make

  • Using a pre-8.0 index with dense_vector fields and expecting approximate kNN to work. Recall this directly — reindex with "index": true first; older mappings don't support this feature.
  • Assuming ELSER is included in every Elasticsearch tier. Recall the explicit licensing requirement directly — the retriever framework itself is broadly available, but ELSER specifically needs Platinum/Enterprise or Elastic Cloud.
  • Merging BM25 and vector results client-side when the rrf retriever already does this server-side. This duplicates work Elasticsearch already handles natively, and reintroduces exactly the score-normalization complexity the hybrid search article warned about.
  • Skipping pre-filtering and filtering results after the fact instead. Recall the mechanical distinction directly — filters inside the knn clause reduce the candidate pool before the HNSW search runs, genuinely different from post-hoc filtering of a fixed top-k result set.
  • Standing up a separate vector database when an existing Elastic Stack deployment already covers the need. Recall the Drupal case study directly — this is real, avoidable operational duplication for teams already running Elasticsearch for other purposes.
  • Relevant Search by Tom Hammerbacher et al. — the practical guide to relevance engineering with Elasticsearch, directly covering the BM25 side of the hybrid equation this article pairs with dense and sparse retrieval.
  • AI-Powered Search by Trey Grainger et al. — modern search engineering including vector, keyword, and learned sparse retrieval, matching the three-way RRF fusion pattern covered here.
  • Designing Machine Learning Systems by Chip Huyen — data and retrieval system design context for deciding when reusing an existing Elastic deployment beats standing up a dedicated vector database.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is Elasticsearch vector search used for?

Elasticsearch vector search powers semantic search, RAG retrieval, and hybrid full-text + meaning-based queries inside an existing Elastic Stack deployment. It combines BM25 keyword matching, dense_vector kNN, and ELSER sparse retrieval in a single request.

Yes. dense_vector fields with index set to true support approximate kNN search (available since Elasticsearch 8.0). Queries use the knn clause with a query vector, k, num_candidates, and optional pre-filtering.

What is ELSER in Elasticsearch?

ELSER (Elastic Learned Sparse EncodeR) is Elasticsearch's built-in learned sparse-vector model for semantic search. It produces explainable sparse embeddings, needs no separate model hosting, and requires a Platinum or Enterprise license (or Elastic Cloud).

How does Elasticsearch hybrid search work?

Elasticsearch's rrf retriever fuses BM25, dense kNN, and ELSER sparse results server-side using Reciprocal Rank Fusion — no client-side merging or score normalization required. All three child retrievers resolve in one request.

Is Elasticsearch better than a dedicated vector database?

If you already run Elastic Stack for logs, observability, or search, adding dense_vector fields avoids standing up a separate vector database. For pure vector-only workloads without an existing Elastic deployment, purpose-built tools like Qdrant or Milvus may be simpler.

What is pre-filtering in Elasticsearch kNN?

Pre-filtering applies term/range filters inside the knn clause before the HNSW graph search runs, reducing the candidate pool up front — generally more efficient than searching everything then discarding non-matching results.

Wrapping This Up

Elasticsearch's vector search story is genuinely distinctive precisely because of its history — a mature full-text search engine that added dense vector kNN in 8.0 and its own sparse semantic model (ELSER) on top, then fused all three retrieval methods together in a single server-side request via the same RRF algorithm the hybrid search article covered earlier in this series. For teams already running Elastic Stack infrastructure, this genuinely means semantic search and RAG retrieval arrive without standing up any of the four dedicated vector databases covered in this series' preceding tutorials.

Remember that ELSER's explainability and license requirement are both genuine, concrete factors worth checking directly against your situation, and that pre-filtering inside the knn clause is mechanically distinct from — and generally more efficient than — filtering results after the fact. FYI, this article genuinely closes the vector database arc running through this series' last five tutorials with the one option that isn't vector-first at all — proof that the hybrid search principles from much earlier in this series apply as cleanly to a search engine that added vectors as to a vector database that added search :)

Now go check whether any RAG pipeline you've built earlier in this series is running against infrastructure that already includes an Elastic Stack deployment for logs or observability — if so, the genuinely fastest path to production-grade hybrid search might be adding a dense_vector field to an index you already operate, rather than standing up a fifth new system alongside it.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles