Sam Austin AI

Qdrant Tutorial: Open-Source Vector Search Engine Getting Started

September 23, 2026 14 min read Sam Austin
Contents

Recall the Weaviate vs Pinecone vs Qdrant comparison from much earlier in this series naming Qdrant as the price-performance pick — fast, self-hostable, genuinely hard to beat on cost-efficiency at scale. This is the hands-on follow-through on that recommendation — the same tutorial treatment the Pinecone and ChromaDB articles got, now for the Rust-built option that comparison article said earns its reputation.

Currently at qdrant-client v1.18.0, and worth knowing about directly: Qdrant Edge, a genuinely new, lightweight variant that runs embedded, in-process, with no separate server at all — supporting Python and Rust — meaning Qdrant now spans the same "local prototype to production cluster" range that made ChromaDB and Pinecone each appealing for different reasons, without forcing you to switch tools as you scale.

By the end of this guide, you'll have Qdrant running locally, a collection created and populated, and a working semantic search query — plus enough of the underlying concepts to know when to reach for hybrid search and payload filtering as your project grows. IMO, the fact that the exact same client code works whether you're running fully in-memory or against a production cluster is genuinely the most practically useful thing about this library :)

Qdrant Tutorial Open-Source Vector Search Engine Getting Started

Figure 1: Qdrant — the Rust-built open-source vector search engine for semantic search and RAG

Image Alt Text: "Qdrant tutorial setup for open-source vector search and RAG applications"

What Qdrant Actually Is

Qdrant is an open-source, Rust-built vector search engine — recall the local LLM tools comparison's point about Rust-based tools directly; the same "written in Rust, and it shows" performance story from that article applies here to vector search specifically, not LLM inference.

  • Apache 2.0 licensed, genuinely free and open-source, with a managed Qdrant Cloud option available once you want someone else operating the infrastructure.
  • Supports REST and gRPC APIs, with official clients for Python, Rust, Go, JavaScript/TypeScript, and .NET — genuinely broad language coverage beyond just the Python ecosystem this series has focused on.
  • Runs in three genuinely distinct modes: fully in-memory for quick experiments, as a local Docker server for realistic development, or as a managed cloud cluster for production — the same client code works against all three.

Installing the Client

pip install qdrant-client

That's the entire installation for local, in-memory experimentation — no Docker, no server, no account required to start writing and testing queries.

Mode One: Fully In-Memory, Zero Setup

Genuinely the fastest way to get a feel for the API, exactly like ChromaDB's in-memory client from the earlier tutorial in this series.

from qdrant_client import QdrantClient

client = QdrantClient(":memory:")

Everything disappears when your script ends — the right mode for a first exploration, a unit test, or a CI pipeline, not for anything you need to persist.

Mode Two: Local Persistent Mode

client = QdrantClient(path="path/to/db")

This persists to disk without running a separate server process at all — genuinely useful for local development and prototyping, letting you restart your script without losing data, similar in spirit to ChromaDB's PersistentClient from earlier in this series, just implemented through Qdrant's own embedded storage engine.

Mode Three: Docker Server, for Realistic Development and Production

docker pull qdrant/qdrant
docker run -p 6333:6333 -p 6334:6334 \
    -v $(pwd)/qdrant_storage:/qdrant/storage \
    qdrant/qdrant
client = QdrantClient(url="http://localhost:6333")

This is genuinely the mode that matters for anything beyond solo prototyping — a real server process, accessible over REST or gRPC, with data persisted to the mounted volume. Notice the client connection code barely changes between this mode and the in-memory or local-persistent modes above — the same collection, upsert, and query calls work identically regardless of which backend is actually running underneath.

Core Concepts: Collections, Points, and Payloads

Recall the Pinecone tutorial's "index" and ChromaDB's "collection" terminology from earlier in this series — Qdrant genuinely uses the same conceptual shape with its own vocabulary.

  • Collection — the named container for your vectors, roughly equivalent to a table, analogous to Pinecone's index or ChromaDB's collection.
  • Point — a single record: a vector, an ID, and an optional payload (Qdrant's term for the metadata dictionary attached to each vector — category, source, timestamp, anything you want to filter or retrieve alongside the vector itself).
  • Distance metric — configured per collection at creation time, typically cosine similarity for text embeddings, matching the same convention from the Pinecone and ChromaDB tutorials.

Creating a Collection

from qdrant_client import models

client.create_collection(
    collection_name="documents",
    vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
)

That size=384 needs to match your embedding model's actual output dimension — recall the Sentence Transformers article's all-MiniLM-L6-v2 model from earlier in this series, which genuinely produces 384-dimensional vectors; mismatching this number is a common first-attempt error worth checking before troubleshooting anything else.

Adding Data Without Handling Embeddings Yourself: FastEmbed

Here's a genuinely convenient feature Qdrant's client ships with directly: FastEmbed, an optional dependency letting you add raw text and have Qdrant handle embedding generation internally — recall this being conceptually similar to Pinecone's integrated-embedding feature from that earlier tutorial.

docs = [
    "Qdrant is a vector search engine written in Rust.",
    "The Eiffel Tower was completed in 1889 in Paris, France.",
    "Sentence Transformers produce dense embeddings for semantic search.",
]

client.add(
    collection_name="documents",
    documents=docs,
    ids=[1, 2, 3],
)

Notice client.add() requires no explicit create_collection() call first — it handles collection creation implicitly using FastEmbed's default embedding model dimensions, genuinely the fastest path from zero to a working semantic search collection.

Adding Points With Explicit Vectors and Payloads

For real production use, you'll typically bring your own embeddings — recall the Sentence Transformers or embedding models articles from earlier in this series — and attach genuine metadata for filtering.

client.upsert(
    collection_name="documents",
    points=[
        models.PointStruct(
            id=1,
            vector=[0.05, 0.61, 0.76, ...],  # your actual 384-dim embedding
            payload={"category": "science", "source": "wikipedia"}
        ),
    ],
)

upsert genuinely means what it says — insert if the ID doesn't exist, update if it does, the same idempotent pattern from the ETL pipeline article's loading-step discussion earlier in this series.

results = client.query_points(
    collection_name="documents",
    query=[0.1, 0.2, 0.3, ...],  # query embedding
    limit=5,
)

for point in results.points:
    print(point.score, point.payload)

Notice this returns both a similarity score and the full payload for each match — genuinely the same result shape from the Pinecone and ChromaDB tutorials, letting you retrieve the original text or metadata alongside the raw similarity ranking.

Searching With FastEmbed-Managed Text Directly

results = client.query(
    collection_name="documents",
    query_text="What engineering material is Qdrant built with?",
    limit=3,
)

This searches using a raw text query, with FastEmbed handling the embedding step transparently — the same "no separate embedding API call needed" convenience from the Pinecone tutorial's integrated embedding feature, just implemented locally instead of through a hosted service.

Payload Filtering: Combining Semantic Search With Hard Constraints

Recall the Pinecone and hybrid search articles' filtering discussions directly — Qdrant's filtering is genuinely one of its most-cited strengths from the earlier vector database comparison article.

results = client.query_points(
    collection_name="documents",
    query=[0.1, 0.2, 0.3, ...],
    query_filter=models.Filter(
        must=[
            models.FieldCondition(
                key="category",
                match=models.MatchValue(value="science")
            )
        ]
    ),
    limit=5,
)

This scopes your semantic search to only records where category equals "science", combining conceptual similarity with a hard metadata constraint — recall the vector database comparison article's direct claim that Qdrant handles queries like "vectors where tenant_id = X" particularly cleanly, exactly the pattern shown here.

Hybrid Search: Combining Keyword and Vector Matching

Recall the hybrid search article's RRF-based fusion of BM25 and vector search from much earlier in this series — Qdrant supports this natively, one of the platforms that article specifically named for built-in hybrid search support.

from qdrant_client.models import Prefetch, FusionQuery, Fusion

results = client.query_points(
    collection_name="documents",
    prefetch=[
        Prefetch(query=dense_vector, using="dense", limit=20),
        Prefetch(query=sparse_vector, using="sparse", limit=20),
    ],
    query=FusionQuery(fusion=Fusion.RRF),
)

Notice Fusion.RRF directly — this is genuinely the exact Reciprocal Rank Fusion algorithm covered in detail in the hybrid search article, available here as a built-in query type rather than something you'd need to implement yourself.

Quantization: Recall the Model Compression Arc, Applied to Vectors

Worth connecting directly to the quantization article from earlier in this series — the same size-versus-precision tradeoff applies to vector storage, not just neural network weights.

client.create_collection(
    collection_name="documents",
    vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
    quantization_config=models.ScalarQuantization(
        scalar=models.ScalarQuantizationConfig(
            type=models.ScalarType.INT8,
            quantile=0.99,
        )
    ),
)

This applies the identical INT8 quantization principle from the model compression article directly to your stored vectors — a meaningful memory reduction with a small, well-understood accuracy cost, genuinely the same tradeoff, just applied to embedding storage rather than model weights.

Where This Fits Against Pinecone and ChromaDB

Recall the earlier vector database tutorials directly for the concrete comparison this series already built:

ChromaDB Pinecone Qdrant
Best for Fast local prototyping Fully managed, zero ops Self-hosted performance + filtering
Local mode Yes, in-process No Yes, in-memory or embedded (Qdrant Edge)
Managed cloud option No Yes, primary offering Yes, optional
Hybrid search No native support Sparse-dense pairs Native, built-in RRF fusion
Payload/metadata filtering Basic where clause Supported Particularly strong, per prior comparison

Recall the earlier comparison article's exact conclusion directly — Qdrant is the pick when you have some infrastructure comfort and genuinely care about cost-efficiency and filtering performance at scale; this tutorial is the concrete "how" behind that recommendation.

Common Mistakes People Make

  • Mismatching size in VectorParams against your actual embedding model's output dimension. This produces an immediate, confusing error — always confirm your embedding model's dimension count before creating a collection.
  • Using :memory: mode and expecting data to persist. Recall this being the exact same caveat from the ChromaDB tutorial — switch to local-persistent or Docker server mode the moment you need data to survive a restart.
  • Forgetting payload filters when scoping search to a specific tenant, category, or date range. Recall the multi-tenancy use case directly from the vector database comparison article — pure semantic search without filtering returns technically-similar-but-practically-wrong results once your dataset grows.
  • Skipping quantization on large collections without checking the actual memory savings. Recall the model compression article's INT8 tradeoff directly — this is close to free in accuracy terms for most use cases and meaningfully reduces memory footprint at scale.
  • Assuming Docker is required for any real project. Recall Qdrant Edge directly — the newer embedded mode genuinely extends how far you can go without standing up a separate server process at all.
  • Designing Machine Learning Systems by Chip Huyen — covers data pipelines, feature storage, and retrieval system design, giving the production context for when a dedicated vector database like Qdrant becomes the right storage layer instead of an in-process store.
  • AI-Powered Search by Trey Grainger et al. — modern search engineering including vector, keyword, and hybrid retrieval, directly matching the hybrid RRF and filtering patterns this tutorial introduces in Qdrant.
  • Machine Learning Engineering by Andriy Burkov — practical reference for shipping ML systems, useful for the operational side of running a self-hosted Qdrant cluster beyond the prototyping modes covered here.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is Qdrant used for?

Qdrant is an open-source vector search engine for semantic search, RAG pipelines, and recommendation systems. It stores embeddings and retrieves similar content using cosine distance or other metrics, with strong payload filtering and native hybrid search.

Is Qdrant free to use?

Yes. Qdrant is Apache 2.0 licensed and fully open-source. You can run it in-memory, embedded, or as a Docker server at no cost. Qdrant Cloud offers a managed option starting around $25/month when you want someone else operating the infrastructure.

How is Qdrant different from Pinecone and ChromaDB?

ChromaDB is fastest for local prototyping, Pinecone is fully managed with zero ops, and Qdrant is the self-hosted performance pick with particularly strong payload filtering and native RRF hybrid search. The same Qdrant client code works across in-memory, embedded, Docker, and cloud modes.

What is FastEmbed in Qdrant?

FastEmbed is an optional dependency bundled with the qdrant-client that generates embeddings from raw text inside the client. It removes the need for a separate embedding API call when adding documents or running text queries.

Yes. Qdrant supports native hybrid search combining dense and sparse vectors with built-in Reciprocal Rank Fusion (RRF) fusion, one of the platforms named for out-of-the-box hybrid support in this series' vector database comparison.

Can Qdrant run without Docker?

Yes. Qdrant supports fully in-memory mode via :memory:, local persistent embedded storage via a path argument, and Qdrant Edge for in-process embedded deployments in Python and Rust. Docker is only needed when you want a real server process for multi-client or production use.

Wrapping This Up

Qdrant delivers on the price-performance and filtering strengths the earlier vector database comparison article attributed to it — collections, points, and payloads following the same conceptual shape as Pinecone and ChromaDB, with genuinely strong native hybrid search (built-in RRF fusion) and payload filtering, plus a Rust-built performance foundation that shows up directly in the benchmarks that comparison article cited. The same client code scaling from :memory: experimentation through local persistence to a full Docker or cloud cluster means you're not locked into a different tool as your project grows.

Remember to match your VectorParams size to your actual embedding model's output dimension, and that payload filtering is genuinely one of Qdrant's most-cited strengths — worth reaching for whenever your search needs to combine semantic similarity with hard metadata constraints. FYI, this tutorial closes the loop on the vector database arc from earlier in this series — you've now got hands-on tutorials for all three platforms that comparison article evaluated, and can genuinely choose based on tested experience rather than a feature table alone :)

Now go take the Sentence Transformers embedding pipeline from earlier in this series and swap its FAISS or in-memory storage for a real Qdrant collection with payload filtering enabled. That's genuinely the fastest way to feel the difference a dedicated vector database makes once your search needs to scope by more than similarity alone.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles