Contents
Recall the Weaviate vs Pinecone vs Qdrant comparison from much earlier in this series naming Qdrant as the price-performance pick — fast, self-hostable, genuinely hard to beat on cost-efficiency at scale. This is the hands-on follow-through on that recommendation — the same tutorial treatment the Pinecone and ChromaDB articles got, now for the Rust-built option that comparison article said earns its reputation.
Currently at qdrant-client v1.18.0, and worth knowing about directly: Qdrant Edge, a genuinely new, lightweight variant that runs embedded, in-process, with no separate server at all — supporting Python and Rust — meaning Qdrant now spans the same "local prototype to production cluster" range that made ChromaDB and Pinecone each appealing for different reasons, without forcing you to switch tools as you scale.
By the end of this guide, you'll have Qdrant running locally, a collection created and populated, and a working semantic search query — plus enough of the underlying concepts to know when to reach for hybrid search and payload filtering as your project grows. IMO, the fact that the exact same client code works whether you're running fully in-memory or against a production cluster is genuinely the most practically useful thing about this library :)
Figure 1: Qdrant — the Rust-built open-source vector search engine for semantic search and RAG
Image Alt Text: "Qdrant tutorial setup for open-source vector search and RAG applications"
What Qdrant Actually Is
Qdrant is an open-source, Rust-built vector search engine — recall the local LLM tools comparison's point about Rust-based tools directly; the same "written in Rust, and it shows" performance story from that article applies here to vector search specifically, not LLM inference.
- Apache 2.0 licensed, genuinely free and open-source, with a managed Qdrant Cloud option available once you want someone else operating the infrastructure.
- Supports REST and gRPC APIs, with official clients for Python, Rust, Go, JavaScript/TypeScript, and .NET — genuinely broad language coverage beyond just the Python ecosystem this series has focused on.
- Runs in three genuinely distinct modes: fully in-memory for quick experiments, as a local Docker server for realistic development, or as a managed cloud cluster for production — the same client code works against all three.
Installing the Client
pip install qdrant-client
That's the entire installation for local, in-memory experimentation — no Docker, no server, no account required to start writing and testing queries.
Mode One: Fully In-Memory, Zero Setup
Genuinely the fastest way to get a feel for the API, exactly like ChromaDB's in-memory client from the earlier tutorial in this series.
from qdrant_client import QdrantClient
client = QdrantClient(":memory:")
Everything disappears when your script ends — the right mode for a first exploration, a unit test, or a CI pipeline, not for anything you need to persist.
Mode Two: Local Persistent Mode
client = QdrantClient(path="path/to/db")
This persists to disk without running a separate server process at all — genuinely useful for local development and prototyping, letting you restart your script without losing data, similar in spirit to ChromaDB's PersistentClient from earlier in this series, just implemented through Qdrant's own embedded storage engine.
Mode Three: Docker Server, for Realistic Development and Production
docker pull qdrant/qdrant
docker run -p 6333:6333 -p 6334:6334 \
-v $(pwd)/qdrant_storage:/qdrant/storage \
qdrant/qdrant
client = QdrantClient(url="http://localhost:6333")
This is genuinely the mode that matters for anything beyond solo prototyping — a real server process, accessible over REST or gRPC, with data persisted to the mounted volume. Notice the client connection code barely changes between this mode and the in-memory or local-persistent modes above — the same collection, upsert, and query calls work identically regardless of which backend is actually running underneath.
Core Concepts: Collections, Points, and Payloads
Recall the Pinecone tutorial's "index" and ChromaDB's "collection" terminology from earlier in this series — Qdrant genuinely uses the same conceptual shape with its own vocabulary.
- Collection — the named container for your vectors, roughly equivalent to a table, analogous to Pinecone's index or ChromaDB's collection.
- Point — a single record: a vector, an ID, and an optional payload (Qdrant's term for the metadata dictionary attached to each vector — category, source, timestamp, anything you want to filter or retrieve alongside the vector itself).
- Distance metric — configured per collection at creation time, typically cosine similarity for text embeddings, matching the same convention from the Pinecone and ChromaDB tutorials.
Creating a Collection
from qdrant_client import models
client.create_collection(
collection_name="documents",
vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
)
That size=384 needs to match your embedding model's actual output dimension — recall the Sentence Transformers article's all-MiniLM-L6-v2 model from earlier in this series, which genuinely produces 384-dimensional vectors; mismatching this number is a common first-attempt error worth checking before troubleshooting anything else.
Adding Data Without Handling Embeddings Yourself: FastEmbed
Here's a genuinely convenient feature Qdrant's client ships with directly: FastEmbed, an optional dependency letting you add raw text and have Qdrant handle embedding generation internally — recall this being conceptually similar to Pinecone's integrated-embedding feature from that earlier tutorial.
docs = [
"Qdrant is a vector search engine written in Rust.",
"The Eiffel Tower was completed in 1889 in Paris, France.",
"Sentence Transformers produce dense embeddings for semantic search.",
]
client.add(
collection_name="documents",
documents=docs,
ids=[1, 2, 3],
)
Notice client.add() requires no explicit create_collection() call first — it handles collection creation implicitly using FastEmbed's default embedding model dimensions, genuinely the fastest path from zero to a working semantic search collection.
Adding Points With Explicit Vectors and Payloads
For real production use, you'll typically bring your own embeddings — recall the Sentence Transformers or embedding models articles from earlier in this series — and attach genuine metadata for filtering.
client.upsert(
collection_name="documents",
points=[
models.PointStruct(
id=1,
vector=[0.05, 0.61, 0.76, ...], # your actual 384-dim embedding
payload={"category": "science", "source": "wikipedia"}
),
],
)
upsert genuinely means what it says — insert if the ID doesn't exist, update if it does, the same idempotent pattern from the ETL pipeline article's loading-step discussion earlier in this series.
Running a Search
results = client.query_points(
collection_name="documents",
query=[0.1, 0.2, 0.3, ...], # query embedding
limit=5,
)
for point in results.points:
print(point.score, point.payload)
Notice this returns both a similarity score and the full payload for each match — genuinely the same result shape from the Pinecone and ChromaDB tutorials, letting you retrieve the original text or metadata alongside the raw similarity ranking.
Searching With FastEmbed-Managed Text Directly
results = client.query(
collection_name="documents",
query_text="What engineering material is Qdrant built with?",
limit=3,
)
This searches using a raw text query, with FastEmbed handling the embedding step transparently — the same "no separate embedding API call needed" convenience from the Pinecone tutorial's integrated embedding feature, just implemented locally instead of through a hosted service.
Payload Filtering: Combining Semantic Search With Hard Constraints
Recall the Pinecone and hybrid search articles' filtering discussions directly — Qdrant's filtering is genuinely one of its most-cited strengths from the earlier vector database comparison article.
results = client.query_points(
collection_name="documents",
query=[0.1, 0.2, 0.3, ...],
query_filter=models.Filter(
must=[
models.FieldCondition(
key="category",
match=models.MatchValue(value="science")
)
]
),
limit=5,
)
This scopes your semantic search to only records where category equals "science", combining conceptual similarity with a hard metadata constraint — recall the vector database comparison article's direct claim that Qdrant handles queries like "vectors where tenant_id = X" particularly cleanly, exactly the pattern shown here.
Hybrid Search: Combining Keyword and Vector Matching
Recall the hybrid search article's RRF-based fusion of BM25 and vector search from much earlier in this series — Qdrant supports this natively, one of the platforms that article specifically named for built-in hybrid search support.
from qdrant_client.models import Prefetch, FusionQuery, Fusion
results = client.query_points(
collection_name="documents",
prefetch=[
Prefetch(query=dense_vector, using="dense", limit=20),
Prefetch(query=sparse_vector, using="sparse", limit=20),
],
query=FusionQuery(fusion=Fusion.RRF),
)
Notice Fusion.RRF directly — this is genuinely the exact Reciprocal Rank Fusion algorithm covered in detail in the hybrid search article, available here as a built-in query type rather than something you'd need to implement yourself.
Quantization: Recall the Model Compression Arc, Applied to Vectors
Worth connecting directly to the quantization article from earlier in this series — the same size-versus-precision tradeoff applies to vector storage, not just neural network weights.
client.create_collection(
collection_name="documents",
vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
quantization_config=models.ScalarQuantization(
scalar=models.ScalarQuantizationConfig(
type=models.ScalarType.INT8,
quantile=0.99,
)
),
)
This applies the identical INT8 quantization principle from the model compression article directly to your stored vectors — a meaningful memory reduction with a small, well-understood accuracy cost, genuinely the same tradeoff, just applied to embedding storage rather than model weights.
Where This Fits Against Pinecone and ChromaDB
Recall the earlier vector database tutorials directly for the concrete comparison this series already built:
| ChromaDB | Pinecone | Qdrant | |
|---|---|---|---|
| Best for | Fast local prototyping | Fully managed, zero ops | Self-hosted performance + filtering |
| Local mode | Yes, in-process | No | Yes, in-memory or embedded (Qdrant Edge) |
| Managed cloud option | No | Yes, primary offering | Yes, optional |
| Hybrid search | No native support | Sparse-dense pairs | Native, built-in RRF fusion |
| Payload/metadata filtering | Basic where clause | Supported | Particularly strong, per prior comparison |
Recall the earlier comparison article's exact conclusion directly — Qdrant is the pick when you have some infrastructure comfort and genuinely care about cost-efficiency and filtering performance at scale; this tutorial is the concrete "how" behind that recommendation.
Common Mistakes People Make
- Mismatching
sizeinVectorParamsagainst your actual embedding model's output dimension. This produces an immediate, confusing error — always confirm your embedding model's dimension count before creating a collection. - Using
:memory:mode and expecting data to persist. Recall this being the exact same caveat from the ChromaDB tutorial — switch to local-persistent or Docker server mode the moment you need data to survive a restart. - Forgetting payload filters when scoping search to a specific tenant, category, or date range. Recall the multi-tenancy use case directly from the vector database comparison article — pure semantic search without filtering returns technically-similar-but-practically-wrong results once your dataset grows.
- Skipping quantization on large collections without checking the actual memory savings. Recall the model compression article's INT8 tradeoff directly — this is close to free in accuracy terms for most use cases and meaningfully reduces memory footprint at scale.
- Assuming Docker is required for any real project. Recall Qdrant Edge directly — the newer embedded mode genuinely extends how far you can go without standing up a separate server process at all.
Recommended Books
- Designing Machine Learning Systems by Chip Huyen — covers data pipelines, feature storage, and retrieval system design, giving the production context for when a dedicated vector database like Qdrant becomes the right storage layer instead of an in-process store.
- AI-Powered Search by Trey Grainger et al. — modern search engineering including vector, keyword, and hybrid retrieval, directly matching the hybrid RRF and filtering patterns this tutorial introduces in Qdrant.
- Machine Learning Engineering by Andriy Burkov — practical reference for shipping ML systems, useful for the operational side of running a self-hosted Qdrant cluster beyond the prototyping modes covered here.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is Qdrant used for?
Qdrant is an open-source vector search engine for semantic search, RAG pipelines, and recommendation systems. It stores embeddings and retrieves similar content using cosine distance or other metrics, with strong payload filtering and native hybrid search.
Is Qdrant free to use?
Yes. Qdrant is Apache 2.0 licensed and fully open-source. You can run it in-memory, embedded, or as a Docker server at no cost. Qdrant Cloud offers a managed option starting around $25/month when you want someone else operating the infrastructure.
How is Qdrant different from Pinecone and ChromaDB?
ChromaDB is fastest for local prototyping, Pinecone is fully managed with zero ops, and Qdrant is the self-hosted performance pick with particularly strong payload filtering and native RRF hybrid search. The same Qdrant client code works across in-memory, embedded, Docker, and cloud modes.
What is FastEmbed in Qdrant?
FastEmbed is an optional dependency bundled with the qdrant-client that generates embeddings from raw text inside the client. It removes the need for a separate embedding API call when adding documents or running text queries.
Does Qdrant support hybrid search?
Yes. Qdrant supports native hybrid search combining dense and sparse vectors with built-in Reciprocal Rank Fusion (RRF) fusion, one of the platforms named for out-of-the-box hybrid support in this series' vector database comparison.
Can Qdrant run without Docker?
Yes. Qdrant supports fully in-memory mode via :memory:, local persistent embedded storage via a path argument, and Qdrant Edge for in-process embedded deployments in Python and Rust. Docker is only needed when you want a real server process for multi-client or production use.
Wrapping This Up
Qdrant delivers on the price-performance and filtering strengths the earlier vector database comparison article attributed to it — collections, points, and payloads following the same conceptual shape as Pinecone and ChromaDB, with genuinely strong native hybrid search (built-in RRF fusion) and payload filtering, plus a Rust-built performance foundation that shows up directly in the benchmarks that comparison article cited. The same client code scaling from :memory: experimentation through local persistence to a full Docker or cloud cluster means you're not locked into a different tool as your project grows.
Remember to match your VectorParams size to your actual embedding model's output dimension, and that payload filtering is genuinely one of Qdrant's most-cited strengths — worth reaching for whenever your search needs to combine semantic similarity with hard metadata constraints. FYI, this tutorial closes the loop on the vector database arc from earlier in this series — you've now got hands-on tutorials for all three platforms that comparison article evaluated, and can genuinely choose based on tested experience rather than a feature table alone :)
Now go take the Sentence Transformers embedding pipeline from earlier in this series and swap its FAISS or in-memory storage for a real Qdrant collection with payload filtering enabled. That's genuinely the fastest way to feel the difference a dedicated vector database makes once your search needs to scope by more than similarity alone.