Contents
Recall the Elasticsearch article's framing directly — a search engine with vectors added versus a vector-first database. Redis inverts that story again, in a third direction: not a search engine, not a purpose-built vector store, but the in-memory cache and data structure server this series has assumed as infrastructure without ever naming — now doing double duty as your vector database too. The genuinely distinctive pitch here is speed, not features: millisecond-scale query responses because search and indexing both happen entirely in RAM, no disk round-trip involved at any point in the query path.
Redis packages this AI-focused capability as Redis Iris, its real-time context engine for agents, with RedisVL as the Python client wiring vector search and semantic caching directly into application code. Worth knowing directly: Redis 8.8 introduced FT.HYBRID, native server-side hybrid search — meaning the same "single request, no client-side merging" story from the Elasticsearch article now applies here too, on infrastructure a genuinely large share of existing web applications already run for caching alone.
By the end of this guide, you'll understand Redis's vector indexing options, RedisVL's higher-level client, and the specific caching capability — semantic caching — that no pure vector database in this series' prior five tutorials offers at all. IMO, "up to 73% lower LLM inference costs" from semantic caching is genuinely the single most financially compelling number in this entire vector database arc :)
Figure 1: Redis as a vector database — in-memory semantic search with RedisVL and semantic caching
Image Alt Text: "Redis vector database tutorial with RedisVL in-memory semantic search and hybrid search"
Why Redis, Specifically: Speed From Being In-Memory by Design
Redis delivers vector search performance by leveraging the same in-memory architecture that's made it the standard caching layer for web applications for over a decade — recall this being genuinely different from every prior tutorial in this series, where in-memory modes (ChromaDB's default client, Qdrant's :memory:, Milvus Lite) were explicitly the prototyping tier, not the production answer.
- RediSearch, the module adding vector capability, supports HNSW and flat (brute-force) indexing — the same two index-type tradeoff from the Elasticsearch article's HNSW-versus-exact-search discussion, here implemented as a Redis module rather than a native field type.
- Redis's own 2026 benchmarking reports Intel Scalable Vector Search (SVS)-based compression delivering up to 144% higher query throughput for FP32 vectors at 0.95 precision — recall the quantization article's precision-versus-speed tradeoff directly; this is that same principle, applied specifically to vector index compression rather than model weights.
- A flat index checks every vector for perfect recall — genuinely useful specifically for small datasets or for generating ground truth when measuring how much recall your HNSW configuration is actually sacrificing for speed, echoing the exact same "start with FlatL2, add approximation once you need it" guidance from the FAISS tutorial much earlier in this series.
Installing and Connecting: RedisVL
pip install redisvl
Requires Python 3.10+. RedisVL is described directly as "the AI-native Redis Python client" — genuinely a purpose-built layer on top of raw Redis commands, comparable in spirit to how pymilvus's MilvusClient or Qdrant's client wrap their respective databases' lower-level protocols.
from redisvl.index import SearchIndex
schema = {
"index": {"name": "documents", "prefix": "doc"},
"fields": [
{"name": "text", "type": "text"},
{"name": "category", "type": "tag"},
{
"name": "embedding",
"type": "vector",
"attrs": {
"dims": 384,
"distance_metric": "cosine",
"algorithm": "hnsw",
"datatype": "float32"
}
}
]
}
index = SearchIndex.from_dict(schema)
index.connect("redis://localhost:6379")
index.create(overwrite=True)
Notice the schema declares text, category, and embedding fields together in one index definition — genuinely the same "one document, multiple field types, one query surface" pattern from the Elasticsearch article's mapping, just expressed through RedisVL's schema dictionary instead of an Elasticsearch mapping JSON.
Loading Data and Running a Vector Search
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
docs = [
{"text": "Redis delivers millisecond-scale in-memory search.", "category": "database"},
{"text": "The Eiffel Tower was completed in 1889 in Paris.", "category": "history"},
]
data = [
{**doc, "embedding": model.encode(doc["text"]).astype("float32").tobytes()}
for doc in docs
]
index.load(data)
Recall all-MiniLM-L6-v2 appearing yet again from the Sentence Transformers article earlier in this series — the same embedding model now feeding its sixth different vector store across this series' tutorials, a genuinely useful reminder that the embedding-generation step stays constant regardless of which storage backend you ultimately choose.
from redisvl.query import VectorQuery
query_vector = model.encode("fast search technology").astype("float32").tobytes()
query = VectorQuery(
vector=query_vector,
vector_field_name="embedding",
return_fields=["text", "category"],
num_results=5
)
results = index.query(query)
Hybrid Filtering: Combining Vector Similarity With Tag Constraints
Recall every prior vector database tutorial's filtering discussion directly — RedisVL's filter expressions genuinely mirror the same pattern.
from redisvl.query.filter import Tag
filter_expression = Tag("category") == "database"
query = VectorQuery(
vector=query_vector,
vector_field_name="embedding",
filter_expression=filter_expression,
num_results=5
)
This applies the tag filter as part of the same query, the same pre-filtering-before-search principle from the Elasticsearch and Qdrant tutorials directly — narrowing the HNSW graph traversal to only relevant candidates rather than filtering a fixed result set after the fact.
Native Hybrid Search: FT.HYBRID
Worth knowing directly, and worth checking your Redis version against: FT.HYBRID, native server-side hybrid search fusing keyword and vector matching, requires Redis 8.8 or later — vector search alone via FT.SEARCH works on earlier versions, but the combined, single-request hybrid capability is genuinely newer.
FT.HYBRID documents_idx
SEARCH "something to listen to music on a run"
VSIM @embedding $query_vector
This is genuinely the same architectural story as the Elasticsearch article's rrf retriever — a shopper searching "something to listen to music on a run" won't type "headphones" or "earbuds," and a pure keyword search misses them entirely; FT.HYBRID fuses the keyword match attempt with the vector similarity match in one server-side call, recall this being precisely the hybrid search article's original motivating example, now shown running natively inside Redis.
The Capability No Prior Tutorial in This Series Has: Semantic Caching
This is genuinely the standout, distinctive Redis capability worth building the rest of this article's recommendation around. Recall the cost optimization article's "batching and caching avoid redundant GPU calls" guidance directly — semantic caching is the concrete mechanism implementing that guidance at the embedding-similarity level, not just exact-string-match caching.
from redisvl.extensions.llmcache import SemanticCache
cache = SemanticCache(
name="llm_cache",
redis_url="redis://localhost:6379",
distance_threshold=0.1
)
response = cache.check(prompt="What is the capital of France?")
if response:
answer = response[0]["response"]
else:
answer = call_expensive_llm("What is the capital of France?")
cache.store(prompt="What is the capital of France?", response=answer)
That distance_threshold parameter is doing genuinely important work — rather than requiring an exact string match to hit the cache, semantic caching matches on embedding similarity, meaning "What's the capital city of France?" and "What is the capital of France?" hit the same cached answer despite being different strings entirely. Redis's own managed LangCache service reports up to 15x faster responses on cache hits and up to 73% lower LLM inference costs in high-repetition workloads — recall this connecting directly to the cost optimization article's "inference eats roughly 80% of AI infrastructure budgets" finding from earlier in this series; semantic caching is genuinely one of the highest-leverage, lowest-effort levers against that specific cost center.
Agent Memory: Session-Scoped and Long-Term, Together
Worth knowing about specifically for anyone building on the local RAG chatbot pattern from earlier in this series: Redis Agent Memory keeps session-scoped working memory (the immediate conversation context) while storing long-term memory as vector embeddings retrieved through the exact same semantic search mechanism covered above — genuinely one system handling both a chatbot's short-term conversational state and its longer-term, retrievable knowledge, rather than two separate infrastructure pieces.
Where Redis Fits Against the Five Prior Vector Database Tutorials
| ChromaDB | Pinecone | Qdrant | Milvus | Elasticsearch | Redis | |
|---|---|---|---|---|---|---|
| Primary identity | Vector-first, lightweight | Vector-first, managed | Vector-first, self-hosted | Vector-first, billion-scale | Full-text search + vectors | In-memory cache/data store + vectors |
| Native hybrid search | No | Sparse-dense pairs | Native RRF | Partial | Native RRF (3-way) | Native (FT.HYBRID, 8.8+) |
| Semantic caching built in | No | No | No | No | No | Yes — genuinely unique to Redis here |
| Existing infra reuse case | N/A | N/A | N/A | N/A | If already on Elastic Stack | If already using Redis for caching |
The pattern across the last two tutorials is genuinely consistent, worth naming directly: both Elasticsearch and Redis offer the strongest case specifically for teams already running that infrastructure for something else — logs and full-text search for Elasticsearch, caching for Redis — where adding vector capability means genuinely one fewer system to operate, not a feature checklist advantage in isolation.
Common Mistakes People Make
- Assuming FT.HYBRID works on any Redis version. Recall the explicit version requirement directly — 8.8 or later; earlier versions support vector search via
FT.SEARCHalone, without the native hybrid fusion. - Setting a semantic cache's distance threshold too loosely. A threshold that's too permissive returns cached answers for genuinely different questions — recall this being conceptually the same precision-recall tradeoff from the FAISS and HNSW discussions throughout this series' vector database tutorials.
- Choosing Redis as a vector database when you're not already using it for caching. Recall the consistent "existing infrastructure reuse" pattern directly — without that existing deployment, a purpose-built vector database (Qdrant, Milvus) may genuinely be the more focused choice.
- Treating flat indexing as obsolete now that HNSW exists. Recall this directly from the FAISS tutorial's own guidance much earlier in this series — flat indexing remains genuinely useful for ground-truth recall measurement and small datasets where exact search costs little.
- Ignoring semantic caching entirely and only using Redis for raw vector storage. This skips the single most distinctive, cost-saving capability covered in this article — recall the 73% inference cost reduction figure directly as the concrete reason to actually use it.
Recommended Books
- Redis in Action by Jos J. L. Carlson — the foundational practical guide to Redis data structures and patterns, useful background for understanding where RediSearch and RedisVL sit on top of core Redis.
- Designing Machine Learning Systems by Chip Huyen — covers inference cost, caching, and retrieval system design, directly matching the semantic-caching cost-lever and existing-infrastructure-reuse arguments in this article.
- AI-Powered Search by Trey Grainger et al. — modern search engineering including hybrid keyword + vector retrieval, relevant to FT.HYBRID's server-side fusion pattern.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
Can Redis be used as a vector database?
Yes. Redis supports vector search through the RediSearch module (HNSW and flat indexing) and the RedisVL Python client. Redis 8.8 added FT.HYBRID for native server-side hybrid keyword + vector search, and Redis also offers semantic caching for LLM workloads.
What is RedisVL?
RedisVL is the AI-native Redis Python client for vector search and semantic caching. It provides SearchIndex schema definitions, VectorQuery with tag filters, and SemanticCache for embedding-similarity LLM response caching on top of a standard Redis server.
What is semantic caching in Redis?
Semantic caching stores LLM responses keyed by embedding similarity rather than exact string match. Similar prompts hit the same cached answer via a distance threshold, reported to deliver up to 15x faster cache-hit responses and up to 73% lower LLM inference costs in high-repetition workloads.
What is FT.HYBRID in Redis?
FT.HYBRID is native server-side hybrid search in Redis 8.8+ that fuses keyword and vector matching in a single request — the same no-client-side-merging story as Elasticsearch's rrf retriever. Earlier versions support vector search via FT.SEARCH but not combined hybrid fusion.
How does Redis compare to Qdrant or Milvus for vector search?
Qdrant and Milvus are purpose-built vector databases; Redis is an in-memory cache/data structure server with vector capability added. Redis is strongest when you already run it for caching — one fewer system to operate — while Qdrant and Milvus fit greenfield dedicated vector deployments better.
Does Redis support HNSW indexing?
Yes. RediSearch supports both HNSW (approximate) and flat (brute-force) vector indexes. Flat search checks every vector for perfect recall and remains useful for ground-truth measurement or small datasets.
Wrapping This Up
Redis's vector search story is genuinely built on speed from architecture rather than features from design — in-memory HNSW and flat indexing through the RediSearch module, RedisVL as a purpose-built AI-native client, and native FT.HYBRID fusion in Redis 8.8+ closing the same gap the Elasticsearch article covered, all running on infrastructure a genuinely large share of existing applications already operate for caching. Semantic caching is the standout capability with no real equivalent in any of the five prior vector database tutorials in this series — a concrete, evidenced lever against the exact inference-cost problem the cost optimization article identified as the dominant AI infrastructure expense.
Remember that FT.HYBRID requires Redis 8.8 or later, and that semantic caching's distance threshold is a genuine precision-recall tradeoff worth tuning deliberately rather than accepting a default. FYI, this article genuinely closes the vector database arc running through this series' last six tutorials with the option most likely to already be running somewhere in your existing stack — Pinecone, ChromaDB, Qdrant, Milvus, and Elasticsearch each asked "which new system should you adopt," and Redis is the one most likely to ask instead "which system you already have could just do this too" :)
Now go check whether the local RAG chatbot from earlier in this series, or any LLM-calling code you've built throughout this arc, could genuinely benefit from a semantic cache layer in front of its most repetitive queries — that single addition, more than any vector-store migration this series has covered, is the one with a direct, quantifiable dollar figure attached to it.