Contents
Your RAG system works great for text-only questions, but the moment someone asks "show me the diagram on page 5" or "find similar product images," it falls apart. That's because traditional RAG treats images as invisible — it can only search what it can read, and images aren't words. Multi-modal RAG fixes this by letting your system search across both text and images using a single unified approach.
I started exploring multi-modal RAG after a client asked why their product search couldn't find items when users uploaded photos instead of typing descriptions. That gap between "what users have" and "what systems can search" is exactly where multi-modal retrieval lives. Ever wished your RAG system could understand both a written question and an uploaded image? This guide shows you how.
By the end, you'll understand how CLIP creates shared embeddings for images and text, how to build a multi-modal retrieval pipeline, and when this approach actually adds value versus when it's overkill. IMO, this is the most exciting frontier in RAG right now :)
Figure 1: Multi-modal RAG retrieves both images and text using shared embeddings
Why Text-Only RAG Has a Blind Spot
Here's the core limitation: traditional RAG systems convert text to vectors and search those vectors. Images are treated as opaque blobs — the system might store a URL to an image, but it can't search the image's content. If your knowledge base includes diagrams, charts, product photos, or technical illustrations, that visual information is effectively invisible to retrieval.
This matters more than most teams realize. Technical documentation often contains critical information in diagrams that text descriptions can't fully capture. Product catalogs need visual search. Medical records include images alongside text. Forcing users to describe images in words before searching is friction that multi-modal RAG eliminates entirely.
What Multi-Modal RAG Actually Does
Multi-modal RAG extends the retrieval pipeline to handle multiple types of content — text, images, and sometimes audio or video — using a unified embedding space. Instead of separate search systems for different content types, everything lives in one index that can be queried from any modality.
- Search for images using a text query ("find diagrams of neural network architectures")
- Search for text using an image query (upload a photo, find matching products)
- Retrieve both modalities together for a single query (get the text explanation AND the accompanying diagram)
- Generate answers that reference both textual and visual sources
The key enabler is a shared embedding space where images and text live as comparable vectors.
CLIP: The Foundation of Multi-Modal Search
CLIP (Contrastive Language-Image Pre-training) from OpenAI is the breakthrough that made multi-modal search practical. It learns to map images and text into the same vector space, so a photo of a dog and the phrase "a photo of a dog" end up close together in embedding space.
How CLIP Works
CLIP trains on millions of image-text pairs, learning which images correspond to which descriptions. The result:
- An image encoder that produces image embeddings
- A text encoder that produces text embeddings
- Both embeddings live in the same 512 or 768-dimensional space
- Cosine similarity between image and text embeddings measures relevance
import open_clip
import torch
from PIL import Image
model, _, preprocess = open_clip.create_model_and_transforms(
'ViT-B-32', pretrained='laion2b_s34b_b79k'
)
tokenizer = open_clip.get_tokenizer('ViT-B-32')
# Encode an image
image = preprocess(Image.open("diagram.png")).unsqueeze(0)
image_embedding = model.encode_image(image)
# Encode text
text = tokenizer(["neural network architecture diagram"])
text_embedding = model.encode_text(text)
# Compute similarity
similarity = torch.cosine_similarity(image_embedding, text_embedding)
That similarity score tells you how well the image matches the text query — even if the text never explicitly describes that specific image.
Choosing the Right CLIP Model
| Model | Dimensions | Speed | Quality |
|-------|------------|-------|---------|
| ViT-B-32 | 512 | Fast | Good |
| ViT-L-14 | 768 | Medium | Better |
| ViT-H-14 | 1024 | Slow | Best |
Start with ViT-B-32 for prototyping. Move to ViT-L-14 when quality matters more than speed.
Building a Multi-Modal Retrieval Pipeline
Here's the architecture for a basic multi-modal RAG system:
Step 1: Embed Both Text and Images
import open_clip
from PIL import Image
model, _, preprocess = open_clip.create_model_and_transforms(
'ViT-L-14', pretrained='laion2b_s32b_b82k'
)
tokenizer = open_clip.get_tokenizer('ViT-L-14')
def embed_text(texts):
tokens = tokenizer(texts)
with torch.no_grad():
return model.encode_text(tokens)
def embed_images(image_paths):
images = [preprocess(Image.open(p)).unsqueeze(0) for p in image_paths]
images = torch.cat(images)
with torch.no_grad():
return model.encode_image(images)
Step 2: Store in Vector Database
from chromadb import Client
client = Client()
collection = client.create_collection("multi_modal_docs")
# Add text documents
text_embeddings = embed_text(["RAG combines retrieval with generation", "CLIP enables cross-modal search"])
collection.add(
embeddings=text_embeddings.tolist(),
documents=["RAG combines retrieval with generation", "CLIP enables cross-modal search"],
metadatas=[{"modality": "text"}, {"modality": "text"}],
ids=["text_1", "text_2"]
)
# Add images
image_embeddings = embed_images(["diagram.png", "chart.jpg"])
collection.add(
embeddings=image_embeddings.tolist(),
documents=["diagram.png", "chart.jpg"],
metadatas=[{"modality": "image"}, {"modality": "image"}],
ids=["img_1", "img_2"]
)
Step 3: Unified Search
def multi_modal_search(query, top_k=5):
# Text query searches across both text and images
query_embedding = embed_text([query])
results = collection.query(
query_embeddings=query_embedding.tolist(),
n_results=top_k
)
return results
# Search for images using text
results = multi_modal_search("neural network architecture")
# Search for text using an image
image_query = embed_images(["upload.jpg"])
results = collection.query(
query_embeddings=image_query.tolist(),
n_results=5
)
Advanced: Image-Text Retrieval with Re-ranking
For production systems, combining CLIP retrieval with a cross-modal reranker improves precision significantly:
def enhanced_multi_modal_search(query, top_k=10, final_k=5):
# Stage 1: CLIP retrieval
initial_results = multi_modal_search(query, top_k=top_k)
# Stage 2: Cross-modal reranking
candidates = initial_results['documents'][0]
reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')
pairs = [[query, doc] for doc in candidates]
scores = reranker.predict(pairs)
ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)
return [doc for doc, score in ranked[:final_k]]
Handling Images in RAG Generation
Retrieving images is half the battle — you also need to generate answers that reference both text and images. Here's how to handle this with GPT-4V:
import base64
from openai import OpenAI
client = OpenAI()
def multi_modal_generate(query, text_context, image_paths):
# Encode images for GPT-4V
image_contents = []
for path in image_paths:
with open(path, "rb") as f:
image_data = base64.b64encode(f.read()).decode()
image_contents.append({
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image_data}"}
})
response = client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": f"Context: {text_context}\n\nQuestion: {query}"},
*image_contents
]
}]
)
return response.choices[0].message.content
When Multi-Modal RAG Actually Matters
Not every RAG system needs multi-modal retrieval. Here's when it adds genuine value:
Strong use cases:
- Product catalogs with images (e-commerce search)
- Technical documentation with diagrams and charts
- Medical records with images and text
- Research papers with figures and tables
- Visual knowledge bases (architecture diagrams, flowcharts)
When to skip it:
- Pure text documentation without images
- Simple FAQ bots
- Systems where images are purely decorative
- When latency requirements are extremely strict (multi-modal adds overhead)
Common Mistakes in Multi-Modal RAG
- Treating images as text descriptions only. Manually describing images loses visual information that CLIP can capture directly.
- Using the same chunking strategy for images and text. Images are atomic units — don't try to "chunk" them like text.
- Ignoring image quality. Low-resolution or blurry images produce poor embeddings. Ensure images are clear and readable.
- Forgetting modality metadata. Always tag whether a result is text or image so the generator knows what it's working with.
Frequently Asked Questions
What is multi-modal RAG?
Multi-modal RAG retrieves both text and images (or other modalities) together based on a query. Instead of searching only text documents, it finds relevant images, diagrams, and visual content alongside textual information.
How does CLIP enable multi-modal search?
CLIP learns shared embeddings for images and text in the same vector space. This allows you to search for images using text queries and vice versa, enabling cross-modal retrieval without separate indices.
What are use cases for multi-modal RAG?
Common use cases include searching product catalogs with images, finding diagrams in technical documentation, visual question answering, and building knowledge bases that combine text and visual information.
How do I index images for RAG?
Use CLIP to generate image embeddings, then store them in a vector database alongside text embeddings. You can use a single index with modality metadata or separate indices for each type.
Can I use GPT-4V for multi-modal RAG?
Yes, GPT-4V can process both images and text. You can pass retrieved images directly to GPT-4V along with text context for generation. This simplifies the pipeline by handling both modalities in one model.
What is the difference between multi-modal and cross-modal retrieval?
Multi-modal retrieves across different types (text + images) together. Cross-modal specifically searches one modality using another (search images using text). Multi-modal is the broader concept that includes cross-modal.
Wrapping This Up
Multi-modal RAG extends your retrieval system beyond text to include images, diagrams, and visual content. CLIP provides the foundation by creating shared embeddings for both modalities, and modern LLMs like GPT-4V can generate answers that reference both text and images together.
Start with CLIP ViT-B-32 for prototyping, store everything in a single vector database with modality metadata, and use GPT-4V for generation when you need to reason about images. FYI, as knowledge bases become increasingly visual, multi-modal RAG is shifting from nice-to-have to essential — teams that build it now will have a real advantage :)
Now go look at your knowledge base and ask: how much important information lives in images that your current RAG system can't see? That's the gap multi-modal RAG fills.