Sam Austin AI

Multi-Modal RAG: Retrieving Images and Text Together

September 1, 2026 12 min read Updated September 2, 2026 Sam Austin
Contents

Your RAG system works great for text-only questions, but the moment someone asks "show me the diagram on page 5" or "find similar product images," it falls apart. That's because traditional RAG treats images as invisible — it can only search what it can read, and images aren't words. Multi-modal RAG fixes this by letting your system search across both text and images using a single unified approach.

I started exploring multi-modal RAG after a client asked why their product search couldn't find items when users uploaded photos instead of typing descriptions. That gap between "what users have" and "what systems can search" is exactly where multi-modal retrieval lives. Ever wished your RAG system could understand both a written question and an uploaded image? This guide shows you how.

By the end, you'll understand how CLIP creates shared embeddings for images and text, how to build a multi-modal retrieval pipeline, and when this approach actually adds value versus when it's overkill. IMO, this is the most exciting frontier in RAG right now :)

Multi-Modal RAG Retrieving Images and Text Together
Multi-Modal RAG Retrieving Images and Text Together

Figure 1: Multi-modal RAG retrieves both images and text using shared embeddings

Why Text-Only RAG Has a Blind Spot

Here's the core limitation: traditional RAG systems convert text to vectors and search those vectors. Images are treated as opaque blobs — the system might store a URL to an image, but it can't search the image's content. If your knowledge base includes diagrams, charts, product photos, or technical illustrations, that visual information is effectively invisible to retrieval.

This matters more than most teams realize. Technical documentation often contains critical information in diagrams that text descriptions can't fully capture. Product catalogs need visual search. Medical records include images alongside text. Forcing users to describe images in words before searching is friction that multi-modal RAG eliminates entirely.

What Multi-Modal RAG Actually Does

Multi-modal RAG extends the retrieval pipeline to handle multiple types of content — text, images, and sometimes audio or video — using a unified embedding space. Instead of separate search systems for different content types, everything lives in one index that can be queried from any modality.

  • Search for images using a text query ("find diagrams of neural network architectures")
  • Search for text using an image query (upload a photo, find matching products)
  • Retrieve both modalities together for a single query (get the text explanation AND the accompanying diagram)
  • Generate answers that reference both textual and visual sources

The key enabler is a shared embedding space where images and text live as comparable vectors.

CLIP (Contrastive Language-Image Pre-training) from OpenAI is the breakthrough that made multi-modal search practical. It learns to map images and text into the same vector space, so a photo of a dog and the phrase "a photo of a dog" end up close together in embedding space.

How CLIP Works

CLIP trains on millions of image-text pairs, learning which images correspond to which descriptions. The result:

  • An image encoder that produces image embeddings
  • A text encoder that produces text embeddings
  • Both embeddings live in the same 512 or 768-dimensional space
  • Cosine similarity between image and text embeddings measures relevance
import open_clip
import torch
from PIL import Image

model, _, preprocess = open_clip.create_model_and_transforms(
    'ViT-B-32', pretrained='laion2b_s34b_b79k'
)
tokenizer = open_clip.get_tokenizer('ViT-B-32')

# Encode an image
image = preprocess(Image.open("diagram.png")).unsqueeze(0)
image_embedding = model.encode_image(image)

# Encode text
text = tokenizer(["neural network architecture diagram"])
text_embedding = model.encode_text(text)

# Compute similarity
similarity = torch.cosine_similarity(image_embedding, text_embedding)

That similarity score tells you how well the image matches the text query — even if the text never explicitly describes that specific image.

Choosing the Right CLIP Model

| Model | Dimensions | Speed | Quality |

|-------|------------|-------|---------|

| ViT-B-32 | 512 | Fast | Good |

| ViT-L-14 | 768 | Medium | Better |

| ViT-H-14 | 1024 | Slow | Best |

Start with ViT-B-32 for prototyping. Move to ViT-L-14 when quality matters more than speed.

Building a Multi-Modal Retrieval Pipeline

Here's the architecture for a basic multi-modal RAG system:

Step 1: Embed Both Text and Images

import open_clip
from PIL import Image

model, _, preprocess = open_clip.create_model_and_transforms(
    'ViT-L-14', pretrained='laion2b_s32b_b82k'
)
tokenizer = open_clip.get_tokenizer('ViT-L-14')

def embed_text(texts):
    tokens = tokenizer(texts)
    with torch.no_grad():
        return model.encode_text(tokens)

def embed_images(image_paths):
    images = [preprocess(Image.open(p)).unsqueeze(0) for p in image_paths]
    images = torch.cat(images)
    with torch.no_grad():
        return model.encode_image(images)

Step 2: Store in Vector Database

from chromadb import Client

client = Client()
collection = client.create_collection("multi_modal_docs")

# Add text documents
text_embeddings = embed_text(["RAG combines retrieval with generation", "CLIP enables cross-modal search"])
collection.add(
    embeddings=text_embeddings.tolist(),
    documents=["RAG combines retrieval with generation", "CLIP enables cross-modal search"],
    metadatas=[{"modality": "text"}, {"modality": "text"}],
    ids=["text_1", "text_2"]
)

# Add images
image_embeddings = embed_images(["diagram.png", "chart.jpg"])
collection.add(
    embeddings=image_embeddings.tolist(),
    documents=["diagram.png", "chart.jpg"],
    metadatas=[{"modality": "image"}, {"modality": "image"}],
    ids=["img_1", "img_2"]
)
def multi_modal_search(query, top_k=5):
    # Text query searches across both text and images
    query_embedding = embed_text([query])
    results = collection.query(
        query_embeddings=query_embedding.tolist(),
        n_results=top_k
    )
    return results

# Search for images using text
results = multi_modal_search("neural network architecture")

# Search for text using an image
image_query = embed_images(["upload.jpg"])
results = collection.query(
    query_embeddings=image_query.tolist(),
    n_results=5
)

Advanced: Image-Text Retrieval with Re-ranking

For production systems, combining CLIP retrieval with a cross-modal reranker improves precision significantly:

def enhanced_multi_modal_search(query, top_k=10, final_k=5):
    # Stage 1: CLIP retrieval
    initial_results = multi_modal_search(query, top_k=top_k)
    
    # Stage 2: Cross-modal reranking
    candidates = initial_results['documents'][0]
    reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')
    
    pairs = [[query, doc] for doc in candidates]
    scores = reranker.predict(pairs)
    ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)
    
    return [doc for doc, score in ranked[:final_k]]

Handling Images in RAG Generation

Retrieving images is half the battle — you also need to generate answers that reference both text and images. Here's how to handle this with GPT-4V:

import base64
from openai import OpenAI

client = OpenAI()

def multi_modal_generate(query, text_context, image_paths):
    # Encode images for GPT-4V
    image_contents = []
    for path in image_paths:
        with open(path, "rb") as f:
            image_data = base64.b64encode(f.read()).decode()
        image_contents.append({
            "type": "image_url",
            "image_url": {"url": f"data:image/png;base64,{image_data}"}
        })
    
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text": f"Context: {text_context}\n\nQuestion: {query}"},
                *image_contents
            ]
        }]
    )
    return response.choices[0].message.content

When Multi-Modal RAG Actually Matters

Not every RAG system needs multi-modal retrieval. Here's when it adds genuine value:

Strong use cases:

  • Product catalogs with images (e-commerce search)
  • Technical documentation with diagrams and charts
  • Medical records with images and text
  • Research papers with figures and tables
  • Visual knowledge bases (architecture diagrams, flowcharts)

When to skip it:

  • Pure text documentation without images
  • Simple FAQ bots
  • Systems where images are purely decorative
  • When latency requirements are extremely strict (multi-modal adds overhead)

Common Mistakes in Multi-Modal RAG

  • Treating images as text descriptions only. Manually describing images loses visual information that CLIP can capture directly.
  • Using the same chunking strategy for images and text. Images are atomic units — don't try to "chunk" them like text.
  • Ignoring image quality. Low-resolution or blurry images produce poor embeddings. Ensure images are clear and readable.
  • Forgetting modality metadata. Always tag whether a result is text or image so the generator knows what it's working with.

Frequently Asked Questions

What is multi-modal RAG?

Multi-modal RAG retrieves both text and images (or other modalities) together based on a query. Instead of searching only text documents, it finds relevant images, diagrams, and visual content alongside textual information.

CLIP learns shared embeddings for images and text in the same vector space. This allows you to search for images using text queries and vice versa, enabling cross-modal retrieval without separate indices.

What are use cases for multi-modal RAG?

Common use cases include searching product catalogs with images, finding diagrams in technical documentation, visual question answering, and building knowledge bases that combine text and visual information.

How do I index images for RAG?

Use CLIP to generate image embeddings, then store them in a vector database alongside text embeddings. You can use a single index with modality metadata or separate indices for each type.

Can I use GPT-4V for multi-modal RAG?

Yes, GPT-4V can process both images and text. You can pass retrieved images directly to GPT-4V along with text context for generation. This simplifies the pipeline by handling both modalities in one model.

What is the difference between multi-modal and cross-modal retrieval?

Multi-modal retrieves across different types (text + images) together. Cross-modal specifically searches one modality using another (search images using text). Multi-modal is the broader concept that includes cross-modal.

Wrapping This Up

Multi-modal RAG extends your retrieval system beyond text to include images, diagrams, and visual content. CLIP provides the foundation by creating shared embeddings for both modalities, and modern LLMs like GPT-4V can generate answers that reference both text and images together.

Start with CLIP ViT-B-32 for prototyping, store everything in a single vector database with modality metadata, and use GPT-4V for generation when you need to reason about images. FYI, as knowledge bases become increasingly visual, multi-modal RAG is shifting from nice-to-have to essential — teams that build it now will have a real advantage :)

Now go look at your knowledge base and ask: how much important information lives in images that your current RAG system can't see? That's the gap multi-modal RAG fills.

Share this article X Facebook LinkedIn Reddit WhatsApp