Sam Austin AI

RAG for PDFs: Extracting and Retrieving from Complex Documents

September 24, 2026 13 min read Sam Austin
Contents

Here's a fun experiment: take the RAG pipeline you've polished through ten tutorials of technique stacking — parent-child chunks, compression, rerankers, all of it — and point it at a 60-page financial report. Then watch your beautiful pipeline feed your LLM sentences like Q3 2023 1,245.6 and See Table 4 and absolutely nothing else. Congratulations, your bot just learned to hallucinate from receipts. PDFs are where naive ingestion goes to die, and today we're fixing the single most underrated step in any document RAG system: extraction.

My scar tissue on this one is deep. I once shipped a "works perfectly" RAG bot over a set of policy PDFs, and two weeks later a user asked why the bot kept citing section numbers that didn't exist. I dug in and found my ingestion pipeline had quietly shredded the table of contents into confetti, split a three-column pricing table into gibberish, and lost every single footer note. The retrieval layer wasn't broken — the documents arriving at it were already garbage. You can't retrieve what you never extracted. :/

Why PDFs Break RAG Pipelines

Let's be clear about the enemy. A PDF isn't a document — it's a rendering instruction file. It tells a viewer where to draw glyphs on a page. There's no paragraph structure, no reading order guarantee, no semantic meaning — just coordinates. That means:

  • No reliable reading order — two-column layouts extract as interleaved sentence soup
  • Tables collapse — a 5-column pricing table becomes a word salad of numbers with no labels
  • Headers and footers pollute chunks — Page 12 — CONFIDENTIAL glued onto every single chunk
  • Scanned PDFs have no text at all — your extractor returns an empty string and a shrug
  • Figures and charts vanish — the data lives in an image your pipeline never sees

FYI, that last one is the silent killer. Ask what was the revenue trend? and your bot has nothing to say, because the trend lives in a chart, and charts are pixels. We'll fix that too.

The Extraction Toolbox

You have options, and picking the wrong one for your document type wastes days. My honest ranking:

  • pypdf / PyPDF2 — pure Python, fast, free. Fine for simple text-heavy PDFs. Falls over on complex layouts.
  • pdfplumber — the sweet spot for tables. It extracts words with coordinates, so you can reconstruct table structure. My go-to for financial docs.
  • PyMuPDF (fitz) — fast and thorough, handles layout better than pypdf. Great default.
  • Unstructured — the batteries-included option: detects sections, tables, and titles. Heavier, but it does a lot for you.
  • docling / visual, LLM-based extractors — parse the page visually, which we'll get to shortly.

Here's the pdfplumber table pattern that saved my financial-doc project:

import pdfplumber

with pdfplumber.open("quarterly_report.pdf") as pdf:
    all_text = []
    for page in pdf.pages:
        # Extract tables first, as structured rows
        for table in page.extract_tables():
            for row in table:
                all_text.append(" | ".join(
                    cell or "" for cell in row
                ))
        # Then extract non-table text
        text = page.extract_text()
        if text:
            all_text.append(text)

Notice the trick: tables become pipe-delimited rows before they ever meet your chunker. Now Enterprise | $2,400/year | SSO included is a retrievable, readable line — not the number-ghost soup I shipped in version one. ;)

Scanned PDFs: OCR Time

If extract_text() returns empty, your PDF is a scan. You need OCR:

import pytesseract
from pdf2image import convert_from_path

pages = convert_from_path("scanned_contract.pdf", dpi=300)
text = "\n".join(pytesseract.image_to_string(p) for p in pages)

Fair warning: OCR at 300 DPI over a big archive is slow, and OCR quality varies wildly with scan quality. Budget for it, run it once during ingestion, and store the results. Never OCR at query time — your latency dashboard will never forgive you.

The Layout-Aware Upgrade: Vision LLMs

Here's where things get genuinely fun. The hardest layouts — nested tables, multi-column reports, figures with captions — defeat rule-based extractors. So we cheat: hand the page image to a vision-capable LLM and let it write structured Markdown.

Tools like docling or a simple vision-LLM prompt do exactly this. The pattern:

from openai import OpenAI
import base64

def extract_page_with_vision(image_path: str) -> str:
    with open(image_path, "rb") as f:
        img_b64 = base64.b64encode(f.read()).decode()
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text": (
                    "Extract all content from this page as "
                    "clean Markdown. Preserve tables as Markdown "
                    "tables. Ignore headers/footers/page numbers."
                )},
                {"type": "image_url",
                 "image_url": {"url": f"data:image/png;base64,{img_b64}"}},
            ],
        }],
    )
    return response.choices[0].message.content

The output is gorgeous: real Markdown tables, section headings that mean something, figure descriptions. But it's expensive — one LLM call per page. My production rule: vision extraction for the gnarly 10% of pages, pdfplumber for the boring 90%. Route by document type, not by vibes. A pipeline that vision-processes 400 pages of plain-text contracts is a charity for your LLM provider. :/

RAG for PDFs Complex Document Extraction Tables Vision LLM

Figure 1: RAG for PDFs — extraction quality caps everything downstream

Image Alt Text: "RAG for PDFs extracting and retrieving from complex documents with pdfplumber and vision LLMs"

Chunking PDFs: Where Our Series Skills Kick In

Now extraction feeds into chunking, and — plot twist — everything we've built applies but needs PDF-specific tweaks:

  • Chunk by heading, not just size. Unstructured and docling give you section boundaries. A chunk that starts at 3.2 Refund Eligibility and ends at the next heading is worth ten arbitrary 500-token slices.
  • Add metadata at extraction time — page number, section title, document title, table detection flag. You'll thank me when debugging: this answer came from page 41, section 'Termination' is a citation your users can actually check.
  • Keep tables as atomic chunks. Never split a table across chunks — the rows without headers are unreadable noise.
  • Parent-child chunking? Still your friend. Section = parent, paragraph = child. PDFs are made for this pattern, since they arrive with natural hierarchy built in.
metadata = {
    "source": "quarterly_report.pdf",
    "page": 41,
    "section": "3.2 Refund Eligibility",
    "doc_title": "Q3 2023 Quarterly Report",
}
chunks = splitter.split_text(section_text)  # then attach metadata

Retrieval Over Tables: The Sneaky Failure Mode

One more PDF-specific trap, and it's brutal. Even with perfect extraction, embedding models don't retrieve table rows well. A query like what does the Enterprise plan cost? doesn't match Enterprise | $2,400/year strongly, because the table row isn't a natural sentence.

Two fixes, and I use both:

  • Generate a text summary per table with an LLM — "This table lists pricing: Starter at $29/mo, Pro at $99/mo, Enterprise at $2,400/yr." Index the summary alongside the raw table. Query matches summary; you return both.
  • Keyword/BM25 hybrid search — remember that exact-match matters for codes and prices. A hybrid search layer (vector + BM25) catches E-4012 and $2,400 that pure embeddings miss.

Measure the Whole Chain — Extraction Included

You know the rule: numbers beat vibes. But here's the PDF-specific twist: when your RAGAS scores look sad, run extraction diagnostics before blaming retrieval. Ask these first:

  • Did the text even extract? — sample 20 pages, check for empty or truncated text
  • Do tables survive? — pick 10 tables, verify their content appears intact in your chunks
  • Is reading order sane? — read a two-column page's extracted text yourself. If you can't follow it, your embedder definitely can't.

Then run your usual RAGAS baseline. If context recall stinks and your extraction diagnostics found garbage, fix extraction first — that's your root cause, and no reranker in the world rescues text that never made it into the index. I spent two weeks tuning chunk sizes on a broken extraction pipeline once. Two. Weeks. Don't be me. :/

Failure Symptom Fix layer
Empty extract_text() No answers at all OCR at ingestion
Shuffled columns Incoherent sentences pdfplumber / vision LLM
Split tables Numbers without labels Atomic table chunks + summaries
Missing chart data Can't answer trend questions Vision LLM page extraction
Junk headers in chunks Citations to page footers Strip headers at extraction

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Why do PDFs break RAG pipelines?

A PDF is a rendering instruction file with coordinates, not semantic structure — no guaranteed reading order, tables collapse into number soup, headers/footers pollute chunks, scans contain no extractable text, and figures/charts are invisible pixels your pipeline never sees.

What is the best Python library for PDF table extraction?

pdfplumber is the sweet spot for tables because it extracts words with coordinates for structure reconstruction. PyMuPDF (fitz) is a fast layout-aware default, pypdf works for simple text-heavy PDFs, and Unstructured detects sections, tables, and titles out of the box.

How do I handle scanned PDFs in RAG?

Run OCR during ingestion — for example pdf2image at 300 DPI plus pytesseract — and store the extracted text. Never OCR at query time; quality varies with scan quality, so budget a one-time ingestion pass and cache results.

When should I use vision LLMs for PDF extraction?

Route the gnarly minority of pages — nested tables, multi-column reports, figures with captions — to a vision LLM that writes structured Markdown. Keep pdfplumber or PyMuPDF for the boring 90% of plain-text pages to control cost.

How should I chunk PDF documents for RAG?

Chunk by heading boundaries rather than fixed size, attach metadata (page, section, doc title, table flag) at extraction time, keep tables as atomic unsplit chunks, and use parent-child chunking with sections as parents and paragraphs as children.

Why are table rows hard to retrieve with embeddings?

Table rows like Enterprise | $2,400/year aren't natural sentences, so embedding similarity to a query like "what does the Enterprise plan cost?" is weak. Fix by indexing an LLM-generated text summary per table and adding BM25/hybrid keyword search for exact codes and prices.

Wrapping Up

Recap: PDFs aren't documents, they're drawing instructions — and extraction quality caps everything downstream. Match your tool to your document: pdfplumber for tables, PyMuPDF as your default, OCR for scans, vision LLMs for the gnarly layouts. Add metadata during extraction, chunk by headings with parent-child patterns, make tables atomic, summarize tables for embeddings, add hybrid search for exact values, and diagnose extraction before blaming retrieval when evals turn sour.

My parting take after this whole series? We've stacked layer after layer — compression, chunking, expansion, agents, graphs, fine-tuning — but every single one operates on whatever extraction hands it. Extraction is the foundation under the entire tower. It's also the least glamorous step, which is exactly why it's the cheapest big win in document RAG: one good ingestion pipeline beats five clever retrieval tricks built on shredded confetti.

So go sample twenty pages of your existing corpus and read what your pipeline actually extracted. I promise it'll be educational — mine read like a ransom note written by a shredder. ;) Fix the front door first. Everything else in this series works better once the documents arrive in one piece.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles