Contents
Every RAG tutorial out there wants you to install six frameworks before writing a single line of actual logic. I get why frameworks exist, but sometimes you just want to understand what's happening under the hood first. That's exactly what we're doing today — no LangChain, no LlamaIndex, just plain Python and enough understanding to know what those frameworks are actually doing for you later.
I built my first RAG pipeline this way specifically because I hate using tools I don't understand. Once you've built one from scratch, every framework afterward feels like a shortcut instead of a black box, and that mental shift is worth the extra hour of setup.
By the end of this, you'll have a working RAG pipeline running locally, and you'll actually know why each piece exists. IMO, that's a far better foundation than copy-pasting a LangChain tutorial and hoping it works :)
What We're Actually Building
Before touching code, let's get the mental model straight. A basic RAG pipeline needs exactly four pieces working together:
- Documents — the source material you want the AI to reference.
- An embedding model — turns text into numerical vectors representing meaning.
- A vector store — holds those vectors and finds the closest matches to a query.
- An LLM — generates the final answer using retrieved context.
That's genuinely it. Everything else is optimization layered on top of this core loop. Ever wondered why some tutorials make this feel like rocket science? Usually it's because they jump straight to production concerns before you've even seen the basic pattern work once.
Figure 1: Building a complete RAG pipeline from scratch with Python
Building a RAG pipeline from scratch teaches you what frameworks do under the hood
Setting Up Your Environment
Let's get the boring part out of the way first. You'll need a few packages, and I'm deliberately keeping this list short.
pip install openai numpy tiktoken
That's the whole dependency list for a minimal working version. No vector database service, no framework — just the OpenAI SDK for embeddings and generation, NumPy for the actual similarity math, and tiktoken for token counting.
A Quick Note on API Keys
You'll need an OpenAI API key set as an environment variable, since we're using their embedding and completion endpoints for this example.
import os
os.environ["OPENAI_API_KEY"] = "your-key-here"
Obviously, don't hardcode your key like that in anything you commit to GitHub. I've made that mistake exactly once, and GitHub's secret scanning caught it within minutes — mildly humiliating, genuinely useful safety net.
Step 1: Preparing Your Documents
Let's start with some sample documents. In a real project, this would be your PDFs, help docs, or whatever knowledge base you're working with.
documents = [
"The Eiffel Tower was completed in 1889 and stands 330 meters tall.",
"Python was created by Guido van Rossum and released in 1991.",
"The Great Wall of China stretches over 21,000 kilometers.",
"RAG combines retrieval systems with generative language models.",
"Vector embeddings represent text as numerical arrays capturing meaning."
]
Chunking Larger Documents
Real documents need to be broken into smaller pieces, since retrieving an entire 50-page PDF as "context" defeats the whole purpose. Here's a simple chunking function that splits text by sentence count.
def chunk_text(text, sentences_per_chunk=3):
sentences = text.split(". ")
chunks = []
for i in range(0, len(sentences), sentences_per_chunk):
chunk = ". ".join(sentences[i:i + sentences_per_chunk])
chunks.append(chunk)
return chunks
This is deliberately basic — production chunking strategies get a lot more sophisticated, accounting for semantic boundaries instead of arbitrary sentence counts. Don't over-engineer this on day one. Get something working first.
Step 2: Generating Embeddings
This is where the actual "understanding meaning" magic happens. We're converting each document chunk into a vector using OpenAI's embedding model.
from openai import OpenAI
client = OpenAI()
def get_embedding(text, model="text-embedding-3-small"):
response = client.embeddings.create(input=text, model=model)
return response.data[0].embedding
document_embeddings = [get_embedding(doc) for doc in documents]
Each call returns a list of numbers — usually 1,536 dimensions for this particular model. You don't need to understand what each number represents individually. What matters is that similar meanings produce similar vectors, and that's the whole trick.
Why Not Just Use Keyword Matching?
Fair question, honestly. Keyword search would find "Eiffel Tower" if you searched "Eiffel Tower," sure, but it completely misses "How tall is that famous tower in Paris?" Embeddings capture that conceptual link even when the exact words don't overlap.
Step 3: Building the Retrieval Function
Now we need a way to find which document chunks are most relevant to a given question. This means comparing the question's embedding against every document embedding using cosine similarity.
import numpy as np
def cosine_similarity(a, b):
a, b = np.array(a), np.array(b)
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
def retrieve(query, documents, document_embeddings, top_k=2):
query_embedding = get_embedding(query)
similarities = [
cosine_similarity(query_embedding, doc_emb)
for doc_emb in document_embeddings
]
ranked_indices = np.argsort(similarities)[::-1][:top_k]
return [documents[i] for i in ranked_indices]
That's genuinely the entire retrieval engine. No FAISS, no Pinecone, no dedicated vector database — just NumPy computing distances between number-lists. At small scale, this brute-force approach works perfectly fine.
When Does This Approach Break Down?
Once you're past a few thousand documents, comparing your query against every single embedding one-by-one gets slow. That's precisely when tools like FAISS, Qdrant, or pgvector earn their keep — they use indexing structures to avoid brute-force comparison. But for learning purposes and small projects, don't reach for that complexity prematurely.
Step 4: Generating the Final Answer
With relevant chunks retrieved, we feed them to the LLM alongside the original question, instructing it to answer based on that context.
def generate_answer(query, retrieved_chunks):
context = "\n".join(retrieved_chunks)
prompt = f"""Answer the question using only the context below.
If the answer isn't in the context, say you don't know.
Context:
{context}
Question: {query}
"""
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}]
)
return response.choices[0].message.content
Notice that instruction telling the model to admit when it doesn't know something. That single line does more to reduce hallucination than almost anything else in this pipeline. Skip it, and the model will happily guess instead of saying "I don't know" — which is exactly the behavior RAG exists to prevent.
Putting It All Together
Here's the complete pipeline running end-to-end.
def rag_pipeline(query, documents, document_embeddings):
retrieved = retrieve(query, documents, document_embeddings)
answer = generate_answer(query, retrieved)
return answer, retrieved
query = "How tall is the Eiffel Tower?"
answer, sources = rag_pipeline(query, documents, document_embeddings)
print("Answer:", answer)
print("Sources used:", sources)
Run this, and you'll see the model correctly pull the height information and cite the relevant chunk as its source. That's a genuine, functioning RAG pipeline in under 50 lines of code. No frameworks, no external services beyond the OpenAI API itself.
Common Mistakes When Building This Yourself
I've made most of these while learning, so treat this as a shortcut past my own mistakes.
- Forgetting to normalize or handle empty documents. An empty string embedding will throw off your similarity math in ways that are annoying to debug.
- Chunking too aggressively or not enough. Tiny chunks lose context; massive chunks dilute retrieval relevance — there's no universal perfect size, only experimentation.
- Not testing with adversarial queries. Try questions with no good answer in your documents and confirm the model actually says "I don't know" instead of confidently inventing one.
- Recomputing embeddings on every run. Cache them. Recalculating embeddings for unchanged documents wastes API calls and money for absolutely no benefit.
When Should You Actually Move to a Framework?
This from-scratch version is genuinely fine for small projects and for learning the concepts properly. But once you need persistent storage, metadata filtering, hybrid search, or you're juggling thousands of documents, hand-rolling everything becomes tedious fast.
That's when tools like LlamaIndex, LangChain, or a dedicated vector database like Qdrant or pgvector start paying for themselves. Understanding this manual version first means you'll actually know what those frameworks are doing instead of trusting a black box. That's the whole point of this exercise.
Frequently Asked Questions
What is a RAG pipeline?
A RAG pipeline combines document retrieval with language model generation. It finds relevant documents using vector similarity, then feeds them to an LLM to generate accurate, grounded answers.
Do I need LangChain to build RAG?
No. A basic RAG pipeline can be built in under 50 lines of Python using just OpenAI API, NumPy, and basic vector math. Frameworks add convenience but aren't required.
What is the best embedding model for RAG?
OpenAI's text-embedding-3-small is a good starting point with 1,536 dimensions. For open-source alternatives, consider sentence-transformers or BAAI/bge-small.
How do I chunk documents for RAG?
Start with simple sentence-based chunking (3-5 sentences per chunk). For production, consider semantic chunking that respects paragraph and topic boundaries.
What is cosine similarity in RAG?
Cosine similarity measures the angle between two vectors to determine how similar their meanings are. It's used to find the most relevant document chunks for a query.
When should I move from custom RAG to a framework?
Move to a framework when you need persistent storage, metadata filtering, hybrid search, or are managing thousands of documents. For learning and small projects, custom code works fine.
Wrapping This Up
Building RAG from scratch boils down to four pieces working in sequence: chunk your documents, embed them, retrieve the closest matches to a query, and generate an answer grounded in that retrieved context. Everything else is optimization on top of that core loop.
Remember to cache your embeddings, test with questions that genuinely have no answer in your data, and don't reach for a heavyweight framework before you've actually outgrown the simple version. FYI, most production systems still run this exact logic underneath — frameworks just wrap it in nicer syntax :)
Now go build something with actual documents instead of my five sample sentences about towers and walls. You'll learn more from your own mistakes than from any tutorial, mine included.