Contents
You built a RAG pipeline. It runs without errors. The answers look good. But here's the uncomfortable question: how do you actually know it works? "I read some outputs and they seemed fine" isn't an evaluation strategy — it's vibes. And vibes don't survive a stakeholder demo. RAGAS (Retrieval Augmented Generation Assessment) turns those vibes into numbers, and today I'll show you how.
True confession: I once shipped a RAG pipeline that I'd tested by asking it ten questions and nodding approvingly at the answers. Two weeks later, a user found a response where the bot confidently cited a document that said the opposite of what the answer claimed. That was the day I stopped eyeballing and started measuring. Let me save you that awkward Slack thread. :)
What Is RAGAS and Why Do You Need It?
RAGAS is an open-source framework that evaluates RAG pipelines automatically using LLMs as judges. Instead of hand-labeling hundreds of answers, you feed RAGAS your questions, retrieved contexts, generated answers, and ground-truth answers — and it scores the whole system for you.
Why does this matter so much? Because RAG pipelines fail in sneaky ways:
- Retrieval misses — the right document exists, but the retriever never found it
- Hallucinations — the generator invents facts the context doesn't support
- Irrelevant answers — the model technically answered, just not the question you asked
You can't fix what you can't measure. And when you swap your embedding model or tweak your chunk size, how do you know you made things better instead of worse? RAGAS gives you a before-and-after scoreboard — the exact missing piece from the Haystack RAG pipeline tutorial we just published.
Installing RAGAS
The install is painless:
pip install ragas
FYI, RAGAS uses an LLM-as-a-judge approach, so it needs an LLM under the hood. OpenAI's models work out of the box, but you can plug in anything from Hugging Face or a local model. More on that later.
The Four Core Metrics
RAGAS measures your pipeline from two angles: retrieval quality and generation quality. Here are the four metrics you'll reach for constantly.
Faithfulness
Faithfulness asks: does the generated answer stick to facts found in the retrieved context? The LLM judge breaks your answer into individual claims and checks each one against the context. A score of 1.0 means every claim traces back to your documents. A score of 0.4 means your bot is writing fan fiction.
IMO, this is the single most important metric. A wrong-but-faithful answer means your retrieval failed; a right-but-unfaithful answer means your generator hallucinated its way to success. You want to know which failure mode you're living with.
Answer Relevancy
Answer relevancy measures whether the answer actually addresses the question. RAGAS generates questions from your answer and compares them to the original question — clever trick, right? This metric catches the classic "technically coherent, completely off-topic" response.
Context Precision
Context precision evaluates the retrieved context. If your retriever returns ten chunks and the useful one sits at position nine, this metric punishes you. It rewards pipelines that rank the genuinely useful chunks near the top.
Context Recall
Context recall measures whether retrieval found all the information needed to answer. You compare retrieved contexts against your ground-truth answer. Low recall? Your retriever is leaving crucial documents on the table — and your generator can't work with context it never received.
Quick cheat sheet:
| Metric | Angle | What a low score means |
|---|---|---|
| Faithfulness | Generation | Answer claims aren't grounded in context |
| Answer relevancy | Generation | Answer is off-topic for the question |
| Context precision | Retrieval | Useful chunks ranked too low |
| Context recall | Retrieval | Needed chunks never retrieved |
Building Your Evaluation Dataset
RAGAS needs four things per test sample: a question, the retrieved contexts, the generated answer, and (for some metrics) the ground-truth answer. Let's build a dataset the modern way, with SingleTurnSample:
from ragas.dataset_schema import SingleTurnSample
sample1 = SingleTurnSample(
user_input="What is RAGAS?",
retrieved_contexts=[
"RAGAS is an open-source framework for evaluating RAG pipelines."
],
response="RAGAS is a framework that evaluates RAG pipelines using LLM-based metrics.",
reference="RAGAS is an open-source evaluation framework for RAG systems.",
)
Wrap multiple samples into an EvaluationDataset:
from ragas.dataset_schema import EvaluationDataset
dataset = EvaluationDataset([sample1, sample2, sample3])
Here's my hard-won advice: aim for 20–50 diverse samples minimum. Three questions won't catch edge cases. Mix easy questions, ambiguous ones, and a few your pipeline should honestly fail at. Testing only the easy stuff is how pipelines look great in demos and crumble in production. Ask me about the demo that crumbled. Actually, don't. :/
Running the Evaluation
Now the fun part — actually scoring things:
from ragas import evaluate
from ragas.llms import LangchainLLMWrapper
from langchain_openai import ChatOpenAI
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o"))
result = evaluate(dataset, llm=evaluator_llm)
print(result)
That's it. RAGAS runs each metric against each sample and hands you a scorecard. The output gives you average scores per metric, and you can convert it to a pandas DataFrame for per-sample inspection:
df = result.to_pandas()
print(df[["user_input", "faithfulness", "answer_relevancy"]])
Pro tip from the trenches: inspect your worst individual samples, not just the averages. A 0.85 average faithfulness could hide one catastrophic sample that cites imaginary documents. Averages lie; individual rows tell stories.
Figure 1: RAGAS scorecards — turning RAG vibes into measurable metrics
Image Alt Text: "RAGAS tutorial automated evaluation for RAG pipelines with faithfulness and context metrics"
Choosing Your Judge LLM
Your evaluator LLM matters a lot. A weak judge produces noisy, inconsistent scores, and then you're optimizing against garbage feedback.
- Strong judges — GPT-4o or Claude-class models give the most reliable verdicts
- Budget option — GPT-4o-mini works surprisingly well for iteration-speed evaluations
- Local models — RAGAS supports local LLMs if you can't send data to external APIs
Trade-off in plain terms: strong judges cost more per run, but they reduce the "wait, is this score real?" debugging sessions. I'd rather pay a few extra cents than distrust my own scoreboard.
Interpreting the Scores (Without Fooling Yourself)
Scores are useless without context, right? Here's how I read them.
Scores Are Relative, Not Absolute
A faithfulness of 0.78 means nothing in isolation. But if you swap your chunking strategy and faithfulness jumps from 0.78 to 0.91? Now you've learned something. Track metrics across changes — that's the whole game.
Diagnose Before You Optimize
Each metric points at a different pipeline component:
- Low context recall → fix retrieval (better embedder, more chunks, metadata filters)
- Low faithfulness → fix generation (better prompts, stricter grounding instructions)
- Low context precision → fix ranking (rerankers help enormously here)
- Low answer relevancy → fix prompts (the model is rambling off-topic)
This is why I love the metric structure: it doesn't just say "bad," it tells you where things went bad. That beats staring at a wrong answer and wondering which layer betrayed you.
Generating a Synthetic Test Set
Hate writing test questions by hand? RAGAS can generate them. The TestsetGenerator takes your knowledge base documents and produces a synthetic evaluation dataset — questions, ground truth, everything:
from ragas.testset import TestsetGenerator
generator = TestsetGenerator(llm=evaluator_llm)
testset = generator.generate_with_langchain_docs(my_documents, testset_size=25)
dataset = testset.to_evaluation_dataset()
The generated questions even span different difficulty types, from simple lookups to multi-hop reasoning. One caveat: synthetic questions skew toward what an LLM finds easy to ask about. Blend synthetic samples with real user questions — real users ask things no generator would ever dream up. My roughest pipeline failures came from user questions, not synthetic ones.
Working RAGAS Into Your Workflow
Here's how I actually use RAGAS day-to-day:
- Baseline first — evaluate before changing anything, or you'll have nothing to compare against
- Evaluate after each change — new embedder, new chunk size, new prompt? Re-run the scoreboard
- Keep a metrics log — a simple CSV of runs and scores beats your memory, trust me
- Watch for regressions — an improvement that tanks another metric isn't an improvement
Ever noticed how a prompt tweak that fixes one failure mode quietly breaks two others? Metrics logs catch those silent regressions. My memory certainly doesn't. ;)
Recommended Books
- Evaluating LLM Applications — practical evaluation methodology for LLM systems, the direct companion discipline to RAGAS-style automated scoring.
- Retrieval-Augmented Generation (RAG) — RAG architecture context for understanding what faithfulness and context recall actually diagnose inside your pipeline.
- Designing Machine Learning Systems by Chip Huyen — evaluation, monitoring, and iteration loops that frame RAGAS scoreboards as part of a larger ML systems practice.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is RAGAS?
RAGAS (Retrieval Augmented Generation Assessment) is an open-source framework that automatically evaluates RAG pipelines using LLMs as judges. You feed it questions, retrieved contexts, generated answers, and ground-truth answers, and it scores retrieval and generation quality.
What are the four core RAGAS metrics?
Faithfulness (does the answer stick to retrieved context), answer relevancy (does the answer address the question), context precision (are useful chunks ranked high), and context recall (did retrieval find all needed information). Two measure generation, two measure retrieval.
Which judge LLM should I use with RAGAS?
GPT-4o or Claude-class models give the most reliable verdicts. GPT-4o-mini is a solid budget option for fast iteration, and RAGAS also supports local LLMs when data cannot leave your environment. Weak judges produce noisy scores you end up optimizing against.
How many samples do I need for a RAGAS evaluation?
Aim for 20–50 diverse samples minimum. Mix easy, ambiguous, and intentionally hard questions. Testing only easy questions makes pipelines look great in demos and fail in production.
How do I generate a synthetic test set with RAGAS?
Use TestsetGenerator with your knowledge base documents to produce questions, ground truth, and an evaluation dataset. Generated questions span difficulty levels, but blend synthetic samples with real user questions since synthetic sets skew toward what LLMs find easy to ask.
How should I interpret RAGAS scores?
Scores are relative, not absolute — track them across changes to a pipeline. Use each metric diagnostically: low context recall points at retrieval, low faithfulness at generation prompts, low context precision at ranking/rerankers, and low answer relevancy at prompt quality.
Wrapping Up
Here's the recap: RAGAS scores your RAG pipeline on faithfulness, answer relevancy, context precision, and context recall — two metrics for generation, two for retrieval. You build a dataset, pick a judge LLM, run evaluate(), and then read the scores diagnostically, letting each metric point you toward the component that needs work. No more nodding approvingly at ten outputs and calling it testing.
My honest take after living with this tool: RAGAS doesn't make your pipeline better — it makes you better at improving your pipeline. The scoreboard changes how you build. You start making decisions with evidence instead of vibes, and your demos stop containing surprise horror moments.
So go run a baseline evaluation on your existing pipeline. I promise the first scorecard will humble you — mine certainly humbled me. But hey, a mediocre number you can improve beats a great vibe you can't verify. Go get your numbers. :D