Contents
Everyone optimizes their RAG pipeline by feel — tweaking chunk sizes, swapping embedding models, adjusting retrieval k — without ever measuring whether any of those changes actually made things better. That's like tuning a car engine by listening to the sound instead of looking at the dyno chart. It might work, but you're guessing when you could know.
I learned this the hard way after spending a week "improving" a RAG system that was secretly getting worse. My manual test queries looked fine, but the automated evaluation revealed I'd broken recall on an entire category of questions I hadn't thought to test manually. Ever wondered why your RAG bot works great in demos but disappoints in production? This is usually why — you're measuring the wrong things, or not measuring at all.
By the end of this guide, you'll know exactly which metrics actually predict whether your RAG system works, which ones are vanity numbers, and how to build an evaluation setup that catches regressions before your users do. IMO, this is the single most neglected skill in the entire RAG stack :)
Figure 1: RAG evaluation metrics help you measure what actually matters for retrieval quality
Why Most RAG Teams Don't Measure Enough
Here's the uncomfortable truth: most teams test their RAG system with five handpicked queries, declare it "working," and move on. Then they're genuinely surprised when users report confidently wrong answers a week later.
The problem isn't that these teams are lazy — it's that RAG evaluation feels harder than it actually is. You need a golden test set, a few automated metrics, and a habit of running them before and after changes. That's genuinely most of it. The rest is knowing which metrics predict real-world behavior and which ones just look impressive on a dashboard.
The Three Layers of RAG Evaluation
RAG systems have three distinct failure points, and each needs its own metric layer:
- Retrieval quality — did you find the right documents?
- Generation quality — did you use those documents correctly?
- End-to-end quality — did the user actually get a good answer?
Most teams only measure layer three intuitively (by asking "does this feel right?"). Production RAG systems need automated metrics at all three layers.
Layer 1: Retrieval Metrics
These measure whether your retriever is finding the right documents before the LLM ever sees them.
Precision@K
Precision@K answers: "Of the top K documents I retrieved, how many are actually relevant?"
Precision@K = (relevant docs in top K) / K
- A system retrieving 3 relevant docs out of 5 has Precision@5 = 0.6.
- High precision means your LLM isn't wasting context window on irrelevant noise.
- Low precision means your LLM sees relevant information buried among irrelevant documents.
Recall@K
Recall@K answers: "Of all the relevant documents in my corpus, how many did I actually retrieve?"
Recall@K = (relevant docs in top K) / (total relevant docs in corpus)
- A system retrieving 3 relevant docs when 5 exist has Recall@K = 0.6.
- High recall means you're not missing important information.
- Low recall means your LLM literally can't access documents it needs to answer correctly.
Mean Reciprocal Rank (MRR)
MRR measures where the first relevant result appears in your ranking. It answers: "How far down does a user have to scroll before finding something useful?"
MRR = 1 / (position of first relevant result)
- If the first result is always relevant, MRR = 1.0.
- If relevant results average position 2, MRR = 0.5.
- MRR is particularly useful when users only look at the first few results.
Normalized Discounted Cumulative Gain (NDCG)
NDCG measures ranking quality, giving more weight to higher-positioned relevant documents. It's the metric search engines care about most because position matters — a relevant document at position 1 is worth more than the same document at position 10.
- NDCG@10 is the standard metric for measuring retrieval ranking quality.
- Scores range from 0 to 1, with higher being better.
- NDCG accounts for graded relevance, not just binary relevant/irrelevant.
Layer 2: Generation Metrics
These measure whether your LLM correctly uses the retrieved context to generate answers.
Faithfulness
Faithfulness measures whether the generated answer is actually supported by the retrieved context. This is the single most important generation metric for RAG.
- A faithful answer only contains information found in the retrieved documents.
- An unfaithful answer hallucinates details that aren't in the source material.
- Faithfulness directly predicts user trust — users who catch one hallucination stop trusting the system entirely.
Answer Relevance
Answer relevance measures whether the generated answer actually addresses the user's question, regardless of whether it's faithful to the context.
- A relevant answer directly addresses what was asked.
- An irrelevant answer might be factually correct but off-topic.
- Low relevance often indicates the retriever found wrong documents or the prompt isn't guiding the LLM correctly.
Context Precision
Context precision measures whether the retrieved context contains the information needed to answer the question, without excessive irrelevant content.
- High context precision means the LLM receives focused, relevant context.
- Low context precision means the LLM wades through irrelevant information before finding the answer.
Context Recall
Context recall measures whether all the information needed to answer the question is present in the retrieved context.
- High context recall means nothing important was missed during retrieval.
- Low context recall means the retriever failed to find documents containing key information.
Layer 3: End-to-End Metrics
These measure the final output quality from the user's perspective.
Answer Correctness
Answer correctness compares the generated answer against a ground truth answer, measuring factual accuracy.
- Requires a golden test set with known correct answers.
- Useful for tracking improvements over time.
- Less useful for production monitoring since ground truth answers don't exist for new queries.
Hallucination Rate
Hallucination rate measures the percentage of generated answers that contain information not supported by the retrieved context.
- Lower is better — ideally zero.
- Directly impacts user trust and system reliability.
- Can be measured automatically using faithfulness scoring.
Building Your Golden Test Set
The foundation of RAG evaluation is a representative test set. Here's how to build one:
- Collect real queries from production logs, support tickets, or user feedback.
- Add edge cases — questions your system should handle but might struggle with.
- Include difficult queries — ambiguous questions, multi-hop reasoning, questions requiring information from multiple documents.
- Label ground truth — for each query, identify which documents are relevant and what the correct answer is.
- Start small, grow over time — 50-100 well-labeled queries is a solid starting point.
My honest take? Don't overthink the size. A small, well-labeled test set beats a large, poorly-labeled one every time. The goal isn't volume — it's coverage of your actual use cases.
Automated Evaluation Tools
Manually evaluating hundreds of queries isn't practical. Here are the tools that automate the process:
RAGAS
RAGAS is the most popular open-source RAG evaluation framework. It computes faithfulness, answer relevance, context precision, and context recall automatically using LLM-based scoring.
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy
result = evaluate(
dataset=test_dataset,
metrics=[faithfulness, answer_relevancy]
)
print(result)
DeepEval
DeepEval provides end-to-end evaluation with metrics for both retrieval and generation, plus built-in test case management.
TruLens
TruLens focuses on feedback functions that score RAG outputs using LLM-based evaluation, with good visualization and tracking.
LangSmith
LangSmith provides tracing and evaluation integrated with LangChain, making it easy to evaluate RAG pipelines built on that framework.
Common Evaluation Mistakes
I've seen these repeated across enough teams to call them patterns.
- Only testing with easy queries. Your system probably handles straightforward questions well — test the hard ones that break it.
- Not having a golden test set. Without ground truth labels, you can't compute recall or measure improvement over time.
- Ignoring retrieval metrics entirely. If retrieval is broken, no amount of prompt engineering fixes the generation.
- Over-optimizing for a single metric. Improving faithfulness by refusing to answer anything uncertain isn't a win — it's a different failure mode.
- Evaluating on toy data. A system that works on three sample PDFs might completely fail on your actual document corpus.
A Practical Evaluation Workflow
Here's the workflow I'd recommend:
- Build a golden test set of 50-100 representative queries.
- Run baseline metrics before making any changes.
- Make one change at a time and re-evaluate.
- Track metrics over time to catch regressions.
- Set thresholds for automated alerts — if faithfulness drops below 0.9, something broke.
The key insight: evaluation isn't a one-time task. It's a habit. Run it before every significant change, and you'll catch problems before your users do.
Frequently Asked Questions
How do you evaluate a RAG system?
Evaluate RAG systems using retrieval metrics (precision, recall, MRR), generation metrics (faithfulness, relevance), and end-to-end metrics (answer correctness). Build a golden test set with representative queries and expected answers.
What is faithfulness in RAG evaluation?
Faithfulness measures whether the generated answer is actually supported by the retrieved context. A faithful answer doesn't hallucinate information beyond what the documents contain. It's the most important generation metric for RAG.
What is the difference between precision and recall in RAG?
Precision measures how many retrieved documents are actually relevant. Recall measures how many relevant documents were successfully retrieved. High precision means fewer irrelevant results. High recall means fewer missed relevant documents.
What is MRR in RAG?
MRR (Mean Reciprocal Rank) measures where the first relevant result appears in your ranking. An MRR of 1.0 means the first result is always relevant. MRR of 0.5 means relevant results appear around position 2 on average.
How many test queries do I need for RAG evaluation?
Start with 50-100 representative queries covering your main use cases. Include edge cases and difficult questions. More queries give more reliable metrics but even a small golden set is better than no evaluation at all.
What tools can I use to evaluate RAG systems?
Popular RAG evaluation tools include RAGAS, DeepEval, TruLens, and LangSmith. These tools automate metric computation and provide dashboards for tracking improvements over time.
Wrapping This Up
RAG evaluation isn't as hard as most teams make it — you need a golden test set, a few key metrics at retrieval and generation layers, and the habit of running them regularly. Faithfulness is the metric that predicts user trust more than any other, and precision@k tells you whether your retriever is actually finding the right documents.
Remember that evaluation isn't a one-time project — it's a continuous practice. FYI, the teams that build evaluation into their workflow from day one consistently outperform the ones that try to bolt it on after things go wrong :)
Now go build that golden test set instead of trusting your gut feeling about whether your RAG system works. Your users deserve better than "I think it's fine."