RAG Evaluation: How to Grade Your Retrieval

RedHub AI Editorialupdated September 20, 20268 min read

A drawer packed edge-on with cream sheets, one red-lit divider standing proud among them.
Jump to a section9

TL;DR

  • The problem: when a RAG app hallucinates, the model usually isn't the culprit — retrieval is. The model can only answer from the chunks it's handed, and when the right chunk never arrives, it improvises confidently.
  • What RAG evaluation is: scoring the retrieval step against a gold set — real questions paired with the chunks that should answer them — using standard metrics like Recall@k and Precision@k.
  • Why deterministic: retrieval scoring doesn't need an LLM judge. Against a gold set, the score is exact, reproducible, free to run, and strong enough to gate CI.
  • The honest part: a great retrieval score doesn't prove the final answer is good — it proves the answer had a fair chance. Grade retrieval first, because nothing downstream can fix what retrieval missed.

What is RAG evaluation?

RAG evaluation is measuring whether your retrieval step actually finds the right chunks. You build a gold set — real questions paired with the passages that should answer them — then score retrieval with metrics like Recall@k, Precision@k, MRR, and nDCG. The result is exact and repeatable, so you can gate releases on it.

Best for: developers running a RAG pipeline who keep hearing "the AI made that up" — the failure the RAG Retrieval Grader is built to measure and catch.


Here's the uncomfortable truth about most RAG problems: the model didn't hallucinate. It answered exactly the question you asked, from exactly the context you gave it. The context was just wrong. Retrieval pulled three chunks about the 2023 pricing page when the user asked about 2025, and the model — doing its job — wrote a fluent answer from stale facts.

This is why RAG evaluation starts at retrieval, not at the answer. The generation step gets all the attention, but it sits downstream of a search problem. If the right chunk never reaches the model, no prompt engineering, no better model, and no amount of "please be accurate" can recover it. Garbage in, confident garbage out.

It's also one of the four quiet failure surfaces we map in the AI reliability pillar guide — quiet because nothing errors. Retrieval always returns something. The question is whether it returned the right thing, and most teams have no way to answer that.

Why RAG feels right but retrieves wrong

Vector search is similarity search. It returns the chunks whose embeddings sit closest to the question's embedding — which is often, but not always, the chunk that answers the question. The gaps are predictable:

  • Near-miss chunks. A chunk that talks about the topic outranks the chunk that answers the question. Similar isn't the same as relevant.
  • Split answers. Chunking cut the key fact in half, so no single chunk contains the whole answer and neither half ranks well.
  • Noisy top-k. The right chunk is there — at position 7 — buried under six loosely related ones that crowd the context window and dilute the model's attention.
  • Silent drift. You swap an embedding model, tweak chunk size, or re-index new documents, and retrieval quality shifts. Nothing fails. Nobody notices until users do.

Every one of these produces a demo that feels fine. You test the questions you'd naturally ask — usually the easy ones the index handles well — and conclude retrieval works. The misses live in the questions you didn't try.

Precision and recall, in plain language

Retrieval quality comes down to two questions, and the standard metrics are just precise ways of asking them.

Recall asks: did the right chunks show up? Of the chunks that should have been retrieved for this question, how many actually appeared in the top k results? Low recall is the killer — it means the answer never reached the model at all. This is the headline metric.

Precision asks: how much junk came with them? Of the chunks you retrieved, what share were actually relevant? Low precision floods the context window with noise, which costs tokens and drags answer quality down even when the right chunk is present.

MetricThe question it answers
Recall@kOf the chunks that should have been retrieved, how many showed up in the top k? Low recall means the answer never reached the model.
Precision@kWhat share of retrieved chunks were actually relevant? Surfaces noisy retrieval that fills the context window with junk.
MRRHow high up does the first correct chunk land, on average? Rewards putting the answer at the top instead of position 8.
nDCG@kIs the ranking itself good — best chunks above weaker ones, not just present somewhere in the list?

You don't need to memorize the formulas. You need the habit: score all of them against a gold set, watch Recall@k first, and treat a drop as a build-breaking event.

Gold sets: the part everyone skips

A gold set is the ground truth that makes grading possible: a list of real questions, each paired with the chunk or chunks that genuinely answer it. Without one, "is retrieval good?" is a matter of opinion. With one, it's arithmetic.

Teams skip this step because it sounds like a labeling project. It doesn't have to be. A useful starter set is small — a few dozen questions — and grows over time:

  1. Start with real questions. Pull them from support tickets, chat logs, and search queries — the things users actually ask, phrased the way they actually ask them.
  2. Label the answering chunks. For each question, find the passage in your corpus that answers it. That pairing is one gold row. (The RAG Retrieval Grader ships a gold-set template, a labeling workflow, and a helper that bootstraps a starter set from your existing documents, so you're not starting from a blank page.)
  3. Add every production miss. Each time a user catches a bad answer, trace it to the retrieval miss and add that question to the gold set. Your test set grows exactly where your system is weakest.

Why not just use an LLM judge? Judge-based frameworks are strong for grading generation quality — faithfulness, groundedness. But for retrieval they're costly, rate-limited, and variable from run to run. Scoring retrieval against a gold set is deterministic: exact, reproducible, free to rerun, and therefore gateable in CI. Use both — the judge for the answer, the gold set for the search.

Grading retrieval and gating it in CI

Once the gold set exists, the workflow is short. Run every gold question through your retriever. Score the results — Recall@k, Precision@k, MRR, nDCG. Compare against your thresholds and get a verdict: ship, hold, or fix. Then wire that same run into CI, so an embedding swap, a chunking tweak, or a re-index that drops retrieval below your bar fails the build instead of shipping quietly.

The verdict matters more than the dashboard. A number on a chart invites debate; a failed build forces a decision. Set a minimum Recall@k you're willing to live with, and let the gate hold the line while you experiment freely above it. This is the same baseline-and-gate discipline that prompt regression testing applies to prompts — applied to search. It's also what you rerun when the change wasn't yours at all: an embedding-model update moves retrieval the same way a deprecated or repointed model moves generation.

And when the grade comes back poor? The fix list is usually short: chunking strategy, embedding model, k value, or index hygiene. The point of measuring first is that you change one thing, re-grade, and know — instead of changing five things and hoping.

The honest boundary

A clean retrieval score does not certify your answers. The model can still misread a correct chunk, and no metric here catches a prompt that mangles good context — that's what output-level LLM evaluation is for. What retrieval grading proves is narrower and foundational: the model had the right material in front of it. Prove that first. It's the failure surface most RAG teams have never measured, and the root cause of most of what gets blamed on the model.

Grade your retrieval this week

The RAG Retrieval Grader scores your retrieval step with Recall@k, Precision@k, MRR, and nDCG against a gold set, returns a ship / hold / fix verdict, and lists every query where retrieval missed. Python and TypeScript engines with identical scores, adapters for Pinecone, Weaviate, Qdrant, Chroma, pgvector, LangChain, and LlamaIndex, a gold-set bootstrapper, and a ready-to-use GitHub Actions workflow to gate CI. Deterministic — no judge tokens burned.

Get the RAG Retrieval Grader — $89 →

Keep going: retrieval is one of four quiet failure surfaces. The full map is in AI reliability: how to know your AI feature actually works. If your pipeline passes retrieval but you still can't say whether v2 beats v1, the next read is LLM evaluation: how to test AI outputs.


Decision Guide

Grade retrieval now if: users report made-up or stale answers from your RAG app, or you're about to change anything upstream — embedding model, chunk size, k, or the index itself.

Wait a beat if: you haven't shipped RAG yet. But collect real user questions from day one — they're your future gold set.

Best first step: take ten real user questions, run them through your retriever, and check by hand whether the answering chunk appears in the top k. If two or more miss, you've found your problem — and your first ten gold rows.

FAQ

What is RAG evaluation?

RAG evaluation is measuring the quality of a retrieval-augmented generation pipeline — most importantly the retrieval step. You score whether the chunks fed to the model are relevant and complete, using metrics like Recall@k, Precision@k, MRR, and nDCG against a gold set of question-to-chunk pairs.

Why does my RAG app hallucinate if the documents are correct?

Usually because retrieval handed the model the wrong chunks — or missed the right one — and the model filled the gap with a confident guess. The documents being correct doesn't help if the answering passage never reached the context window. Grade retrieval before blaming the model.

What is a gold set?

A gold set is your retrieval ground truth: real questions, each paired with the chunk or chunks that should be retrieved to answer them. It's what makes retrieval scoring exact instead of subjective. A few dozen rows built from real user questions is enough to start, and every production miss you add makes it stronger.

What's a good Recall@k score?

There's no honest universal number — it depends on your corpus, your chunking, and what a miss costs you. The useful move is to baseline your current score, set a minimum you're willing to ship at, and gate CI on it. The trend against your own baseline matters more than anyone else's benchmark.

Do I need an LLM judge to evaluate RAG?

Not for retrieval. Against a gold set, retrieval scoring is deterministic — exact, reproducible, and free to run, which is what lets it gate a build. LLM judges earn their cost on generation quality (faithfulness, groundedness), where the grading is genuinely subjective. The two are complementary, not competing.

Can retrieval quality drop without any code change?

Yes. Re-indexing new documents shifts what ranks where, and swapping an embedding model changes the whole similarity space. Both can drop retrieval quality with zero code diff — which is exactly why the grade belongs in CI, run on every change, rather than as a one-time audit.

What does the RAG Retrieval Grader include?

Metrics engines in Python and TypeScript (identical scores), a gold-set template plus a bootstrapper that builds a starter set from your documents, adapters for common vector DBs and frameworks, a CLI, a ship / hold / fix verdict with every miss listed, and a GitHub Actions workflow to gate retrieval in CI. One-time $89 at redhub.ai.