LLM Evaluation: How to Test AI Outputs
RedHub AI Editorialupdated September 20, 20268 min read

Jump to a section9
TL;DR
- The problem: most teams judge LLM output by reading a few responses and nodding. That's a demo, not a test — it leaves no record, covers no edge cases, and can't answer "is v2 better than v1?"
- What an eval is: three parts — a test set of real cases with expected outcomes, a scorer that grades each output, and a threshold that turns scores into a pass/fail verdict.
- How to work: eval-driven development. Write the eval before you tune the prompt, run it on every change, and let the score — not the vibe — decide what ships.
- The honest part: an eval can't prove your AI is always right. It proves your changes don't make things worse, and it turns quality into a number you can trend. That's the whole game.
What is LLM evaluation?
LLM evaluation is testing your AI's outputs systematically instead of eyeballing them. You build a test set of real inputs with expected outcomes, run your prompt and model against it, score every output with a defined scorer, and compare against a threshold. The result is a repeatable verdict you can track across every prompt edit and model swap.
Best for: dev teams shipping LLM features who need evals in their own repo, on their own infra — the setup the Prompt Evaluation & Versioning System ships ready to drop in.
Ask a team how they know their AI feature works and you'll usually get some version of "we tested it." Press on what that means and it comes down to: someone typed in a handful of inputs, read the outputs, and they looked good.
"It looked good" is the most expensive sentence in AI development. It's not a lie — the outputs probably did look good. But it measured nothing, recorded nothing, and covered almost nothing. Two weeks later someone edits the prompt, the outputs shift, and there is no way to say whether things got better or worse — because "good" was never defined and never written down. LLM evaluation is how you replace that sentence with a number.
This is the fourth quiet failure surface in our AI reliability guide — un-measured output — and it's the one that makes the other three unfixable. You can't debug retrieval, catch regressions, or tune an agent if you can't score the result.
One scope note before we start: this post is about evaluating your own system's outputs — the engineering discipline. Choosing which AI vendor or tool to buy is a different problem with different criteria, and it's not what "eval" means here.
An eval has exactly three parts
Strip away the tooling and every eval is the same machine:
- A test set. Real inputs paired with expected outcomes. Sometimes the expectation is exact ("this support ticket should be classified billing"), sometimes it's a property ("the summary must mention the deadline and stay under 100 words"). Either way, it's written down before the run.
- A scorer. A function that grades one output against one expectation and returns a score. Scorers range from exact match to rule checks to embedding similarity to an LLM judge — more on the tradeoffs below.
- A threshold. The line that turns scores into a decision: pass rate at or above the bar, ship; below it, hold. Without a threshold, an eval is a report nobody acts on.
That's it. No platform required, no PhD required. A folder of test cases, a scoring script, and a number your CI can compare against a bar.
Building your first test set
The test set is where most teams stall, because it sounds like a research project. It isn't. The working recipe:
- Start from production, not imagination. Pull real inputs from logs, tickets, and user sessions. Real inputs are messier, shorter, more misspelled, and more hostile than anything you'd invent — which is exactly why they belong in the set.
- Start small and real. A few dozen well-chosen cases beat a thousand synthetic ones. Cover the common cases first, then the edges: ambiguous inputs, empty inputs, adversarial phrasing, and the longest input you legitimately expect.
- Label the expectation, not the exact wording. For generative tasks, define what a passing output must contain, must not contain, and how long it can be. You're grading properties, not prose style.
- Feed it every failure. Each production miss becomes a permanent test case. This is the compounding move: your eval grows exactly where your system is weakest, and a bug fixed once can never quietly return.
Choosing scorers: deterministic first, judge last
| Scorer | Best for | The tradeoff |
|---|---|---|
| Exact match | Classification, extraction, structured fields | Cheap, perfectly reproducible — but only fits tasks with one right answer |
| Rule checks | Format, required content, forbidden content, length, latency | Deterministic and free; catches breakage, not eloquence |
| Embedding similarity | "Close enough in meaning" comparisons | Handles paraphrase; the threshold takes tuning |
| LLM judge | Genuinely subjective quality — tone, helpfulness | Flexible but costs per run and varies between runs; audit it, don't gate on it alone |
The ordering matters. Deterministic scorers — exact match and rules — are reproducible and free, which makes them the foundation you can gate a build on. An LLM judge grading another LLM adds cost and its own variance; use it where subjective quality genuinely matters, and spot-check its verdicts against human judgment before you trust it. A useful default: deterministic checks decide the gate, the judge adds color.
Eval-driven development: the loop that replaces vibes
Once the eval exists, the working rhythm flips. Instead of tune-then-eyeball, you eval-then-tune:
- Write the eval first. Before touching the prompt, define what passing looks like in test cases. This forces the question "what is this feature supposed to do?" to get a written answer.
- Baseline the current behavior. Run the eval, record the score. This number is now the thing to beat.
- Change one thing, rerun, compare. Prompt edit, model swap, parameter change — the eval says whether it helped, hurt, or did nothing. No more arguing from three cherry-picked outputs.
- Gate the merge. Wire the eval into CI so a change that drops the score below your bar can't ship. This is where evals meet prompt regression testing — the baseline diff that catches what a score summary can miss.
- Version everything. Prompts get versions, changelogs, and owners, like code — because a prompt with an eval score attached is finally a managed artifact instead of a string someone edits at midnight.
Why "it looked good" fails as a test, in one line: it has no test set (you tried the inputs you thought of), no scorer (your mood that afternoon), and no threshold (nodding). An eval is the same activity with all three written down — which is what makes it repeatable, comparable, and honest.
The honest boundary
An eval cannot prove your AI is always right. The test set is a sample, not the universe; the model stays probabilistic; and a pass rate is a measurement, not a warranty. What an eval proves is narrower and far more useful: this change didn't make things worse, this version beats that version on the cases we care about, and quality is trending in a direction you can see. Teams that have that stop debating outputs in Slack and start reading a chart. That's the entire trade — and it's a good one.
Skip the week of eval plumbing
The Prompt Evaluation & Versioning System is the drop-in starting point: a TypeScript + Python eval framework with adapters for the OpenAI, Anthropic, and Google AI SDKs, a golden-dataset runner with pluggable scorers, CI integration, a dashboard template for trending runs over time, prompt versioning conventions (semver, changelogs, rollback), and a Notion war-room for the non-engineers in the loop. Runs in your repo with your API keys — no telemetry, no platform lock-in.
Get the Prompt Eval System — $49 →Ready to cover every surface — retrieval, regressions, security, and agents — in one pass? Step up to the AI Reliability Bundle — $329 for the four CI-ready tools plus the connective playbook.
Keep going: the full four-surface program is in AI reliability: how to know your AI feature actually works. To put a hard baseline-diff gate behind your eval, read prompt regression testing. The same eval is what you rerun when the provider moves the model rather than when you move the prompt — that case is model deprecation. And if your outputs depend on retrieval, score that layer with its own metrics: RAG evaluation.
Decision Guide
Build an eval now if: an LLM feature is in front of users, more than one person edits prompts, or a model migration is coming. Any of those without an eval means changes ship on faith.
Wait a beat if: you're still prototyping and the feature changes shape daily. But save every interesting input you hit — that pile becomes your test set the week you stabilize.
Best first step: write twenty test cases from real inputs, define pass/fail for each in one sentence, and run them against your current prompt. The score you get is your baseline — and probably the first real number your AI feature has ever had.
FAQ
What is LLM evaluation?
LLM evaluation is systematically testing your AI system's outputs: a test set of real inputs with expected outcomes, a scorer that grades each output, and a threshold that turns scores into a pass/fail verdict. Run on every change, it replaces "it looked good" with a number you can compare across versions.
What's the difference between an eval and just testing manually?
Manual testing has no fixed test set, no defined scorer, and no threshold — so it can't be repeated, compared, or trended. An eval writes all three down. The same cases run every time, the same rules grade them, and the same bar decides the verdict, which is what makes version-to-version comparison possible.
How big does a test set need to be?
Smaller than you think to start: a few dozen well-chosen real cases covering common inputs, edge cases, and past failures. What matters more than size is growth — every production failure becomes a permanent test case, so the set compounds exactly where your system is weakest.
Should I use an LLM to judge my LLM's outputs?
Sparingly. LLM judges are useful for genuinely subjective qualities like tone, but they cost money per run and can disagree with themselves between runs. Build the foundation on deterministic scorers — exact match, rule checks — which are free and reproducible, and let those gate the build. Add a judge for color, and audit it against human judgment.
What is eval-driven development?
Writing the eval before tuning the prompt, then letting the score drive every change: baseline current behavior, change one thing, rerun, compare, and gate the merge in CI if the score drops. It's test-driven development adapted for systems whose outputs vary — the eval defines "working" so changes stop being arguments.
Does an eval catch model updates breaking my prompts?
Yes — that's one of its highest-value moments. When a provider updates a model or you migrate to a new one, rerunning the eval shows exactly which cases shifted, instead of discovering the breakage from users. Pair it with a committed baseline diff (see prompt regression testing) so any regression fails the build.
Do I need an eval platform, or can this live in my repo?
It can live in your repo. The core machine — test cases, scorers, thresholds, CI wiring — is files and scripts, not infrastructure. The Prompt Evaluation & Versioning System ($49) ships that setup with SDK adapters, a dashboard, and versioning conventions, with no telemetry and no data leaving your stack. Platforms earn their keep later, at scale — and the kit doesn't block that path.