LLM Regression Testing: Prove a New Model Is Better
RedHub AI Editorialupdated October 2, 20267 min read

Jump to a section8
Before a new model takes over a workflow, LLM regression testing runs it, behind the same prompts, against a frozen set of your own representative tasks, and compares the results with the model you have now. It turns "the new model seems better" into evidence. The question it answers is narrower than the one vendors answer. A vendor can tell you a model is stronger in general. Only your own task set can tell you whether it is better at your work, and whether it got worse anywhere that matters.
TL;DR: Freeze a versioned set of real tasks, run the current model and the candidate through it with the same prompts, and score more than answer quality: format, tools, latency, cost, safety and consistency across runs. Then apply release gates per metric. A higher overall pass rate is not a pass if the candidate breaks a high-impact task the old model handled. For agents, the Agent Reliability Harness ($149) scores the whole run and gates CI on ship, hold or fix. Start with the pillar: AI Model Lifecycle: Your AI Stack Has an Expiration Date.
Same prompts, new model
Our post on prompt regression testing covers the case where you edit a prompt and hold the model fixed. This post is the opposite case: you hold the prompts, tools and workflow fixed and change only the model. That is the test you run when a provider retires your model, ships a stronger one, or cuts a price. If you have never built a test set at all, start with our guide to LLM evaluation, then come back.
The model swap has one property the prompt edit does not. When you edit a prompt, you know what you changed. When you change the model, everything the model does changed at once, and the test set is the only map you have of where to look.
Build a golden task set
A golden task set is a fixed collection of real tasks with known good outcomes. Start with real work, not demos. Pull from normal volume, valuable edge cases, past failures and high-risk scenarios. For each task, record:
- The input and any context the model sees
- The expected outcome and the required format
- The tools it may use and the actions it must not take
- How it is scored
Version the set and freeze it for the comparison. If the tasks change every time the model changes, you lose the ability to compare results over time. New tasks go into the next version, and you compare models only within one version.
Measure more than answer quality
- Task success: did the workflow reach the correct business outcome?
- Format validity: did the output meet the schema, JSON, template or field rules?
- Tool reliability: did the agent pick valid tools with valid parameters?
- Latency: how long until an acceptable result?
- Cost: tokens, tool calls, retries and human review time.
- Safety: did the system follow policy and avoid prohibited actions?
- Stability: does it give the same quality of answer across repeated runs?
Stability is the one teams skip. A model that passes a task once and fails it on the second run has not passed it. Running each task more than once is the cheapest way to find out which kind you have.
A testing loop you can repeat
- Freeze the baseline. Current model, prompt and workflow, recorded as they run today.
- Run both. The golden set through the baseline and through the candidate.
- Score the results. Automatically where a rule can decide, and with a person where judgment is needed.
- Classify failures. By type, severity and business impact.
- Fix and rerun. Adjust prompts, routing, schemas, tools or policies, then run again.
- Pilot the candidate. Send a controlled slice of production traffic to the candidate.
- Feed it back. Watch for drift, and add each new real failure to the next version of the set.
Use gates, not gut feeling
A release gate is a threshold the candidate has to clear before it ships. Each metric gets its own gate, set before testing starts. These are examples; the right numbers depend on the workflow.
| Metric | Example release gate |
|---|---|
| Structured-output validity | At least 99% for critical schemas |
| Task completion | No regression on high-value representative tasks |
| Tool-action error rate | Below a defined safety threshold |
| Latency | Within the workflow service-level objective |
| Cost per accepted result | At or below target after review cost |
| Safety violations | Zero tolerance for defined high-impact violations |
A service-level objective is the response-time target the workflow already promises its users. Gates matter because an average can hide the one result that should stop the release.
The next move is step five of the loop: find out why the candidate misses deadline questions, fix the prompt or retrieval, and rerun. If the fix holds and the 2 high-impact tasks pass, the candidate has earned a pilot. If not, the old model stays until a candidate can clear every gate.
Who grades the grader
One common way to score open-ended answers is a second model, often called an LLM judge, because a person cannot read every output. That judge is a model too. It has a version, quirks and its own retirement date. If you swap the judge at the same time as the model under test, you no longer know which change moved the scores.
Two habits help, and neither fully solves it. Score with deterministic checks first (schema valid, field present, tool allowed, number matches) and use a judge only for what a rule cannot decide. And pin the judge's version, checking a sample of its scores against a person's whenever the judge itself has to change. How far to trust a judge on your tasks is a question only data from your tasks can answer.
The timing problem gets worse when a provider retires the model on a fixed date. The judge, the baseline and the candidate all have to be pinned before the old model stops answering, because a baseline you did not capture cannot be rebuilt later. Our post on model deprecation covers what changes when a retirement sets the schedule.
When the workflow is an agent, test the whole run
The Agent Reliability Harness evaluates AI agents at the trajectory level, meaning the whole sequence of steps: tool choice, argument validity, step efficiency, cost and policy. It gates CI on a ship, hold or fix verdict. Six deterministic evaluators, framework-agnostic, zero-dependency Python.
Get the Agent Reliability Harness — $149Pairs well with
The Prompt Regression Lab ($89) diffs every prompt change against a saved baseline, which is the check you want at step five, when the prompt edits you make for the new model need their own regression test. The Prompt Evaluation & Versioning System ($49) pairs an eval framework and a dashboard with a Notion war-room, for teams that treat prompts like code. When the candidate is a cheaper model, the Model Swap tab in the AI Unit-Economics & Token-Shock Exposure Kit ($59) takes the pass rates from your own test and computes cost per finished task for both models, returning ROUTE DOWN, KEEP or UNMEASURED.
More in this guide
What is LLM regression testing?
It is running a proposed change, such as a new model, prompt, tool or policy, against a frozen set of representative tasks before it reaches production, and comparing the results with the current setup.
How is this different from prompt regression testing?
Prompt regression testing holds the model fixed and checks a prompt edit. Testing a model swap holds the prompts and workflow fixed and changes the model. The mechanics are similar, but a model swap can change everything the model does at once.
How many tasks should a golden set have?
Enough to cover normal volume, valuable edge cases, past failures and high-risk scenarios for that workflow. Coverage counts for more than the raw number. Keep the set versioned and frozen for each comparison.
If the new model has a higher pass rate, should I switch?
Not on the pass rate alone. Check which tasks it broke as well as which it fixed. A candidate can score higher overall and still fail a high-impact task the current model handled, which a release gate should stop.
What should I measure besides accuracy?
Format validity, tool reliability, latency, cost including retries and review time, safety, and stability across repeated runs. A model that passes a task once and fails it on the next run has not passed it.
Can I use an AI model to score the results?
Yes, for open-ended answers, but the judge is a model with its own versions and quirks. Use deterministic checks first, pin the judge's version, and compare a sample of its scores with a person's when the judge changes.


The gate this post refers to, drawn from the tool’s own logic. See the tool.