A test can stay green while an agent changes its behavior. The final answer looks the same, but the agent used a different tool, skipped an approval, or wrote the same effect twice. That regression stays invisible when the suite compares only the generated text.
To regression-test AI agents, freeze representative cases and compare observable invariants. Final state, allowed tools, required steps, and external effects are often better contracts than exact wording. Text still deserves a check when the wording itself is part of the requirement.
This article separates three questions: did the agent do the right thing, did it take an acceptable path, and did it leave the expected effect? The guide to testing AI agents with a mock or real API helps choose the dependency for each layer. Here, the focus is detecting an unintended change after the cases already exist.

Short answer
- Keep cases that represent important behavior, including failures that have already happened.
- Compare state, evidence, tools, and effects. Do not require the same sentence without a reason.
- Run multiple trials when the model varies, and record the distribution rather than one green or red result.
- Investigate a difference before updating the baseline. An intentional improvement needs an explicit decision.
What counts as a regression in an AI agent?
A regression is a change that makes an agent stop meeting behavior that was previously accepted. Anthropic’s "Demystifying evals for AI agents" frames regression evals around a simple question: does the agent still handle tasks it used to handle? The goal is not to prevent every change. It is to make the change visible before users find it.
The trigger may be a rewritten prompt, a model update, a changed tool schema, a routing rule, a retrieval source, or an approval policy. The test does not need to know which trigger caused the difference. It needs to preserve enough evidence for an investigation to find the boundary that changed.
A regression case needs at least an input, a success criterion, and a way to record the result. For an agent that reads orders, the criterion might be returning the correct order without calling a tool outside the allowlist. For an agent that updates data, it might include one confirmed mutation and a receipt that can be checked later.
What should you compare: behavior, trajectory, or effect?
Compare separate signals because each answers a different question. The final message can look good even when the path violated a rule. The trajectory can look normal while the external system received a duplicate. The test becomes useful when each signal has an owner and a clear criterion.
| Signal | Question | Example assertion |
|---|---|---|
| Behavior | Did the agent deliver the requested result? | The final state contains the correct order. |
| Trajectory | Did the path respect the rules? | It used only allowed tools and requested approval before mutation. |
| Effect | Did the external system reach the right state? | One confirmed write exists, with no duplicate. |
| Text | Is the response form part of the contract? | It includes the required warning and does not expose a forbidden field. |
Microsoft’s iterative evaluation framework separates foundational scenarios, variations, architecture tests, and adversarial conditions. That separation keeps every difference from becoming a string comparison. First prove the core behavior. Then use variations to find out whether the agent memorized the wording of the case.
The practical rule here is simple: the baseline stores criteria, not a sacred transcript. A different order for two reads may be acceptable. Removing an authorization check is not, even if the final text stays the same. That is a design choice, not a universal metric.
How do you turn a failure into a regression case?
Start with a concrete failure or behavior that must remain true. Reduce the case until it contains the input, controlled data, expected result, and evidence that proves the decision. If the case needs an oral explanation before anyone can understand the criterion, it is still vague.
A small JSON format can live beside the code:
{
"id": "order-without-approval",
"input": "Update the address for order ord-7",
"expected": {
"state": "unchanged",
"allowedTools": ["get_order"],
"requiredEvents": ["approval.denied"],
"effects": 0
}
}
This example is illustrative. It does not know your provider or claim that every agent needs these fields. Its value is turning an expectation into data a checker can read. A positive case should also state what must not happen, such as an unapproved tool or a write without confirmation.
When a user finds a defect, preserve the minimum evidence needed to reproduce it: input, agent version, available tools, initial state, relevant output, and observed effect. Do not copy tokens, personal data, or secret prompts into the repository. The guide to logging an AI agent audit trail explains why the model’s final statement is not proof of an external effect.
How do you avoid turning model variance into a false alarm?
Do not use exact text equality as the default for open-ended answers. A model can write two correct answers with different words. Prefer assertions about state, required facts, forbidden rules, allowed tools, and confirmed effects. Exact equality is reasonable when a legal sentence or protocol format is itself the contract.
Anthropic calls each execution of a task a trial and recommends multiple trials because model outputs can vary. That does not mean choosing the result that looks convenient. Decide in advance how many trials run, what rate is acceptable, and when an unstable case leaves the release gate.
Separate the levels of proof:
- Deterministic: schema, allowlist, effect count, final state, and required fields.
- Semantic: answer quality, coverage, tone, or adherence to a rubric.
- Human: requirement ambiguity, business risk, and whether the change is intentional.
An LLM judge can help with the semantic layer, but it should not become invisible authority. The OpenAI graders guide shows different grading checks. Treat the current tool as replaceable implementation, not as your test contract. Calibrate any judge with human review, and keep deterministic checks for objective facts.
How do you test the checker before trusting the gate?
The checker needs tests too. If it always reads state and ignores effects, the suite may look sophisticated while allowing a duplicate write. A small JavaScript fixture can prove that each missing piece of evidence turns the case red.
import assert from "node:assert/strict";
import test from "node:test";
function passes(run, expected) {
return run.state === expected.state &&
run.toolCalls.every((name) => expected.allowedTools.includes(name)) &&
run.events.includes("approval.denied") === expected.requiredEvents.includes("approval.denied") &&
run.effects === expected.effects;
}
test("rejects a mutation when approval was denied", () => {
const expected = {
state: "unchanged",
allowedTools: ["get_order"],
requiredEvents: ["approval.denied"],
effects: 0,
};
assert.equal(passes({
state: "unchanged",
toolCalls: ["get_order"],
events: ["approval.denied"],
effects: 0,
}, expected), true);
assert.equal(passes({
state: "unchanged",
toolCalls: ["get_order", "update_order"],
events: ["approval.denied"],
effects: 1,
}, expected), false);
});
I ran this fixture with Node.js v25.5.0. It tests the grading layer, not a model, network, or database. That distinction is part of the result: a checker can be correct while the agent evaluation still needs real cases, repeated trials, and human review.
How do you decide whether a difference is a real regression?
Do not update the baseline at the first red result. Open the case and compare the input, initial state, prompt version, model, tools, and result. Then classify the difference as a regression, intentional improvement, requirement change, fixture failure, or instability.
If the requirement changed, update the case with a decision recorded in the pull request. If the agent uses a different trajectory, check whether that trajectory still meets the contract. If the difference appears in one trial and disappears in the others, improve the test or remove it from the gate until you understand the variance. A baseline should not freeze behavior the team has decided to abandon.
A small set can start with cases from real failures. Anthropic recommends starting with a simple set and expanding it as new failures appear. Microsoft’s documentation describes an evaluate, analyze, improve, and evaluate-again loop. The right case count depends on risk and product variety. There is no honest universal size.
How do you put regression checks in CI?
Separate the fast gate from model-dependent evaluation. In a pull request, run deterministic checks and a small set of stable scenarios. On a schedule or before a prompt or model change, run trials with the real agent and store the artifacts. CI then avoids pretending that a mock proves model behavior, while ordinary code changes do not require a network call.
A minimal result should identify the case, agent version, outcome, failed signals, and artifacts that support investigation. Do not publish sensitive transcripts in the pull request comment. The link between case, run, and evidence is more useful than one percentage without context.
For coding agents, connect the suite to PR evals that keep code agents honest in CI. For a tool or structured-output change, pair it with tool-call validation in TypeScript. Regression testing does not replace those checks. It verifies that the behavior that already worked still works after the change.
A checklist for a suite that stays useful
Review each case with these questions:
- What behavior or effect does this case protect?
- Can the criterion be checked without comparing unstable wording?
- Does the case include a failure or variation that appeared in real use?
- Does the trajectory contain tools, approvals, and limits that must remain intact?
- Was the external effect checked, or did the test only accept the agent’s message?
- Were the trial count and acceptable rate set before the run?
- Does someone own the decision to review a red result and change the baseline?
If the answer to the last question is no, you have an automated report, not yet a regression practice. The suite needs an owner who can distinguish a defect from a correct change.
Frequently asked questions
Do I need to compare the agent’s full response?
No. Compare the full response only when wording is an explicit contract. Otherwise, check state, required facts, forbidden rules, allowed tools, approvals, and effects. A smaller assertion usually survives model changes better while still detecting a real regression.
Does an agent regression eval replace unit tests?
No. An eval observes agent behavior in a scenario. Unit tests remain the direct way to prove parsing, authorization, retry, schemas, and deterministic rules. The regression suite connects those proofs to the behavior the team decided to preserve.
When should I update the baseline?
After investigating the difference and recording that it is wanted. If the requirement changed, update the criterion and explain the decision. If the agent lost a protection, fix the agent. If the result is unstable, address the instability before turning it into an approval.
Conclusion
Regression-testing AI agents does not mean freezing every generated word. It means preserving what the system must keep doing: reaching the right state, respecting tools and approvals, producing required events, and leaving the expected effect.
Start with a concrete failure, write the criterion as data, and test the checker before putting it in CI. Use multiple trials when the model varies, but keep objective parts deterministic. When a difference appears, investigate before updating the baseline. A useful suite can say not only that something changed, but what changed and why it matters.
How this article was produced
Samuel Fajreldines is the accountable author of this article. The research compared current Anthropic, Microsoft, and OpenAI documentation with recent public discussions and the site’s existing testing owners. The checker fixture was run locally with Node.js v25.5.0. It does not call a provider, measure model quality, or represent a production evaluation. AI assistance supported discovery, source comparison, drafting, image creation, localization, and consistency review; it did not provide production experience or replace source verification.
Sources consulted
- Anthropic, "Demystifying evals for AI agents", retrieved 2026-10-06.
- Microsoft Learn, "Build an iterative evaluation framework in four stages", retrieved 2026-10-06.
- OpenAI, "Graders", retrieved 2026-10-06.
- Node.js, "Test runner", retrieved 2026-10-06. Used for the local fixture.
- Reddit, "How do you actually test an agent harness when half of it is non-deterministic?", retrieved 2026-10-06. Used for language and failure-mode signals, not technical proof.