BehavTest

BehavTest › Learn

AI regression testing

Regression testing answers one question: did a change break something that used to work? In traditional software the answer comes from rerunning tests whose expected results never change. AI applications keep the question but break most of the machinery behind the answer. This page contrasts the two, lists where regressions in AI applications actually come from, and shows what an AI regression test has to do differently.

Regression testing in traditional software

A conventional regression suite rests on three assumptions:

Under those assumptions a single red test is strong evidence, and the tooling (unit test runners, snapshot files, CI that fails on any failure) follows naturally.

What changes for AI applications

Traditional softwareAI application
Same input, same output?YesNo: sampled outputs vary between calls
Expected resultExact value or snapshotA property of the answer: a fact, a rubric, a tool call, grounding
Who decides pass/failAn equality checkA deterministic check where possible, otherwise a model (an LLM judge)
A test that sometimes failsA bug in the testOften the true behavior of the system: a case with a pass rate below 100%
Evidence neededOne runSeveral attempts per case, compared statistically
A change "breaks" something whenAny test turns redThe pass rate drops by more than the noise

The last row is the core difference. An AI regression test cannot treat one failed attempt as proof, because a correct system fails some attempts by chance, and it cannot treat one passed attempt as proof either. It has to estimate how often each behavior holds, before and after, and decide whether the difference is real. Behavioral regression testing explains the statistics.

Where regressions in AI applications come from

Many regressions arrive without anyone touching the code that looks responsible:

A useful AI regression suite is the set of behaviors you'd be embarrassed to lose when any of these happens.

Behavioral assertions

Instead of asserting an exact output, an AI regression test asserts properties. Common ones, from cheapest to most expensive:

Each assertion turns one attempt into pass or fail; repeating the case turns those into a pass rate.

Evaluation versus regression testing

The two are easy to confuse because they share the same parts (datasets, scorers, judges):

A system can score 72% on an evaluation and pass its regression tests every day, as long as it keeps scoring 72%. LLM evaluation covers the difference in detail.

Practical examples

Switching to a cheaper model. Run the same suite through the current and the candidate model with enough attempts, and compare pass rates case by case, plus latency and cost. With BehavTest, matrix runs do this in one command: each model is a variant, and every variant is compared with a reference using the same statistics.

Editing a system prompt. Record a baseline before the edit, rerun after, and compare. Cases that regressed significantly are the ones to read; cases that only wobble are noise. See check whether a prompt or model change made things worse.

Changing RAG chunking. List, for each case, the document ids a correct retrieval returns, and score retrieval deterministically (recall at k) alongside a grounding check on the answer. A drop in retrieval recall with stable answers usually means the model is compensating from its own knowledge, which is its own risk. See RAG.

Refactoring an agent's tools. Assert the tool calls, not the final text: which tool, which arguments, how many steps. See traces.

Where BehavTest fits

BehavTest is a regression-testing tool for this setting: it runs your cases through the real application (an HTTP endpoint, an OpenAI-compatible or Anthropic model, or a TypeScript function), repeats them, scores behavior with deterministic checks, trace checks, RAG checks or an LLM judge, and compares runs with significance tests, locally or in CI. It is not a benchmark suite, a red-teaming tool or a production monitor, and it doesn't replace ordinary unit tests for the deterministic code around your model. When to use BehavTest lists the fit in detail.