BehavTest

BehavTest › Learn

Behavioral regression testing

A behavioral regression test checks that a system still does what it is supposed to do after a change, judged by properties of what it does rather than by the exact output it produces. When the system is nondeterministic, as every LLM application is, "still does it" becomes a rate: the test asks whether the behavior now happens less often than before, by more than chance would explain.

This page builds the idea up from the ground: what counts as behavior, why exact outputs are the wrong thing to compare, why one run of each version can't settle the question, and what a statistically sound answer looks like.

Output versus behavior

Take a support assistant and one question: "Can I return an item after 40 days?" The shop accepts returns for 30 days. Here are three answers the assistant might give:

AnswerSame wording as before?Correct behavior?
"No: returns are accepted within 30 days."yes (the old answer)yes
"Our policy allows returns within 30 days, so 40 days is too late."noyes
"Sure, just bring it back with the receipt!"nono

A test that compares exact output (a snapshot) treats the second and third answers the same way: both differ from the stored text. The first is a harmless rephrasing and the second is a real regression, but the snapshot can't tell them apart. The behavior you care about is a property of the answer: it states the 30-day limit and doesn't promise an exception. A behavioral test checks that property, so the rephrasing passes and the wrong promise fails.

Behavior isn't only about the words. For an AI application it usually includes:

Four terms, defined

Why a single run of each version can't tell you

Suppose a case passes on the old version and fails on the new one. Is that a regression? With a nondeterministic system you can't know from one attempt each. If a case passes 80% of the time, two versions that behave identically will disagree on a single attempt about a third of the time.

The noise is large even with several attempts. In BehavTest's nondeterministic example, a bot that answers correctly 90% of the time was run twice, with nothing changed, 10 attempts per case:

  ✗ regressed opening-hours 10/10 → 7/10 100% → 70%  p=0.211
      not statistically significant at this sample size
  ✓ improved  refund-window 7/10 → 10/10 70% → 100%  p=0.211
      not statistically significant at this sample size

One case "dropped" 30 points and another "rose" 30 points, and both movements are pure chance. Anyone comparing single runs would spend the afternoon chasing them.

Repeated testing: behavior as a pass rate

The fix is to measure each case several times and treat its behavior as a pass rate with an uncertainty range. A 95% Wilson score interval shows how much a pass rate can be trusted:

Attempts passedPass rate95% interval
1 of 1100%21% to 100%
3 of 3100%44% to 100%
7 of 1070%40% to 89%
9 of 1090%60% to 98%
45 of 5090%79% to 96%

One passing attempt is consistent with a true pass rate as low as 21%. The intervals narrow as attempts grow, which is exactly the trade-off you manage in practice: more attempts give sharper answers and cost more model calls.

Repeats also produce a category that single runs can't: a flaky case, one that passes on some attempts and fails on others. Flakiness is information. It tells you the behavior is not reliable even before anything changes.

Statistical confidence: is the drop bigger than the noise?

With pass rates on both sides of a change, the question becomes a standard statistical one. Two levels matter:

On the same example, when the bot's accuracy really dropped from 90% to 60%, most individual cases were still "not significant" at 10 attempts each, but the overall test was not in doubt:

  ✗ regressed shipping-time 10/10 → 5/10 100% → 50%  p=0.033 significant
  ✗ regressed support-email 9/10 → 5/10 90% → 50%  p=0.141
      not statistically significant at this sample size
  ...
  attempt pass rate  85% [76%–91%] → 56% [45%–67%]  (8 comparable cases; descriptive)
  overall change     mean per case -28.7 pts, 95% CI [-41.3 pts, -16.3 pts], p=<0.0001 → significant regression

Compare the unchanged run from the previous section: its overall change was -3.7 points with an interval from -15 to +7.5 points and p ≈ 0.67, which is noise. That contrast is the whole point of behavioral regression testing: the same kind of per-case wobble, two very different verdicts, decided by evidence rather than by eye.

The workflow

  test cases + expected behaviors
              |
   run each case N times (baseline)        run each case N times (after the change)
              |                                          |
      pass rate per case                         pass rate per case
              \__________________  __________________/
                                 \/
         compare: per-case exact test, overall permutation test
                                 |
       regressed / improved / flaky / not significant  ->  pass or fail the check

A baseline can be the last good run, a run of the main branch, or a small file committed to the repository so that every pull request is compared with the same reference.

Doing it with BehavTest

BehavTest implements exactly this loop. You describe cases and their expected behaviors in a suite, it runs each case --repeat times through your application, and behavtest compare applies the per-case and overall tests:

behavtest run suite.json --repeat 10 --label before
# change the prompt, the model or the retrieval settings
behavtest run suite.json --repeat 10 --label after
behavtest compare --fail-on-regression --significant-only

Reproduce the numbers on this page with the keyless example: behavtest run examples/nondeterministic/suite.mjs --label before, then SEED=2 behavtest run examples/nondeterministic/suite.mjs (same bot) or BOT_ACCURACY=0.6 SEED=3 behavtest run examples/nondeterministic/suite.mjs (worse bot), then behavtest compare. The overall p-values are Monte Carlo estimates, so their last digits vary between runs. See how BehavTest works, how compare decides and the full statistical reference.

What behavioral regression testing does not tell you