BehavTest

BehavTest › Reference

Compare runs: what regressed, and is it real?

behavtest run suite.json --repeat 5 --label prompt-v6     # before your change
# ...edit the prompt / swap the model...
behavtest run suite.json --repeat 5 --label prompt-v7     # after
behavtest compare                                         # latest run vs the one before it
behavtest compare · support-bot
  base  f033e1c8  2026-09-21 14:53  prompt-v6
  head  22ec5145  2026-09-21 14:53  prompt-v7

  ✗ regressed author-of-hamlet    5/5 → 0/5   100% → 0%  p=0.008 significant
  ✗ regressed symbol-for-gold     5/5 → 2/5   100% → 40%  p=0.167
      not statistically significant at this sample size
  ✓ improved  is-pluto-a-planet   0/5 → 5/5   0% → 100%  p=0.008 significant
  ~ flaky     largest-ocean       3/5 → 3/5   60% → 60%
  6 unchanged cases hidden (use --all to list them)

  attempt pass rate  88% [78%–94%] → 67% [54%–77%]  (12 comparable cases; descriptive)
  overall change     mean per case -21.7 pts, 95% CI [-28.3 pts, -15.0 pts], p=0.0015 → significant regression

behavtest compare takes [base] [head]: run ids (unique prefixes work) or run files. With one run id it compares that run with the run before it; with one run file, that file (as the baseline) with the latest run of its suite; with none, the latest two. Add --fail-on-regression to make it a CI gate (exit 1), --json / --md to write the result, and --all to list unchanged cases. behavtest report <run> --against <base> --out report.html writes the same comparison as a single-file HTML report.

How it decides. Model outputs are random, so one run each is rarely enough to call a regression. BehavTest is explicit about what it knows:

--fail-on-regression fails on any regressed case (significant or not, because single-attempt suites can't do better), on any case that errored in the head run, and on a significant overall drop. --significant-only ignores regressions that aren't statistically significant. Every method, assumption and limit is documented in the statistical reference. The tests check the statistics against textbook reference values and, by simulation, that the overall test rejects under 9% of the time when nothing changed and over 95% of the time for a real drop.