Compare runs: what regressed, and is it real?
behavtest run suite.json --repeat 5 --label prompt-v6 # before your change
# ...edit the prompt / swap the model...
behavtest run suite.json --repeat 5 --label prompt-v7 # after
behavtest compare # latest run vs the one before it
behavtest compare · support-bot
base f033e1c8 2026-09-21 14:53 prompt-v6
head 22ec5145 2026-09-21 14:53 prompt-v7
✗ regressed author-of-hamlet 5/5 → 0/5 100% → 0% p=0.008 significant
✗ regressed symbol-for-gold 5/5 → 2/5 100% → 40% p=0.167
not statistically significant at this sample size
✓ improved is-pluto-a-planet 0/5 → 5/5 0% → 100% p=0.008 significant
~ flaky largest-ocean 3/5 → 3/5 60% → 60%
6 unchanged cases hidden (use --all to list them)
attempt pass rate 88% [78%–94%] → 67% [54%–77%] (12 comparable cases; descriptive)
overall change mean per case -21.7 pts, 95% CI [-28.3 pts, -15.0 pts], p=0.0015 → significant regression
behavtest compare takes [base] [head]: run ids (unique prefixes work) or run files. With one run id it compares that run with the run before it; with one run file, that file (as the baseline) with the latest run of its suite; with none, the latest two. Add --fail-on-regression to make it a CI gate (exit 1), --json / --md to write the result, and --all to list unchanged cases. behavtest report <run> --against <base> --out report.html writes the same comparison as a single-file HTML report.
How it decides. Model outputs are random, so one run each is rarely enough to call a regression. BehavTest is explicit about what it knows:
- Per case it compares pass rates with Wilson 95% intervals, and runs Fisher's exact test. A change is significant only when p < 0.05. That takes several attempts per case: with 3 attempts per side even 3/3 → 0/3 is p = 0.1. Changes on a single attempt are still listed, flagged "could be noise, re-run with
--repeat". - Overall it runs a paired permutation test, stratified by case, on the mean change in pass rate, with a within-case bootstrap for the interval. The question a gate asks is "on this suite, did the pass rate move by more than the pipeline's sampling noise?", so the randomness that matters is within each case, not which cases happen to exist. With one attempt per case this reduces to an exact sign test on the cases that flipped: six one-way flips are significant (p = 0.031), five are not (p = 0.063).
- Never compared: a case whose definition changed between the runs (
modified, which includes a different judge model for judged cases), a case in only one run (new/removed), and a case with an errored attempt (errored: no verdict). They are listed, never counted as regressions.
--fail-on-regression fails on any regressed case (significant or not, because single-attempt suites can't do better), on any case that errored in the head run, and on a significant overall drop. --significant-only ignores regressions that aren't statistically significant. Every method, assumption and limit is documented in the statistical reference. The tests check the statistics against textbook reference values and, by simulation, that the overall test rejects under 9% of the time when nothing changed and over 95% of the time for a real drop.