Check whether a prompt or model change made things worse
Run the suite before and after the change, with a few attempts per case, then compare:
behavtest run suite.json --repeat 5 --label before
# change the prompt, the model, the retrieval settings...
behavtest run suite.json --repeat 5 --label after
behavtest compare --fail-on-regression
compare lists regressed, improved, flaky and changed cases, with pass rates and p-values, and an overall verdict. See compare.