Non-determinism: repeat your cases
behavtest run suite.json --repeat 5
Each case runs 5 times as separate attempts. A case where every attempt passes is passed, none failed, and a mix is flaky, which exits non-zero. One green run of a stochastic pipeline proves little; repeated attempts show you the real pass rate. Every attempt is stored, and behavtest compare uses them to tell a real regression from noise.
--min-pass-rate 0.9 replaces "every case must pass" with "at least 90% of attempts must pass" (errored attempts count as not passed), for suites where some flakiness is acceptable. Then use compare to catch it getting worse.