BehavTest

BehavTest › Reference

Statistical testing in BehavTest

behavtest compare decides whether a change made an AI application behave worse. Because each test case passes only some of the time, that is a statistical question, and this page documents exactly how BehavTest answers it: what is computed, why each method was chosen, and what the numbers can and cannot tell you. For the intuition first, read behavioral regression testing.

The data

Each run stores every attempt: one execution of one case, with a status of passed (every scorer passed), failed (the pipeline answered and a scorer said no) or errored (no verdict was possible: the pipeline was down, timed out, or a judge could not answer). With --repeat N, a case has N attempts. A case's pass rate is its passed attempts divided by its attempts.

Cases in the two runs are matched by id. Before any statistics, three kinds of case are set aside:

Everything below applies to the remaining comparable cases.

Per case: pass rates and Wilson intervals

Each side's pass rate is reported with a 95% Wilson score interval. Wilson is used rather than the textbook normal approximation because it stays sensible at the extremes that are common here: 10 of 10 passing gives 72% to 100%, not a zero-width interval at 100%.

The case is then classified by the direction of its pass rate: lower is regressed, higher is improved, equal is unchanged, or flaky when equal and the new run's attempts were mixed.

Per case: Fisher's exact test

Whether a case's change is significant comes from Fisher's exact test on the 2×2 table of passed and not-passed attempts before and after, two-sided, at α = 0.05. It is exact (no large-sample approximation), which matters with 3 to 20 attempts per side. The consequence is a floor on what one case can show:

Attempts per sideMost extreme changepCan be significant?
11/1 → 0/11.000no
33/3 → 0/30.100no
55/5 → 0/50.008yes
1010/10 → 5/100.033yes

A change that is not significant is still listed, with the note "not statistically significant at this sample size" (or, with one attempt per side, "could be noise, re-run with --repeat to confirm"). Per-case tests are not corrected for multiple comparisons: in a suite of 100 unchanged cases, a few may come out significant by chance. Treat a single significant case as a lead to investigate, and the overall test as the verdict on the suite.

Overall: a case-stratified permutation test

The overall question is "on this suite, did the pass rate move by more than the pipeline's sampling noise?" The statistic is the mean change in pass rate per case (each comparable case weighted equally, so a case with more attempts doesn't dominate).

The test is a paired permutation test stratified by case: within each case, the before/after labels of that case's own attempts are shuffled, the mean change is recomputed, and the p-value is the share of shuffles at least as extreme as the observed change (two-sided, with the usual +1 correction so it is never exactly zero). Cases are held fixed because the randomness that matters is within each case (the same input gives different outputs), not which cases happen to be in the suite. With one attempt per side, this reduces to an exact sign test on the cases whose outcome flipped: six one-way flips are significant (p = 0.031), five are not (p = 0.063).

A drop with p < 0.05 is a significant regression, a rise a significant improvement, and anything else not significant.

The interval printed next to the mean change is a 95% percentile bootstrap interval that resamples attempts within each case (before and after separately), again holding the cases fixed, so its width reflects the same within-case noise.

Monte Carlo, and why the last digits vary

The permutation test and the bootstrap use random resampling: between 2,000 and 10,000 resamples, depending on the number of attempts. The random seed is derived from the two run ids, so comparing the same two runs always gives the same numbers, and a report can be regenerated exactly. Two different runs of an unchanged pipeline will give slightly different p-values, as in this repeat of the no-change experiment from the nondeterministic example: p = 0.678, 0.668 and 0.671 on three fresh pairs of runs. The smallest reportable p-value is about 0.0001, printed as <0.0001.

The descriptive pass rate

The line attempt pass rate 85% [76%–91%] → 56% [45%–67%] pools all attempts of all comparable cases on each side, with Wilson intervals. It is labelled descriptive because it is not the test: pooling attempts ignores which case they came from, and cases with more attempts weigh more. The decision comes from the permutation test on the per-case mean.

Validation

The statistics are checked against textbook reference values in the test suite, and the overall test is checked by simulation: with nothing changed it rejects less than 9% of the time, and it rejects more than 95% of the time for a real drop. The nondeterministic example has an end-to-end test that keeps the numbers quoted in this documentation true.

Assumptions and limits