BehavTest

BehavTest › Reference

How it works

A test run moves through the same stages every time:

 Test definition      suite: cases, scorers, pipeline (JSON or TypeScript)
        |
 Execution            each case sent to your application: HTTP, OpenAI, Anthropic or a function
        |
 Repeated evaluation  --repeat N attempts per case
        |
 Result collection    every attempt, score and trace saved (SQLite, or a portable run file)
        |
 Behavioral scoring   scorers pass or fail each attempt; each case gets a verdict
        |             (passed, failed, flaky, errored)
        |
 Statistical analysis compare with the baseline: per-case pass rates with Wilson intervals and
        |             Fisher's exact test; overall change with a case-stratified permutation test
        |
 Regression decision  regressed / improved / flaky / not significant, per case and overall
        |
 Test result          report + exit code (0 passed, 1 regression or failure, 2 configuration error)
  1. Run every test case through your real pipeline several times (--repeat), before and after a change.
  2. Score every attempt: exact match, an LLM judge, latency and cost limits, the tool calls the agent made, or what a RAG pipeline retrieved and whether the answer is grounded in it.
  3. Save every run, attempt, score and trace to a local SQLite file, or to a portable run file you commit as the baseline.
  4. Compare the candidate run with the baseline case by case: Wilson intervals on each pass rate, Fisher's exact test per case, and a case-stratified permutation test (with a bootstrap interval) on the overall change.
  5. Decide: a case that passes only sometimes is reported as flaky, a change within the noise is labelled not significant, and a real drop is a regression. --fail-on-regression, or the GitHub Action, fails the build.

A minimal suite, for a support bot served over HTTP. Each case names the behavior its answer must show:

{
  "name": "support-bot",
  "defaults": { "repeat": 5, "judge": "openai:gpt-4.1-nano" },
  "pipeline": { "adapter": "http", "config": { "url": "${PIPELINE_URL:-http://localhost:4000/pipeline}" } },
  "cases": [
    { "id": "refund-window", "input": "Can I return an item after 40 days?",
      "scorers": ["llmJudge"],
      "scorerConfig": { "llmJudge": { "rubric": "Does the answer state the 30-day limit and avoid promising an exception?" } } },
    { "id": "capital", "input": "What is the capital of France?", "expected": "Paris", "scorers": ["exactMatch"] }
  ]
}

behavtest run suite.json sends each input to the endpoint five times, scores every answer and saves the run; after a change, behavtest compare runs the statistics against the previous run (or a committed baseline file).

What the decision looks like, from the nondeterministic example after the bot's accuracy dropped from 90% to 60% (10 attempts per case; abridged; the overall p-value and interval are Monte Carlo estimates, so their last digits vary from run to run):

behavtest compare · support-bot
  ✗ regressed shipping-time 10/10 → 5/10 100% → 50%  p=0.033 significant
  ✗ regressed support-email 9/10 → 5/10 90% → 50%  p=0.141
      not statistically significant at this sample size
  ...
  attempt pass rate  85% [76%–91%] → 56% [45%–67%]  (8 comparable cases; descriptive)
  overall change     mean per case -28.7 pts, 95% CI [-41.3 pts, -16.3 pts], p=<0.0001 → significant regression

Each case on its own has only 10 attempts per side, so most per-case drops are "not significant"; the overall test pools the evidence across cases and is sure. Cases whose definition, scorer or judge changed between the two runs are reported as modified and never counted as regressions, so changing a test is not mistaken for a change in behavior. The details: compare runs and the statistical reference. The concept, from first principles: behavioral regression testing.

Next: try it in the quickstart, connect your stack with an integration, gate pull requests with the GitHub Action, or browse the examples.