BehavTest

BehavTest › Learn

LLM regression testing

How do you regression test an LLM application when its outputs change between runs? Keep a fixed set of test cases, check each answer for the behavior you need rather than its exact wording, run every case several times, and compare the pass rates with a baseline using a significance test. Fail the build only when a change is larger than the run-to-run noise.

This page walks through each of those steps, why each one is needed, and how to wire the result into CI. The concepts behind it are on behavioral regression testing.

Why can't I just snapshot LLM outputs?

Snapshot testing (store the output once, fail when it changes) works for deterministic code. For an LLM it fails in both directions:

Snapshots are still useful for the deterministic parts around the model (prompt templates, parsers), which LLM testing covers.

Step 1: write cases that check behavior

Each case is an input plus the behavior its answer must show. Pick the cheapest check that captures the behavior:

Behavior to protectHow to check itIn BehavTest
States a required fact or keywordDeterministic: substring or regular expressionA custom scorer in a code suite (a one-line function)
A fixed answer (a label, "OK", a number)Deterministic: normalized equalityexactMatch
Correct and complete in meaningAn LLM judge with a rubricllmJudge
Calls the right tool, doesn't loopCheck the agent's recorded stepstoolCalled, maxSteps
Retrieves the right documents, answers from themRetrieval metrics, a judge on groundingretrieval, faithfulness, contextRelevance
Fast and cheap enoughThresholds on latency and costlatencyCost

Prefer deterministic checks: they cost nothing and add no noise of their own. Use an LLM judge where meaning matters, and check that the judge agrees with you.

A minimal suite that tests a prompt on an OpenAI model:

{
  "$schema": "https://unpkg.com/behavtest/schema/suite.schema.json",
  "name": "support-prompt",
  "defaults": { "judge": "openai:gpt-4.1-nano", "repeat": 5 },
  "pipeline": {
    "adapter": "openai",
    "config": { "model": "gpt-6-luna", "system": "You are a concise support agent. Returns are accepted within 30 days." }
  },
  "cases": [
    {
      "id": "refund-window",
      "input": "Can I return an item after 40 days?",
      "expected": "No: returns are accepted within 30 days.",
      "scorers": ["llmJudge"],
      "scorerConfig": { "llmJudge": { "rubric": "Does the answer state the 30-day limit and avoid promising an exception?" } }
    },
    { "id": "one-word", "input": "Reply with only the word OK.", "expected": "OK", "scorers": ["exactMatch"] }
  ]
}

Your application doesn't need to be JavaScript: BehavTest can also call any HTTP endpoint, or a function in a TypeScript suite. See adapters and the integrations for OpenAI, Anthropic, Ollama, FastAPI, LangChain and the Vercel AI SDK.

Step 2: record a baseline

Run the suite on the version you trust, with several attempts per case, and keep the result:

export OPENAI_API_KEY=sk-...
npx behavtest run prompt.suite.json --repeat 5 --label baseline --export behavtest.baseline.json --compact

--export ... --compact writes a small run file with pass/fail per attempt and a hash of each case, but no inputs or outputs, so it is safe to commit. The hash matters: if you later edit a case, the comparison reports it as modified instead of mistaking your edit for a regression.

Step 3: change something and compare

Edit the prompt, swap the model or change retrieval, then run again and compare:

npx behavtest run prompt.suite.json --repeat 5 --label new-prompt
npx behavtest compare behavtest.baseline.json

compare lists every case that regressed, improved or is flaky, with its pass rate before and after, a p-value from Fisher's exact test, and an overall verdict from a case-stratified permutation test. Read how compare decides for the details.

Step 4: choose how many attempts

The number of attempts per case decides what the comparison can detect. For a single case, these are the p-values of the most extreme result possible (the case went from always passing to never passing, or to half):

Attempts per sideChangep (Fisher's exact test)Significant at 0.05?
11/1 → 0/11.000never
33/3 → 0/30.100never
55/5 → 0/50.008yes
1010/10 → 5/100.033yes
2020/20 → 14/200.020yes

So a single case needs about five attempts per side before even a total collapse can be significant, and more to detect partial drops. The overall test pools all cases, so across a suite fewer attempts can still reveal a broad regression. A practical starting point is --repeat 3 for fast feedback and 5 to 10 for the suite that gates merges; the cost is cases × attempts model calls (plus judge calls), and behavtest run prints the count before it starts.

Step 5: decide what fails the build

This is where nondeterminism bites hardest. BehavTest's nondeterministic example runs a bot that is right 90% of the time; running it twice with nothing changed (10 attempts per case) produced five "regressed" cases, none significant, and an overall change of -3.7 points with p ≈ 0.67. The gate you choose decides what that means:

GateFails onResult for the unchanged bot
compare --fail-on-regressionany case whose pass rate dropped, significant or not, any errored case, or a significant overall dropfails (exit 1): noise counts
compare --fail-on-regression --significant-onlyonly statistically significant case drops, an errored case, or a significant overall droppasses (exit 0)
run --min-pass-rate 0.8this run alone falling below 80% of attempts passedpasses (exit 0)

When the same bot's accuracy really fell to 60%, the significant-only gate failed as it should (overall p < 0.001). The rule of thumb: use --fail-on-regression for suites that are nearly deterministic (mostly deterministic scorers on stable behavior, where any drop is suspicious), and add --significant-only for genuinely nondeterministic behavior, with enough attempts for significance to be reachable.

Step 6: run it in CI

On GitHub, the BehavTest Action runs the suite on every pull request, compares it with the committed baseline, writes the comparison to the job summary and fails the check according to gate (regression, significant, cases or none):

name: BehavTest
on:
  pull_request:

permissions:
  contents: read
  pull-requests: write   # only for comment: true

jobs:
  behavtest:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: dhrumilbhut/behavtest@v0
        with:
          suite: prompt.suite.json
          repeat: 5
          gate: significant
          comment: true
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

Any other CI system can run the same two commands (behavtest run, then behavtest compare behavtest.baseline.json --fail-on-regression); exit code 1 means the gate failed and 2 means a configuration error. Both recipes are in baselines and CI. Update the baseline in the same pull request when a change is intended, so moving it is a reviewed decision.

Limitations