BehavTest
v0.8.1 Open source · MIT · No account

Behavioral regression testing for AI applications

AI output is nondeterministic, so one run proves little. BehavTest runs your LLM app, agent or RAG test cases repeatedly, scores every answer, and uses statistics to tell a real change in behavior from random noise, then fails the pull request that made things worse.

npx behavtest init --ts && npx behavtest run behavtest/suite.mts

No API key needed to try it. Requires Node.js 24 or newer.

behavtest compare
$ npx behavtest compare --fail-on-regression
behavtest compare · support-bot
  base  f033e1c8  prompt-v6
  head  22ec5145  prompt-v7

  ✗ regressed author-of-hamlet   5/5 → 0/5  p=0.008 significant
  ✗ regressed symbol-for-gold    5/5 → 2/5  p=0.167
  ✓ improved  is-pluto-a-planet  0/5 → 5/5  p=0.008 significant
  ~ flaky     largest-ocean      3/5 → 3/5

  overall  -21.7 pts  95% CI [-28.3, -15.0]  p=0.0015
  → significant regression · gate failed
Tests Any HTTP servicePython · FastAPI · LangChainOpenAIAnthropic ClaudeAzure · Ollama · vLLM · OpenRouterTypeScript functions

What it does

Know whether a change made your AI behave worse

A prompt tweak, a model swap or a new retrieval setting can quietly change how your application behaves. BehavTest turns that into a behavioral regression test you run locally and in CI.

Test the real pipeline

An HTTP endpoint in any language, an OpenAI-compatible or Anthropic model, or a function in your own process.

Score every answer

Exact match, a prompt-injection-hardened LLM judge, latency and cost limits, or your own scorers in TypeScript.

Tell regressions from noise

Repeat each case, then compare runs with Wilson intervals, Fisher's exact test and a case-stratified permutation test.

Fail the pull request

A GitHub Action compares every pull request with a committed baseline and fails the check when quality drops.

Check what the agent did

Store each run's tool calls and LLM steps, and test that the agent called the right tool without looping.

Zero infrastructure

One CLI, one local SQLite file, and a local dashboard to browse it. No hosted service, no account, no telemetry, no default provider. MIT licensed.

How it works

Three steps from prompt change to confident merge

Suites are plain JSON you can commit, or TypeScript when you want to call your agent directly.

Write a suite

List your pipeline and the cases that matter, with the scorers for each.

{
  "name": "support-bot",
  "pipeline": { "adapter": "openai",
    "config": { "model": "gpt-6-luna" } },
  "cases": [{
    "id": "refund-window",
    "input": "Return after 40 days?",
    "scorers": ["llmJudge"] }]
}

Run, change, compare

Run before and after your change, a few attempts per case, and see what moved.

behavtest run suite.json --repeat 5
# edit the prompt or swap the model
behavtest run suite.json --repeat 5
behavtest compare

Gate every pull request

Commit a baseline once; CI fails the pull request that makes results worse.

behavtest run suite.json \
  --repeat 3 --compact \
  --export behavtest.baseline.json
# then in CI:
behavtest compare \
  behavtest.baseline.json \
  --fail-on-regression

Learn

Testing AI applications, from first principles

What regression testing means when outputs are nondeterministic, how it differs from evaluation, and how to do it in practice. Useful whether or not you use BehavTest.

Blog

Notes from building BehavTest

Articles on behavioral regression testing and on the decisions behind BehavTest, newest first.

All posts · RSS feed

Integrations

Test the stack you already have

Step-by-step setups for the providers and frameworks BehavTest works with, each with a working example.

All integrations

Comparisons

How BehavTest relates to other tools

Neutral, sourced comparisons with other LLM evaluation and testing tools, and when each approach fits.

All comparisons

Try it in one command

A working suite with a stand-in agent. No API key, no server, no account.

npx behavtest init --ts && npx behavtest run behavtest/suite.mts