Test the real pipeline
An HTTP endpoint in any language, an OpenAI-compatible or Anthropic model, or a function in your own process.
AI output is nondeterministic, so one run proves little. BehavTest runs your LLM app, agent or RAG test cases repeatedly, scores every answer, and uses statistics to tell a real change in behavior from random noise, then fails the pull request that made things worse.
npx behavtest init --ts && npx behavtest run behavtest/suite.mtsNo API key needed to try it. Requires Node.js 24 or newer.
$ npx behavtest compare --fail-on-regression behavtest compare · support-bot base f033e1c8 prompt-v6 head 22ec5145 prompt-v7 ✗ regressed author-of-hamlet 5/5 → 0/5 p=0.008 significant ✗ regressed symbol-for-gold 5/5 → 2/5 p=0.167 ✓ improved is-pluto-a-planet 0/5 → 5/5 p=0.008 significant ~ flaky largest-ocean 3/5 → 3/5 overall -21.7 pts 95% CI [-28.3, -15.0] p=0.0015 → significant regression · gate failed
What it does
A prompt tweak, a model swap or a new retrieval setting can quietly change how your application behaves. BehavTest turns that into a behavioral regression test you run locally and in CI.
An HTTP endpoint in any language, an OpenAI-compatible or Anthropic model, or a function in your own process.
Exact match, a prompt-injection-hardened LLM judge, latency and cost limits, or your own scorers in TypeScript.
Repeat each case, then compare runs with Wilson intervals, Fisher's exact test and a case-stratified permutation test.
A GitHub Action compares every pull request with a committed baseline and fails the check when quality drops.
Store each run's tool calls and LLM steps, and test that the agent called the right tool without looping.
One CLI, one local SQLite file, and a local dashboard to browse it. No hosted service, no account, no telemetry, no default provider. MIT licensed.
How it works
Suites are plain JSON you can commit, or TypeScript when you want to call your agent directly.
List your pipeline and the cases that matter, with the scorers for each.
{
"name": "support-bot",
"pipeline": { "adapter": "openai",
"config": { "model": "gpt-6-luna" } },
"cases": [{
"id": "refund-window",
"input": "Return after 40 days?",
"scorers": ["llmJudge"] }]
}Run before and after your change, a few attempts per case, and see what moved.
behavtest run suite.json --repeat 5
# edit the prompt or swap the model
behavtest run suite.json --repeat 5
behavtest compareCommit a baseline once; CI fails the pull request that makes results worse.
behavtest run suite.json \
--repeat 3 --compact \
--export behavtest.baseline.json
# then in CI:
behavtest compare \
behavtest.baseline.json \
--fail-on-regressionLearn
What regression testing means when outputs are nondeterministic, how it differs from evaluation, and how to do it in practice. Useful whether or not you use BehavTest.
Blog
Articles on behavioral regression testing and on the decisions behind BehavTest, newest first.
Integrations
Step-by-step setups for the providers and frameworks BehavTest works with, each with a working example.
Comparisons
Neutral, sourced comparisons with other LLM evaluation and testing tools, and when each approach fits.
How-to guides
Short, task-first guides with copy-paste examples.
Reference
The suite format, every adapter and scorer, the statistics, the CLI and the library API.
A working suite with a stand-in agent. No API key, no server, no account.
npx behavtest init --ts && npx behavtest run behavtest/suite.mts