# BehavTest > BehavTest is an open-source (MIT) command-line tool and Node.js library for behavioral regression testing of AI applications: LLM apps, AI agents and RAG pipelines. AI output is nondeterministic, so BehavTest runs each test case through your pipeline several times (an HTTP endpoint in any language, an OpenAI-compatible or Anthropic model, or a TypeScript/JavaScript function), scores every attempt (exact match, LLM-as-a-judge, latency and cost limits, tool-call and step-count checks on the agent's trace, RAG checks on retrieval and grounding, or custom functions), stores every run in a local SQLite file, and compares it with a baseline using Wilson intervals, Fisher's exact test and a case-stratified permutation test, so a real change in behavior is told apart from random variation and CI fails when a prompt, model or code change makes results worse. It was called Regrade until version 0.8.0. Key facts: - Install: `npx behavtest` or `npm install --global behavtest`. Requires Node.js 24+. npm package: `behavtest`. - Start without an API key: `npx behavtest init --ts && npx behavtest run behavtest/suite.mts`. - Suites are JSON files (schema: https://unpkg.com/behavtest/schema/suite.schema.json) or TypeScript/JavaScript modules with inline scorers and in-process pipelines. - Adapters: `http` (POST `{"input": ...}`, expects `{"output": "..."}`), `openai` (also Azure, Ollama, vLLM, OpenRouter via `baseUrl`), `anthropic`. - Built-in scorers: `exactMatch`, `llmJudge`, `latencyCost`, `toolCalled`, `maxSteps`, and for RAG `retrieval` (deterministic: hit/recall/precision/MRR against a case's `expectedDocs`), `faithfulness` and `contextRelevance` (LLM judge). - Judge calibration: `behavtest calibrate` measures how often the LLM judge agrees with human pass/fail labels (Cohen's kappa, false-pass rate); labels are saved from the dashboard, or read from a JSONL file with `--labels` (exported from the HTML report). - Dashboard: `behavtest serve` opens a local web UI on the results database (127.0.0.1:4800): runs, pass-rate trends per suite, comparing any two runs, labelling judge verdicts, and judge calibration. - Matrix runs: `variants` on a suite (each changes the pipeline config, e.g. the model or prompt) make one run per variant; `behavtest matrix` compares them side by side with Wilson intervals, cost, latency and a permutation test against a reference. - Commands: `run`, `compare`, `matrix`, `runs`, `show`, `report` (single-file HTML), `serve` (local dashboard), `calibrate`, `export` / `import` (run files), `init`, `schema`. - CI: commit a compact baseline (`behavtest run suite.json --export behavtest.baseline.json --compact`). On GitHub use the Action `dhrumilbhut/behavtest@v0` (inputs `suite`, `repeat`, `gate`, `comment`); elsewhere run `behavtest compare behavtest.baseline.json --fail-on-regression`. - Exit codes: 0 passed, 1 failed/flaky/errored or gate failed, 2 configuration error, 130 interrupted. - Formerly Regrade (npm `regrade`, now deprecated): the commands are the same with `behavtest` in place of `regrade`, and `.regrade/results.db`, `regrade.baseline.json`, `REGRADE_JUDGE` and Regrade run files are still read. See https://github.com/dhrumilbhut/behavtest#migrating-from-regrade. - No telemetry, no hosted service, no default provider. Secrets are referenced as `${ENV_VAR}` and never stored. ## Docs - [README (full documentation)](https://github.com/dhrumilbhut/behavtest#readme): quickstart, how-to guides, concepts, reference, FAQ - [Full documentation as plain text](https://dhrumilbhut.github.io/behavtest/llms-full.txt): the README in one file - [Quickstart](https://github.com/dhrumilbhut/behavtest#quickstart): no-key start, testing an OpenAI or Anthropic prompt, testing your own HTTP service - [GitHub Action](https://github.com/dhrumilbhut/behavtest#github-action): block the pull request that makes results worse (inputs, outputs, baseline refresh) - [Matrix runs](https://github.com/dhrumilbhut/behavtest#matrix-runs-compare-models-and-prompts-side-by-side): compare models and prompt versions side by side - [Baselines and CI](https://github.com/dhrumilbhut/behavtest#baselines-and-ci-fail-the-pull-request-that-made-things-worse): the same with CLI steps, for any CI system - [Traces](https://github.com/dhrumilbhut/behavtest#traces-check-what-the-agent-did-not-just-what-it-said): reporting an agent's steps and checking its tool calls - [The LLM judge](https://github.com/dhrumilbhut/behavtest#the-llm-judge): prompt-injection hardening, structured verdicts, choosing a judge model - [RAG](https://github.com/dhrumilbhut/behavtest#rag-test-retrieval-and-grounded-answers): retrieval, faithfulness and context-relevance scorers - [Judge calibration](https://github.com/dhrumilbhut/behavtest#judge-calibration-does-the-judge-agree-with-you): measuring the judge against your own labels - [Compare runs](https://github.com/dhrumilbhut/behavtest#compare-runs-what-regressed-and-is-it-real): how regressions are separated from noise - [Dashboard](https://github.com/dhrumilbhut/behavtest#dashboard-browse-compare-and-label-runs): `behavtest serve`, the local web UI for browsing, comparing and labelling runs - [Migrating from Regrade](https://github.com/dhrumilbhut/behavtest#migrating-from-regrade): the former name, and what changed with the rename - [For AI coding assistants](https://github.com/dhrumilbhut/behavtest#for-ai-coding-assistants): the rules for setting BehavTest up correctly ## Learn - [Behavioral regression testing](https://dhrumilbhut.github.io/behavtest/behavioral-regression-testing/): the concept, from first principles, with real numbers - [LLM regression testing](https://dhrumilbhut.github.io/behavtest/llm-regression-testing/): how to regression test an LLM application whose outputs change between runs, including CI - [AI regression testing](https://dhrumilbhut.github.io/behavtest/ai-regression-testing/): traditional versus AI regression testing - [LLM testing](https://dhrumilbhut.github.io/behavtest/llm-testing/): unit tests, integration tests, evaluation, regression tests, production checks - [AI application testing](https://dhrumilbhut.github.io/behavtest/ai-application-testing/): what to test in prompts, retrieval, agents, dependencies - [LLM evaluation](https://dhrumilbhut.github.io/behavtest/llm-evaluation/): methods, LLM-as-a-judge, and evaluation versus regression testing - [Statistical testing in BehavTest](https://dhrumilbhut.github.io/behavtest/docs/statistics/): exactly what compare computes, and its limits - [Troubleshooting](https://dhrumilbhut.github.io/behavtest/docs/troubleshooting/): common error messages and fixes ## Integrations - [OpenAI](https://dhrumilbhut.github.io/behavtest/integrations/openai/) · [Anthropic](https://dhrumilbhut.github.io/behavtest/integrations/anthropic/) · [Ollama and OpenAI-compatible servers](https://dhrumilbhut.github.io/behavtest/integrations/ollama/) · [HTTP services (Python, FastAPI)](https://dhrumilbhut.github.io/behavtest/integrations/http/) · [LangChain](https://dhrumilbhut.github.io/behavtest/integrations/langchain/) · [Vercel AI SDK](https://dhrumilbhut.github.io/behavtest/integrations/vercel-ai-sdk/) ## Comparisons - [BehavTest vs Promptfoo](https://dhrumilbhut.github.io/behavtest/comparisons/promptfoo/) · [vs DeepEval](https://dhrumilbhut.github.io/behavtest/comparisons/deepeval/) · [vs Ragas](https://dhrumilbhut.github.io/behavtest/comparisons/ragas/) · [vs LangSmith](https://dhrumilbhut.github.io/behavtest/comparisons/langsmith/) ## Examples - [Documentation website](https://dhrumilbhut.github.io/behavtest/): one page per how-to guide and reference topic - [Live sample report](https://dhrumilbhut.github.io/behavtest/sample/): a healthy pipeline compared with a degraded one - [Example suite and mock pipeline](https://github.com/dhrumilbhut/behavtest/tree/main/examples/qa-http) - [Example RAG pipeline and suite](https://github.com/dhrumilbhut/behavtest/tree/main/examples/rag) - [Demo repository](https://github.com/dhrumilbhut/behavtest-demo/pulls): the GitHub Action passing one pull request and blocking another - [Example matrix suite](https://github.com/dhrumilbhut/behavtest/tree/main/examples/matrix): four OpenAI models side by side ## Blog All posts: https://dhrumilbhut.github.io/behavtest/blog/ (RSS: https://dhrumilbhut.github.io/behavtest/blog/feed.xml) - [AI Output Is Nondeterministic. Here's How to Test It Anyway.](https://dhrumilbhut.github.io/behavtest/blog/ai-output-is-nondeterministic-how-to-test-it-anyway/) (2026-10-01): Why pass or fail testing breaks for AI applications, and the statistics bug that shaped how BehavTest tells a real regression from noise. - [My RAG Eval Kept Failing the Cheap Model. The Model Wasn't the Problem.](https://dhrumilbhut.github.io/behavtest/blog/my-rag-eval-kept-failing-the-cheap-model/) (2026-10-01): A real debugging story: why a cheap LLM judge scored 7 out of 12 on a RAG faithfulness check, and why the fix was the prompt, not the model. ## Optional - [Changelog](https://github.com/dhrumilbhut/behavtest/blob/main/CHANGELOG.md) - [Suite JSON Schema](https://unpkg.com/behavtest/schema/suite.schema.json) - [npm package](https://www.npmjs.com/package/behavtest) - [Security policy](https://github.com/dhrumilbhut/behavtest/blob/main/SECURITY.md)