BehavTest

BehavTest › Reference

FAQ

What is BehavTest? An open-source (MIT) CLI and Node.js library for behavioral regression testing of AI applications: LLM apps, AI agents and RAG pipelines. It runs test cases through your pipeline several times, scores the answers, stores every run, and tells you whether a change made results worse, with statistics that separate a real change in behavior from nondeterministic noise. It was called Regrade until version 0.8.0 (see migrating from Regrade).

Does it need an API key? No, not to start: behavtest init --ts and behavtest init run without one. You need a provider key only for the openai / anthropic adapters or for llmJudge.

Does it work with Python, LangChain or LlamaIndex? Yes, through the HTTP adapter: expose one endpoint that takes { "input": ... } and returns { "output": "..." } (example). BehavTest itself runs on Node.js 24+, which CI runners already have.

Can I use local or self-hosted models (Ollama, vLLM)? Yes: use the openai adapter with baseUrl pointing at any OpenAI-compatible server, for the pipeline or for the judge (OPENAI_BASE_URL).

Which LLM should I use as the judge? A different model from the one being tested, ideally one that accepts temperature 0, chosen by measured cost per verdict. gpt-4.1-nano is a cheap choice that worked well in our tests; see the LLM judge.

How many repeats do I need? For a single case to show a significant drop, about 5 attempts per side (5/5 → 0/5 gives p = 0.008; 3/3 → 0/3 is only p = 0.1). Across many cases, fewer attempts can still show a significant overall drop. Start with --repeat 3 and raise it for important suites.

Where are my results stored, and does BehavTest send data anywhere? In .behavtest/results.db on your machine. BehavTest contacts only the pipeline and providers you configure. There is no telemetry and no update check.

How do I run it in GitHub Actions or another CI system? On GitHub, use the GitHub Action (uses: dhrumilbhut/behavtest@v0) with a committed baseline. In any other CI that runs Node.js, run behavtest run … --export and behavtest compare behavtest.baseline.json --fail-on-regression (baselines and CI).

Can I compare several models or prompts? Yes: add variants to a suite and behavtest run makes one run per variant; behavtest matrix shows them side by side with confidence intervals, cost and latency, and tests each against a reference (matrix runs).

How is BehavTest different from Promptfoo, DeepEval, Inspect AI or Ragas? Those are mature evaluation tools, several with broader feature sets or hosted options. BehavTest focuses narrowly on behavioral regression testing: repeated attempts per case, significance tests on the change between two runs, run files as CI baselines, and zero infrastructure, with no default provider. See prior art.

Is the LLM judge reliable? It is hardened (prompt-injection fencing, structured output, fail-closed, pre-run check), but it is still one model's opinion. Measure it: label some answers yourself and run behavtest calibrate for its agreement with you (Cohen's kappa). Prefer exactMatch, toolCalled, retrieval or your own programmatic scorers where possible, and repeat cases.

How do I test a RAG pipeline? Report the retrieved documents as a retrieval trace step, list the right document ids per case in expectedDocs, and use retrieval (did it find them), faithfulness (is the answer grounded in them) and contextRelevance (was the context on topic). See RAG and examples/rag.

Is there a UI? Yes, a local one: behavtest serve opens a dashboard on your results database (runs, pass-rate trends, comparisons, labelling and judge calibration), and behavtest report writes a single-file HTML report you can share. Neither needs an account or a hosted service.

Is it free? Yes, MIT-licensed. You pay only your model providers for the calls your suites make.