Troubleshooting
The messages below are quoted from BehavTest (without the backticks some of them contain); … stands for the part that names your file, case or variable.
… is not set. The openai adapter reads its API key from the environment. Export the key in the shell or CI job that runs BehavTest (OPENAI_API_KEY, ANTHROPIC_API_KEY, or the variable named by apiKeyEnv). Keys are never read from the suite file.
Missing environment variable referenced in pipeline.config: … The suite uses ${VAR} and VAR is empty or unset. Set it, or give a default: ${PIPELINE_URL:-http://localhost:4000/pipeline}.
case "…" uses llmJudge but no judge model is configured. Judged scorers need a model: set "defaults": { "judge": "openai:gpt-4.1-nano" } in the suite, pass --judge provider:model, or export BEHAVTEST_JUDGE.
The judge … does not work: … Before any case runs, BehavTest makes one tiny call to each judge. The message after the colon is the provider's answer: usually a wrong model name, a missing key, or a model without structured output. Fix the judge, or skip the check with --no-judge-check.
Failed to load suite module …: … Node is treating this file as CommonJS … A .ts or .js suite is an ES module only when package.json says "type": "module". Rename the suite to .mts / .mjs (what behavtest init --ts does), or set "type": "module". Local files it imports need the same treatment.
Failed to load suite module …: … If it imports another TypeScript file, write the extension in the import … Node's type stripping needs the extension in local imports (import { answer } from "./agent.ts"). If the suite imports values such as tracer or defineSuite from behavtest, install it in the project (npm i -D behavtest); import type needs no install.
Cannot load …: this Node.js (…) cannot import TypeScript files. BehavTest needs Node.js 24 or newer. Upgrade Node, or write the suite as .mjs or .json.
No results database at "…". Run behavtest run <suite> first, or pass --db <path>. runs, show, compare, report, matrix, calibrate and serve read the database a previous run wrote in the current directory. Run from the same directory, or point --db at it. In CI, compare against a committed run file instead: behavtest compare behavtest.baseline.json.
There is no earlier run of "…" to compare run … against. compare without arguments needs two runs of the same suite (and the same variant). Run the suite again, or name both runs or files: behavtest compare <base> <head>.
… is not a BehavTest run file (expected "kind": "behavtest.run"). compare and report --against take run ids or run files written by --export or behavtest export. A --json report is not a run file: it has no case hashes. (Run files written by Regrade are still read.)
Every run exits 1, but nothing looks broken. Exit 1 means at least one case failed, was flaky or errored. With a nondeterministic pipeline, flaky cases are expected: gate on the change instead (behavtest compare … --fail-on-regression --significant-only), or accept a pass rate (--min-pass-rate 0.9). See deal with flaky, non-deterministic outputs.
Cases show as modified instead of regressed or improved. The case's definition, its scorer code or its judge model changed between the two runs, so their results aren't comparable. This is intended: update the baseline in the same change.
Attempts are errored with network error calling … or HTTP 5xx from …. The pipeline could not answer, which is different from answering wrongly. Network errors, HTTP 429 and 5xx are retried (honoring Retry-After); raise retries in pipeline.config, check the service, or lower --concurrency if it is rate-limiting you.
Using .regrade/results.db (from Regrade, BehavTest's former name). Not an error: results from before the rename are still used. Rename the .regrade folder to .behavtest to make the notice go away.
Port 4800 is in use. Pick another with --port <n>. Another program (or another behavtest serve) is using the port: behavtest serve --port 5000.
Still stuck? Open an issue with the command, the full message and behavtest --version.