BehavTest

BehavTest › Reference

For AI coding assistants

If you are an AI assistant setting up BehavTest for a user, these facts are enough to do it correctly:

  1. Requires Node.js 24+. Run it with npx behavtest <command>; npm i -D behavtest only if the suite imports values (tracer, defineSuite) from "behavtest".
  2. Scaffold with npx behavtest init --ts (code suite, no server, no key) or npx behavtest init (JSON suite plus mock HTTP pipeline). In a CommonJS project use .mts / .mjs for code suites.
  3. A JSON suite needs name, pipeline: { adapter, config } (http | openai | anthropic) and cases[] with a stable id, an input and scorers. Add "$schema": "https://unpkg.com/behavtest/schema/suite.schema.json". Unknown keys are errors.
  4. Never put secrets in suite files: use ${ENV_VAR} placeholders in pipeline.config. Keys come from OPENAI_API_KEY / ANTHROPIC_API_KEY.
  5. exactMatch requires expected. llmJudge needs a judge (defaults.judge: "provider:model"), preferably not the pipeline's own model. toolCalled / maxSteps need the pipeline to return steps.
  6. Use --repeat 3 or more for LLM pipelines. Exit codes: 0 all passed, 1 failures/flaky/errored or gate failed, 2 configuration error, 130 interrupted.
  7. To gate CI: create behavtest.baseline.json with run --export behavtest.baseline.json --compact, commit it, and in CI run behavtest run … || test $? -eq 1 then behavtest compare behavtest.baseline.json --fail-on-regression.
  8. Add .behavtest/ to .gitignore (init does this): the database holds raw inputs and outputs.
  9. RAG: the pipeline reports retrieved documents as a kind: "retrieval" step whose output lists { id, text }; cases list expectedDocs for the retrieval scorer. faithfulness and contextRelevance need a judge.
  10. To check the judge: the user labels judged answers in behavtest serve (saved to the database), then behavtest calibrate --min-kappa 0.6. A labels file (JSONL { run, case, attempt, scorer, label: "pass" | "fail" }, exported from the HTML report) works with --labels <file>.
  11. On GitHub, prefer the Action: uses: dhrumilbhut/behavtest@v0 with suite: (and repeat: equal to the baseline's). It needs a committed compact behavtest.baseline.json; comment: true needs permissions: pull-requests: write.
  12. To compare models or prompts, add variants: [{ name, pipeline: { config: {...} } }] to the suite (at least two; config is merged over pipeline.config), run it, then behavtest matrix.
  13. behavtest serve is for a person to look at results; it is not needed in CI. It binds to 127.0.0.1; do not suggest --host 0.0.0.0 on shared machines.

A machine-readable summary is at dhrumilbhut.github.io/behavtest/llms.txt, and this README as plain text at llms-full.txt.