For AI coding assistants
If you are an AI assistant setting up BehavTest for a user, these facts are enough to do it correctly:
- Requires Node.js 24+. Run it with
npx behavtest <command>;npm i -D behavtestonly if the suite imports values (tracer,defineSuite) from"behavtest". - Scaffold with
npx behavtest init --ts(code suite, no server, no key) ornpx behavtest init(JSON suite plus mock HTTP pipeline). In a CommonJS project use.mts/.mjsfor code suites. - A JSON suite needs
name,pipeline: { adapter, config }(http|openai|anthropic) andcases[]with a stableid, aninputandscorers. Add"$schema": "https://unpkg.com/behavtest/schema/suite.schema.json". Unknown keys are errors. - Never put secrets in suite files: use
${ENV_VAR}placeholders inpipeline.config. Keys come fromOPENAI_API_KEY/ANTHROPIC_API_KEY. exactMatchrequiresexpected.llmJudgeneeds a judge (defaults.judge: "provider:model"), preferably not the pipeline's own model.toolCalled/maxStepsneed the pipeline to returnsteps.- Use
--repeat 3or more for LLM pipelines. Exit codes: 0 all passed, 1 failures/flaky/errored or gate failed, 2 configuration error, 130 interrupted. - To gate CI: create
behavtest.baseline.jsonwithrun --export behavtest.baseline.json --compact, commit it, and in CI runbehavtest run … || test $? -eq 1thenbehavtest compare behavtest.baseline.json --fail-on-regression. - Add
.behavtest/to.gitignore(initdoes this): the database holds raw inputs and outputs. - RAG: the pipeline reports retrieved documents as a
kind: "retrieval"step whoseoutputlists{ id, text }; cases listexpectedDocsfor theretrievalscorer.faithfulnessandcontextRelevanceneed a judge. - To check the judge: the user labels judged answers in
behavtest serve(saved to the database), thenbehavtest calibrate --min-kappa 0.6. A labels file (JSONL{ run, case, attempt, scorer, label: "pass" | "fail" }, exported from the HTML report) works with--labels <file>. - On GitHub, prefer the Action:
uses: dhrumilbhut/behavtest@v0withsuite:(andrepeat:equal to the baseline's). It needs a committed compactbehavtest.baseline.json;comment: trueneedspermissions: pull-requests: write. - To compare models or prompts, add
variants: [{ name, pipeline: { config: {...} } }]to the suite (at least two; config is merged overpipeline.config), run it, thenbehavtest matrix. behavtest serveis for a person to look at results; it is not needed in CI. It binds to 127.0.0.1; do not suggest--host 0.0.0.0on shared machines.
A machine-readable summary is at dhrumilbhut.github.io/behavtest/llms.txt, and this README as plain text at llms-full.txt.