CLI reference
behavtest run <suite> [options] Run a suite (.json, or a code suite: .ts .mts .js .mjs), score outputs, save the run
--db <path> SQLite file (default .behavtest/results.db)
--json <file> also write a JSON report
--md <file> also write a Markdown summary
--export <file> [--compact] also write a run file (--compact: only what a comparison needs)
--min-pass-rate <0-1> pass if at least this fraction of attempts pass
--concurrency <n> attempts in flight (default 4)
--repeat <n> attempts per case (overrides the suite)
--timeout <ms> per-attempt timeout (default 30000)
--tag <tag> only cases with this tag (repeatable)
--case <id> only this case (repeatable)
--label <text> label the run (e.g. a prompt version)
--variant <name> matrix suites: only this variant (repeatable)
--judge <provider:model> LLM judge model
--no-judge-check skip the one tiny call that checks the judge before any case runs
--no-trace do not store the steps pipelines report
--prices <file> extra/override model prices
--no-color plain output (also honours NO_COLOR; set BEHAVTEST_ASCII=1 for ASCII symbols)
behavtest runs [--suite <name>] [--limit <n>] List saved runs, newest first
behavtest show <run> [case] [--full] A run's summary, or one case's input, outputs, scores and trace
behavtest compare [base] [head] [options] What regressed, improved, or is just flaky; runs are ids or run files
--fail-on-regression | --significant-only exit 1 when the gate fails
--all --json <file> --md <file> --suite <name>
behavtest report <run> [--against <base>] [--out <file>] Single-file HTML report (runs are ids or run files)
behavtest export <run> [--out <file>] [--compact] Write a run file (a baseline to commit, or to compare or import elsewhere)
behavtest import <file> Load a full run file into the database
behavtest calibrate [--labels <file>] [--min-kappa <k>] [--json <file>] [--md <file>] How often the judge agrees with your labels
(default: the labels saved from the dashboard)
behavtest serve [--port <n>] [--host <host>] [--open] Local dashboard: runs, trends, compare, matrices, labels, calibration
behavtest matrix [id] [--reference <v>] [--list] [--md|--json|--out <file>] Variants of a matrix side by side
behavtest init [--dir <dir>] [--force] [--ts] Scaffold an example suite (--ts: a code suite, no server needed)
behavtest schema [--out <file>] Print the suite JSON Schema
Environment variables: OPENAI_API_KEY, OPENAI_BASE_URL, ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL, BEHAVTEST_JUDGE (default judge), NO_COLOR, BEHAVTEST_ASCII, plus any ${VAR} your suite references.