BehavTest

BehavTest › Reference

CLI reference

behavtest run <suite> [options]     Run a suite (.json, or a code suite: .ts .mts .js .mjs), score outputs, save the run
  --db <path>                     SQLite file (default .behavtest/results.db)
  --json <file>                   also write a JSON report
  --md <file>                     also write a Markdown summary
  --export <file> [--compact]     also write a run file (--compact: only what a comparison needs)
  --min-pass-rate <0-1>           pass if at least this fraction of attempts pass
  --concurrency <n>               attempts in flight (default 4)
  --repeat <n>                    attempts per case (overrides the suite)
  --timeout <ms>                  per-attempt timeout (default 30000)
  --tag <tag>                     only cases with this tag (repeatable)
  --case <id>                     only this case (repeatable)
  --label <text>                  label the run (e.g. a prompt version)
  --variant <name>                matrix suites: only this variant (repeatable)
  --judge <provider:model>        LLM judge model
  --no-judge-check                skip the one tiny call that checks the judge before any case runs
  --no-trace                      do not store the steps pipelines report
  --prices <file>                 extra/override model prices
  --no-color                      plain output (also honours NO_COLOR; set BEHAVTEST_ASCII=1 for ASCII symbols)
behavtest runs [--suite <name>] [--limit <n>]        List saved runs, newest first
behavtest show <run> [case] [--full]                 A run's summary, or one case's input, outputs, scores and trace
behavtest compare [base] [head] [options]            What regressed, improved, or is just flaky; runs are ids or run files
  --fail-on-regression | --significant-only        exit 1 when the gate fails
  --all  --json <file>  --md <file>  --suite <name>
behavtest report <run> [--against <base>] [--out <file>]   Single-file HTML report (runs are ids or run files)
behavtest export <run> [--out <file>] [--compact]    Write a run file (a baseline to commit, or to compare or import elsewhere)
behavtest import <file>                              Load a full run file into the database
behavtest calibrate [--labels <file>] [--min-kappa <k>] [--json <file>] [--md <file>]   How often the judge agrees with your labels
                                                   (default: the labels saved from the dashboard)
behavtest serve [--port <n>] [--host <host>] [--open] Local dashboard: runs, trends, compare, matrices, labels, calibration
behavtest matrix [id] [--reference <v>] [--list] [--md|--json|--out <file>]   Variants of a matrix side by side
behavtest init [--dir <dir>] [--force] [--ts]        Scaffold an example suite (--ts: a code suite, no server needed)
behavtest schema [--out <file>]                      Print the suite JSON Schema

Environment variables: OPENAI_API_KEY, OPENAI_BASE_URL, ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL, BEHAVTEST_JUDGE (default judge), NO_COLOR, BEHAVTEST_ASCII, plus any ${VAR} your suite references.