Dashboard: browse, compare and label runs
behavtest serve # http://127.0.0.1:4800/
behavtest serve --open --port 5000 --db path/to/results.db
A local web dashboard on your results database, for looking around rather than gating CI. Runs still start from the CLI or CI; the dashboard reads what they saved.
- Runs: every run with its outcome, label, git commit, cases passed, attempt pass rate, flaky cases and cost, filterable by suite, and a pass-rate trend per suite (each run's attempt pass rate with its 95% interval; hover or use the arrow keys for details, click to open a run).
- Run: the same drill-down as the HTML report: summary, each case's input, attempts, outputs, scores, judge reasoning and traces (loaded when you open a case).
- Compare: pick any two runs for the full comparison: what regressed, improved or is flaky, with the significance tests from
behavtest compare. - Labels and calibration: mark judged answers Pass or Fail; labels are saved to the database as you click,
behavtest calibratereads them, and the Calibration page shows each judge's agreement, kappa, confusion matrix and the answers where it disagreed with you.
It is one plain page with no external assets, in light and dark themes, served by Node's own HTTP server (no extra dependencies). The only thing it writes is your labels. Its JSON API (/api/v1/runs, /api/v1/runs/<id>, /api/v1/compare?base=&head=, /api/v1/trend?suite=, /api/v1/calibration, /api/v1/labels) is available to scripts on the same machine.
Security: it listens on 127.0.0.1 only by default. It refuses requests whose Host is not localhost, an IP address or the host you started it with (so a web page cannot reach it through DNS rebinding), refuses label changes sent from other sites, and serves a strict Content-Security-Policy. There is no login: --host 0.0.0.0 makes your runs (inputs, outputs, traces) readable by anyone who can reach the port, and BehavTest prints a warning when you do it.