BehavTest

BehavTest › Reference

Judge calibration: does the judge agree with you?

An LLM judge's pass rate is only as good as the judge. behavtest calibrate compares its verdicts with your own labels on the same answers.

  1. Label. Run behavtest serve and open a run: every judge verdict has Your label: Pass / Fail buttons, and each click is saved to the results database. Or, without a server, use the same buttons in an HTML report (behavtest report <run> --out report.html), where labels stay in your browser until you Export labels to behavtest-labels-<run>.jsonl. You can also write that file yourself: one { "run": "<id or prefix>", "case": "<id>", "attempt": 1, "scorer": "llmJudge", "label": "pass" | "fail" } per line (attempt defaults to 1, scorer to llmJudge).
  2. Measure. behavtest calibrate reads the labels saved in the database; --labels <file> reads a file instead. For example, with 40 labels on one rubric:
behavtest calibrate
behavtest calibrate · 40 labels, 40 matched

  llmJudge · judge openai:gpt-4.1-nano · "Cites the policy?"
    labels 40   agreement 90% [77%–96%]   kappa 0.80 [0.59, 0.95]  (almost perfect agreement)
    judge passed 2 of 20 answers you failed (false pass 10%) · failed 2 of 20 you passed (false fail 10%)
                 you: pass  you: fail
    judge pass          18          2
    judge fail           2         18