Judge calibration: does the judge agree with you?
An LLM judge's pass rate is only as good as the judge. behavtest calibrate compares its verdicts with your own labels on the same answers.
- Label. Run
behavtest serveand open a run: every judge verdict has Your label: Pass / Fail buttons, and each click is saved to the results database. Or, without a server, use the same buttons in an HTML report (behavtest report <run> --out report.html), where labels stay in your browser until you Export labels tobehavtest-labels-<run>.jsonl. You can also write that file yourself: one{ "run": "<id or prefix>", "case": "<id>", "attempt": 1, "scorer": "llmJudge", "label": "pass" | "fail" }per line (attemptdefaults to 1,scorertollmJudge). - Measure.
behavtest calibratereads the labels saved in the database;--labels <file>reads a file instead. For example, with 40 labels on one rubric:
behavtest calibrate
behavtest calibrate · 40 labels, 40 matched
llmJudge · judge openai:gpt-4.1-nano · "Cites the policy?"
labels 40 agreement 90% [77%–96%] kappa 0.80 [0.59, 0.95] (almost perfect agreement)
judge passed 2 of 20 answers you failed (false pass 10%) · failed 2 of 20 you passed (false fail 10%)
you: pass you: fail
judge pass 18 2
judge fail 2 18
- Per scorer, judge model and rubric: a judge is calibrated for one rubric, not in general.
- Cohen's kappa is agreement beyond chance (1 = perfect, 0 = chance); the interval is a bootstrap. The false-pass rate is how often the judge lets through an answer you would fail.
- Gate:
--min-kappa 0.6exits 1 unless every group has at least 30 labels and kappa at or above 0.6. Fewer than 30 labels is reported as too few, never as a pass. - Labels that match no stored verdict, labels on verdicts where the judge errored, and duplicates (the last one counts) are reported.
--json,--md. - The dashboard's Calibration page shows the same numbers live, and lists the verdicts where the judge disagreed with you, each linked to the answer.