Check that the LLM judge agrees with you
Label some judged answers yourself (Pass/Fail buttons in the dashboard, saved to the results database), and measure the agreement:
behavtest serve --open # open a run, mark judged answers Pass or Fail
behavtest calibrate --min-kappa 0.6 # reads the labels you saved
See judge calibration.