Scorers
| Scorer | What it does |
|---|---|
exactMatch | Deterministic equality with expected. By default trimmed, whitespace-normalised and case-insensitive. Options: caseSensitive, trim, normalizeWhitespace. |
llmJudge | Asks an LLM to judge the output against a rubric (scorerConfig.llmJudge.rubric, default: "does the output correctly and completely address the input, matching the intent of the expected answer?"). Model: provider:model from scorerConfig.llmJudge.judge, --judge, defaults.judge, or BEHAVTEST_JUDGE. |
latencyCost | Records latency and fails if maxLatencyMs or maxCostUsd is exceeded. With no thresholds it always passes. If maxCostUsd is set but the cost is unknown it reports an error, not a silent pass. |
toolCalled | Checks the pipeline's trace: was a tool called (with these arguments, this many times), or not called. |
maxSteps | Checks the trace: did the attempt finish within a step budget (optionally of one kind)? |
retrieval | RAG: were the case's expectedDocs retrieved? hit, recall, precision or mrr at k. Deterministic. |
faithfulness | RAG, LLM judge: is the answer supported by the retrieved documents? One verdict, or claim by claim. |
contextRelevance | RAG, LLM judge: were the retrieved documents relevant to the question? |
| your own | Any function in a code suite, or a scorer registered through the library. |
The LLM judge
Judge scores are useful, but they are not ground truth. Studies find raw judge agreement overstates real accuracy, and judges can be talked into passing bad answers. BehavTest takes these precautions:
- Prompt-injection resistant. The pipeline output is untrusted text. It is fenced inside a per-call random delimiter, and the judge is told everything inside is data, never instructions.
- Structured verdicts. The judge must return schema-validated JSON (
{reasoning, verdict}, reasoning first) using the provider's native structured output, at temperature 0. Models that accept only their default temperature (such as current OpenAI reasoning models) reject that; BehavTest then asks again without it, so those judges run at their default temperature, and it warns you, because their verdicts can vary more between runs. - Checked before the run. Before any case runs, BehavTest asks each judge one trivial question. If the judge cannot answer with a valid verdict (unknown model, bad key, no structured output), the run stops with exit 2 and says why, instead of erroring every attempt. It costs one tiny call per judge;
--no-judge-checkskips it. - Fail closed. A malformed, refused or failed judge response makes the attempt errored, never an implicit pass.
- Recorded. Judge spend is recorded separately from pipeline cost, and every verdict records which judge model produced it and at what temperature (
behavtest showand the HTML report display it). - Self-preference warning. BehavTest warns when the judge model is the same as the pipeline model (judges favour their own output).
- Changing the judge is a change, not a regression. The judge model is part of each judged case's identity, so
behavtest comparereports those cases asmodifiedwhen two runs used different judges.
Choosing a judge model. Pick by measured cost per verdict, not list price: reasoning models can spend hundreds of hidden tokens on one verdict. In a small test (2026-09-23, two to four verdicts per model), gpt-5-nano (the lowest list price) used 376 to 888 output tokens per verdict, mostly hidden reasoning, and cost 5 to 12 times as much per verdict as gpt-4.1-nano or gpt-6-luna, which used 40 to 60.
It is still a single LLM making a judgment. Use an exact or programmatic check where you can, treat judge results as one signal, and measure how often the judge agrees with you before relying on it.