BehavTest

BehavTest › Reference

Scorers

ScorerWhat it does
exactMatchDeterministic equality with expected. By default trimmed, whitespace-normalised and case-insensitive. Options: caseSensitive, trim, normalizeWhitespace.
llmJudgeAsks an LLM to judge the output against a rubric (scorerConfig.llmJudge.rubric, default: "does the output correctly and completely address the input, matching the intent of the expected answer?"). Model: provider:model from scorerConfig.llmJudge.judge, --judge, defaults.judge, or BEHAVTEST_JUDGE.
latencyCostRecords latency and fails if maxLatencyMs or maxCostUsd is exceeded. With no thresholds it always passes. If maxCostUsd is set but the cost is unknown it reports an error, not a silent pass.
toolCalledChecks the pipeline's trace: was a tool called (with these arguments, this many times), or not called.
maxStepsChecks the trace: did the attempt finish within a step budget (optionally of one kind)?
retrievalRAG: were the case's expectedDocs retrieved? hit, recall, precision or mrr at k. Deterministic.
faithfulnessRAG, LLM judge: is the answer supported by the retrieved documents? One verdict, or claim by claim.
contextRelevanceRAG, LLM judge: were the retrieved documents relevant to the question?
your ownAny function in a code suite, or a scorer registered through the library.

The LLM judge

Judge scores are useful, but they are not ground truth. Studies find raw judge agreement overstates real accuracy, and judges can be talked into passing bad answers. BehavTest takes these precautions:

Choosing a judge model. Pick by measured cost per verdict, not list price: reasoning models can spend hundreds of hidden tokens on one verdict. In a small test (2026-09-23, two to four verdicts per model), gpt-5-nano (the lowest list price) used 376 to 888 output tokens per verdict, mostly hidden reasoning, and cost 5 to 12 times as much per verdict as gpt-4.1-nano or gpt-6-luna, which used 40 to 60.

It is still a single LLM making a judgment. Use an exact or programmatic check where you can, treat judge results as one signal, and measure how often the judge agrees with you before relying on it.