BehavTest

BehavTest › Reference

Concepts

TermMeaning
SuiteA JSON file (or a TypeScript/JavaScript module) listing the pipeline to test and the test cases
CaseOne input, an optional expected answer, and the scorers to run. Its id must stay stable across runs
AttemptOne execution of a case. With --repeat 5, each case has 5 attempts
ScorerA check on an attempt's output or trace; returns pass or fail (or an error if it could not evaluate)
VerdictPer case: passed (all attempts pass), failed (none pass), flaky (a mix), errored (no verdict possible, e.g. pipeline down)
RunOne execution of a suite, saved with every attempt, score and trace
Run fileA portable JSON copy of a run; compact run files hold only what comparisons need and are safe to commit
BaselineThe run you compare against, usually a committed compact run file
TraceThe steps an attempt took (LLM calls, tool calls, retrievals), reported by the pipeline
Expected docsThe ids of the documents a RAG case should retrieve (expectedDocs), for the retrieval scorer
LabelYour own pass/fail on a judged answer, used by behavtest calibrate to measure the judge
Variant / matrixA variant changes the pipeline's config (a model, a prompt); a matrix is one run per variant of the same cases, compared side by side
ModifiedA case whose definition, scorer code or judge model changed between two runs; listed, never counted as a regression