| Suite | A JSON file (or a TypeScript/JavaScript module) listing the pipeline to test and the test cases |
| Case | One input, an optional expected answer, and the scorers to run. Its id must stay stable across runs |
| Attempt | One execution of a case. With --repeat 5, each case has 5 attempts |
| Scorer | A check on an attempt's output or trace; returns pass or fail (or an error if it could not evaluate) |
| Verdict | Per case: passed (all attempts pass), failed (none pass), flaky (a mix), errored (no verdict possible, e.g. pipeline down) |
| Run | One execution of a suite, saved with every attempt, score and trace |
| Run file | A portable JSON copy of a run; compact run files hold only what comparisons need and are safe to commit |
| Baseline | The run you compare against, usually a committed compact run file |
| Trace | The steps an attempt took (LLM calls, tool calls, retrievals), reported by the pipeline |
| Expected docs | The ids of the documents a RAG case should retrieve (expectedDocs), for the retrieval scorer |
| Label | Your own pass/fail on a judged answer, used by behavtest calibrate to measure the judge |
| Variant / matrix | A variant changes the pipeline's config (a model, a prompt); a matrix is one run per variant of the same cases, compared side by side |
| Modified | A case whose definition, scorer code or judge model changed between two runs; listed, never counted as a regression |