BehavTest

BehavTest › Learn

LLM testing

Testing an LLM application is not one activity but several, each answering a different question. Some of them are ordinary software tests that happen to sit next to a model; others only make sense because the model is nondeterministic. Mixing them up is how teams end up with either a flaky test suite nobody trusts or no tests at all. This page lays out the layers, what each one catches, and where a regression-testing tool such as BehavTest belongs among them.

The layers

  unit tests          deterministic code around the model         every commit, milliseconds, no API calls
  integration tests   the real model call works end to end        every commit or PR, a few calls
  evaluation          how good is it, on a dataset                when choosing or reporting, many calls
  regression tests    did this change make behavior worse         every PR that touches the AI path
  production checks   what is happening with real traffic         continuously, after release
LayerQuestion it answersDeterministic?BehavTest's role
Unit testsDoes my code (templates, parsers, tools) do what I wrote?YesNone: use your usual test runner
Integration testsDoes the whole path run: keys, network, schemas, tool wiring?MostlyPartly: a run errors (exit 1) when the pipeline fails
EvaluationHow good are the answers, in absolute terms?NoPartly: scorers and pass rates, but no benchmark library
Regression testsDid a change make behavior worse than a baseline?NoIts core purpose
Production checksIs live traffic going wrong right now?NoNone: use an observability tool

Unit tests: the code around the model

Most of an LLM application is ordinary code: building prompts, parsing responses, implementing tools, routing between steps. Test it the ordinary way, with the model replaced by a fixed response. These tests are fast, free and deterministic, and they catch the bugs that have nothing to do with the model:

// parseAnswer.test.ts: an ordinary unit test (vitest); no model is called
import { expect, test } from "vitest";
import { parseRefundDays } from "./parseAnswer.ts";

test("extracts the refund window from the model's answer", () => {
  expect(parseRefundDays("Returns are accepted within 30 days.")).toBe(30);
  expect(parseRefundDays("We don't accept returns on sale items.")).toBeNull();
});

If a test here is flaky, it is a real bug, exactly as in any other codebase.

Integration tests: the real call works

An integration test sends a real request through the real stack and checks that it completes: the key is valid, the model name exists, the response has the expected shape, the tool definitions are accepted. It usually runs once per case and asserts structure, not quality. Failures are mostly infrastructure: an expired key, a renamed model, a schema the provider now rejects.

BehavTest separates this kind of failure from a quality failure: an attempt where the pipeline could not produce an answer is errored, not failed, and errors make the run exit non-zero. Before any case runs it also checks that the configured LLM judge works, and stops with a configuration error (exit 2) if it doesn't.

Evaluation: how good is it?

Evaluation measures quality on a dataset: accuracy against references, rubric scores from an LLM judge, groundedness for RAG, task success for agents. It answers questions like "which model should we use?" or "how good is this assistant?", and produces numbers rather than a pass/fail decision. LLM evaluation covers methods, LLM-as-a-judge, and offline versus online evaluation.

Regression tests: did this change make it worse?

A regression test compares the behavior after a change with a baseline and decides pass or fail. Because the outputs are nondeterministic, it needs behavioral assertions instead of exact outputs, several attempts per case, and a statistical comparison instead of a single red/green result. That is what BehavTest does:

npx behavtest init --ts                      # a code suite with a stand-in agent, no API key needed
npx behavtest run behavtest/suite.mts --repeat 5 --label before
# edit the prompt or the agent
npx behavtest run behavtest/suite.mts --repeat 5 --label after
npx behavtest compare --fail-on-regression

LLM regression testing walks through the full workflow, including how many attempts to use and which gate to put in CI.

Nondeterminism changes what "a failing test" means

In the first two layers a failing test means something is broken. In the last three a single failed attempt may be the system's normal behavior: a case that is right 90% of the time fails one attempt in ten. Two consequences:

Behavioral assertions

Every nondeterministic layer checks properties rather than exact text. The cheapest property that captures what you care about is the best one:

BehavTest's built-in scorers cover these (exactMatch, toolCalled, maxSteps, retrieval, faithfulness, contextRelevance, llmJudge, latencyCost), and any other check is a function in a code suite.

CI/CD: what runs where

Model calls cost money and time, so not every layer belongs on every commit:

Where BehavTest fits, and where it doesn't

BehavTest covers the regression layer and part of the evaluation layer: it runs cases through the real application, repeats them, scores behavior, and compares runs statistically, locally and in CI. It does not replace unit tests for your own code, does not ship a benchmark or a large metric library, does not red-team for safety, and does not monitor production traffic. Pair it with the tools that do.