Traces: check what the agent did, not just what it said
An agent can give the right answer for the wrong reason, or the same answer after twice as many steps. If your pipeline reports its steps (LLM calls, tool calls, retrievals), BehavTest stores them with each attempt, shows them, and can score them.
From an HTTP pipeline, add steps to the response:
{
"output": "Your order shipped on Monday.",
"steps": [
{ "kind": "agent", "name": "order-agent", "startOffsetMs": 0, "durationMs": 78, "children": [
{ "kind": "retrieval", "name": "search", "durationMs": 9, "input": { "query": "order 123" } },
{ "kind": "tool", "name": "lookup_order", "durationMs": 25, "input": { "orderId": 123 }, "output": { "status": "shipped" } },
{ "kind": "llm", "name": "answer", "durationMs": 40 }
] }
]
}
kind is one of llm, tool, retrieval, agent, other; everything except kind and name is optional.
From a function pipeline, record steps with tracer(). Steps started inside another step become its children, and errors are recorded on the step:
import { tracer } from "behavtest"; // a value import: npm i -D behavtest
async function run(question: string) {
const t = tracer();
const docs = await t.step("retrieval", "search", () => search(question), { input: { query: question } });
const order = await t.step("tool", "lookup_order", () => lookupOrder(123), { input: { orderId: 123 } });
const text = await t.step("llm", "answer", () => answer(question, docs, order));
return { output: text, steps: t.steps };
}
Score the trace with two built-in scorers (a case using them errors, never passes, when the pipeline reported no trace):
| Scorer | Config | Passes when |
|---|---|---|
toolCalled | tool; optional argsInclude (the call's input contains these values; objects match partially), times (exact count), not | the tool was called (with those arguments, that many times), or with not: true, was not |
maxSteps | max; optional kind | the attempt took at most max steps (of that kind): catches loops and runaway retries |
"scorers": ["llmJudge", "toolCalled", "maxSteps"],
"scorerConfig": {
"toolCalled": { "tool": "lookup_order", "argsInclude": { "orderId": 123 } },
"maxSteps": { "max": 5, "kind": "retrieval" }
}
Your own scorers receive the full trace as trace in their arguments.
See it: behavtest show <run> <case> prints the step tree with durations (--full adds each step's input and output), and the HTML report has a collapsible trace with timing bars under each attempt. Run files include traces, except compact ones.
What is stored. Values under secret-looking keys (authorization, api_key, token, password...) are masked, step inputs and outputs longer than 20,000 characters are clipped, and at most 1,000 steps are kept per attempt; anything cut is marked. behavtest run --no-trace stores none; scorers still see them.