Library API and custom scorers
The easy way to add a scorer is an inline scorer in a code suite. To reuse scorers across projects, or to run BehavTest from your own program, register them on a registry and call the library:
import { createRegistry, loadSuite, runSuite, SqliteStore } from "behavtest";
const registry = createRegistry().registerScorer({
name: "mentionsParis",
async score({ output }) {
const pass = /paris/i.test(output);
return { pass, value: pass ? 1 : 0, reasoning: pass ? undefined : "never mentions Paris" };
},
});
const suite = loadSuite("suite.json", registry); // suites can now list "mentionsParis"
const store = new SqliteStore(".behavtest/results.db");
const outcome = await runSuite({ suite, registry, store, behavtestVersion: "custom" });
process.exitCode = outcome.exitCode;
score receives { input, expected, output, config, meta: { latencyMs, costUsd, usage, ... }, trace, runtime } and returns { pass, value, reasoning?, costUsd?, error?, metadata? }. Return error (rather than pass: false) when you couldn't evaluate, so the attempt is recorded as errored. Custom adapters work the same way through registerAdapter. The adapter and scorer interfaces are the library's stable contracts and change only additively.
Other exports include calibrate, cohensKappa, retrievedDocs, compareRuns, regressionGate, renderHtmlReport, renderRunMarkdown, renderCompareMarkdown, buildRunFile, readRunFile, tracer, defineSuite and the statistics helpers (wilsonInterval, fisherExact, stratifiedPermutationTest). Type definitions ship with the package.