BehavTest

BehavTest › Comparisons

BehavTest vs Promptfoo

Both are open-source command-line tools that run test cases against LLMs and applications and fit into CI. They differ in emphasis: Promptfoo is a broad evaluation and red-teaming toolkit with a large set of assertions and providers; BehavTest is a narrow tool for deciding, with statistics, whether a change made behavior worse than a baseline.

What is Promptfoo?

Promptfoo is an open-source (MIT) CLI and library, written in TypeScript, for evaluating and red-teaming LLM applications. Tests are declared in YAML configuration: prompts, providers (OpenAI, Anthropic, Google, Azure, local models, custom HTTP and code providers and more) and test cases with assert lists. Its documented assertions include deterministic checks (equals, contains, regex, is-json, latency, cost, tool-call validation), model-graded checks (llm-rubric, g-eval, factuality), RAG checks (context-faithfulness, context-recall, context-relevance, answer-relevance) and agent trajectory checks (trajectory:tool-used, trajectory:tool-sequence, trajectory:step-count and others); assertions can have thresholds and weights. promptfoo eval --repeat <n> runs each test several times, promptfoo view opens a local web viewer with side-by-side results, and a GitHub Action posts a before/after comparison on pull requests that change prompts. It runs locally, and it also has a substantial red-teaming feature set for security testing.

What is BehavTest?

BehavTest is an open-source (MIT) CLI and Node.js library for behavioral regression testing. Suites are JSON or TypeScript; the application under test is an HTTP endpoint (any language), an OpenAI-compatible or Anthropic model, or a function. It repeats every case, scores each attempt (deterministic checks, trace checks on tool calls, three RAG scorers, a hardened LLM judge, latency and cost limits, or custom functions), stores every run in SQLite, and compares runs case by case with Fisher's exact test and overall with a case-stratified permutation test. A comparison can fail CI only on statistically significant regressions. It also measures how often the LLM judge agrees with your own labels (Cohen's kappa).

Feature comparison

CapabilityBehavTestPromptfoo
ConfigurationJSON or TypeScript/JavaScript suitesYAML configuration (plus code for custom providers and assertions)
Repeated executionYes: --repeat per run, per-case repeatYes: --repeat <n>
Pass rate per case with a confidence intervalYes (Wilson 95%)Not documented
Statistical significance of a changeYes: Fisher's exact test per case, permutation test overallNot documented
Comparing two runsbehavtest compare against a previous run or a committed baseline fileSide-by-side results in the web viewer; the GitHub Action compares before and after on a pull request
CI gateExit code from compare (--fail-on-regression, --significant-only); GitHub Action with gateExit code on failed tests; GitHub Action comments with the comparison
Deterministic assertionsexactMatch, latencyCost, and custom functionsMany built in (equality, contains, regex, JSON, SQL, similarity metrics, latency, cost…)
LLM-as-a-judgellmJudge with rubric; injection-fenced, schema-validated, pre-run checkllm-rubric, g-eval, factuality and other model-graded assertions
Judge calibration against human labelsYes: behavtest calibrate (Cohen's kappa, false-pass rate)Not documented
RAG evaluationretrieval (hit, recall, precision, MRR vs expected document ids), faithfulness, contextRelevancecontext-faithfulness, context-recall, context-relevance, answer-relevance
Agent / tool-call evaluationtoolCalled, maxSteps on the steps the application reportsTool-call validation assertions and trajectory:* assertions
Red teaming / security testingNoYes, a major feature
Local results and viewerSQLite file; single-file HTML report; behavtest serve dashboardLocal web viewer (promptfoo view); sharing via promptfoo.app or a self-hosted server
LicenseMITMIT

Promptfoo facts from its introduction, command line, assertions, web viewer and GitHub Action documentation, checked 2026-09-29.

When each approach makes sense

Promptfoo fits when you want a broad toolkit: many assertion types without writing code, many providers side by side in one configuration, a web viewer to browse outputs, and red teaming for security testing. It is a natural choice for prompt engineering (comparing prompts and models on the same inputs) and for teams already standardized on YAML-configured evals.

BehavTest fits when the question is specifically "did this change make my application behave worse?" and your outputs are nondeterministic enough that a single run can't answer it: you want per-case pass rates with intervals, a significance test before failing a build, a committed baseline in the repository, and evidence that your LLM judge agrees with you. It is also a fit when you want to test the real application over HTTP in any language rather than prompts.

Together: nothing stops you using Promptfoo for exploration and red teaming and BehavTest as the statistical regression gate on pull requests.