BehavTest

BehavTest › Reference

Quickstart

Requires Node.js 24 or newer (see installation). Every command below also works with npx behavtest.

1. Try it with no API key

npx behavtest init --ts            # writes behavtest/suite.mts: a small suite with a stand-in agent
npx behavtest run behavtest/suite.mts

Or the JSON version, which tests a local mock HTTP service:

npx behavtest init                     # writes behavtest/suite.json and behavtest/mock-pipeline.mjs
node behavtest/mock-pipeline.mjs &     # start the mock pipeline (or use a second terminal)
npx behavtest run behavtest/suite.json
behavtest 0.8.0 · my-first-suite · http → localhost:4000/pipeline
  2 cases · concurrency 4

  ✓ capital-of-france     177 ms  exactMatch ✓  latencyCost ✓
  ✓ simple-math           175 ms  exactMatch ✓

  cases 2 · passed 2 · failed 0 · flaky 0 · errored 0
  latency avg 176 ms · p95 177 ms
  cost pipeline unknown

  All 2 cases passed.
  run 1219f529 saved → .behavtest/results.db

2. Test a prompt on OpenAI or Anthropic

Save as prompt.suite.json:

{
  "$schema": "https://unpkg.com/behavtest/schema/suite.schema.json",
  "name": "support-prompt",
  "defaults": { "judge": "openai:gpt-4.1-nano", "repeat": 3 },
  "pipeline": {
    "adapter": "openai",
    "config": { "model": "gpt-6-luna", "system": "You are a concise support agent. Returns are accepted within 30 days." }
  },
  "cases": [
    {
      "id": "refund-window",
      "input": "Can I return an item after 40 days?",
      "expected": "No: returns are accepted within 30 days.",
      "scorers": ["llmJudge"]
    },
    { "id": "one-word", "input": "Reply with only the word OK.", "expected": "OK", "scorers": ["exactMatch"] }
  ]
}
export OPENAI_API_KEY=sk-...           # PowerShell: $env:OPENAI_API_KEY="sk-..."
npx behavtest run prompt.suite.json --label prompt-v1
# edit the system prompt, then:
npx behavtest run prompt.suite.json --label prompt-v2
npx behavtest compare                    # what changed between the two runs, and is it real?

For Claude, use "adapter": "anthropic", a model such as "claude-haiku-4-5", and ANTHROPIC_API_KEY.

3. Test your own service (any language)

Expose one endpoint that takes { "input": ... } and returns { "output": "..." }, then point a suite at it: see test a Python, LangChain or other HTTP service.