BehavTest

BehavTest › Integrations

OpenAI

BehavTest's built-in openai adapter sends each test case to an OpenAI model through the Chat Completions API, with your system prompt, and scores the answers. Use it to check that a prompt edit or a model swap (say, to a cheaper model) doesn't make behavior worse. OpenAI models can also be the LLM judge for any suite.

If what you want to test is your whole application (its retrieval, tools and code), not just a prompt, call the application instead: see HTTP services or Vercel AI SDK.

Installation

BehavTest needs Node.js 24 or newer. Nothing else is required: it calls the API over HTTPS itself, without the OpenAI SDK.

npx behavtest --version
export OPENAI_API_KEY=sk-...          # PowerShell: $env:OPENAI_API_KEY="sk-..."

Setup

Save as support.suite.json:

{
  "$schema": "https://unpkg.com/behavtest/schema/suite.schema.json",
  "name": "support-prompt-openai",
  "defaults": { "repeat": 5, "judge": "openai:gpt-4.1-nano" },
  "pipeline": {
    "adapter": "openai",
    "config": {
      "model": "gpt-4.1-mini",
      "system": "You are a concise support agent. Returns are accepted within 30 days.",
      "maxTokens": 300
    }
  },
  "cases": [
    {
      "id": "refund-window",
      "input": "Can I return an item after 40 days?",
      "expected": "No: returns are accepted within 30 days.",
      "scorers": ["llmJudge", "latencyCost"],
      "scorerConfig": {
        "llmJudge": { "rubric": "Does the answer state the 30-day limit and avoid promising an exception?" },
        "latencyCost": { "maxLatencyMs": 5000, "maxCostUsd": 0.001 }
      }
    },
    { "id": "one-word", "input": "Reply with only the word OK.", "expected": "OK", "scorers": ["exactMatch"] }
  ]
}

Options of the openai adapter (pipeline.config):

OptionMeaning
modelThe model id, e.g. gpt-4.1-mini (required)
systemThe system prompt: usually the thing you are testing
maxTokensOutput token limit, sent as max_completion_tokens (or max_tokens with maxTokensParam)
temperatureSent only if you set it; several current OpenAI reasoning models accept only their default
inputTemplateTurns an object input into a prompt, e.g. "{{question}}"
baseUrl, apiKeyEnvAnother endpoint or key variable (defaults: OPENAI_BASE_URL or https://api.openai.com/v1, and OPENAI_API_KEY)
retries, retryBaseDelayMsRetries for network errors, HTTP 429 and 5xx (honoring Retry-After)

A case's input can be a string (one user message) or { "messages": [...] } (a conversation history).

Running the test

npx behavtest run support.suite.json --label gpt-4.1-mini

Every case runs five times (defaults.repeat). Before any case runs, BehavTest makes one tiny call to the judge to check it works; a wrong model name or key stops the run with exit code 2 instead of erroring every attempt. Token usage from each response is priced with a bundled table (input, cached input and output separately), so the report shows pipeline and judge cost; an unknown model's cost is reported as unknown, never guessed.

Regression detection

Change something, run again, and compare:

# edit the system prompt, or change "model" to the candidate
npx behavtest run support.suite.json --label candidate
npx behavtest compare --fail-on-regression --significant-only

compare shows each case's pass rate before and after, whether the change is statistically significant, and an overall verdict. What it catches here: the answer no longer states the 30-day limit or promises an exception (judge), the one-word format breaks (exact match), or responses get slower or more expensive than the limits (latency and cost).

To compare several models at once, add variants to the suite and read them side by side with behavtest matrix: see matrix runs.

Choosing the judge

Use a different model from the one under test (BehavTest warns when they are the same: judges favor their own output), and pick by measured cost per verdict rather than list price: reasoning models can spend hundreds of hidden tokens on a single verdict. Before relying on a judge, measure how often it agrees with you.

CI/CD

Store the key as a repository secret and use the GitHub Action:

      - uses: dhrumilbhut/behavtest@v0
        with:
          suite: support.suite.json
          repeat: 5
          gate: significant
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

Every CI run makes real API calls: cases × attempts, plus judge calls. behavtest run prints the planned count before it starts.

Limitations