BehavTest

BehavTest › Reference

Adapters: what to test

HTTP (any language, any framework)

BehavTest POSTs { "input": <case input> } and expects { "output": "<string>" }:

{ "adapter": "http", "config": { "url": "http://localhost:8000/answer", "headers": { "X-Team": "search" } } }

Options: url, method (POST/PUT/PATCH), headers, outputField (default output), retries, retryBaseDelayMs.

The pipeline can also report costUsd, usage, steps (a trace) and metadata in its response. Cost and usage are used, and steps are stored and can be scored (see traces).

OpenAI (and anything OpenAI-compatible)

{ "adapter": "openai", "config": { "model": "gpt-6-luna", "system": "Be concise.", "maxTokens": 300 } }

Reads OPENAI_API_KEY. Set baseUrl (or OPENAI_BASE_URL) to use Ollama, vLLM, OpenRouter, Azure, or a local stub. temperature is sent only if you set it (current OpenAI reasoning models accept only their default).

Anthropic

{ "adapter": "anthropic", "config": { "model": "claude-haiku-4-5", "system": "Be concise.", "maxTokens": 300 } }

Reads ANTHROPIC_API_KEY (and ANTHROPIC_BASE_URL). maxTokens defaults to 1024 because the API requires it.

Both LLM adapters accept apiKeyEnv (to name a different env var), inputTemplate (e.g. "{{question}}", to turn an object input into a prompt), and retries. Retries happen only for network errors, HTTP 429 and 5xx (honouring Retry-After); latency is that of the final successful attempt.

A function in your own process

Write the suite in TypeScript or JavaScript and give it pipeline: { run: async (input) => ... }: see code suites.