BehavTest

BehavTest › Integrations

Ollama and OpenAI-compatible servers

Ollama serves local models through an OpenAI-compatible API, so BehavTest's openai adapter can test them: point baseUrl at the local server and name the model. The same works for other servers that implement the OpenAI Chat Completions API (for example vLLM or OpenRouter). Use it to check a prompt on a local model, to compare a local model with a hosted one, or to test without API costs.

Installation

Install Ollama, pull a model, and check that the server answers:

ollama pull codellama:7b
curl http://localhost:11434/v1/models
npx behavtest --version        # Node.js 24 or newer

Setup

Save as ollama.suite.json:

{
  "name": "local-model",
  "defaults": { "repeat": 3 },
  "pipeline": {
    "adapter": "openai",
    "config": {
      "baseUrl": "http://localhost:11434/v1",
      "apiKeyEnv": "OLLAMA_API_KEY",
      "model": "codellama:7b",
      "system": "Answer with a single word, without punctuation.",
      "maxTokens": 20
    }
  },
  "cases": [
    { "id": "capital", "input": "What is the capital of France?", "expected": "Paris",
      "scorers": ["exactMatch", "latencyCost"], "scorerConfig": { "latencyCost": { "maxLatencyMs": 60000 } } },
    { "id": "color", "input": "What color is a clear daytime sky?", "expected": "Blue", "scorers": ["exactMatch"] }
  ]
}

Running the test

OLLAMA_API_KEY=ollama npx behavtest run ollama.suite.json
# PowerShell: $env:OLLAMA_API_KEY="ollama"; npx behavtest run ollama.suite.json

A real run of this suite against codellama:7b on a laptop (abridged):

  ✓ capital #2/3   9,568 ms  exactMatch ✓  latencyCost ✓
  ✓ capital #3/3   9,637 ms  exactMatch ✓  latencyCost ✓
  ✗ color #2/3       287 ms  exactMatch ✗ (expected "Blue", got "Creamy")
  ✓ color #3/3       264 ms  exactMatch ✓

  cases 2 · passed 1 · failed 0 · flaky 1 · errored 0
  attempts 6 · passed 5 · failed 1 · errored 0
  latency avg 6,449 ms · p95 9,637 ms
  cost pipeline unknown

  ~ flaky: color (2/3 attempts passed)

This is the nondeterminism BehavTest is built for: the same question got "Blue" twice and "Creamy" once, so the case is flaky, not simply passed or failed. Cost is reported as unknown because a local model has no price; BehavTest never guesses one. The first request is slow while the model loads, which is why the latency limit here is generous.

Regression detection

Change the prompt or the model (for example a newer local model), run again, and compare:

OLLAMA_API_KEY=ollama npx behavtest run ollama.suite.json --repeat 10 --label candidate
npx behavtest compare

Small local models are often quite random, so use enough attempts (10 or more per case) for the comparison to separate a real drop from noise, and gate on significant changes. See LLM regression testing.

A local model can also be the judge ("judge": "openai:<model>" with OPENAI_BASE_URL pointing at the server), but the judge requires JSON-schema structured output, which not every local server or model supports. BehavTest checks the judge before the run and stops with a clear message if it can't produce a valid verdict.

CI/CD

A GitHub-hosted runner has no local model server; run the suite on a machine that does (a self-hosted runner), or start the server in the job before running BehavTest. The commands are the same: behavtest run, then behavtest compare against a committed baseline. See baselines and CI.

Limitations