BehavTest

BehavTest › Integrations

HTTP services: Python, FastAPI and any language

BehavTest's http adapter tests your real application, whatever it is written in: it POSTs each test case to an endpoint you expose and scores the answer. Because the whole application runs (its prompts, retrieval, tools and business logic), this is the most faithful way to catch regressions, and the way to test frameworks BehavTest has no dedicated support for (LlamaIndex, CrewAI, Pydantic AI, Mastra and others).

The contract

BehavTest sends, for every attempt:

{ "input": "Where is order 123?" }

input is the case's input as written in the suite: a string, { "messages": [...] }, or any JSON object you choose. Your endpoint replies with:

{
  "output": "Order 123 was shipped on Monday.",
  "usage": { "inputTokens": 52, "outputTokens": 9 },
  "costUsd": 0.00012,
  "steps": [
    { "kind": "tool", "name": "lookup_order", "input": { "orderId": 123 }, "output": "shipped on Monday" }
  ]
}

Only output (a string) is required. The rest is optional and unlocks more checks:

FieldUsed for
outputEvery scorer. Another field name can be set with outputField
usageToken counts (inputTokens, outputTokens, cachedInputTokens, cacheWriteTokens, reasoningTokens), stored and shown
costUsdYour own cost figure for the attempt, used by cost limits
stepsA trace: kind (llm, tool, retrieval, agent, other), name, optional input, output, durationMs, startOffsetMs, attributes, children. Needed for the tool-call and RAG scorers
metadataAnything else you want stored with the attempt

A non-2xx status, invalid JSON or a missing output makes the attempt errored, not failed; network errors, 429 and 5xx responses are retried first.

Installation

npx behavtest --version                  # Node.js 24 or newer on the machine that runs the tests
pip install fastapi uvicorn              # for the Python example below

Setup: a FastAPI endpoint

Add one endpoint that calls your application and reports what it did. support_agent stands in for your code:

from fastapi import FastAPI

app = FastAPI()

ORDERS = {123: "shipped on Monday"}


def lookup_order(order_id: int) -> str:
    return ORDERS.get(order_id, "not found")


def support_agent(question: str) -> dict:
    # stand-in for your agent: call your model and tools here
    status = lookup_order(123)
    return {
        "text": f"Order 123 was {status}.",
        "tool_calls": [{"name": "lookup_order", "args": {"orderId": 123}, "result": status}],
        "usage": {"input_tokens": 52, "output_tokens": 9},
    }


@app.post("/behavtest")
def behavtest_endpoint(body: dict):
    result = support_agent(body["input"])
    return {
        "output": result["text"],
        "usage": {"inputTokens": result["usage"]["input_tokens"], "outputTokens": result["usage"]["output_tokens"]},
        "steps": [
            {"kind": "tool", "name": call["name"], "input": call["args"], "output": call["result"]}
            for call in result["tool_calls"]
        ],
    }

And a suite, agent.suite.json, that checks the answer's behavior through its steps:

{
  "$schema": "https://unpkg.com/behavtest/schema/suite.schema.json",
  "name": "support-agent",
  "defaults": { "repeat": 3 },
  "pipeline": { "adapter": "http", "config": { "url": "${PIPELINE_URL:-http://localhost:8000/behavtest}" } },
  "cases": [
    {
      "id": "order-status",
      "input": "Where is order 123?",
      "scorers": ["toolCalled", "maxSteps", "latencyCost"],
      "scorerConfig": {
        "toolCalled": { "tool": "lookup_order", "argsInclude": { "orderId": 123 } },
        "maxSteps": { "max": 5 },
        "latencyCost": { "maxLatencyMs": 3000 }
      }
    }
  ]
}

HTTP adapter options: url (required), method (POST, PUT or PATCH), headers (e.g. { "Authorization": "Bearer ${SERVICE_TOKEN}" }: secrets come from environment variables, never the file), outputField, retries, retryBaseDelayMs.

Running the test

uvicorn main:app --port 8000 &
npx behavtest run agent.suite.json
  ✓ order-status #1/3      24 ms  toolCalled ✓  maxSteps ✓  latencyCost ✓
  ✓ order-status #2/3      33 ms  toolCalled ✓  maxSteps ✓  latencyCost ✓
  ✓ order-status #3/3      33 ms  toolCalled ✓  maxSteps ✓  latencyCost ✓

Regression detection

Once the endpoint reports its steps, a suite can check far more than the final text: that the agent called the right tool with the right arguments (toolCalled), didn't loop (maxSteps), retrieved the right documents (retrieval, with expectedDocs on the case), stayed grounded in them (faithfulness), and stayed within latency and cost limits. Add an LLM judge with a rubric for the answer itself. After a change, behavtest compare reports which of those behaviors regressed and whether the change is larger than the noise. For a retrieval example, see LangChain or RAG.

CI/CD

Start the service in the CI job, then run BehavTest against it; with the GitHub Action:

      - run: pip install -r requirements.txt && (uvicorn main:app --port 8000 &) && sleep 3
      - uses: dhrumilbhut/behavtest@v0
        with:
          suite: agent.suite.json
          repeat: 3

Limitations