BehavTest

BehavTest

Integrations

BehavTest tests your application where it runs. There are three ways to connect it, and every integration below is one of them:

Each page below has a working example, the command to run it, what regressions it catches, and its limits. The examples were run against the current versions of each library.

Model providers

IntegrationHow it connectsWhat you can test
OpenAIopenai adapter (Chat Completions)Prompts and models; OpenAI models as the LLM judge
Anthropic (Claude)anthropic adapter (Messages API)Prompts and models; Claude models as the LLM judge
Ollama and OpenAI-compatible serversopenai adapter with baseUrlLocal and self-hosted models

Applications and frameworks

IntegrationHow it connectsWhat you can test
HTTP services: Python, FastAPI, any languagehttp adapterYour real service end to end, including retrieval and tool steps
LangChainHTTP (Python) or a code suite (LangChain.js)Chains, retrievers (LangChain documents are read directly in JavaScript)
Vercel AI SDKA code suite calling generateTextAnswers, token usage, and the tool calls of multi-step agents

CI

IntegrationHow it connects
GitHub ActionsThe dhrumilbhut/behavtest@v0 Action: runs the suite on each pull request, compares it with a committed baseline, writes the job summary, fails the check
Any other CI systemTwo CLI commands and a committed run file; exit code 1 fails the job

Not listed here

Frameworks without a page (LlamaIndex, CrewAI, Pydantic AI, Mastra and others) can still be tested through the HTTP adapter: expose one endpoint that takes { "input": ... } and returns { "output": "..." }. BehavTest has no framework-specific support for them, so they have no page of their own.

To use BehavTest from your own program, see the library API.