When to use BehavTest
Use BehavTest when:
- You changed a prompt, a model or a retrieval setting and want to know whether anything got worse before users notice.
- You are switching models (for example from GPT to Claude, or to a cheaper model) and need evidence that answers, latency and cost stay acceptable.
- You want a CI check that fails a pull request when it breaks your LLM feature, the way unit tests do for code.
- Your agent calls tools, and you need to test that it calls the right one with the right arguments, without looping.
- You run a RAG pipeline, and need to know whether it still retrieves the right documents and answers only from them.
- You rely on an LLM judge, and want evidence that it agrees with a human before you trust its scores.
- Your outputs are nondeterministic, so a single pass/fail is a coin flip and you need repeated attempts and a verdict on whether a change is real.
- You want to stay vendor-neutral and local: no hosted platform, no account, results in a file you own.
Something else may fit better if you need a hosted evaluation platform with a team UI, production observability and tracing of live traffic, or an extensive library of ready-made RAG metrics today. See prior art.