BehavTest

BehavTest › Reference

Why AI applications need behavioral regression tests

What problem does BehavTest solve?

You change a prompt, swap a model or tune retrieval, and something that used to work stops working: the bot no longer states the refund window, the agent calls the wrong tool, the RAG pipeline answers from the wrong document. Nothing crashes, so ordinary tests stay green, and you find out when a user complains. BehavTest turns "did this change make the application behave worse?" into a test you run before merging: the same cases, run the same way, compared with a known-good baseline.

Why single runs and snapshots are not enough

Traditional regression tests assume the same input gives the same output. LLM applications break that assumption twice:

So BehavTest repeats each case, treats its behavior as a pass rate, and asks whether the rate moved by more than the noise. On that same example, the unchanged bot scored 85% then 81% (overall change p ≈ 0.67: not significant), while a bot whose accuracy really dropped to 60% scored 56% (p < 0.001: a significant regression).

Who it is for

Developers and small teams who ship an LLM feature (a support bot, RAG search, an agent) and want a local, vendor-neutral check they can run on every change and in CI. It is not a hosted evaluation platform or production monitoring; see prior art for tools that are.