Roadmap
Shipped: suites (JSON and code), HTTP / OpenAI / Anthropic / function pipelines, eight built-in scorers including RAG (retrieval, faithfulness, contextRelevance), repeats and flakiness, compare with significance tests, judge calibration against your labels, HTML / Markdown / JSON reports, a local dashboard (behavtest serve), run files and CI baselines, a GitHub Action, matrix runs across models and prompts, traces. Next: turning production failures into test cases, and a Python client. See CHANGELOG.md for what changed in each release.