Compare models or prompt versions side by side
Add variants to the suite, each changing the pipeline's config, and run it once:
"variants": [
{ "name": "gpt-4.1-nano" },
{ "name": "gpt-5.4-nano", "pipeline": { "config": { "model": "gpt-5.4-nano" } } }
]
behavtest run suite.json --repeat 3 # one run per variant
behavtest matrix --out matrix.html # side by side: pass rate with intervals, cost, latency, each case
See matrix runs.