Matrix runs: compare models and prompts side by side
A matrix runs the same cases through several pipeline setups, a model, a prompt version or a temperature, and compares them with the same statistics as compare. Add variants to a suite:
{
"name": "model-shootout",
"pipeline": { "adapter": "openai", "config": { "model": "gpt-4.1-nano", "system": "Answer with only the final answer." } },
"variants": [
{ "name": "gpt-4.1-nano" },
{ "name": "gpt-4o-mini", "pipeline": { "config": { "model": "gpt-4o-mini" } } },
{ "name": "gpt-5.4-nano", "pipeline": { "config": { "model": "gpt-5.4-nano" } } },
{ "name": "gpt-6-luna", "pipeline": { "config": { "model": "gpt-6-luna" } } }
],
"defaults": { "repeat": 3 },
"cases": [ ... ]
}
behavtest run prints how many calls the matrix will make, then runs each variant as an ordinary run (labelled with the variant) and prints the comparison. This is examples/matrix/suite.json, run for real (12 questions × 3 attempts × 4 models, about $0.002):
behavtest matrix · model-shootout · 4 variants · matrix b4812ebe
variant pipeline attempt pass rate [95% CI] cases passed flaky p95 latency cost vs gpt-4.1-nano
gpt-4.1-nano * openai:gpt-4.1-nano 92% [78%–97%] 11/12 0 2,382 ms $0.0002 reference
gpt-4o-mini openai:gpt-4o-mini 100% [90%–100%] 12/12 0 2,877 ms $0.0003 +8.3 pts [+8.3 pts, +8.3 pts] p=0.101 not significant
gpt-5.4-nano openai:gpt-5.4-nano 89% [75%–96%] 10/12 2 2,697 ms $0.0005 -2.8 pts [-8.3 pts, +2.8 pts] p=1.000 not significant
gpt-6-luna openai:gpt-6-luna 100% [90%–100%] 12/12 0 4,500 ms $0.0008 +8.3 pts [+8.3 pts, +8.3 pts] p=0.106 not significant
case gpt-4.1-nano gpt-4o-mini gpt-5.4-nano gpt-6-luna
letters-strawberry 0/3 ✗ 3/3 ✓ 3/3 ✓ 3/3 ✓
bat-and-ball 3/3 ✓ 3/3 ✓ 1/3 ~ 3/3 ✓
decimal-compare 3/3 ✓ 3/3 ✓ 1/3 ~ 3/3 ✓
...
Twelve questions cannot tell these models apart with confidence: every difference is "not significant". That is the point of the intervals. Add cases and attempts until the differences you care about are significant, or until you are confident they are small.
- How variants combine: a variant's
configis merged overpipeline.config(nested objects merged, other values replaced); a variant with a differentadapterreplaces the config instead. In a code suite a variant'spipelinecan be a function. - Same cases, same judge: variants cannot change cases or the judge, so every variant is judged the same way and cases compare one to one. Case hashes do not include the pipeline, so nothing shows as
modified. behavtest matrix [id]shows the latest matrix (or one by id prefix;--listlists them).--reference <variant>picks what the others are compared with (default: the first).--md,--json, and--out report.html(a single-file report with a dot-and-interval chart and the case grid). The dashboard has a Matrix page.--variant <name>(repeatable) runs only some variants.behavtest compareand the dashboard's trends compare a run with earlier runs of the same variant.- Cost: a matrix multiplies calls (variants × cases × attempts); the count is printed before anything runs. Variants run one after another.