BehavTest

BehavTest › Reference

Matrix runs: compare models and prompts side by side

A matrix runs the same cases through several pipeline setups, a model, a prompt version or a temperature, and compares them with the same statistics as compare. Add variants to a suite:

{
  "name": "model-shootout",
  "pipeline": { "adapter": "openai", "config": { "model": "gpt-4.1-nano", "system": "Answer with only the final answer." } },
  "variants": [
    { "name": "gpt-4.1-nano" },
    { "name": "gpt-4o-mini", "pipeline": { "config": { "model": "gpt-4o-mini" } } },
    { "name": "gpt-5.4-nano", "pipeline": { "config": { "model": "gpt-5.4-nano" } } },
    { "name": "gpt-6-luna", "pipeline": { "config": { "model": "gpt-6-luna" } } }
  ],
  "defaults": { "repeat": 3 },
  "cases": [ ... ]
}

behavtest run prints how many calls the matrix will make, then runs each variant as an ordinary run (labelled with the variant) and prints the comparison. This is examples/matrix/suite.json, run for real (12 questions × 3 attempts × 4 models, about $0.002):

behavtest matrix · model-shootout · 4 variants · matrix b4812ebe

  variant         pipeline             attempt pass rate [95% CI]  cases passed  flaky  p95 latency     cost  vs gpt-4.1-nano
  gpt-4.1-nano *  openai:gpt-4.1-nano  92% [78%–97%]                      11/12      0     2,382 ms  $0.0002  reference
  gpt-4o-mini     openai:gpt-4o-mini   100% [90%–100%]                    12/12      0     2,877 ms  $0.0003  +8.3 pts [+8.3 pts, +8.3 pts] p=0.101 not significant
  gpt-5.4-nano    openai:gpt-5.4-nano  89% [75%–96%]                      10/12      2     2,697 ms  $0.0005  -2.8 pts [-8.3 pts, +2.8 pts] p=1.000 not significant
  gpt-6-luna      openai:gpt-6-luna    100% [90%–100%]                    12/12      0     4,500 ms  $0.0008  +8.3 pts [+8.3 pts, +8.3 pts] p=0.106 not significant

  case                gpt-4.1-nano  gpt-4o-mini  gpt-5.4-nano  gpt-6-luna
  letters-strawberry         0/3 ✗        3/3 ✓         3/3 ✓       3/3 ✓
  bat-and-ball               3/3 ✓        3/3 ✓         1/3 ~       3/3 ✓
  decimal-compare            3/3 ✓        3/3 ✓         1/3 ~       3/3 ✓
  ...

Twelve questions cannot tell these models apart with confidence: every difference is "not significant". That is the point of the intervals. Add cases and attempts until the differences you care about are significant, or until you are confident they are small.