GitHub Action
Block the pull request that makes your LLM app worse. The Action installs BehavTest, runs your suite, compares it with a committed baseline, writes the comparison to the job summary, uploads the HTML report, and fails the check according to gate. See it on a demo repository: one pull request passes, the other is blocked with the two cases it broke.
name: BehavTest
on:
pull_request:
permissions:
contents: read
pull-requests: write # only for comment: true
jobs:
behavtest:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
# Start your pipeline here if the suite calls it over HTTP.
- uses: dhrumilbhut/behavtest@v0
with:
suite: behavtest/suite.json
repeat: 3
comment: true
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
Create the baseline once from a run you accept, and commit it: npx behavtest run behavtest/suite.json --repeat 3 --export behavtest.baseline.json --compact. Without a baseline the Action still runs and its summary says how to make one.
| Input | Default | Meaning |
|---|---|---|
suite | (required) | Suite file, relative to working-directory |
baseline | behavtest.baseline.json | The committed baseline run file |
gate | regression | regression: fail if any case regressed or errored, or the pass rate dropped significantly. significant: only significant regressions. cases: fail if any case failed, was flaky or errored in this run (with min-pass-rate, if too few attempts passed). none: never fail (configuration errors still do) |
repeat | suite's | Attempts per case; use the same number as the baseline |
min-pass-rate | With gate: cases, the fraction of attempts that must pass | |
judge | suite's | LLM judge model, e.g. openai:gpt-5.4-nano |
args | Extra behavtest run arguments, e.g. --tag smoke | |
comment | false | Keep one pull request comment up to date with the result (needs pull-requests: write; skipped outside pull requests, a warning if refused) |
report | true | Upload the HTML report, run file and summaries as an artifact |
working-directory, artifact-name, node-version, github-token | As named |
Outputs: result (pass, fail or error), regressed (number of regressed cases), run-id, report-path.
Choosing a gate. regression fails on any case whose pass rate dropped, which suits suites that are close to deterministic. If your pipeline is genuinely random, some cases will drop by chance on an unchanged branch: use gate: significant with enough attempts per case (5 or more) that a real drop can reach significance, or gate: cases with min-pass-rate. LLM regression testing shows the difference on real numbers.
The Action runs the BehavTest release that matches its tag (@v0 follows the latest 0.x release; pin @v0.8.0 for a fixed version). API keys come from your workflow's env, as for any step.
Updating the baseline is a reviewed change. A manual workflow that opens a pull request with a fresh baseline:
name: Update BehavTest baseline
on: workflow_dispatch
permissions:
contents: write
pull-requests: write
jobs:
baseline:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v7
with:
node-version: 24
- run: npx behavtest@0.8 run behavtest/suite.json --repeat 3 --export behavtest.baseline.json --compact || test $? -eq 1
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
- env:
GH_TOKEN: ${{ github.token }}
run: |
git switch -c behavtest-baseline-${{ github.run_id }}
git -c user.name=github-actions -c user.email=github-actions@users.noreply.github.com commit -am "Update BehavTest baseline"
git push -u origin HEAD
gh pr create --fill --title "Update BehavTest baseline"