<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>BehavTest blog</title>
    <link>https://dhrumilbhut.github.io/behavtest/blog/</link>
    <description>Articles on behavioral regression testing for AI applications, and on the decisions behind BehavTest, newest first.</description>
    <language>en</language>
    <atom:link href="https://dhrumilbhut.github.io/behavtest/blog/feed.xml" rel="self" type="application/rss+xml"/>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>AI Output Is Nondeterministic. Here's How to Test It Anyway.</title>
      <link>https://dhrumilbhut.github.io/behavtest/blog/ai-output-is-nondeterministic-how-to-test-it-anyway/</link>
      <guid isPermaLink="true">https://dhrumilbhut.github.io/behavtest/blog/ai-output-is-nondeterministic-how-to-test-it-anyway/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Dhrumil Bhut</dc:creator>
      <category>testing</category>
      <category>statistics</category>
      <category>ai-agents</category>
      <category>regression-testing</category>
      <description>Why pass or fail testing breaks for AI applications, and the statistics bug that shaped how BehavTest tells a real regression from noise.</description>
      <content:encoded><![CDATA[<p>Most teams shipping an AI feature test it by eye. Someone types a few questions into the chatbot, reads the answers, decides they look right, and ships. Then a prompt gets tweaked, or the model gets swapped for something cheaper, and nobody notices anything changed until a user complains.</p>
<p>The problem isn't that testing AI is hard because the code is hard. It's that the output refuses to sit still. Ask the same question twice and you can get two different answers, both reasonable, neither wrong. A test suite built for deterministic code, where the same input always produces the same output, has no idea what to do with that.</p>
<p>I found out how badly a testing tool can get this wrong by building one and watching it lie to me.</p>
<h2 id="one-pass-or-one-fail-tells-you-almost-nothing">One Pass or One Fail Tells You Almost Nothing</h2>
<p>If you run a test case against an LLM pipeline once and it passes, you've learned very little. Run it again and it might fail. Neither result tells you whether the pipeline is good, bad, or just noisy that day.</p>
<p>The only way around this is to stop treating a test case as a single question with a single answer, and start treating it as a question you ask several times, then look at the pattern. Five attempts at the same case, scored individually, gives you an actual pass rate instead of one lucky or unlucky roll. This sounds obvious once you say it out loud. It is not how most people test AI features, because it means every test run costs five times as many model calls, and most testing tools weren't built with that cost in mind from day one.</p>
<p>That decision, attempts as first class data instead of an afterthought, ended up mattering more than almost anything else in the design. It's also what made the next bug possible.</p>
<figure><img src="https://dhrumilbhut.github.io/behavtest/blog/ai-output-is-nondeterministic-how-to-test-it-anyway/ui-runs-dark.png" alt="The runs list, showing several saved runs with their pass rates and a flaky run flagged (demo data)" width="1280" height="900" loading="lazy" decoding="async"><figcaption>The runs list, showing several saved runs with their pass rates and a flaky run flagged (demo data)</figcaption></figure>
<h2 id="the-bug-in-my-own-statistics">The Bug in My Own Statistics</h2>
<p>Early in building BehavTest, I wired up a comparison feature: run a suite twice, once against a baseline configuration and once against a change, and tell me whether the change made things better or worse. The mechanism was straightforward. Run each case five times in both configurations, compare the pass rates, and use a statistical test to decide whether the difference was real or just noise.</p>
<p>I tested it against a deliberately bad change. Three test cases that used to pass every single time, five out of five attempts, started failing every single time, zero out of five. The suite's overall pass rate dropped from 88 percent to 67 percent. This is about as unambiguous a regression as you can construct on purpose.</p>
<p>My own tool looked at that and said: not significant.</p>
<p>It was wrong, and it was wrong for an interesting reason. My first implementation bootstrapped over cases: it resampled which test cases were in the suite, as if the suite were a random sample of every question anyone might ask. That answers a different question, &quot;will this be worse in general?&quot;, and with twelve cases the answer drowns in case-to-case variation. A regression gate asks something narrower: on this exact suite, did the pass rate move by more than the pipeline's own randomness? That randomness lives inside each case, in its five attempts, not in which cases happen to be in the suite.</p>
<p>The fix was a case stratified paired permutation test, with a separate within case bootstrap for the confidence interval. Each case is treated as its own paired observation, base versus head, and the test asks how likely a shuffle of labels would produce a gap this large by chance. I checked it against known reference values (Wilson intervals, Fisher's exact test on small samples) and then ran it against simulated data thousands of times: under a true null hypothesis it flags a false positive under 9 percent of the time at a 0.05 significance level, and it catches a real 90 percent to 40 percent drop more than 95 percent of the time.</p>
<p>That one bug reshaped how I think about the whole product. A testing tool that can be fooled by its own statistics is worse than no testing tool, because it hands you false confidence instead of an honest &quot;I don't know.&quot;</p>
<figure><img src="https://dhrumilbhut.github.io/behavtest/blog/ai-output-is-nondeterministic-how-to-test-it-anyway/compare-dark.png" alt="The compare view showing a statistically significant improvement, with the case-stratified permutation test result and per-case breakdown (demo data)" width="1280" height="1300" loading="lazy" decoding="async"><figcaption>The compare view showing a statistically significant improvement, with the case-stratified permutation test result and per-case breakdown (demo data)</figcaption></figure>
<h2 id="never-guess-and-never-go-green-on-a-broken-pipeline">Never Guess, and Never Go Green on a Broken Pipeline</h2>
<p>Once that lesson landed, it turned into a rule I apply everywhere else in the tool, not just in the comparison logic.</p>
<p>If a pipeline is down, or a judge model can't produce a verdict, the run errors. It does not silently pass, and it does not silently skip the case. Early real API testing against OpenAI found a model that rejected the temperature setting the judge asked for, and every judge call on it was failing outright. The correct behavior in that moment isn't to catch the error and move on quietly. It's to surface it loudly, because a testing tool that goes green on a broken pipeline is lying about the one thing it exists to tell you the truth about.</p>
<p>The same principle shows up in smaller places. If a model's pricing isn't confidently known, the reported cost is null, not zero. A model's output, the thing it actually produces, is treated as untrusted text, the same way you'd treat user input in a web app, because it's entirely possible for a chatty or manipulated model to try to talk its way past whatever is grading it. None of these are exotic ideas. They're the same defensive habits any backend engineer already applies to unreliable systems, just pointed at a part of the stack most eval tools treat as trustworthy by default.</p>
<h2 id="why-this-has-to-work-before-it-can-gate-a-pull-request">Why This Has to Work Before It Can Gate a Pull Request</h2>
<p>None of the statistics matter if the only place you can see them is a terminal output you have to remember to run by hand. The actual goal was always to get this into CI: commit a baseline once, and let a pull request fail automatically the moment a prompt or model change makes behavior measurably worse, the same way a broken unit test blocks a merge today.</p>
<p>That only works if the tool underneath it is honest. A CI gate built on top of statistics that can call an 88 percent to 67 percent collapse &quot;not significant&quot; would either train a team to ignore it, because it cries wolf, or worse, let real regressions through with a green checkmark attached. Getting the statistics right wasn't a nice to have on the way to a CI integration. It was the prerequisite for one being worth building at all.</p>
<p>The teams that end up finding out about a regression from an angry user in production are the ones treating AI output as fundamentally untestable, something you can only eyeball and hope. The teams that catch it in a pull request are the ones who accepted that the output is noisy, and built the statistics to tell a real change from noise anyway. That's a solvable problem. It just isn't solved by pass or fail.</p>
<p>If you're testing something similarly nondeterministic, I'd be curious how you're approaching it.</p>]]></content:encoded>
    </item>
    <item>
      <title>My RAG Eval Kept Failing the Cheap Model. The Model Wasn't the Problem.</title>
      <link>https://dhrumilbhut.github.io/behavtest/blog/my-rag-eval-kept-failing-the-cheap-model/</link>
      <guid isPermaLink="true">https://dhrumilbhut.github.io/behavtest/blog/my-rag-eval-kept-failing-the-cheap-model/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Dhrumil Bhut</dc:creator>
      <category>rag</category>
      <category>llm-judge</category>
      <category>evaluation</category>
      <category>prompt-engineering</category>
      <description>A real debugging story: why a cheap LLM judge scored 7 out of 12 on a RAG faithfulness check, and why the fix was the prompt, not the model.</description>
      <content:encoded><![CDATA[<p>RAG feels like magic in a demo and breaks in ways nobody can quite explain in production. The usual first instinct when an answer goes wrong is to ask whether the model retrieved the right document. The second instinct, once retrieval looks fine, is to blame the model for generating a bad answer anyway.</p>
<p>I built RAG specific scorers for BehavTest recently: one that checks whether an answer stays faithful to what was actually retrieved, and one that checks whether the retrieved documents were even relevant to the question in the first place. Both rely on an LLM judge to make the call. Judges are themselves a thing you have to test, and the first real test run taught me something I didn't expect: my own judge prompt was the bug, and the model I was about to write off as too weak had done nothing wrong.</p>
<h2 id="what-a-faithfulness-judge-is-actually-supposed-to-check">What a Faithfulness Judge Is Actually Supposed to Check</h2>
<p>A faithfulness scorer answers one question: is this answer supported by the text that was retrieved, or did the model add something that isn't actually in there? That sounds simple, but there are two ways to implement it, and the difference matters for cost as much as for correctness.</p>
<p>The straightforward version asks the judge for one verdict per answer: supported or not. The more useful version, claim mode, asks the judge to first break the answer into individual factual claims, then mark each one supported or unsupported against the retrieved text, with the source document cited per claim. Structured this way, claim mode costs roughly one to two times a single verdict call instead of one call per claim, because the judge does the decomposition and the scoring in one structured response rather than one call per fact.</p>
<p>A second scorer, context relevance, checks the retrieval side directly: for each document that came back, was it actually relevant to the question, independent of whether the final answer used it well. Between the two, you get a real answer to the question that matters most when a RAG pipeline goes wrong: was this a retrieval failure or a generation failure. That distinction is the entire reason claim level attribution is worth building in the first place, rather than a single collapsed faithfulness score that can't tell you which half of the pipeline actually broke.</p>
<h2 id="the-weak-model-that-wasnt">The Weak Model That Wasn't</h2>
<p>I ran the first real test against a small set of prompts, using a handful of different judge models so I could compare them against each other. One of the cheapest models on the list, the kind of model you'd pick specifically to keep judging costs down, scored 7 out of 12. The obvious read was that a model this cheap simply isn't strong enough to be trusted as a judge, and the honest move would have been to write that conclusion into the documentation and move on with a more expensive default.</p>
<p>Something about a 7 out of 12 result on a task this simple, judging whether a short answer is or isn't supported by a short passage, didn't sit right. It's not a hard reasoning problem. It's closer to reading comprehension. So instead of downgrading the model, I went looking at what the judge was actually being asked, and found three separate problems, all in the prompt, none in the model.</p>
<p>First, the faithfulness judge was seeing the original question alongside the answer and the retrieved context, and it was quietly scoring relevance to the question instead of faithfulness to the context. An answer that was completely supported by the retrieved text but slightly off topic relative to the question was getting marked unfaithful, which isn't what faithfulness means at all.</p>
<p>Second, in claim mode, the judge was occasionally listing the question itself as one of the answer's claims, then correctly noting the question wasn't supported by anything, which dragged the score down for a reason that had nothing to do with the actual answer.</p>
<p>Third, the context relevance scorer had a subtler bug: it was picking up the anti tampering nonce used to fence untrusted text in the prompt, and returning that nonce string as if it were a document id, instead of the id of an actual retrieved document.</p>
<p>None of these are model failures. They're all cases where the harness asked an ambiguous or slightly wrong question and then graded the answer to that wrong question as if it were the real one.</p>
<figure><img src="https://dhrumilbhut.github.io/behavtest/blog/my-rag-eval-kept-failing-the-cheap-model/ui-run-light.png" alt="The run detail view, showing a case's judge verdict alongside a &quot;check the judge&quot; pass/fail labeling control (demo data)" width="1280" height="900" loading="lazy" decoding="async"><figcaption>The run detail view, showing a case's judge verdict alongside a &quot;check the judge&quot; pass/fail labeling control (demo data)</figcaption></figure>
<h2 id="fixing-the-prompt-not-the-model">Fixing the Prompt, Not the Model</h2>
<p>Each fix was small on its own. Faithfulness judging was changed so the judge never sees the original question at all, only the answer and the retrieved context, which removes any path for relevance to leak into a faithfulness call. Document ids are now listed explicitly outside the fenced, untrusted portion of the prompt, and the judge's output schema enforces that any id it returns has to be one of the real ids offered, via an enum constraint rather than free text.</p>
<p>After those three changes, I reran the same comparison across models: 216 verdicts total. Two of the models, including a genuinely small one, went 36 out of 36 correct across both faithfulness modes. The model that had scored 7 out of 12 went 36 out of 36 on answer level faithfulness and 34 out of 36 in claim mode, with the two misses being one wrong claim call and one timeout, not a pattern of the model failing to understand the task.</p>
<p>The lesson generalizes past this one scorer: when a cheap model looks unreasonably bad at a task that shouldn't be hard, check what you're actually asking it before you conclude the model can't do the job. I'd told myself a clean story about model capability, and the real explanation was three separate, fixable, entirely mundane prompt bugs.</p>
<h2 id="retrieved-documents-are-untrusted-input-too">Retrieved Documents Are Untrusted Input Too</h2>
<p>There's a second failure mode worth testing for once the faithfulness prompt itself is fixed: what happens when one of the retrieved documents is actively trying to manipulate the judge. I ran a test where one retrieved document contained an instruction telling the judge to return a passing verdict regardless of the actual answer, the same category of attack as prompt injection against a chatbot, just aimed at the evaluator instead of the end user facing model.</p>
<p>Two of the models tested were never fooled by it, across every attempt. The cheapest model in the set was fooled once, out of two attempts in claim mode, and cited the injected document as its supporting source when it happened. That's a genuinely useful data point if you're choosing a judge model for a system where retrieved content isn't fully trusted, which in most real deployments, it isn't. Retrieved documents come from wherever your retrieval system points, and treating them as safe input by default is the same mistake as treating a chatbot's raw output as safe to render without escaping.</p>
<figure><img src="https://dhrumilbhut.github.io/behavtest/blog/my-rag-eval-kept-failing-the-cheap-model/ui-cal-dark.png" alt="The judge calibration view, showing Cohen's kappa, agreement percentage, and a confusion matrix of judge verdicts against human labels (demo data)" width="1280" height="900" loading="lazy" decoding="async"><figcaption>The judge calibration view, showing Cohen's kappa, agreement percentage, and a confusion matrix of judge verdicts against human labels (demo data)</figcaption></figure>
<h2 id="trusting-a-judge-requires-measuring-it-not-assuming-it">Trusting a Judge Requires Measuring It, Not Assuming It</h2>
<p>None of this replaces actually checking whether a judge agrees with a human. Pass rates from a judge can look reasonable and still be systematically wrong in a direction nobody notices, which is exactly what almost happened here before the prompt bugs were found. The real check is comparing the judge's verdicts against a small set of your own labeled examples, and reporting agreement with something like Cohen's kappa rather than trusting a raw percentage. A judge that agrees with itself ninety percent of the time and a judge that agrees with a human ninety percent of the time are very different claims, and only one of them is the one that actually matters.</p>
<p>If you're building judge prompts for a RAG evaluation, I'd be glad to compare notes on what you've found.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
