RAG: test retrieval and grounded answers
A retrieval-augmented pipeline can fail in two places: it retrieves the wrong documents, or it answers with something the documents do not say. BehavTest scores both, from the documents the pipeline reports in its trace.
Report what was retrieved as a retrieval step. Its output is a list of documents: { "id": "refunds", "text": "..." }, plain strings (text only), or LangChain documents ({ pageContent, metadata: { id | source } }). Several retrieval steps are read in order, each id once:
{
"output": "A refund is issued within 5 business days.",
"steps": [
{ "kind": "retrieval", "name": "search", "output": [
{ "id": "refunds", "text": "Refunds go back to the original payment method. A refund is issued within 5 business days...", "score": 6.2 },
{ "id": "gift-cards", "text": "Gift cards never expire...", "score": 1.6 }
] },
{ "kind": "llm", "name": "answer" }
]
}
| Scorer | Config | Value / passes when |
|---|---|---|
retrieval | metric: hit (default), recall, precision, mrr; optional k, min (default 1; required for precision) | the metric at k for the case's expectedDocs, at least min. Deterministic: no model |
faithfulness | mode: answer (default) or claims; min (claims, default 1); judge | answer: every statement is supported by the retrieved text (1/0). claims: the supported fraction of the answer's claims, all checked in the same single judge call |
contextRelevance | min (default: at least one relevant document); judge | the fraction of retrieved documents the judge rates relevant to the question |
- Errors, never passes, when there is nothing to compare: no retrieval step, no
expectedDocs(forretrieval), or documents without ids (retrieval) or text (the judge scorers). - Retrieved documents are untrusted input. A poisoned document can carry prompt injection, so the judge sees documents fenced like the answer and is told never to follow them. Long documents are clipped (4,000 characters each, 24,000 in total).
- Judge choice, measured. On the example pipeline (2026-09-28: 36 answers per judge and mode with known right verdicts, 12 of them with an invented claim),
gpt-4.1-miniandgpt-5.4-nanowere right every time in both modes;gpt-4.1-nanowas right 36/36 in answer mode and 34/36 in claims mode (one wrong verdict, one timeout). With a retrieved document that told the judge to pass everything,gpt-5.4-nanoandgpt-6-lunastill failed the invented claim every time, butgpt-4.1-nanowas fooled once in two claims-mode attempts (it cited the injected document as the source). Prefer a capable judge, check the sources in the reasoning, and measure your own data withbehavtest calibrate. faithfulnessandcontextRelevanceshare the judge's safeguards: structured output, fail-closed, the pre-run check, and the recorded judge model and temperature.- Try it:
examples/raghas a store-policy RAG pipeline withhealthy,degraded(retrieval breaks) andhallucinate(adds an unsupported claim) modes, and a suite for it:node examples/rag/server.mjs, thenbehavtest run examples/rag/suite.json.