BehavTest

BehavTest › Reference

RAG: test retrieval and grounded answers

A retrieval-augmented pipeline can fail in two places: it retrieves the wrong documents, or it answers with something the documents do not say. BehavTest scores both, from the documents the pipeline reports in its trace.

Report what was retrieved as a retrieval step. Its output is a list of documents: { "id": "refunds", "text": "..." }, plain strings (text only), or LangChain documents ({ pageContent, metadata: { id | source } }). Several retrieval steps are read in order, each id once:

{
  "output": "A refund is issued within 5 business days.",
  "steps": [
    { "kind": "retrieval", "name": "search", "output": [
      { "id": "refunds", "text": "Refunds go back to the original payment method. A refund is issued within 5 business days...", "score": 6.2 },
      { "id": "gift-cards", "text": "Gift cards never expire...", "score": 1.6 }
    ] },
    { "kind": "llm", "name": "answer" }
  ]
}
ScorerConfigValue / passes when
retrievalmetric: hit (default), recall, precision, mrr; optional k, min (default 1; required for precision)the metric at k for the case's expectedDocs, at least min. Deterministic: no model
faithfulnessmode: answer (default) or claims; min (claims, default 1); judgeanswer: every statement is supported by the retrieved text (1/0). claims: the supported fraction of the answer's claims, all checked in the same single judge call
contextRelevancemin (default: at least one relevant document); judgethe fraction of retrieved documents the judge rates relevant to the question