BehavTest

BehavTest › How-to guides

Test a RAG pipeline: retrieval and grounded answers

Report the retrieved documents as a retrieval step (output: a list of { id, text }, plain strings, or LangChain documents), list the ids each case should retrieve in expectedDocs, and combine the RAG scorers:

{
  "id": "refund-time",
  "input": "How long does a refund take?",
  "expectedDocs": ["refunds"],
  "scorers": ["retrieval", "faithfulness", "contextRelevance"],
  "scorerConfig": { "retrieval": { "metric": "recall", "k": 3 }, "faithfulness": { "mode": "claims" } }
}

retrieval is deterministic (no model); faithfulness and contextRelevance use the judge. A complete, runnable example with a small store-policy corpus is in examples/rag. See RAG.