Test a RAG pipeline: retrieval and grounded answers
Report the retrieved documents as a retrieval step (output: a list of { id, text }, plain strings, or LangChain documents), list the ids each case should retrieve in expectedDocs, and combine the RAG scorers:
{
"id": "refund-time",
"input": "How long does a refund take?",
"expectedDocs": ["refunds"],
"scorers": ["retrieval", "faithfulness", "contextRelevance"],
"scorerConfig": { "retrieval": { "metric": "recall", "k": 3 }, "faithfulness": { "mode": "claims" } }
}
retrieval is deterministic (no model); faithfulness and contextRelevance use the judge. A complete, runnable example with a small store-policy corpus is in examples/rag. See RAG.