RAGStack / Evaluation

Every component must
earn its place.

Two evaluation surfaces ship in the box — a golden-set quality harness and a deterministic self-benchmark — plus the published research that justified each architectural bet.

Golden-set harness

E1 · EVAL

Write cases as YAML: the question, which sources should be retrieved (substring match on path/title), keywords that must appear, and optionally whether an LLM judge should score faithfulness.

cases:
  - question: "What storage engine does the vector index use?"
    expected_docs: ["lance"]
    expected_keywords: ["lancedb"]
    judge: true          # LLM-scores answer faithfulness 0-5

Run it against any mode's retriever:

terminal
ragstack eval my_golden.yaml -k 8

Reports hit@k (did an expected source appear), MRR (how high it ranked), keyword pass-rate and mean faithfulness. Exit code signals pass/fail so it can gate CI.

Deterministic self-benchmark

E2 · BENCH

ragstack bench --docs 50 generates a fixed synthetic corpus across five topics with known ground truth, then measures indexing throughput (chunks/s), retrieval latency (p50/p95) and retrieval quality (hit@k, MRR). It exists to catch regressions between versions on your own machine — the numbers are deliberately not comparable across machines or to other systems' published benchmarks.

The research behind the design

E3 · EVIDENCE
DecisionEvidence

Bounded agentic loop + forced final answer

A-RAG / SoK agentic-RAG (2026): unbounded loops compound failure and cost without quality gains.

Hybrid fusion + reranking as default

Anthropic contextual-retrieval study: contextual embeddings −35% failed retrievals; +contextual BM25 −49%; +reranker −67%.

Contextual enrichment chosen over late chunking

Late chunking (arXiv:2409.04701) helps dense-only pipelines; blurbs improve BM25, rerankers and generation too — three consumers for one call.

CRAG evidence grading + strip refinement

Yan et al. 2024: grading retrieved documents before generation adds ~3 accuracy points over RAG baselines.

Query decomposition for compound asks

IRCoT (ACL 2023): interleaved retrieval improves retrieval by up to 21 points and QA by up to 15 on multi-hop benchmarks.

Semantic cache with dynamic thresholds

Higress production report: ~50 ms recurrent-query latency at 0.95 similarity; stricter threshold under hedged phrasing.

Honesty rule: these are the studies that shaped choices here, not claims that RAGStack reproduces their headline numbers on your corpus. Run the harness on your own golden set before trusting any configuration.

Evaluation discipline

E4 · DISCIPLINE
  • Build gold sets from real queries plus hand-made hard cases (no-answer, adversarial, identifier-shaped).
  • Use deterministic metrics (hit/MRR/keyword) for gating; reserve the LLM judge for faithfulness, never as the only signal.
  • Add regression cases whenever the agent answers wrongly in practice — the golden file is the memory of failures.
  • Run evals with --no-cache semantics in mind: cached responses bypass retrieval and would mask regressions.