Two evaluation surfaces ship in the box — a golden-set quality harness and a deterministic self-benchmark — plus the published research that justified each architectural bet.
Write cases as YAML: the question, which sources should be retrieved (substring match on path/title), keywords that must appear, and optionally whether an LLM judge should score faithfulness.
cases:
- question: "What storage engine does the vector index use?"
expected_docs: ["lance"]
expected_keywords: ["lancedb"]
judge: true # LLM-scores answer faithfulness 0-5
Run it against any mode's retriever:
ragstack eval my_golden.yaml -k 8
Reports hit@k (did an expected source appear), MRR (how high it ranked), keyword pass-rate and mean faithfulness. Exit code signals pass/fail so it can gate CI.
ragstack bench --docs 50 generates a fixed synthetic corpus across five topics with known ground truth, then measures indexing throughput (chunks/s), retrieval latency (p50/p95) and retrieval quality (hit@k, MRR). It exists to catch regressions between versions on your own machine — the numbers are deliberately not comparable across machines or to other systems' published benchmarks.
| Decision | Evidence |
|---|---|
Bounded agentic loop + forced final answer | A-RAG / SoK agentic-RAG (2026): unbounded loops compound failure and cost without quality gains. |
Hybrid fusion + reranking as default | Anthropic contextual-retrieval study: contextual embeddings −35% failed retrievals; +contextual BM25 −49%; +reranker −67%. |
Contextual enrichment chosen over late chunking | Late chunking (arXiv:2409.04701) helps dense-only pipelines; blurbs improve BM25, rerankers and generation too — three consumers for one call. |
CRAG evidence grading + strip refinement | Yan et al. 2024: grading retrieved documents before generation adds ~3 accuracy points over RAG baselines. |
Query decomposition for compound asks | IRCoT (ACL 2023): interleaved retrieval improves retrieval by up to 21 points and QA by up to 15 on multi-hop benchmarks. |
Semantic cache with dynamic thresholds | Higress production report: ~50 ms recurrent-query latency at 0.95 similarity; stricter threshold under hedged phrasing. |
--no-cache semantics in mind: cached responses bypass retrieval and would mask regressions.