Two pipelines carry everything: an ingestion pipeline that runs once per document, and a query pipeline that runs per question. This page walks both, end to end, with honest status markers for what is built versus planned.
Every file passes through five stages. Results are cached and fingerprinted, so re-indexing a changed corpus only touches what changed.
PDF, DOCX, PPTX, XLSX and scanned images go through Docling (layout-aware, OCR included). Markdown, logs and text are read as-is; code files are wrapped in language-tagged blocks; HTML main-content is extracted; CSVs become header plus sample rows; web pages are fetched politely (robots.txt aware, same-domain) and cleaned.
Structure-aware splitting keeps heading paths attached to every chunk ([Title > Section > Subsection]) and never cuts a code block. Chunks target ~512 tokens with sentence-level overlap so no idea straddles a boundary unnoticed.
With --enrich, one LLM sentence per chunk situates it inside its document ("this chunk defines the 2026 termination notice period"). The blurb is prepended before indexing, which improves keyword matching, embeddings, reranking and generation simultaneously — one fix, four beneficiaries. Cached by content fingerprint.
The same chunks land in three stores: dense vectors in LanceDB; pages and chunks in a tantivy BM25 index (the vectorless path — whole pages catch exact identifiers); entities and relations into the knowledge graph.
Louvain clustering groups tightly-connected graph entities; the LLM writes a summary per community. That is what lets "what are the main themes?" work at corpus scale.
| Index | Answers best | Engine | Status |
|---|---|---|---|
lexical | Exact names, IDs, error codes, statute numbers | tantivy BM25, pages + chunks | shipped |
vector | Meaning-level questions, paraphrases | LanceDB + BGE embeddings, normalized | shipped |
graph | Entity relationships, multi-hop connections | SQLite default · Neo4j optional, provenance edges | shipped |
communities | Corpus-wide themes and summaries | Louvain + LLM summaries | shipped |
recall | "Have we answered this before?" | LanceDB over stored Q&A pairs | shipped |
sql | Counts, totals, records, aggregates | Read-only SQLAlchemy catalog, schema-guarded | shipped |
sparse-learned | Vocabulary-mismatch expansion | SPLADE-class model | roadmap |
late-interaction | Fine-grained passage and visual-page scoring | ColBERT / ColPali style multivectors | roadmap |
visual | Charts, diagrams, page screenshots | Page-image + region embeddings | roadmap |
In auto mode, one cheap classification call produces a structured understanding of the question — intent, complexity, ambiguity, suggested strategy and dynamic top_k. The router then picks the minimum sufficient strategy. Explicit modes always win over suggestions.
| Route | When | Behaviour |
|---|---|---|
clarify | Ambiguous entity / time / document set | Asks a clarifying question instead of guessing |
conversational | No retrieval needed | Answers directly from history, tools withheld |
lexical | Exact identifiers, codes | BM25 single pass + synthesis, cited |
vector | Meaning-based lookup | Dense search single pass, cited |
hybrid | Ordinary factual questions | Dense + BM25 → RRF → cross-encoder rerank |
graph / global | Relationships / themes | Neighborhood walk / community map-reduce |
sql | Numbers when databases registered | Agent restricted to read-only SQL + chunk context |
agentic | Multi-hop, comparisons, research | Bounded ReAct loop across all eleven tools |
If the LLM is unavailable, a deterministic heuristic classifier keeps routing alive — compare-questions go agentic, aggregate questions go SQL when databases exist, identifier-shaped strings go lexical.
The agent is a bounded loop: understand → plan → retrieve → read → assess → repeat or synthesize. State carried between hops includes completed sub-questions, retrieved candidates, the citation registry and remaining step budget.
recall_memory.After generation, the answer is decomposed into atomic claims. Each claim is checked against the evidence it cites (supported / partial / unsupported), producing a machine-readable manifest. Confidence blends the retrieval-side evidence grade with the claim support ratio. When support falls below abstain_threshold, the answer is replaced by an explicit insufficiency notice — sources stay listed so you can judge yourself.
Every query writes one replayable trace to .ragstack/traces/<trace_id>.json: the route chosen, each tool call with arguments, verification results, citations and the final answer. start and done SSE events both carry the trace_id. Retention is bounded (default 200 traces) and can be disabled entirely.