RAGStack / Architecture

How a document becomes four indexes,
and a question becomes a verified answer.

Two pipelines carry everything: an ingestion pipeline that runs once per document, and a query pipeline that runs per question. This page walks both, end to end, with honest status markers for what is built versus planned.

The ingestion pipeline

§01 · INDEXING

Every file passes through five stages. Results are cached and fingerprinted, so re-indexing a changed corpus only touches what changed.

1 · Parse

PDF, DOCX, PPTX, XLSX and scanned images go through Docling (layout-aware, OCR included). Markdown, logs and text are read as-is; code files are wrapped in language-tagged blocks; HTML main-content is extracted; CSVs become header plus sample rows; web pages are fetched politely (robots.txt aware, same-domain) and cleaned.

2 · Chunk

Structure-aware splitting keeps heading paths attached to every chunk ([Title > Section > Subsection]) and never cuts a code block. Chunks target ~512 tokens with sentence-level overlap so no idea straddles a boundary unnoticed.

3 · Enrich (optional)

With --enrich, one LLM sentence per chunk situates it inside its document ("this chunk defines the 2026 termination notice period"). The blurb is prepended before indexing, which improves keyword matching, embeddings, reranking and generation simultaneously — one fix, four beneficiaries. Cached by content fingerprint.

4 · Index ×3

The same chunks land in three stores: dense vectors in LanceDB; pages and chunks in a tantivy BM25 index (the vectorless path — whole pages catch exact identifiers); entities and relations into the knowledge graph.

5 · Communities

Louvain clustering groups tightly-connected graph entities; the LLM writes a summary per community. That is what lets "what are the main themes?" work at corpus scale.

The index layer

§02 · INDEXES
IndexAnswers bestEngineStatus
lexical

Exact names, IDs, error codes, statute numbers

tantivy BM25, pages + chunks

shipped
vector

Meaning-level questions, paraphrases

LanceDB + BGE embeddings, normalized

shipped
graph

Entity relationships, multi-hop connections

SQLite default · Neo4j optional, provenance edges

shipped
communities

Corpus-wide themes and summaries

Louvain + LLM summaries

shipped
recall

"Have we answered this before?"

LanceDB over stored Q&A pairs

shipped
sql

Counts, totals, records, aggregates

Read-only SQLAlchemy catalog, schema-guarded

shipped
sparse-learned

Vocabulary-mismatch expansion

SPLADE-class model

roadmap
late-interaction

Fine-grained passage and visual-page scoring

ColBERT / ColPali style multivectors

roadmap
visual

Charts, diagrams, page screenshots

Page-image + region embeddings

roadmap

Query intelligence & adaptive routing

§03 · ROUTER

In auto mode, one cheap classification call produces a structured understanding of the question — intent, complexity, ambiguity, suggested strategy and dynamic top_k. The router then picks the minimum sufficient strategy. Explicit modes always win over suggestions.

RouteWhenBehaviour
clarify

Ambiguous entity / time / document set

Asks a clarifying question instead of guessing

conversational

No retrieval needed

Answers directly from history, tools withheld

lexical

Exact identifiers, codes

BM25 single pass + synthesis, cited

vector

Meaning-based lookup

Dense search single pass, cited

hybrid

Ordinary factual questions

Dense + BM25 → RRF → cross-encoder rerank

graph / global

Relationships / themes

Neighborhood walk / community map-reduce

sql

Numbers when databases registered

Agent restricted to read-only SQL + chunk context

agentic

Multi-hop, comparisons, research

Bounded ReAct loop across all eleven tools

If the LLM is unavailable, a deterministic heuristic classifier keeps routing alive — compare-questions go agentic, aggregate questions go SQL when databases exist, identifier-shaped strings go lexical.

The agent state machine

§04 · AGENT

The agent is a bounded loop: understand → plan → retrieve → read → assess → repeat or synthesize. State carried between hops includes completed sub-questions, retrieved candidates, the citation registry and remaining step budget.

  • Stop conditions: step budget exhausted; on the final hop tools are removed so the model must answer from evidence gathered — or say what is missing.
  • Evidence grading: every retrieval result is graded correct / ambiguous / incorrect; "incorrect" injects a corrective hint telling the agent to change strategy rather than force an answer.
  • Strip refinement: graded-correct evidence is split into sentences and only query-relevant strips continue onward.
  • Session continuity: follow-ups are rewritten into standalone queries using recent turns; the original phrasing still reaches the generator.
  • Cross-session recall: answered Q&A pairs are embedded into a recall store searchable by later sessions via recall_memory.

Verification & confidence

§05 · VERIFY

After generation, the answer is decomposed into atomic claims. Each claim is checked against the evidence it cites (supported / partial / unsupported), producing a machine-readable manifest. Confidence blends the retrieval-side evidence grade with the claim support ratio. When support falls below abstain_threshold, the answer is replaced by an explicit insufficiency notice — sources stay listed so you can judge yourself.

Design rule: generated context blurbs help retrieval but are never authoritative evidence; citations always resolve to original source text.

Audit traces

§06 · TRACES

Every query writes one replayable trace to .ragstack/traces/<trace_id>.json: the route chosen, each tool call with arguments, verification results, citations and the final answer. start and done SSE events both carry the trace_id. Retention is bounded (default 200 traces) and can be disabled entirely.