RAGStack / Components

Every part, in plain English.

Where each module lives, what it does, and the knobs you can turn. File paths are relative to src/ragstack/.

Ingestion

ingestion/
ComponentWhat it does
parsers.py

File → Document. Docling for PDF/DOCX/PPTX/XLSX/images (layout + OCR); native readers for MD/TXT/code/CSV; trafilatura for HTML. Skips junk dirs (.git, node_modules…), oversized files, duplicate paths.

chunker.py

Structure-aware splitting: heading paths carried per chunk, code fences kept whole, token budgets with sentence-level overlap, small chunks merged under a size cap.

enricher.py

Optional contextual retrieval: one LLM sentence situating each chunk, prepended before indexing. Threaded, JSONL-cached by fingerprint — re-indexing costs nothing.

crawler.py

Same-domain crawler: robots.txt respected, delay between fetches, depth/page caps, main-content extraction.

Stores

stores/
StoreWhat it holds
lexical.py

tantivy BM25 index over whole pages (kind=page) and chunks (kind=chunk). Query sanitization with graceful fallbacks; delete-by-document for re-indexing.

vector.py

LanceDB fixed-size float32 vectors plus text/metadata. Dimension pinned via meta file — switching embedding models without resetting fails loudly instead of silently corrupting search.

graph/sqlite_graph.py

Embedded default graph: entities keyed by (normalized name, type), weighted relation rows with chunk provenance, extraction cache table, communities table. WAL mode, thread-locked.

graph/neo4j_graph.py

Drop-in Neo4j backend: same interface, MERGE-based upserts, constraint + index bootstrap. Enable with one config line.

sql_catalog.py

Registered databases → SQLAlchemy engines (cached). Schema introspection feeds the agent's prompts; queries are hard-guarded to SELECT/WITH/EXPLAIN with row caps.

GraphRAG engine

graphrag/
ModuleRole
extract.py

Per-chunk LLM extraction of typed entities + relationships as strict JSON; aggressive normalization (unknown types → other, dangling relations dropped); cached per chunk fingerprint so interrupted indexing resumes.

communities.py

Builds a NetworkX graph from aggregated relation weights, runs Louvain clustering, writes an LLM summary + keywords per community back into the store.

search.py

Local: seed entities matched by name tokens, walk up to max_hops, return entities + relations + their source chunks. Global: rank community summaries by keyword overlap, map-reduce synthesize.

Retrieval & quality layer

retrieval/ · cache.py · memory.py
ModuleRole
retrieval/vector_rag.py

Dense search, hybrid search (RRF fusion of dense + BM25) and reciprocal-rank fusion itself. Reranking hooks after both single-leg and fused passes.

retrieval/lexical_rag.py

BM25 helpers over pages and chunks with optional reranking.

retrieval/evaluator.py

CRAG evidence grader (heuristic or LLM judge) returning correct / ambiguous / incorrect with corrective hints; knowledge-strip refinement keeps only query-relevant sentences from graded evidence.

cache.py

Semantic response cache in LanceDB. Cosine threshold 0.95, tightened to 0.98 when questions contain uncertainty markers ("maybe", "possibly"). Bypass with --no-cache.

memory.py

SQLite conversation turns per session + co-reference query rewriting ("and its pricing?" → standalone query). RecallStore embeds answered Q&A for cross-session reuse.

The agent & its eleven tools

agent/

runner.py executes the bounded ReAct loop and streams typed events (start, route, rewrite, thought, tool_start/end, verification, answer, error, done). planner.py decomposes compound questions into parallel sub-searches. verifier.py produces the claim manifest and abstention decision.

ToolPurpose
hybrid_search

Fused dense+keyword search, reranked — the default first move.

decomposed_search

Splits compound questions, searches parts in parallel, fuses results.

search_chunks · search_pages

Pure BM25 over chunks / whole pages — exact terms and identifiers.

semantic_search

Dense-only meaning search.

graph_search

Entity neighborhood walk with linked source chunks.

community_overview

Corpus-wide thematic synthesis from community summaries.

recall_memory

Semantic search over past answered questions.

find_contradicting_evidence

Negation-probed search surfacing passages that may conflict with a claim.

sql_query

Read-only SELECT against registered databases.

fetch_url

Live page extraction (only when explicitly requested).

Mode scoping: tools are subsets per mode — e.g. sql exposes only sql_query + search_chunks; graph exposes graph + community + semantic. auto/agentic expose all eleven.

Interfaces & operations

web/ · cli.py · eval/
ModuleRole
web/app.py

FastAPI: SSE streaming on POST /api/query; /api/index, /api/crawl, /api/status, /api/modes; CORS for the hosted site; optional bearer auth on mutating routes.

cli.py

Typer app: index, crawl, query, watch, serve, modes, status, sessions, forget, db, eval, bench, reset.

eval/harness.py

Golden-set runner: hit@k, MRR, keyword pass-rate, LLM-judged faithfulness.

eval/bench.py

Deterministic synthetic corpus generator behind ragstack bench.