DATA β CHUNK β EMBED β INDEX β UNDERSTAND QUERY β RETRIEVE β FUSE β RERANK
β BUILD CONTEXT β GENERATE β VERIFY β CITE β RESPOND
Memory, caching, security, evaluation, and observability wrap this pipeline.
RAG-MultiFile-QA/
βββ backend/
β βββ rag/ # Core RAG engine modules
β β βββ builtin/ # System built-in documents (HOW_TO_USE.md)
β β βββ chunking.py # Structure-aware parent-child chunker
β β βββ ingestion.py # Multi-format doc parsing & sanitizer
β β βββ store.py # FAISS dense + BM25 sparse hybrid index
β β βββ query.py # Intent classifier & query rewriter
β β βββ retrieval.py # Hybrid retrieval + RRF + Cross-Encoder reranker
β β βββ context.py # Context assembler & token budget optimizer
β β βββ generation.py # LLM generator with streaming gate
β β βββ verification.py # Claim-level NLI verifier & citation checker
β β βββ memory.py # Short-term / long-term memory management
β β βββ cache.py # SQLite-backed semantic & KV caching
β β βββ security.py # Injection detection, PII redaction, rate-limiter
β β βββ evaluation.py # Golden dataset evaluation metrics
β β βββ observability.py # Distributed tracing & metrics logging
β β βββ pipeline.py # End-to-end RAG pipeline coordinator
β β βββ cli.py # Backend CLI interface
β β βββ config.py # Centralized configuration dataclass
β βββ data/ # Tenant data, indices, and cache persistence
β βββ logs/ # Observability traces and metrics logs
β βββ requirements.txt # Backend-specific dependencies
β βββ __init__.py
βββ frontend/
β βββ app.py # Streamlit frontend application
β βββ ui/ # UI stylesheet and custom assets
β β βββ style.css
β βββ legacy/ # Legacy single-file application archive
β β βββ main_legacy.py
β βββ requirements.txt # Frontend dependencies
β βββ __init__.py
βββ tests/
β βββ test_rag.py # Unit and integration test suite
β βββ conftest.py # Pytest configuration and path resolution
β βββ test_data/ # Fixtures and evaluation datasets
β β βββ eval.jsonl
β β βββ widgets.md
β βββ __init__.py
βββ docs/
β βββ rag_end2end_notes.md # Architecture & engineering notes
β βββ rag_end2end_notes.pdf # PDF export of engineering notes
βββ .github/
β βββ workflows/
β βββ rag_test.yml # CI workflow (pytest across Python versions)
βββ main.py # Unified launcher entry point
βββ requirements.txt # Unified project dependencies
βββ README.md
Using uv (recommended):
uv syncOr using pip:
pip install -r requirements.txt# Using uv:
uv run streamlit run frontend/app.py
# Or using python3 / virtualenv:
python3 -m streamlit run frontend/app.py
# Offline demo mode (model-free):
RAG_OFFLINE=1 uv run streamlit run frontend/app.py# Ingest documents
uv run python -m backend.rag.cli --offline ingest tests/test_data/widgets.md
# Ask questions
uv run python -m backend.rag.cli --offline ask "How long is the warranty period?"
# Run evaluation suite
uv run python -m backend.rag.cli --offline eval tests/test_data/eval.jsonl --name baseuv run pytest -v| Layer | Module | Description |
|---|---|---|
| 1 Ingestion | ingestion.py, chunking.py, embeddings.py, store.py |
PDF/DOCX/TXT/MD/CSV parsing, boilerplate stripping, metadata extraction. Structure-aware parent-child chunking: sections form parents, sentence-packed children embedded with heading paths. FAISS dense + BM25 sparse hybrid index. |
| 2 Query Understanding | query.py, retrieval.py |
Normalization, conversational query rewriting, sub-query decomposition, metadata filter extraction, RRF (Reciprocal Rank Fusion), and Cross-Encoder reranking. |
| 3 Context Engine | context.py |
Deduplication (Exact / Cosine / Jaccard) β MMR diversity β parent expansion β extractive compression β token budget optimization β [S1]..[Sn] citation labeling. |
| 4 Generation | generation.py, llm.py |
XML-delimited prompt formatting, canary token protection, NO_ANSWER abstention, and token-streaming gate. |
| 5 Verification | verification.py |
Atomic claim extraction, evidence matching, NLI (Natural Language Inference) entailment checking, numeric consistency check, and groundedness scoring. |
| 6 Memory | memory.py |
Short-term conversation history, semantic recall of older turns, rolling summarization, and PII-redacted long-term fact extraction. |
| 7 Caching | cache.py |
SQLite-backed LRU cache for embeddings, retrieval results, and LLM responses, plus semantic answer caching. |
| 8 Security | security.py |
Upload file validation (magic bytes, active content, zip bomb checks), prompt injection neutralization, PII redaction, and rate-limiting. |
| 9 Evaluation | evaluation.py, cli.py |
Recall, Precision, MRR, NDCG, groundedness, citation accuracy, and automated regression comparison. |
| 10 Observability | observability.py |
Per-request tracing (spans, events, timings), stage latency metrics, and PII-safe JSONL log export. |
All configuration settings are defined in backend/rag/config.py (RAGConfig) and can be dynamically overridden via environment variables prefixed with RAG_ (e.g. RAG_LLM_MODEL, RAG_MAX_CONTEXT_ITEMS, RAG_DENSE_K).