Skip to content

About

A RAG (Retrieval-Augmented Generation) AI chatbot that allows users to upload multiple document types (PDF, DOCX, TXT, CSV) and ask questions about the content. Built using LangChain, Hugging Face embeddings, and Streamlit, it enables efficient document search and question answering using vector-based retrieval. πŸš€

Topics

Resources

Stars

6 stars

Watchers

2 watching

Forks

Latest commit

Β 

History

18 Commits

Folders and files

Repository files navigation

Document Q&A: End-to-End Multi-Tenant RAG Pipeline

DATA β†’ CHUNK β†’ EMBED β†’ INDEX β†’ UNDERSTAND QUERY β†’ RETRIEVE β†’ FUSE β†’ RERANK
     β†’ BUILD CONTEXT β†’ GENERATE β†’ VERIFY β†’ CITE β†’ RESPOND

Memory, caching, security, evaluation, and observability wrap this pipeline.


πŸ“ Project Structure

RAG-MultiFile-QA/
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ rag/                       # Core RAG engine modules
β”‚   β”‚   β”œβ”€β”€ builtin/               # System built-in documents (HOW_TO_USE.md)
β”‚   β”‚   β”œβ”€β”€ chunking.py            # Structure-aware parent-child chunker
β”‚   β”‚   β”œβ”€β”€ ingestion.py           # Multi-format doc parsing & sanitizer
β”‚   β”‚   β”œβ”€β”€ store.py               # FAISS dense + BM25 sparse hybrid index
β”‚   β”‚   β”œβ”€β”€ query.py               # Intent classifier & query rewriter
β”‚   β”‚   β”œβ”€β”€ retrieval.py           # Hybrid retrieval + RRF + Cross-Encoder reranker
β”‚   β”‚   β”œβ”€β”€ context.py             # Context assembler & token budget optimizer
β”‚   β”‚   β”œβ”€β”€ generation.py          # LLM generator with streaming gate
β”‚   β”‚   β”œβ”€β”€ verification.py        # Claim-level NLI verifier & citation checker
β”‚   β”‚   β”œβ”€β”€ memory.py              # Short-term / long-term memory management
β”‚   β”‚   β”œβ”€β”€ cache.py               # SQLite-backed semantic & KV caching
β”‚   β”‚   β”œβ”€β”€ security.py            # Injection detection, PII redaction, rate-limiter
β”‚   β”‚   β”œβ”€β”€ evaluation.py          # Golden dataset evaluation metrics
β”‚   β”‚   β”œβ”€β”€ observability.py       # Distributed tracing & metrics logging
β”‚   β”‚   β”œβ”€β”€ pipeline.py            # End-to-end RAG pipeline coordinator
β”‚   β”‚   β”œβ”€β”€ cli.py                 # Backend CLI interface
β”‚   β”‚   └── config.py              # Centralized configuration dataclass
β”‚   β”œβ”€β”€ data/                      # Tenant data, indices, and cache persistence
β”‚   β”œβ”€β”€ logs/                      # Observability traces and metrics logs
β”‚   β”œβ”€β”€ requirements.txt           # Backend-specific dependencies
β”‚   └── __init__.py
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ app.py                     # Streamlit frontend application
β”‚   β”œβ”€β”€ ui/                        # UI stylesheet and custom assets
β”‚   β”‚   └── style.css
β”‚   β”œβ”€β”€ legacy/                    # Legacy single-file application archive
β”‚   β”‚   └── main_legacy.py
β”‚   β”œβ”€β”€ requirements.txt           # Frontend dependencies
β”‚   └── __init__.py
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_rag.py                # Unit and integration test suite
β”‚   β”œβ”€β”€ conftest.py                # Pytest configuration and path resolution
β”‚   β”œβ”€β”€ test_data/                 # Fixtures and evaluation datasets
β”‚   β”‚   β”œβ”€β”€ eval.jsonl
β”‚   β”‚   └── widgets.md
β”‚   └── __init__.py
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ rag_end2end_notes.md       # Architecture & engineering notes
β”‚   └── rag_end2end_notes.pdf      # PDF export of engineering notes
β”œβ”€β”€ .github/
β”‚   └── workflows/
β”‚       └── rag_test.yml           # CI workflow (pytest across Python versions)
β”œβ”€β”€ main.py                        # Unified launcher entry point
β”œβ”€β”€ requirements.txt               # Unified project dependencies
└── README.md

πŸš€ Quick Start

1. Installation

Using uv (recommended):

uv sync

Or using pip:

pip install -r requirements.txt

2. Run the Streamlit Web Application

# Using uv:
uv run streamlit run frontend/app.py

# Or using python3 / virtualenv:
python3 -m streamlit run frontend/app.py

# Offline demo mode (model-free):
RAG_OFFLINE=1 uv run streamlit run frontend/app.py

3. Backend CLI

# Ingest documents
uv run python -m backend.rag.cli --offline ingest tests/test_data/widgets.md

# Ask questions
uv run python -m backend.rag.cli --offline ask "How long is the warranty period?"

# Run evaluation suite
uv run python -m backend.rag.cli --offline eval tests/test_data/eval.jsonl --name base

4. Running Tests

uv run pytest -v

πŸ—οΈ Architecture & Layer Map

Layer Module Description
1 Ingestion ingestion.py, chunking.py, embeddings.py, store.py PDF/DOCX/TXT/MD/CSV parsing, boilerplate stripping, metadata extraction. Structure-aware parent-child chunking: sections form parents, sentence-packed children embedded with heading paths. FAISS dense + BM25 sparse hybrid index.
2 Query Understanding query.py, retrieval.py Normalization, conversational query rewriting, sub-query decomposition, metadata filter extraction, RRF (Reciprocal Rank Fusion), and Cross-Encoder reranking.
3 Context Engine context.py Deduplication (Exact / Cosine / Jaccard) β†’ MMR diversity β†’ parent expansion β†’ extractive compression β†’ token budget optimization β†’ [S1]..[Sn] citation labeling.
4 Generation generation.py, llm.py XML-delimited prompt formatting, canary token protection, NO_ANSWER abstention, and token-streaming gate.
5 Verification verification.py Atomic claim extraction, evidence matching, NLI (Natural Language Inference) entailment checking, numeric consistency check, and groundedness scoring.
6 Memory memory.py Short-term conversation history, semantic recall of older turns, rolling summarization, and PII-redacted long-term fact extraction.
7 Caching cache.py SQLite-backed LRU cache for embeddings, retrieval results, and LLM responses, plus semantic answer caching.
8 Security security.py Upload file validation (magic bytes, active content, zip bomb checks), prompt injection neutralization, PII redaction, and rate-limiting.
9 Evaluation evaluation.py, cli.py Recall, Precision, MRR, NDCG, groundedness, citation accuracy, and automated regression comparison.
10 Observability observability.py Per-request tracing (spans, events, timings), stage latency metrics, and PII-safe JSONL log export.

βš™οΈ Configuration

All configuration settings are defined in backend/rag/config.py (RAGConfig) and can be dynamically overridden via environment variables prefixed with RAG_ (e.g. RAG_LLM_MODEL, RAG_MAX_CONTEXT_ITEMS, RAG_DENSE_K).

About

A RAG (Retrieval-Augmented Generation) AI chatbot that allows users to upload multiple document types (PDF, DOCX, TXT, CSV) and ask questions about the content. Built using LangChain, Hugging Face embeddings, and Streamlit, it enables efficient document search and question answering using vector-based retrieval. πŸš€

Topics

Resources

Stars

6 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages