A full-stack vector-database-free RAG system for document search, evidence retrieval, and grounded question answering.
Vectorless RAG is a full-stack Retrieval-Augmented Generation (RAG) system designed to perform document search and question answering without using embeddings or a vector database.
The system uses a custom BM25 retrieval engine instead of vector similarity search. Documents are extracted into page- and section-aware chunks, indexed locally, and retrieved using lexical relevance with deterministic query and section routing.
A Groq-hosted LLM then generates the final answer strictly from the retrieved evidence, with source citations.
Upload a PDF, ask questions about the document, and get grounded answers based on the retrieved evidence.
- No vector database
- Custom BM25 lexical retrieval
- PDF document indexing
- Page- and section-aware chunking
- Deterministic query/section routing
- Section-aware reranking
- Groq-powered answer generation
- Grounded answers with source citations
- React + Vite frontend
- FastAPI backend
- Deployed using Vercel + Render
flowchart TD
subgraph Client[" Frontend Client (React + Vite + TypeScript)"]
UI["Modern UI / Upload & Chat Interface"]
Viewer["Evidence Viewer & Citation Inspector"]
end
subgraph Backend["Backend API (FastAPI)"]
API["FastAPI REST Controller"]
Endpoints["Endpoints:<br/><code>/api/upload</code> · <code>/api/query</code> · <code>/api/stats</code> · <code>/api/clear</code>"]
end
subgraph Ingestion["Document Ingestion & Indexing Pipeline"]
EXT["Document Extraction<br/>(PyPDF / TXT)"]
SEC["Section Detection & Header Boundary Analysis"]
CHUNK["Page & Section-Aware Chunking"]
TOK["Lexical Tokenization & Normalization"]
INDEX[("Local Inverted Index<br/>(Term Frequencies & Document Lengths)")]
end
subgraph Retrieval[" Vectorless Retrieval Engine (BM25)"]
INTENT["Query Intent Classifier & Section Router"]
BM25["BM25 Lexical Retrieval Engine<br/><code>k1 = 1.5, b = 0.75</code>"]
RERANK["Section-Aware Reranking & Score Booster"]
TOPK["Top Evidence Chunks Extractor"]
end
subgraph Generation[" LLM Generation & Citation Validation"]
PROMPT["Strict Grounded Prompt Construction"]
GROQ["Groq LLM Inference API"]
VAL["Citation & Grounding Validator"]
end
%% Interactions
UI -->|"Upload Document (PDF/TXT)"| API
UI -->|"Submit Question"| API
API --- Endpoints
Endpoints -->|"Document Bytes"| EXT
EXT --> SEC --> CHUNK --> TOK --> INDEX
Endpoints -->|"Query String"| INTENT
INTENT --> BM25
INDEX -.->|"Index Lookup"| BM25
BM25 --> RERANK
RERANK --> TOPK
TOPK -->|"Retrieved Context Chunks"| PROMPT
PROMPT --> GROQ
GROQ --> VAL
VAL -->|"Grounded Answer + Citations"| API
API -->|"JSON Response"| Viewer
flowchart LR
Q(["User Query"]) --> ID["1. Intent Detection<br/>& Section Routing"]
ID --> BM["2. BM25 Lexical<br/>Candidate Search"]
BM --> RR["3. Section-Aware<br/>Reranking"]
RR --> EV["4. Top Evidence<br/>Chunks Extraction"]
EV --> LLM["5. Grounded LLM<br/>Generation (Groq)"]
LLM --> ANS([" Verified Answer<br/>+ Source Citations"])
Vectorless_RAG/
│
├── backend/
│ ├── api.py # FastAPI REST API
│ ├── rag_engine.py # Custom BM25 retrieval engine
│ ├── groq_rag.py # Grounded Groq generation + citation validation
│ ├── test_rag.py # Retrieval/backend tests
│ ├── requirements.txt # Backend dependencies
│ └── .env # Local API credentials (not committed)
│
├── frontend/
│ ├── src/
│ │ ├── App.tsx # Main application UI
│ │ ├── rag_ui.tsx # RAG interface components
│ │ ├── index.css # Global styling
│ │ └── main.tsx # React entry point
│ ├── public/
│ ├── package.json
│ ├── vite.config.ts
│ └── index.html
│
├── README.md
└── .gitignore
The retrieval layer is intentionally vectorless.
The custom engine uses an inverted index and BM25 scoring with:
k1 = 1.5
b = 0.75
The retrieval pipeline:
- Extract document text.
- Detect document sections.
- Split content into retrieval chunks.
- Tokenize chunks.
- Build an inverted index.
- Calculate BM25 relevance.
- Detect query intent.
- Apply section-aware reranking.
- Return the strongest evidence chunks.
No embedding model or vector database is required for retrieval.
Short questions can have very little lexical overlap with the actual document text. The engine therefore detects likely section intent.
Query Detected Intent
What are his projects? Projects
What is his qualification? Education
What did he study? Education
What are his technical skills? Technical Skills
Has he published a paper? Publications
Where did he work? Work Experience
What does he do? Work Experience
Tell me what you analyzed from the PDF Document Summary
backend/groq_rag.py sends only retrieved document evidence to the Groq
API.
The generation layer is instructed to:
- answer only from retrieved evidence;
- avoid unsupported facts;
- handle short/conversational questions;
- include source citations;
- avoid fabricated citations.
Expected citation format:
[Source: Resume.pdf | Page: 1 | Section: Education]
Generated citations are validated against the retrieved document metadata. If the answer cannot be citation-verified, the backend can fall back to retrieved evidence rather than inventing a citation.
The frontend uses React, TypeScript, and Vite.
The UI communicates with:
/api/stats
/api/upload
/api/query
/api/clear
The interface displays document statistics, upload state, generated answers, retrieved evidence, and source/page/section information.
- Python 3.10+
- FastAPI
- Uvicorn
- PyMuPDF
- Requests
- python-dotenv
- Pydantic
Install:
cd backend
pip install -r requirements.txt- Node.js
- npm
Install:
cd frontend
npm installCreate:
backend/.env
and add:
GROQ_API_KEY=gsk_your_api_key_hereNever commit .env or API keys to GitHub.
From backend:
python -m uvicorn api:app --reloadBackend:
http://127.0.0.1:8000
Health check:
http://127.0.0.1:8000/health
Expected:
{"status": "ok"}In another terminal:
cd frontend
npm install
npm run devVite normally serves the frontend at:
http://localhost:5173
Returns the current index statistics.
Uploads and indexes a PDF or TXT file using multipart form data:
file=<document>
Request:
{
"query": "What are his projects?"
}Response:
{
"answer": "...",
"grounded": true,
"citations": [],
"reason": null,
"results": [
{
"source": "Resume.pdf",
"page": 1,
"section": "Projects",
"chunk": 1,
"score": 5.8,
"text": "..."
}
]
}Clears the current in-memory document index.
What is the qualification?
What did he study?
What are his projects?
What are his technical skills?
Where did he work?
Has he published any paper?
What does he do?
Tell me what you analyzed from the PDF.
flowchart LR
subgraph Trad[" Traditional Vector RAG"]
direction TB
TD1[" Raw Document"] --> TD2["Dense Embedding Model<br/>(OpenAI / HuggingFace)"]
TD2 --> TD3[("Vector Database<br/>(Pinecone / Chroma / FAISS)")]
TD3 --> TD4["Cosine Similarity Search<br/>(Approximate Nearest Neighbor)"]
TD4 --> TD5["LLM Generation"]
end
subgraph VLess[" Vectorless RAG (This Project)"]
direction TB
VD1[" Raw Document"] --> VD2["Page & Section Extraction<br/>(PyPDF / TXT)"]
VD2 --> VD3["Lexical Tokenization<br/>& Normalization"]
VD3 --> VD4[("Local Inverted Index<br/>(Term Frequency & Document Length)")]
VD4 --> VD5["BM25 Lexical Retrieval<br/>+ Section Intent Reranking"]
VD5 --> VD6["Grounded LLM Generation<br/>(Groq + Verified Citations)"]
end
- No vector database.
- No embedding model required for retrieval.
- Minimal storage overhead.
- Fast indexing for small/medium document collections.
- Transparent retrieval scores.
- Easy to inspect retrieved evidence.
- Exact source/page/section metadata can be preserved.
Lexical retrieval depends on terminology overlap and query wording. Section-aware routing and query expansion are used to improve conversational queries with weak direct keyword overlap.
