Skip to content

About

A full-stack, vector-database-free RAG system for uploading documents, retrieving relevant evidence, and generating grounded answers.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Vectorless RAG - Semantic Document Search

Vectorless RAG Demo

Try Live Demo

A full-stack vector-database-free RAG system for document search, evidence retrieval, and grounded question answering.


Overview

Vectorless RAG is a full-stack Retrieval-Augmented Generation (RAG) system designed to perform document search and question answering without using embeddings or a vector database.

The system uses a custom BM25 retrieval engine instead of vector similarity search. Documents are extracted into page- and section-aware chunks, indexed locally, and retrieved using lexical relevance with deterministic query and section routing.

A Groq-hosted LLM then generates the final answer strictly from the retrieved evidence, with source citations.

Try It Live

- Open Vectorless RAG

Upload a PDF, ask questions about the document, and get grounded answers based on the retrieved evidence.

Key Features

  • No vector database
  • Custom BM25 lexical retrieval
  • PDF document indexing
  • Page- and section-aware chunking
  • Deterministic query/section routing
  • Section-aware reranking
  • Groq-powered answer generation
  • Grounded answers with source citations
  • React + Vite frontend
  • FastAPI backend
  • Deployed using Vercel + Render

Architecture

flowchart TD
    subgraph Client[" Frontend Client (React + Vite + TypeScript)"]
        UI["Modern UI / Upload & Chat Interface"]
        Viewer["Evidence Viewer & Citation Inspector"]
    end

    subgraph Backend["Backend API (FastAPI)"]
        API["FastAPI REST Controller"]
        Endpoints["Endpoints:<br/><code>/api/upload</code> · <code>/api/query</code> · <code>/api/stats</code> · <code>/api/clear</code>"]
    end

    subgraph Ingestion["Document Ingestion & Indexing Pipeline"]
        EXT["Document Extraction<br/>(PyPDF / TXT)"]
        SEC["Section Detection & Header Boundary Analysis"]
        CHUNK["Page & Section-Aware Chunking"]
        TOK["Lexical Tokenization & Normalization"]
        INDEX[("Local Inverted Index<br/>(Term Frequencies & Document Lengths)")]
    end

    subgraph Retrieval[" Vectorless Retrieval Engine (BM25)"]
        INTENT["Query Intent Classifier & Section Router"]
        BM25["BM25 Lexical Retrieval Engine<br/><code>k1 = 1.5, b = 0.75</code>"]
        RERANK["Section-Aware Reranking & Score Booster"]
        TOPK["Top Evidence Chunks Extractor"]
    end

    subgraph Generation[" LLM Generation & Citation Validation"]
        PROMPT["Strict Grounded Prompt Construction"]
        GROQ["Groq LLM Inference API"]
        VAL["Citation & Grounding Validator"]
    end

    %% Interactions
    UI -->|"Upload Document (PDF/TXT)"| API
    UI -->|"Submit Question"| API
    API --- Endpoints

    Endpoints -->|"Document Bytes"| EXT
    EXT --> SEC --> CHUNK --> TOK --> INDEX

    Endpoints -->|"Query String"| INTENT
    INTENT --> BM25
    INDEX -.->|"Index Lookup"| BM25
    BM25 --> RERANK
    RERANK --> TOPK

    TOPK -->|"Retrieved Context Chunks"| PROMPT
    PROMPT --> GROQ
    GROQ --> VAL
    VAL -->|"Grounded Answer + Citations"| API
    API -->|"JSON Response"| Viewer
Loading

Core Query & Retrieval Pipeline

flowchart LR
    Q(["User Query"]) --> ID["1. Intent Detection<br/>& Section Routing"]
    ID --> BM["2. BM25 Lexical<br/>Candidate Search"]
    BM --> RR["3. Section-Aware<br/>Reranking"]
    RR --> EV["4. Top Evidence<br/>Chunks Extraction"]
    EV --> LLM["5. Grounded LLM<br/>Generation (Groq)"]
    LLM --> ANS([" Verified Answer<br/>+ Source Citations"])
Loading

Project Structure

Vectorless_RAG/
│
├── backend/
│   ├── api.py                 # FastAPI REST API
│   ├── rag_engine.py          # Custom BM25 retrieval engine
│   ├── groq_rag.py            # Grounded Groq generation + citation validation
│   ├── test_rag.py            # Retrieval/backend tests
│   ├── requirements.txt       # Backend dependencies
│   └── .env                   # Local API credentials (not committed)
│
├── frontend/
│   ├── src/
│   │   ├── App.tsx            # Main application UI
│   │   ├── rag_ui.tsx         # RAG interface components
│   │   ├── index.css          # Global styling
│   │   └── main.tsx           # React entry point
│   ├── public/
│   ├── package.json
│   ├── vite.config.ts
│   └── index.html
│
├── README.md
└── .gitignore

Retrieval Engine

The retrieval layer is intentionally vectorless.

The custom engine uses an inverted index and BM25 scoring with:

k1 = 1.5
b  = 0.75

The retrieval pipeline:

  1. Extract document text.
  2. Detect document sections.
  3. Split content into retrieval chunks.
  4. Tokenize chunks.
  5. Build an inverted index.
  6. Calculate BM25 relevance.
  7. Detect query intent.
  8. Apply section-aware reranking.
  9. Return the strongest evidence chunks.

No embedding model or vector database is required for retrieval.

Section-Aware Retrieval

Short questions can have very little lexical overlap with the actual document text. The engine therefore detects likely section intent.

Query Detected Intent


What are his projects? Projects What is his qualification? Education What did he study? Education What are his technical skills? Technical Skills Has he published a paper? Publications Where did he work? Work Experience What does he do? Work Experience Tell me what you analyzed from the PDF Document Summary

Grounded Generation

backend/groq_rag.py sends only retrieved document evidence to the Groq API.

The generation layer is instructed to:

  • answer only from retrieved evidence;
  • avoid unsupported facts;
  • handle short/conversational questions;
  • include source citations;
  • avoid fabricated citations.

Expected citation format:

[Source: Resume.pdf | Page: 1 | Section: Education]

Generated citations are validated against the retrieved document metadata. If the answer cannot be citation-verified, the backend can fall back to retrieved evidence rather than inventing a citation.

Frontend

The frontend uses React, TypeScript, and Vite.

The UI communicates with:

/api/stats
/api/upload
/api/query
/api/clear

The interface displays document statistics, upload state, generated answers, retrieved evidence, and source/page/section information.

Requirements

Backend

  • Python 3.10+
  • FastAPI
  • Uvicorn
  • PyMuPDF
  • Requests
  • python-dotenv
  • Pydantic

Install:

cd backend
pip install -r requirements.txt

Frontend

  • Node.js
  • npm

Install:

cd frontend
npm install

Environment Variables

Create:

backend/.env

and add:

GROQ_API_KEY=gsk_your_api_key_here

Never commit .env or API keys to GitHub.

Run the Project

1. Start FastAPI

From backend:

python -m uvicorn api:app --reload

Backend:

http://127.0.0.1:8000

Health check:

http://127.0.0.1:8000/health

Expected:

{"status": "ok"}

2. Start React

In another terminal:

cd frontend
npm install
npm run dev

Vite normally serves the frontend at:

http://localhost:5173

API Endpoints

GET /api/stats

Returns the current index statistics.

POST /api/upload

Uploads and indexes a PDF or TXT file using multipart form data:

file=<document>

POST /api/query

Request:

{
  "query": "What are his projects?"
}

Response:

{
  "answer": "...",
  "grounded": true,
  "citations": [],
  "reason": null,
  "results": [
    {
      "source": "Resume.pdf",
      "page": 1,
      "section": "Projects",
      "chunk": 1,
      "score": 5.8,
      "text": "..."
    }
  ]
}

POST /api/clear

Clears the current in-memory document index.

Example Queries

What is the qualification?
What did he study?
What are his projects?
What are his technical skills?
Where did he work?
Has he published any paper?
What does he do?
Tell me what you analyzed from the PDF.

Why Vectorless?

flowchart LR
    subgraph Trad[" Traditional Vector RAG"]
        direction TB
        TD1[" Raw Document"] --> TD2["Dense Embedding Model<br/>(OpenAI / HuggingFace)"]
        TD2 --> TD3[("Vector Database<br/>(Pinecone / Chroma / FAISS)")]
        TD3 --> TD4["Cosine Similarity Search<br/>(Approximate Nearest Neighbor)"]
        TD4 --> TD5["LLM Generation"]
    end

    subgraph VLess[" Vectorless RAG (This Project)"]
        direction TB
        VD1[" Raw Document"] --> VD2["Page & Section Extraction<br/>(PyPDF / TXT)"]
        VD2 --> VD3["Lexical Tokenization<br/>& Normalization"]
        VD3 --> VD4[("Local Inverted Index<br/>(Term Frequency & Document Length)")]
        VD4 --> VD5["BM25 Lexical Retrieval<br/>+ Section Intent Reranking"]
        VD5 --> VD6["Grounded LLM Generation<br/>(Groq + Verified Citations)"]
    end
Loading

Advantages

  • No vector database.
  • No embedding model required for retrieval.
  • Minimal storage overhead.
  • Fast indexing for small/medium document collections.
  • Transparent retrieval scores.
  • Easy to inspect retrieved evidence.
  • Exact source/page/section metadata can be preserved.

Trade-off

Lexical retrieval depends on terminology overlap and query wording. Section-aware routing and query expansion are used to improve conversational queries with weak direct keyword overlap.

About

A full-stack, vector-database-free RAG system for uploading documents, retrieving relevant evidence, and generating grounded answers.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages