PolicyProof is an evidence-first Retrieval-Augmented Generation (RAG) and citation-verification system for public AI-governance and regulatory documents.
Instead of treating retrieval as a hidden implementation detail, PolicyProof evaluates retrieval quality explicitly, estimates whether retrieved evidence is sufficient to support an answer, returns source-derived citations, and abstains when evidence is weak.
The project is designed around evaluation, provenance, reproducibility, failure-aware AI behavior, and production deployment rather than unrestricted LLM generation.
PolicyProof is deployed on AWS using EC2, Docker, Nginx, private S3 artifact storage, and a least-privilege IAM instance role.
Application: http://3.128.255.20/
Health endpoint: http://3.128.255.20/api/health
The current deployment uses an EC2 public IPv4 address. The address may change if the instance is stopped and restarted unless a stable IP or domain is configured.
A detailed description of the offline evaluation pipeline, runtime retrieval
path, provenance model, and deployment boundaries is available in
docs/architecture.md.
flowchart LR
U[Browser] -->|HTTP :80| N[Nginx on EC2]
N -->|127.0.0.1:10000| D[PolicyProof Docker Container]
G[GitHub Repository] --> E[EC2 Deployment Host]
S[Private S3 Artifacts] -->|IAM Instance Role| E
D --> R[BM25 Retrieval]
R --> Q[Evidence Sufficiency Gate]
Q --> A[Answer + Citations or Abstain]
The deployed runtime uses:
- Amazon EC2 for compute
- Docker for application packaging
- Nginx as the public reverse proxy
- Amazon S3 for private deployment-artifact storage
- IAM instance roles for temporary AWS credentials
- loopback-only application binding on port
10000 - public HTTP traffic exposed through Nginx on port
80
No long-lived AWS access keys are stored in the repository or on the EC2 host.
Many RAG demos stop at:
documents -> vector database -> LLM -> answer
That structure does not answer several important engineering questions:
- Did retrieval actually find the relevant evidence?
- How does one retrieval strategy compare with another?
- Can evaluation results be reproduced later?
- Does the system know when retrieved evidence is insufficient?
- Can citations be traced back to exact source material?
- Can benchmark, dataset, and model artifacts be tied to immutable versions?
- Can deployment preserve the same artifact contracts used during evaluation?
PolicyProof treats those questions as first-class system requirements.
The resulting workflow is closer to:
controlled corpus
|
v
deterministic ingestion
|
v
provenance-preserving passages
|
v
retrieval benchmarking
|
v
evidence-sufficiency evaluation
|
v
frozen evaluation artifacts
|
v
retrieval + sufficiency gate
|
+------ sufficient ------> cited answer
|
+------ insufficient ----> abstain
The repository currently includes:
- four authoritative AI-governance source documents
- 707 token-safe passages with source provenance
- BM25 retrieval evaluation
- dense BGE-small retrieval evaluation
- hybrid candidate generation
- MiniLM cross-encoder reranking evaluation
- 80 evidence questions
- 160 evidence-sufficiency cases
- query-grouped train, validation, and test partitions
- construction-derived sufficient and insufficient evidence labels
- a frozen evidence-sufficiency baseline
- explicit abstention behavior
- browser demo and JSON CLI
- deterministic deployment artifacts
- SHA-256 artifact bindings
- AWS EC2 deployment
- private S3 artifact storage
- least-privilege IAM access
- Dockerized runtime
- Nginx reverse proxy
- 891 passing automated tests
PolicyProof compares multiple retrieval strategies against the same frozen evaluation corpus.
| Method | Recall@10 | MRR@10 | Direct Evidence Hit@10 | nDCG@10 |
|---|---|---|---|---|
| BM25 | 0.7760 | 0.7433 | 0.9375 | 0.6555 |
| Dense BGE-small | 0.9688 | 0.9062 | 1.0000 | 0.8866 |
| MiniLM reranker | 0.9271 | 0.8250 | 1.0000 | 0.7893 |
Dense retrieval remains the strongest benchmarked ranking method in the current evaluation.
The deployed portable demo intentionally uses deterministic BM25 retrieval rather than bundling the dense ONNX asset into the public runtime.
This distinction is explicit:
best evaluated retrieval:
Dense BGE-small
portable deployed runtime:
Deterministic BM25
The deployed runtime therefore does not imply that BM25 was the strongest retrieval method in the benchmark.
Retrieval alone does not guarantee that the retrieved evidence is sufficient to support an answer.
PolicyProof therefore includes a separate evidence-sufficiency stage.
The evaluation dataset contains complete, incomplete, and distractor evidence cases grouped by query identity.
train: 48 query groups
validation: 16 query groups
test: 16 query groups
Query-group isolation prevents related evidence variants from leaking between training and evaluation partitions.
| Metric | Result |
|---|---|
| Accuracy | 0.9024 |
| Balanced Accuracy | 0.8750 |
| F1 | 0.8571 |
These metrics use construction-derived silver labels.
They are engineering evaluation results and are not presented as independently human-adjudicated gold-label performance.
A runtime query follows this path:
user question
|
v
BM25 retrieval over 707 accepted passages
|
v
candidate evidence selection
|
v
evidence-sufficiency model
|
+------ evidence sufficient
| |
| v
| source-derived answer
| + citations
| + passage IDs
| + document IDs
| + retrieval metadata
|
+------ evidence insufficient
|
v
abstain
Responses expose the evidence and system metadata required to inspect how the decision was reached.
Depending on the query, the response can include:
answerorabstain- sufficiency probability
- frozen decision threshold
- source-derived citation excerpts
- document IDs
- passage IDs
- source labels
- BM25 scores
- ranking metadata
- label-provenance disclosures
Start the local demo and ask:
What risks does unauthorized voice generation create, and how does GPT-4o mitigate them?
A supported response can contain:
- an evidence-backed answer
- sufficiency probability
- source-derived excerpts
- document provenance
- passage provenance
- retrieval scores
An unsupported query should trigger explicit abstention instead of unsupported generation.
- Python 3.12
- Git
Clone the repository:
git clone https://github.com/Bad33/policyproof.git
cd policyproofCreate a virtual environment:
python3.12 -m venv .venv
source .venv/bin/activateInstall:
python -m pip install --upgrade pip
python -m pip install -e .Start the browser demo:
python -m policyproof.demo serve --openThen open:
http://127.0.0.1:8000/
Run a terminal query:
python -m policyproof.demo query \
"What risks does unauthorized voice generation create, and how does GPT-4o mitigate them?"The repository includes a production Dockerfile.
Build:
docker build -t policyproof:latest .Run:
docker run --rm \
-p 10000:10000 \
-e PORT=10000 \
policyproof:latestThen open:
http://127.0.0.1:10000/
Health check:
curl http://127.0.0.1:10000/api/healthThe current public deployment runs on Amazon EC2.
Detailed deployment documentation is available in:
GitHub repository
|
v
Amazon EC2
|
+--------------------------+
| |
v v
private S3 Docker build
artifacts |
^ v
| PolicyProof container
| |
| v
IAM instance role 127.0.0.1:10000
|
v
Nginx
|
v
Internet
The private deployment bucket contains:
retrieval-passages-v1.1.jsonl.gz
evidence-sufficiency-silver-baseline-v0.1.0.json
The bucket has public access blocked.
EC2 retrieves the files using an IAM instance role rather than long-lived access keys.
The EC2 deployment role is restricted to the PolicyProof deployment bucket.
Required permissions are limited to:
s3:ListBucket
s3:GetObject
No account-wide S3 administrative permission is required.
The PolicyProof Docker container is started with:
-p 127.0.0.1:10000:10000The Python application is therefore not directly exposed to the internet.
Nginx accepts public requests and forwards them internally to:
127.0.0.1:10000
The deployment security group permits:
22 SSH administrator IP only
80 HTTP public
443 HTTPS reserved for future TLS configuration
Application port 10000 is not publicly exposed.
The AWS deployment script is located at:
deploy/aws/deploy.sh
Set the deployment bucket:
export POLICYPROOF_BUCKET="your-private-policyproof-bucket"Then run:
./deploy/aws/deploy.shThe script:
- retrieves the accepted deployment artifacts from S3
- builds the Docker image
- removes the previous application container
- launches the new PolicyProof container
- performs a health check
- fails the deployment if the application does not become healthy
This keeps deployment behavior repeatable rather than relying on an undocumented sequence of manual commands.
Reproducibility is a core system requirement.
Published datasets and evaluation outputs are versioned and SHA-256 bound.
Tests enforce contracts around:
- corpus identity
- source provenance
- benchmark identity
- dataset construction
- train/validation/test splits
- model contracts
- frozen evaluation outputs
- deployment artifacts
- byte stability
- retrieval behavior
- evidence-sufficiency behavior
- end-to-end runtime behavior
The repository currently contains:
891 passing tests
Generated artifacts use no-overwrite behavior where appropriate so previously accepted results cannot silently be replaced.
PolicyProof maintains provenance across multiple stages.
A runtime evidence passage retains relationships to:
source document
|
v
source coordinates
|
v
logical retrieval unit
|
v
token-safe passage
|
v
retrieval result
|
v
citation
The system separates:
- retrieval text
- citation text
- source identity
- source coordinates
- passage identity
- retrieval score
This prevents retrieval-specific formatting from being confused with source-derived citation evidence.
policyproof/
├── data/
│ ├── deployment/
│ ├── evaluation/
│ ├── processed/
│ └── results/
│
├── deploy/
│ └── aws/
│ ├── deploy.sh
│ └── nginx-policyproof.conf
│
├── docs/
│ ├── architecture.md
│ ├── deployment.md
│ ├── engineering-decisions.md
│ └── assets/
│
├── evaluation/
│
├── scripts/
│
├── src/
│ └── policyproof/
│
├── tests/
│
├── Dockerfile
├── pyproject.toml
└── README.md
Key locations:
src/policyproof/— ingestion, retrieval, evaluation, sufficiency, and demodata/deployment/— deterministic transport artifacts used for deploymentdata/evaluation/— evidence cases, labels, splits, and frozen baselinesdata/results/— retrieval and reranking evaluation resultsdeploy/aws/— AWS deployment automation and Nginx configurationdocs/architecture.md— detailed system architecturedocs/deployment.md— AWS deployment architecture and proceduredocs/engineering-decisions.md— accepted technical decisions and limitationstests/— regression, integrity, artifact-binding, and end-to-end tests
PolicyProof deliberately makes several conservative engineering choices.
The deployed demo uses BM25 because it is:
- deterministic
- portable
- inexpensive
- reproducible
- independent of external model APIs
Dense retrieval remains benchmarked separately.
The system can refuse to answer when retrieved evidence is insufficient rather than attempting to generate a plausible unsupported response.
Accepted evaluation artifacts are versioned and cryptographically bound to reduce accidental benchmark drift.
Evidence variants belonging to the same query remain within the same dataset partition to reduce leakage.
Deployment artifacts are stored in private S3 rather than being exposed as public object URLs.
EC2 obtains S3 permissions through an IAM instance role instead of permanent AWS access keys.
PolicyProof is a research and compliance-support application.
It does not:
- provide legal advice
- determine legal compliance
- replace qualified legal or policy review
- guarantee that its source corpus is current or comprehensive
- claim independently human-adjudicated evidence-sufficiency accuracy
The current evidence-sufficiency evaluation uses construction-derived silver labels.
Runtime outputs should therefore be treated as evidence-support tooling rather than authoritative legal conclusions.
Current limitations include:
- the deployed runtime uses BM25 rather than the strongest evaluated dense retriever
- evidence-sufficiency labels are construction-derived
- the corpus is intentionally frozen and limited in scope
- the system does not perform live web retrieval
- the public deployment currently uses HTTP rather than a custom-domain HTTPS endpoint
- the EC2 public IPv4 address may change after instance stop/start
- PolicyProof is designed for research and compliance support, not autonomous legal decision-making
- Retrieval-Augmented Generation concepts
- BM25 retrieval
- BGE-small dense retrieval evaluation
- MiniLM cross-encoder reranking evaluation
- evidence-sufficiency modeling
- abstention
- provenance-aware citation handling
- Python 3.12
- NumPy
- ONNX Runtime
- deterministic HTTP demo server
- Recall@10
- MRR@10
- nDCG@10
- direct-evidence hit rate
- accuracy
- balanced accuracy
- F1
- query-group-isolated evaluation
- Amazon EC2
- Amazon S3
- AWS IAM
- Docker
- Nginx
- Ubuntu Linux
- SHA-256 artifact binding
- deterministic builds
- frozen evaluation artifacts
- regression testing
- health checks
- reproducible deployment
- least-privilege cloud access
PolicyProof currently has:
707 accepted passages
80 evidence questions
160 evidence-sufficiency cases
891 passing tests
Dense retrieval Recall@10: 0.9688
Dense retrieval MRR@10: 0.9062
Evidence sufficiency:
Accuracy: 0.9024
Balanced Accuracy: 0.8750
F1: 0.8571
Deployment:
AWS EC2 + Docker + Nginx + private S3 + IAM
The system is actively maintained as an applied AI engineering and reproducibility project.
