Pause. Check. Then Act.
Privacy Brake is a human-centered cybersecurity safety layer that helps people review detected risk indicators before they share a file, open a link, or follow a QR destination. It presents context and a safer next step while leaving the decision with the user.
ShareGuard uses local OCR, sensitive-pattern matching, and opaque masking of selected detected regions. QRGuard decodes locally and sends URL payloads through the same deterministic LinkGuard service used for pasted links. LinkGuard checks URL structure without visiting the destination. An optional Gemini AI Analyst explains sanitized deterministic results; it cannot change findings or risk scores. These checks may miss indicators and do not guarantee safety. External threat intelligence is disabled and no provider adapter is currently implemented.
Sensitive details in screenshots and documents, misleading URLs, malicious QR destinations, and social-engineering prompts are easy to act on before a person notices the risk.
Three focused checks—ShareGuard, LinkGuard, and QRGuard—separate scanner signals from deterministic risk scoring and contextual explanation. QRGuard decodes uploaded codes locally, classifies the content, and displays basic URL indicators before the user chooses an action. The opt-in AI Analyst uses a minimized finding summary to explain possible consequences; deterministic analysis remains authoritative. Privacy Brake makes the evidence and next-step recommendation visible before the user acts. It is a safety layer, not a generic AI chatbot.
User action → local/deterministic detection → authoritative risk engine
→ optional sanitized Gemini context → explanation → user action
The React/TypeScript frontend calls a FastAPI JSON API. ShareGuard performs local Tesseract OCR, OpenCV QR-presence detection, deterministic sensitive-pattern matching, and selected-region masking. QRGuard uses OpenCV to decode uploaded QR images in memory and hands decoded URLs to the shared LinkGuard analyzer. LinkGuard normalizes and parses URL text, calculates deterministic indicators, and never visits submitted destinations. SQLAlchemy persists privacy-safe scan metadata in PostgreSQL. See docs/architecture.md, docs/privacy.md, docs/api.md, docs/testing.md, docs/deployment.md, and docs/demo.md.
- ShareGuard: local OCR and deterministic checks for common sensitive details, with normalized finding boxes.
- Safe Version: user-selected findings are covered with opaque masks; original upload bytes are held temporarily in process memory for this workflow.
- QRGuard: local QR decoding and URL classification; decoded links are never automatically opened.
- LinkGuard: deterministic URL structure checks without fetching destinations or following redirects.
- Risk Engine: transparent scanner-specific weights and consistent LOW / CAUTION / HIGH / STOP levels; AI cannot change the score or findings.
- AI Analyst: optional Gemini-assisted explanation, disabled by default; validated output is contextual, not authoritative detection.
- Risk Path: evidence-grounded display of the analysis steps and possible, not guaranteed, consequences.
- Frontend: React, TypeScript, Vite, Tailwind CSS, React Router
- Backend: Python 3.11+, FastAPI, Pydantic, SQLAlchemy
- Database: PostgreSQL
- Local development: Docker Compose,
.env, Git
-
Copy
.env.exampleto.env. Its database credentials are local-demo placeholders; replace them before sharing or using the configuration outside a local demo. Keep the Postgres andDATABASE_URLcredentials aligned. Threat intelligence and Gemini AI are disabled by default. -
From the project root, run:
docker compose up --build -d
-
Open the frontend at http://localhost:5173. The API is available at http://localhost:8000; interactive API docs are at http://localhost:8000/docs. PostgreSQL is published on
127.0.0.1:5432for local database tools only. -
Check service and database availability:
docker compose ps docker compose exec postgres pg_isready -U privacy_brake -d privacy_brake Invoke-RestMethod http://localhost:8000/api/v1/health
The backend waits for PostgreSQL health and creates missing tables at startup using SQLAlchemy create_all. This is suitable for a fresh demo database, not a replacement for production schema migrations. Uploaded files are not permanently stored. To stop services, run docker compose down. Do not add -v unless you intend to remove the local demo database volume.
AI Analyst is disabled by default. To explicitly enable the Gemini explanation provider, set AI_ENABLED=true, AI_PROVIDER=gemini, and GEMINI_API_KEY in .env. Optionally set GEMINI_MODEL and AI_TIMEOUT_SECONDS. The key stays in the backend environment. When enabled, only sanitized scan categories and bounded metadata are sent to Google's Gemini service; see docs/privacy.md before enabling.
Start PostgreSQL locally and set DATABASE_URL from .env.example. Then:
cd backend
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
uvicorn app.main:app --reload --no-access-logIn a second terminal:
cd frontend
npm install
npm run devFor a backend running on the host, set DATABASE_URL to a host-local PostgreSQL URL in the backend environment. Vite proxies /api to http://localhost:8000 by default; set VITE_API_PROXY_TARGET to change the development proxy. The frontend can also use the public VITE_API_BASE_URL build-time value. It defaults to same-origin /api/v1; it must never contain secrets.
| Variable | Purpose | Default / requirement |
|---|---|---|
DATABASE_URL |
Backend PostgreSQL connection. Use hostname postgres between Compose containers. |
Local Compose value in .env.example; required for production. |
POSTGRES_DB, POSTGRES_USER, POSTGRES_PASSWORD |
PostgreSQL initialization. Keep these aligned with DATABASE_URL. |
Local-only sample values; production must set secure values. |
APP_ENV, DEBUG |
Backend environment and debug mode. | development, false; production enforces DEBUG=false. |
CORS_ORIGINS |
Comma-separated frontend origins allowed by FastAPI. | Localhost in development; explicit HTTPS origins in production. |
VITE_API_BASE_URL |
Public frontend API base URL compiled into the frontend. | /api/v1; may be a deployed API URL. Never put a key or credential here. |
VITE_API_PROXY_TARGET |
Vite development-only proxy destination. | http://localhost:8000 outside Docker; Compose uses http://backend:8000. |
AI_ENABLED, AI_PROVIDER, GEMINI_API_KEY, GEMINI_MODEL, AI_TIMEOUT_SECONDS |
Optional server-side AI explanation. | Disabled by default. GEMINI_API_KEY stays backend-only. |
THREAT_INTEL_ENABLED, THREAT_INTEL_PROVIDER, VIRUSTOTAL_API_KEY |
Reserved configuration for optional reputation providers. | Disabled; no external provider adapter is implemented. |
A multi-stage frontend Dockerfile builds static assets and serves them with Nginx; the backend image runs as a non-root user and includes a database-aware health check. The deployment-target-neutral docker-compose.production.yml keeps PostgreSQL and the API off the public host ports and exposes the frontend container on HTTP_PORT (default 8080). Put a TLS-terminating HTTPS reverse proxy or hosting-provider ingress in front of it before public use. See docs/deployment.md for required production configuration and commands.
docker compose -f docker-compose.production.yml --env-file .env up --build -dProduction mode rejects debug, loopback database URLs, wildcard CORS, and non-HTTPS CORS origins. A placeholder .env.example is not production configuration and must not be deployed unchanged.
Follow docs/demo.md to generate a synthetic ShareGuard image and QR image, demonstrate selected-region sanitization, hand a QR URL to LinkGuard, and show the AI-disabled fallback. No real personal data or live suspicious destinations are needed.
With the Docker stack running:
docker compose cp backend/tests backend:/tmp/privacy-brake-tests
docker compose exec backend python -m unittest discover -s /tmp/privacy-brake-tests -vThe Phase 8 regression and synthetic-evaluation results are documented in docs/testing.md. The evaluation corpus contains synthetic text only; it is not a real-world accuracy benchmark.
All endpoints are under /api/v1.
| Method | Endpoint | Purpose |
|---|---|---|
| GET | /health |
Health and PostgreSQL connectivity |
| POST | /scans/image |
Analyze an uploaded PNG or JPEG using local ShareGuard OCR and detectors |
| POST | /scans/url |
Analyze a JSON { "url": "..." } request with local LinkGuard checks |
| POST | /scans/qr |
Decode an uploaded QR image locally and return classified content and deterministic indicators |
| GET | /scans |
List scan summaries (limit, offset) |
| GET | /scans/{scan_id} |
Get scan details and findings |
| POST | /scans/{scan_id}/sanitize |
Return a sanitized PNG for selected ShareGuard findings |
File uploads are limited to 10 MiB. ShareGuard accepts JPEG and PNG; QRGuard accepts JPEG, PNG, and WebP. The backend validates file signatures and decodes images in memory, with a 25-megapixel limit. FastAPI may spool request bodies and pytesseract may create a short-lived local OCR input file that is removed after processing; uploaded files are not permanently stored. QRGuard's decoded content is returned only in the immediate upload response for user review; raw QR content is not persisted in scan history or application signals. For sanitization, the original ShareGuard upload bytes are held in a bounded, process-local memory cache for up to 30 minutes (maximum 128 MiB total); entries are periodically expired and cleared on backend shutdown. Restarting the backend invalidates pending safe-version requests, so users must upload again. The database stores normalized detections and masked previews, not image bytes or raw OCR text. Sanitized images are returned to the browser and never stored in PostgreSQL.
All three scan endpoints include an ai_analysis object. When AI is disabled or unavailable, the object reports that state and still includes a locally generated explanation, evidence-linked points, possible risk path, and recommended action. The provider-generated summary is stored with provider/model/status metadata in an existing threat_checks row; prompts and raw user inputs are not stored. The AI cannot affect deterministic scores, levels, findings, or signals.
The application shell provides responsive Dashboard, ShareGuard, QRGuard, LinkGuard, and Scan History navigation, accessible keyboard focus, and visible privacy disclaimers. Scan results consistently pair the risk level with a numeric score/meter and the wording “risk score based on detected indicators.” Risk paths use the returned analysis steps and label possible consequences as uncertain.
- ShareGuard: upload → local analysis → review labeled findings/bounds → select regions → create and review the Safe Version → download only the sanitized image.
- QRGuard: upload → decode locally → review non-clickable payload and indicators → optionally use Analyze Destination to open the internal LinkGuard route.
- LinkGuard: paste → analyze as text → review redacted URL, deterministic indicators, Risk Path, and AI-assisted explanation/fallback.
No demo link opens an external URL. The app footer notes that detection is not guaranteed and AI explanations may be imperfect.
ShareGuard defaults high/critical findings to selected and lets the user deselect any region. Creating a Safe Version requires at least one selection. The backend accepts finding IDs only, validates ownership against the scan, retrieves its authoritative stored bounding boxes, and masks selected regions with opaque black pixels. It does not rerun OCR or trust browser-supplied boxes. The operation records a SANITIZE / SUCCESS safe-action audit entry with the mask method, selected count, and finding categories; it does not store image bytes or raw PII. The user receives a PNG download and before/after comparison. Review the safe version before sharing; automated detection may miss sensitive content.
ShareGuard OCR requires English Tesseract; Docker installs it and its OpenCV runtime dependencies. For a backend outside Docker, install Tesseract OCR with English language data before starting FastAPI.
- ShareGuard OCR/pattern and QR-presence checks are heuristic; OCR may miss or misread content, arbitrary government ID documents are not guaranteed to be recognized, and no sanitizer can mask regions that were not detected.
- LinkGuard checks URL structure and contextual indicators only; it does not fetch destinations, resolve DNS, follow redirects, expand shorteners, or execute content. It identifies risk indicators; it does not guarantee that a URL is malicious or safe.
- Threat-intelligence provider interfaces and status reporting are present, but no external provider adapter is implemented. Threat intelligence is disabled by default; a configured API key is not used by this phase.
- Gemini AI explanations are optional and disabled by default. The model sees only sanitized categories and bounded metadata; generated summaries remain untrusted contextual explanations and cannot change security results.
- QRGuard locally decodes supported QR images and passes URL payloads through the shared LinkGuard deterministic analyzer. The immediate QR upload response can contain the decoded payload for review; the payload is not persisted in scan history.
- Safe-version masking is limited to the detected bounding boxes; boxes may include nearby text or miss pixels, so review the output. The process-local original cache is lost on backend restart and is not shared across multiple backend workers.
- No user authentication or authorization; this is a single-user demo environment.
- There is no built-in one-click demo mode; browser testing uses synthetic fixtures so the actual upload and analysis flows remain the demo experience.
- Database schema setup uses SQLAlchemy
create_all; there are no migrations yet. - A low score means only that the automated checks found few weighted indicators; it is not a safety guarantee.
- Add automated tests, database migrations, and user-level data isolation.
- Improve ShareGuard detector coverage and evaluate local OCR across supported image types and languages.
- Evaluate risk scoring against documented test cases and expose signal provenance.
- Evaluate an explicitly opt-in threat-intelligence adapter with documented disclosure and privacy-safe handling.
- Expand evidence-grounded explanation evaluation while preserving deterministic risk scores.
No external AI or threat-intelligence integration is included in this phase. ShareGuard's local OCR and detectors are heuristic and should not be treated as exhaustive.