Skip to content

About

Privacy Brake — Pause. Check. Then Act. A cybersecurity safety layer for safer digital actions.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Privacy Brake

Pause. Check. Then Act.

Privacy Brake is a human-centered cybersecurity safety layer that helps people review detected risk indicators before they share a file, open a link, or follow a QR destination. It presents context and a safer next step while leaving the decision with the user.

ShareGuard uses local OCR, sensitive-pattern matching, and opaque masking of selected detected regions. QRGuard decodes locally and sends URL payloads through the same deterministic LinkGuard service used for pasted links. LinkGuard checks URL structure without visiting the destination. An optional Gemini AI Analyst explains sanitized deterministic results; it cannot change findings or risk scores. These checks may miss indicators and do not guarantee safety. External threat intelligence is disabled and no provider adapter is currently implemented.

Problem

Sensitive details in screenshots and documents, misleading URLs, malicious QR destinations, and social-engineering prompts are easy to act on before a person notices the risk.

Solution

Three focused checks—ShareGuard, LinkGuard, and QRGuard—separate scanner signals from deterministic risk scoring and contextual explanation. QRGuard decodes uploaded codes locally, classifies the content, and displays basic URL indicators before the user chooses an action. The opt-in AI Analyst uses a minimized finding summary to explain possible consequences; deterministic analysis remains authoritative. Privacy Brake makes the evidence and next-step recommendation visible before the user acts. It is a safety layer, not a generic AI chatbot.

Architecture

User action → local/deterministic detection → authoritative risk engine
             → optional sanitized Gemini context → explanation → user action

The React/TypeScript frontend calls a FastAPI JSON API. ShareGuard performs local Tesseract OCR, OpenCV QR-presence detection, deterministic sensitive-pattern matching, and selected-region masking. QRGuard uses OpenCV to decode uploaded QR images in memory and hands decoded URLs to the shared LinkGuard analyzer. LinkGuard normalizes and parses URL text, calculates deterministic indicators, and never visits submitted destinations. SQLAlchemy persists privacy-safe scan metadata in PostgreSQL. See docs/architecture.md, docs/privacy.md, docs/api.md, docs/testing.md, docs/deployment.md, and docs/demo.md.

Core features

  • ShareGuard: local OCR and deterministic checks for common sensitive details, with normalized finding boxes.
  • Safe Version: user-selected findings are covered with opaque masks; original upload bytes are held temporarily in process memory for this workflow.
  • QRGuard: local QR decoding and URL classification; decoded links are never automatically opened.
  • LinkGuard: deterministic URL structure checks without fetching destinations or following redirects.
  • Risk Engine: transparent scanner-specific weights and consistent LOW / CAUTION / HIGH / STOP levels; AI cannot change the score or findings.
  • AI Analyst: optional Gemini-assisted explanation, disabled by default; validated output is contextual, not authoritative detection.
  • Risk Path: evidence-grounded display of the analysis steps and possible, not guaranteed, consequences.

Tech stack

  • Frontend: React, TypeScript, Vite, Tailwind CSS, React Router
  • Backend: Python 3.11+, FastAPI, Pydantic, SQLAlchemy
  • Database: PostgreSQL
  • Local development: Docker Compose, .env, Git

Local setup

Docker Compose

  1. Copy .env.example to .env. Its database credentials are local-demo placeholders; replace them before sharing or using the configuration outside a local demo. Keep the Postgres and DATABASE_URL credentials aligned. Threat intelligence and Gemini AI are disabled by default.

  2. From the project root, run:

    docker compose up --build -d
  3. Open the frontend at http://localhost:5173. The API is available at http://localhost:8000; interactive API docs are at http://localhost:8000/docs. PostgreSQL is published on 127.0.0.1:5432 for local database tools only.

  4. Check service and database availability:

    docker compose ps
    docker compose exec postgres pg_isready -U privacy_brake -d privacy_brake
    Invoke-RestMethod http://localhost:8000/api/v1/health

The backend waits for PostgreSQL health and creates missing tables at startup using SQLAlchemy create_all. This is suitable for a fresh demo database, not a replacement for production schema migrations. Uploaded files are not permanently stored. To stop services, run docker compose down. Do not add -v unless you intend to remove the local demo database volume.

AI Analyst is disabled by default. To explicitly enable the Gemini explanation provider, set AI_ENABLED=true, AI_PROVIDER=gemini, and GEMINI_API_KEY in .env. Optionally set GEMINI_MODEL and AI_TIMEOUT_SECONDS. The key stays in the backend environment. When enabled, only sanitized scan categories and bounded metadata are sent to Google's Gemini service; see docs/privacy.md before enabling.

Run services separately

Start PostgreSQL locally and set DATABASE_URL from .env.example. Then:

cd backend
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
uvicorn app.main:app --reload --no-access-log

In a second terminal:

cd frontend
npm install
npm run dev

For a backend running on the host, set DATABASE_URL to a host-local PostgreSQL URL in the backend environment. Vite proxies /api to http://localhost:8000 by default; set VITE_API_PROXY_TARGET to change the development proxy. The frontend can also use the public VITE_API_BASE_URL build-time value. It defaults to same-origin /api/v1; it must never contain secrets.

Environment variables

Variable Purpose Default / requirement
DATABASE_URL Backend PostgreSQL connection. Use hostname postgres between Compose containers. Local Compose value in .env.example; required for production.
POSTGRES_DB, POSTGRES_USER, POSTGRES_PASSWORD PostgreSQL initialization. Keep these aligned with DATABASE_URL. Local-only sample values; production must set secure values.
APP_ENV, DEBUG Backend environment and debug mode. development, false; production enforces DEBUG=false.
CORS_ORIGINS Comma-separated frontend origins allowed by FastAPI. Localhost in development; explicit HTTPS origins in production.
VITE_API_BASE_URL Public frontend API base URL compiled into the frontend. /api/v1; may be a deployed API URL. Never put a key or credential here.
VITE_API_PROXY_TARGET Vite development-only proxy destination. http://localhost:8000 outside Docker; Compose uses http://backend:8000.
AI_ENABLED, AI_PROVIDER, GEMINI_API_KEY, GEMINI_MODEL, AI_TIMEOUT_SECONDS Optional server-side AI explanation. Disabled by default. GEMINI_API_KEY stays backend-only.
THREAT_INTEL_ENABLED, THREAT_INTEL_PROVIDER, VIRUSTOTAL_API_KEY Reserved configuration for optional reputation providers. Disabled; no external provider adapter is implemented.

Docker production image

A multi-stage frontend Dockerfile builds static assets and serves them with Nginx; the backend image runs as a non-root user and includes a database-aware health check. The deployment-target-neutral docker-compose.production.yml keeps PostgreSQL and the API off the public host ports and exposes the frontend container on HTTP_PORT (default 8080). Put a TLS-terminating HTTPS reverse proxy or hosting-provider ingress in front of it before public use. See docs/deployment.md for required production configuration and commands.

docker compose -f docker-compose.production.yml --env-file .env up --build -d

Production mode rejects debug, loopback database URLs, wildcard CORS, and non-HTTPS CORS origins. A placeholder .env.example is not production configuration and must not be deployed unchanged.

Synthetic demo workflow

Follow docs/demo.md to generate a synthetic ShareGuard image and QR image, demonstrate selected-region sanitization, hand a QR URL to LinkGuard, and show the AI-disabled fallback. No real personal data or live suspicious destinations are needed.

Backend tests

With the Docker stack running:

docker compose cp backend/tests backend:/tmp/privacy-brake-tests
docker compose exec backend python -m unittest discover -s /tmp/privacy-brake-tests -v

The Phase 8 regression and synthetic-evaluation results are documented in docs/testing.md. The evaluation corpus contains synthetic text only; it is not a real-world accuracy benchmark.

API endpoints

All endpoints are under /api/v1.

Method Endpoint Purpose
GET /health Health and PostgreSQL connectivity
POST /scans/image Analyze an uploaded PNG or JPEG using local ShareGuard OCR and detectors
POST /scans/url Analyze a JSON { "url": "..." } request with local LinkGuard checks
POST /scans/qr Decode an uploaded QR image locally and return classified content and deterministic indicators
GET /scans List scan summaries (limit, offset)
GET /scans/{scan_id} Get scan details and findings
POST /scans/{scan_id}/sanitize Return a sanitized PNG for selected ShareGuard findings

File uploads are limited to 10 MiB. ShareGuard accepts JPEG and PNG; QRGuard accepts JPEG, PNG, and WebP. The backend validates file signatures and decodes images in memory, with a 25-megapixel limit. FastAPI may spool request bodies and pytesseract may create a short-lived local OCR input file that is removed after processing; uploaded files are not permanently stored. QRGuard's decoded content is returned only in the immediate upload response for user review; raw QR content is not persisted in scan history or application signals. For sanitization, the original ShareGuard upload bytes are held in a bounded, process-local memory cache for up to 30 minutes (maximum 128 MiB total); entries are periodically expired and cleared on backend shutdown. Restarting the backend invalidates pending safe-version requests, so users must upload again. The database stores normalized detections and masked previews, not image bytes or raw OCR text. Sanitized images are returned to the browser and never stored in PostgreSQL.

All three scan endpoints include an ai_analysis object. When AI is disabled or unavailable, the object reports that state and still includes a locally generated explanation, evidence-linked points, possible risk path, and recommended action. The provider-generated summary is stored with provider/model/status metadata in an existing threat_checks row; prompts and raw user inputs are not stored. The AI cannot affect deterministic scores, levels, findings, or signals.

Interface and demo flows

The application shell provides responsive Dashboard, ShareGuard, QRGuard, LinkGuard, and Scan History navigation, accessible keyboard focus, and visible privacy disclaimers. Scan results consistently pair the risk level with a numeric score/meter and the wording “risk score based on detected indicators.” Risk paths use the returned analysis steps and label possible consequences as uncertain.

  • ShareGuard: upload → local analysis → review labeled findings/bounds → select regions → create and review the Safe Version → download only the sanitized image.
  • QRGuard: upload → decode locally → review non-clickable payload and indicators → optionally use Analyze Destination to open the internal LinkGuard route.
  • LinkGuard: paste → analyze as text → review redacted URL, deterministic indicators, Risk Path, and AI-assisted explanation/fallback.

No demo link opens an external URL. The app footer notes that detection is not guaranteed and AI explanations may be imperfect.

ShareGuard defaults high/critical findings to selected and lets the user deselect any region. Creating a Safe Version requires at least one selection. The backend accepts finding IDs only, validates ownership against the scan, retrieves its authoritative stored bounding boxes, and masks selected regions with opaque black pixels. It does not rerun OCR or trust browser-supplied boxes. The operation records a SANITIZE / SUCCESS safe-action audit entry with the mask method, selected count, and finding categories; it does not store image bytes or raw PII. The user receives a PNG download and before/after comparison. Review the safe version before sharing; automated detection may miss sensitive content.

ShareGuard OCR requires English Tesseract; Docker installs it and its OpenCV runtime dependencies. For a backend outside Docker, install Tesseract OCR with English language data before starting FastAPI.

Current limitations

  • ShareGuard OCR/pattern and QR-presence checks are heuristic; OCR may miss or misread content, arbitrary government ID documents are not guaranteed to be recognized, and no sanitizer can mask regions that were not detected.
  • LinkGuard checks URL structure and contextual indicators only; it does not fetch destinations, resolve DNS, follow redirects, expand shorteners, or execute content. It identifies risk indicators; it does not guarantee that a URL is malicious or safe.
  • Threat-intelligence provider interfaces and status reporting are present, but no external provider adapter is implemented. Threat intelligence is disabled by default; a configured API key is not used by this phase.
  • Gemini AI explanations are optional and disabled by default. The model sees only sanitized categories and bounded metadata; generated summaries remain untrusted contextual explanations and cannot change security results.
  • QRGuard locally decodes supported QR images and passes URL payloads through the shared LinkGuard deterministic analyzer. The immediate QR upload response can contain the decoded payload for review; the payload is not persisted in scan history.
  • Safe-version masking is limited to the detected bounding boxes; boxes may include nearby text or miss pixels, so review the output. The process-local original cache is lost on backend restart and is not shared across multiple backend workers.
  • No user authentication or authorization; this is a single-user demo environment.
  • There is no built-in one-click demo mode; browser testing uses synthetic fixtures so the actual upload and analysis flows remain the demo experience.
  • Database schema setup uses SQLAlchemy create_all; there are no migrations yet.
  • A low score means only that the automated checks found few weighted indicators; it is not a safety guarantee.

Roadmap

  1. Add automated tests, database migrations, and user-level data isolation.
  2. Improve ShareGuard detector coverage and evaluate local OCR across supported image types and languages.
  3. Evaluate risk scoring against documented test cases and expose signal provenance.
  4. Evaluate an explicitly opt-in threat-intelligence adapter with documented disclosure and privacy-safe handling.
  5. Expand evidence-grounded explanation evaluation while preserving deterministic risk scores.

No external AI or threat-intelligence integration is included in this phase. ShareGuard's local OCR and detectors are heuristic and should not be treated as exhaustive.

About

Privacy Brake — Pause. Check. Then Act. A cybersecurity safety layer for safer digital actions.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages