Skip to content

About

Streamline — An autonomous personal productivity OS unifying multi-account Gmail, Calendar, semantic memory (pgvector RAG), and a ReAct AI agent, with a human-in-the-loop policy engine gating every mutating action. Built with Next.js 15, Express 5, Gemini, Neon Postgres, and BullMQ.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Streamline — Autonomous AI Personal Productivity OS

License: MIT Node.js Version Next.js Database Vector Search AI Engine Queues Observability Security Evals

An autonomous, multi-tenant personal operating system unifying multi-account Gmail, Google Calendar, semantic memory, and proactive AI agents — where every mutating action (send_email, create_calendar_event, create_task) requires human approval through a policy engine that holds independently of what the model decides (0/21 model-attempted policy bypasses in the most recent live adversarial run — see evals/results/live/).

Explore Documentation · Architecture Overview · REST API Specs · ADRs · Developer Setup


⚡ Executive Overview

Streamline transforms personal productivity by aggregating fragmented Google Workspace accounts (Work, Personal, University) into a single, high-performance command center.

Beyond standard email clients, Streamline acts as an autonomous personal operating system:

  • ReAct Agent Runtime & Tool Orchestrator: An interactive multi-turn agent capable of scheduling events, drafting emails, managing tasks, and recalling memories using explicit thought signatures and multi-step tool execution loops.
  • OAuth 2.0 Token Lifecycle & Multi-Account Credential Vault: Proactive 5-minute expiry buffer checks, an in-process single-flight mutex plus a Redis-backed distributed lock across replicas (fails open — best-effort, no hang — if Redis itself is unavailable), AES-256-GCM encryption, and automated invalid_grant revocation handling.
  • Dual-Boundary Policy Engine & Human-in-the-Loop Shield: Automated interception of state-mutating actions (send_email, create_calendar_event, create_task) requiring human approval before execution.
  • Policy-Layer Prompt Injection Containment: Untrusted external content is tagged as untrusted at the data layer, but the actual guarantee is downstream — every mutating tool call is queued for human approval regardless of what the model decides. In the most recent live adversarial run (21 real attack/control scenarios, real Gemini model, no scripting), the policy layer held 21/21, and the model itself never attempted a prohibited action either (0/21) — those are two separately measured things, see ADR-0012.
  • pgvector Hybrid RAG & Semantic Memory Engine: Continuous semantic recall across 3 memory classes (User Preferences, Confirmed Decisions, Project Facts) via Neon pgvector HNSW cosine distance fused with PostgreSQL tsvector full-text search via Reciprocal Rank Fusion ($k=60$).
  • OpenTelemetry-Compliant Observability: Real-time waterfall trace profiler tracking per-step latencies, model vs. tool overhead, and exact token/USD cost attribution via live SSE streams.
  • DAG Task Dependency Scheduler & Topological Urgency Engine: Topological sorting with cycle detection (Kahn's / DFS algorithm) and exponential urgency decay scoring ($e^{-\Delta t / 48}$) to deterministically select your next best task against Google Calendar free slots.
  • AI Cost Guard & Circuit Breaker: Autonomous per-user daily token budgets and USD spend limits with automatic tripping mechanisms.

🏛️ System Architecture

Tier 1: 5-Box Executive Architecture (10-Second High-Level Scan)

flowchart LR
    subgraph B1 ["1. Client & Streaming UI"]
        NextJS["Next.js 15 (React 19)\nSSE Waterfall & Drafter Stream"]
    end

    subgraph B2 ["2. Security & Token Gateway"]
        Gateway["Express 5 REST API\nToken Lifecycle Mutex\nAES-256-GCM Vault"]
    end

    subgraph B3 ["3. ReAct Agent Core"]
        AgentCore["ReAct Execution Runtime\nGemini Cascade Fallback\nHITL Dual-Boundary Shield"]
    end

    subgraph B4 ["4. Distributed Async Queues"]
        AsyncQ["BullMQ 5.x + Redis Cluster\nSync, Triage & Digest Workers"]
    end

    subgraph B5 ["5. Hybrid Storage & Vectors"]
        DBStore["Neon Serverless Postgres\npgvector HNSW Cosine Search\nFull-Text tsvector Engine"]
    end

    B1 <-->|HTTP / SSE| B2
    B2 <-->|Enqueued Jobs| B4
    B2 <-->|Read / Write| B5
    B4 <-->|Persistence & Vectors| B5
    B2 <-->|Agent Sessions| B3
    B3 <-->|Context Recall & Mutex| B5
Loading

Tier 2: Deep-Dive Distributed Systems Topology & Dataflows

graph TD
    subgraph MultiAccount ["Multi-Account Ingestion Layer"]
        G1["Work Gmail & Calendar"]
        G2["Personal Gmail & Calendar"]
        G3["University / Org Mail"]
    end

    subgraph TokenLifecycle ["OAuth 2.0 Token Lifecycle & Credential Vault"]
        TokenManager["GoogleTokenManager\n(Single-Flight Concurrency Mutex)"]
        AES["AES-256-GCM Encrypted Vault\n(Unique IV per secret)"]
        ExpiryGuard["Proactive Expiry Guard\n(<5m buffer + invalid_grant handling)"]
    end

    subgraph SecurityCore ["Policy & Isolation Layer"]
        PolicyEngine["Dual-Boundary Policy Engine"]
        Shield["Pending Action Approval Shield\n(Authenticated Human Approval API)"]
        InjectionShield["Untrusted-Content Tagging\n(_contentWarning field + system prompt)"]
    end

    subgraph AsyncEngine ["Distributed Async Engine"]
        BullMQ["BullMQ 5.x Job Queues"]
        RedisBus["Redis Cache & Event Bus"]
        SyncWorker["Account Sync Worker"]
        TriageWorker["AI Email Triage Worker"]
        DigestWorker["Daily Digest Worker"]
    end

    subgraph StorageEngine ["Hybrid Persistence & Vector Engine"]
        NeonDB[("Neon Serverless Postgres")]
        PGVector[("pgvector HNSW Vector Store\n(768-dim cosine distance)")]
        TSVector[("PostgreSQL tsvector Index\n(Sparse BM25 Keyword Search)")]
    end

    subgraph GeminiAI ["AI Intelligence & Orchestration Runtime"]
        AgentRuntime["ReAct Agent Runtime\n(Multi-Turn Reasoning & Tools)"]
        Cascade["Multi-Model Cascade Fallback\n(Gemini 3.5 Lite → 3.6 Flash)"]
        CostGuard["AI Spend Circuit Breaker"]
        Tools["Tool Registry\n(Email, Calendar, Tasks, Memory)"]
        OTel["OpenTelemetry GenAI Span Profiler"]
    end

    subgraph UnifiedOS ["Streamline Personal OS Dashboard"]
        Inbox["Unified Multi-Account Inbox"]
        Drafter["Streaming Contextual Drafter"]
        CalendarUI["Integrated Agenda & Slot Finder"]
        TaskDAG["DAG Task Dependency Scheduler"]
        RAGInspector["Hybrid RAG Retrieval Inspector"]
        SecDash["Security & Injection Test Lab"]
        TraceUI["Live Decision Trace Waterfall"]
    end

    G1 & G2 & G3 --> TokenManager
    TokenManager <--> ExpiryGuard
    TokenManager <--> AES
    TokenManager --> BullMQ
    BullMQ --> RedisBus
    RedisBus --> SyncWorker & TriageWorker & DigestWorker
    SyncWorker --> NeonDB
    TriageWorker & DigestWorker --> Cascade
    TriageWorker --> InjectionShield --> NeonDB
    NeonDB <--> PGVector & TSVector

    NeonDB & PGVector --> AgentRuntime & Cascade
    AgentRuntime --> Tools --> PolicyEngine
    PolicyEngine -->|Safe / Read-Only| Tools
    PolicyEngine -->|Mutating Action| Shield -->|User Approved| Tools
    AgentRuntime --> OTel --> TraceUI
    Cascade --> CostGuard
    Tools & Cascade --> Inbox & Drafter & CalendarUI & TaskDAG & RAGInspector & SecDash
Loading

✨ Core Features

📬 1. Unified Multi-Account Inbox & Mailbox Management

  • Consolidated Account Feed: Aggregate emails from multiple Google accounts into a synchronized view, with every query scoped by userId/accountId at the repository layer.
  • Smart Account Badging: Color-coded mailbox indicators identifying account sources and sender domains.
  • L1 Redis Caching: Sub-millisecond response times for cached mailbox queries with automatic invalidation on mutations.
  • Rich Thread Viewer: Sanitized HTML rendering with attachments, inline participant badges, and full thread history.

🧠 2. Gemini Autonomous Intelligence Pipeline

  • Real-Time Priority Triage (P1–P4):
    • 🔥 P1 Action: Urgent deadlines, critical action items, and executive requests.
    • 💬 P2 Direct: 1-on-1 human interpersonal correspondence.
    • 🔔 P3 Updates: Automated notifications, security alerts, and system policies.
    • 📰 P4 News: Subscriptions, newsletters, and digests.
  • Automated Topic Clustering: Smart semantic pills (🎓 Academics, 🚀 Tech & AI, 💼 DevClub, 👥 Community).
  • Streaming Reply Drafter (SSE): Real-time contextual reply generation supporting tone modulation (Professional, Friendly, Concise, Custom Prompt).
  • Daily Executive Digest: Synthesized morning briefing consolidating newsletters, pending tasks, and upcoming meetings into an actionable dashboard.

🤖 3. ReAct Agent Runtime & Tool Orchestrator

  • Multi-Turn Decision Agent: Natural language assistant equipped with specialized operational tools:
    • get_email, send_email, draft_email
    • create_calendar_event, find_free_slots
    • create_task, get_tasks
    • save_memory, search_memory
  • Thought Signatures: Transparent agent reasoning with explicit internal reasoning steps before executing actions.
  • Dual-Boundary Policy Interception: State-mutating operations are held in a secure pending_actions queue, displaying impact previews and requiring one-click user authorization before touching Google APIs.

🛡️ 4. Durable Memory Vault & pgvector Hybrid RAG Engine

  • 3-Tier Durable Memory Architecture (ADR-0011):
    • preference: Working styles, communication habits, and scheduling preferences.
    • decision: Confirmed agreements, policy rules, and explicit directives.
    • project_fact: System configurations, architectural constraints, and organizational context.
  • Hybrid Search Engine: Neon PostgreSQL pgvector HNSW cosine similarity search combined with full-text keyword indexing.
  • Hybrid RAG Retrieval Inspector: Live inspection interface to execute real-time vector queries, inspect retrieval latency, examine HNSW cosine distance (<=>), and preview exact prompt context injection.

🔒 5. Policy-Layer Security & Prompt Injection Containment

  • Policy-Layer Containment, Not Model Resistance (ADR-0012): Untrusted external email bodies are tagged as untrusted at the data layer (_contentWarning field) and in the system prompt, but the real guarantee doesn't depend on the model respecting either — every mutating tool call is independently intercepted by the policy engine and held for human approval regardless of what the model decides. Most recent live adversarial run (21 real attack/control scenarios, real Gemini model, no scripting): policy-layer containment held 21/21; the model itself also never attempted the prohibited action (0/21) — two separately measured numbers, see evals/results/live/.
  • AES-256-GCM Token Encryption: Google OAuth access and refresh tokens are encrypted at rest with unique initialization vectors and cryptographic authentication tags.
  • OAuth 2.0 with PKCE & CSRF Protection: Strict state parameter validation, double-submit cookie verification, and httpOnly session cookies so tokens are never readable by client-side JavaScript.
  • In-App Security Console: Interactive UI with preset attack payloads (jailbreaks, fake system delimiters, roleplay DAN) for exploring injection framing — note this in-app simulator does simple keyword pattern-matching on the pasted text, it does not route the payload through the real agent/model/policy pipeline; the real adversarial evidence above comes from evals/injection-resistance.eval.ts run live against the actual system, not from this console.

📊 6. OpenTelemetry Decision Tracing & Cost Breakdown

  • Full-Trace Waterfall Profiler: Visual Gantt-chart timeline mapping every sub-step of agent sessions conforming to OpenTelemetry GenAI semantic conventions.
  • Latency & Cost Decomposition: Separate measurements for model generation latency vs. tool execution overhead, with exact per-step token counts and USD expenditure tracking.
  • Live SSE Trace Streaming: Real-time streaming updates as the agent reasons, queries tools, and checks policy guardrails.

📅 7. Calendar Scheduling & DAG Task Planner

  • Two-Way Google Calendar Sync: Bidirectional sync supporting event creation, recurrence parsing, and participant conflict detection.
  • Smart Free-Slot Discovery: Automatic scanning of calendar gaps to recommend optimal task execution windows.
  • DAG Task Dependency Engine (ADR-0009): Topological sorting, cycle detection, and exponential urgency decay scoring across configurable presets (Balanced, Deadline Driven, Deep Work, Quick Wins).

💰 8. AI Cost Guard & Circuit Breaker

  • Autonomous Budget Enforcement: Configurable daily user token limits and USD cost thresholds.
  • Circuit Breaker System: Automatically trips and falls back to deterministic local handlers if daily thresholds are exceeded or provider error rates spike.

🗺️ Dashboard & Application Routes

Route Page Name Primary Capabilities
/inbox Unified Inbox Multi-account email stream, priority badges, thread viewer, streaming reply drafter.
/agent ReAct Agent Orchestrator Multi-turn reasoning agent, tool execution runtime, pending approval cards, thought signature inspection.
/agent/traces/[id] Trace Profiler OpenTelemetry span waterfall, step latencies, token usage, and USD cost decomposition.
/memory Memory Vault & Hybrid RAG Semantic memory records, category filters, live pgvector HNSW + tsvector Hybrid RAG retrieval inspector.
/security Security Guardrails Threat posture monitor, untrusted content shield status, in-app injection payload simulator (keyword-pattern demo — see note in the security section above for how this differs from the real evals/injection-resistance.eval.ts adversarial suite).
/calendar Calendar & Agenda Synchronized multi-calendar view, free-slot finder, direct event scheduling modal.
/tasks DAG Task Scheduler Action item radar, dependency blocker tracking, exponential urgency score ranking.
/digest Executive Digest Daily synthesized morning briefing, newsletter summaries, actionable highlights.
/settings System Settings Google account connections, OAuth token lifecycle status, sync intervals, and AI cost guard limits.

🏗 Technology Stack

Layer Technologies & Libraries
Frontend UI Next.js 15 (App Router), React 19, TypeScript, Tailwind CSS, Radix UI, TanStack Query, Zustand, Lucide Icons
Backend API Express.js 5, Node.js 18+, TypeScript, Zod, Pino Logger, Helmet, CORS, Cookie-Parser
AI & LLM Orchestration Google Gen AI SDK (@google/genai), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, OpenAI-compatible provider fallback
Vector Engine & RAG Neon PostgreSQL pgvector (HNSW indexing, cosine similarity), text embeddings
Database & ORM Neon Serverless PostgreSQL, Drizzle ORM, Drizzle Kit
Queues & Async Jobs BullMQ 5.x, Redis (ioredis), node-cron schedulers
Security & Auth AES-256-GCM symmetric encryption, Bcrypt, JWT (httpOnly), Double-Submit CSRF, Rate Limiting, Secret Redactor
Observability Custom OpenTelemetry GenAI Span Assembler, Server-Sent Events (SSE) trace streaming
Testing & Evals Vitest test runner, Scenario-based AI eval harness (evals/run-all.ts)

📚 Complete Documentation Suite

Comprehensive architectural deep-dives and guides are available in the /docs directory:

📋 Architecture Decision Records (ADRs)

  • ADR-0001: BullMQ & Redis for Async Job Queues
  • ADR-0002: Application-Level AES-256-GCM Token Encryption
  • ADR-0003: Multi-Model Cascade Fallback vs Single-Model Exponential Retry
  • ADR-0004: Neon Serverless PostgreSQL with Drizzle ORM
  • ADR-0005: Server-Sent Events (SSE) over WebSockets for AI Streaming
  • ADR-0006: pgvector on Neon over Dedicated Vector Database
  • ADR-0007: Scenario-Based Eval Harness over Ad-Hoc Manual Testing
  • ADR-0008: Vercel + Railway Deployment Topology
  • ADR-0009: Exponential Urgency Decay and DAG Prioritization
  • ADR-0010: Dual-Boundary Policy Engine and Human-in-the-Loop Safeguards
  • ADR-0011: Three Durable Memory Types and Read-Classified Storage
  • ADR-0012: Policy-Layer Structural Enforcement for Prompt-Injection Defense

🚀 Quick Start Guide

1. Prerequisites

  • Node.js: v18.0.0 or higher
  • Redis: Local or cloud Redis instance (e.g., Upstash / Redis Cloud)
  • PostgreSQL: Neon Serverless Postgres with pgvector enabled
  • Google Cloud Console: OAuth 2.0 Web Client ID with Gmail and Calendar scopes
  • Gemini API Key: From Google AI Studio

2. Clone & Install

git clone https://github.com/PiyushY111/Streamline.git
cd Streamline
npm install

3. Environment Configuration

Generate a cryptographically secure 256-bit encryption key and configure the environment:

node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"

Copy and fill in .env:

cp .env.example .env

Ensure required variables are populated:

# Database & Cache
DATABASE_URL=postgresql://user:password@ep-xyz.neon.tech/streamline?sslmode=require
REDIS_URL=redis://localhost:6379

# Security & Encryption
JWT_SECRET=your_super_secret_jwt_key
ENCRYPTION_KEY=your_64_character_hex_encryption_key

# Google OAuth 2.0
GOOGLE_CLIENT_ID=your-google-client-id.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=your-google-client-secret
GOOGLE_REDIRECT_URI=http://localhost:5001/api/auth/google/callback

# AI Intelligence
GEMINI_API_KEY=your-gemini-api-key
GEMINI_MODEL=gemini-3.5-flash-lite

4. Database Push & Migrations

Push Drizzle ORM schemas into your Neon database:

npm --prefix server run db:push

5. Launch Development Environment

Start both the Express backend and Next.js frontend concurrently:

npm run dev

Or start individual services independently:

# Backend Server (Port 5001)
npm run dev:server

# Frontend Client (Port 3000)
npm run dev:client

Open http://localhost:3000 in your browser to access the Streamline dashboard.


🧪 Automated Eval Harness & Benchmarks

Streamline includes a dedicated scenario-based evaluation suite independent of standard unit tests:

# Run Vitest unit & integration tests
npm test

# Run the complete scenario-based AI evaluation harness
npm run eval

# Run TypeScript type verification across client and server
npm run type-check

Evaluation Benchmark Matrix

npm run eval runs all four suites in fast, deterministic mock mode (the CI default). npm run eval:live runs them against the real configured provider at real (small) API cost — see server/evals/README.md for exactly what --live changes per suite (it's a no-op for two of the four). Most recent live run: 2026-09-21, gemini-3.5-flash-lite, 78/78 scenarios passed.

  • Deterministic Priority Engine (evals/priority.eval.ts): A pure deterministic task-scoring engine (urgency decay, dependency graph) — no AI model call at all, despite living in an "AI eval" directory. 10/10 scenarios.
  • Policy Engine Boundary Tests (evals/tool-selection.eval.ts): Calls enforcePolicy() directly with a hand-specified tool call — tests the policy engine's allow/deny decisions, not an LLM's tool selection. 22/22 scenarios.
  • Hybrid RAG Retrieval Precision (evals/retrieval-precision.eval.ts): The real searchMemory hybrid-RRF path against seeded pgvector rows. 25/25 (100%) on every live run so far; MRR has ranged 0.96–1.00 across three live runs (real embedding calls vary slightly run to run). A three-way ablation (vector-only vs. keyword-only vs. hybrid) lives in evals/retrieval-ablation.eval.ts — see ADR-0011 for the honest result: hybrid does not clearly beat vector-only on this query set.
  • Prompt-Injection Containment & Resistance (evals/injection-resistance.eval.ts): In live mode, the real model runs the full agent loop against 21 real attack/control emails with no scripting. Two separate numbers are measured: policy-layer containment (write/send tools never execute without approval) is a structural guarantee, held 21/21; the model's own resistance (whether it avoided attempting the prohibited action at all) is a real, non-guaranteed behavioral measurement — 0/21 attempted in the most recent run.

📄 License

Distributed under the MIT License. See LICENSE for details.

About

Streamline — An autonomous personal productivity OS unifying multi-account Gmail, Calendar, semantic memory (pgvector RAG), and a ReAct AI agent, with a human-in-the-loop policy engine gating every mutating action. Built with Next.js 15, Express 5, Gemini, Neon Postgres, and BullMQ.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages