Skip to content

About

AI for SRE

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agentic SRE

Open-source incident response agents that learn from every failure.

Build your own resolve.ai in a weekend. LLM agents investigate production incidents, store root causes and playbooks in persistent memory (SenseLab), and resolve recurring failures instantly.

┌─────────────────────────────────────────────────────┐
│  Incident 1 (novel)     →  Full investigation       │
│  Incident 2 (recurring) →  Pattern recognized,      │
│                             playbook applied         │
│                                                     │
│  The difference? Agents that remember.              │
└─────────────────────────────────────────────────────┘

How It Works

A multi-agent Council investigates incidents through coordinated phases:

flowchart TD
    subgraph council [Agent Council]
        T[Triage Agent] --> M[Memory Specialist]
        M --> S1[MCP Specialist 1]
        M --> S2[MCP Specialist 2]
        S1 --> V[Verification Agent]
        S2 --> V
        V --> L[Lead Synthesizer]
    end

    subgraph memory [SenseLab Memory]
        Patterns[Patterns]
        Playbooks[Playbooks]
        Traces[Decision Traces]
    end

    subgraph infra [Your Infrastructure — MCP]
        GCP[Google Cloud]
        DD[Datadog]
        GH[GitHub]
        Jira[Jira]
    end

    infra --> S1
    infra --> S2
    memory <--> M
    L -->|writes findings| memory
    L -->|suggests actions| Action[Action Executor]
    Action --> infra
Loading

Novel incidents: Agents investigate from scratch, find the root cause, and store knowledge in SenseLab — patterns, playbooks, decision traces.

Recurring incidents: The Memory Specialist recognizes patterns from past investigations, specialists confirm with live data, resolution is fast and confident.


Get Started

Prerequisites

  • Node.js 18+ (for MCP servers via npx)
  • Python 3.11+ with uv (for backend)

1. Clone

git clone https://github.com/raia-live/sre-sample.git && cd sre-sample

2. Run

# Install + start (frontend + backend)
cd backend && uv sync && cd ..
cd frontend && npm install && cd ..
make dev

Or with Docker:

docker compose up --build

The onboarding wizard will appear on first launch and guide you through:

  • Selecting your LLM provider (Anthropic, OpenAI) and entering your API key
  • Entering your SenseLab API key (sign up here)
  • Connecting MCP tools (your infrastructure data sources)

All configuration is saved automatically — no need to manually edit .env files.

For advanced users: You can still pre-configure via environment variables if you prefer. Copy .env.example to .env and fill in your keys before starting. The onboarding wizard will detect existing config and skip completed steps.


Connecting Your Tools (MCP)

Agents access your infrastructure via MCP (Model Context Protocol). You can add connectors through the Connectors page in the UI, or configure them manually in mcp-servers.json.

Supported out of the box:

Tool Package What agents can do
Google Cloud @google-cloud/gcloud-mcp Cloud Run, Logging, Monitoring, GKE, BigQuery
GitHub @modelcontextprotocol/server-github Repos, PRs, Issues, Actions, deployments
Datadog datadog-mcp Metrics, logs, traces, dashboards, monitors
Jira mcp-jira-scoped Search issues, create tickets, manage sprints
Notion @notionhq/notion-mcp-server Search pages, create docs, query databases
Sentry @sentry/mcp-server Errors, performance, releases
PagerDuty pagerduty-mcp Incidents, on-call schedules, escalations

Any MCP-compatible server works — just add it via the UI or in mcp-servers.json.

Manual MCP configuration (advanced)

If you prefer config files over the UI, create mcp-servers.json in the project root:

cp mcp-servers.example.json mcp-servers.json

Example configuration:

{
  "gcp": {
    "command": "npx",
    "args": ["-y", "@google-cloud/gcloud-mcp"],
    "env": { "CLOUDSDK_CORE_PROJECT": "your-project-id" }
  },
  "github": {
    "command": "npx",
    "args": ["-y", "@modelcontextprotocol/server-github"],
    "env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "ghp_..." }
  },
  "jira": {
    "command": "npx",
    "args": ["-y", "mcp-jira-scoped"],
    "env": {
      "JIRA_INSTANCE": "your-org",
      "JIRA_USER_EMAIL": "you@company.com",
      "JIRA_API_TOKEN": "ATATT...",
      "JIRA_SCOPES": "read:jira-work,write:jira-work"
    }
  }
}

Supports both stdio (local subprocess via npx) and HTTP (remote MCP servers).


Features

Investigation Engine

  • Multi-Agent Council: Triage, Memory, Specialists, Verification, Synthesis — working together
  • Real-time Activity Feed: Watch agents reason, call tools, and build knowledge live
  • Steering: Redirect agents mid-investigation with follow-up messages
  • Stop & Resume: Halt investigations at any time; partial results are saved
  • Parallel Threads: Fork multiple lines of investigation simultaneously

Action Execution

  • Suggested Actions: After analysis, agents suggest concrete next steps based on connected tools
  • One-click Execution: Create Jira tickets, Notion pages, GitHub issues directly from findings
  • Safety Guards: Actions only use explicitly connected MCP tools — never the wrong tool

Skills System

  • Custom Skills (YAML): Define investigation behaviors, prompts, and tool access
  • Self-Improvement: Agents propose skill updates after investigations
  • Versioning & Rollback: Track skill evolution, revert if needed
  • Effectiveness Rating: See how skills perform over time

Memory-Powered Intelligence (SenseLab)

  • Pattern Recognition: Known failure patterns matched instantly
  • Playbook Reuse: Validated remediation steps applied with confidence
  • Cross-Investigation Learning: What one investigation discovers benefits all future ones
  • Knowledge Graph: Incidents, Patterns, and Playbooks linked and traversable
  • Outcome Propagation: Successful resolutions boost confidence; failures trigger re-evaluation

Connectors

  • MCP-First Architecture: Any MCP server works out of the box
  • Managed Connectors: Pre-configured setups for GCP, GitHub, Datadog, Jira, Notion, Sentry, PagerDuty
  • Connection Validation: Tools are tested before being marked as connected

Architecture

Component Purpose
Backend (FastAPI + Python) Council orchestration, agent execution, skills engine, WebSocket events
Frontend (Next.js + React) Dashboard, incident workbench, skills UI, history, connectors, settings
SenseLab (cloud) Persistent agent memory — patterns, playbooks, decision traces, knowledge graph
Anthropic API LLM reasoning (Claude Sonnet) with tool_use
MCP Servers Connect to any infrastructure tool via Model Context Protocol

Development

# Full local setup
make setup    # Creates .env from template (optional — wizard handles this)
make dev      # Starts backend + frontend

# Individual services
make dev-backend    # Backend only (port 8000)
make dev-frontend   # Frontend only (port 3000)

# Docker
make run      # docker compose up --build
make stop     # docker compose down
make clean    # Remove all containers, volumes, and build artifacts

Environment Variables (advanced)

These are only needed if you want to pre-configure the app without using the onboarding wizard.

Variable Required Description
AMFS_API_KEY Yes SenseLab API key from senselab.ai
ANTHROPIC_API_KEY Yes Anthropic API key
MODEL No claude-sonnet-4-6 (default)
AMFS_HTTP_URL No SenseLab endpoint (default: https://amfs-login.sense-lab.ai)
LOG_LEVEL No INFO (default), DEBUG for verbose

Why SenseLab?

Without persistent memory, every agent starts from zero. At scale with hundreds of services:

  • Reliability: Known patterns are confirmed, not re-investigated
  • Confidence: Playbooks validated by past outcomes get higher confidence
  • Institutional Knowledge: What one agent learns, all agents benefit from
  • Compounding Value: Each incident makes the system smarter
  • Self-Improvement: Skills evolve automatically based on investigation outcomes

Contributing

  1. Fork the repo
  2. Create a feature branch (git checkout -b feature/my-feature)
  3. Make your changes
  4. Run the backend tests: cd backend && uv run pytest
  5. Submit a PR

License

MIT

About

AI for SRE

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages