Open-source incident response agents that learn from every failure.
Build your own resolve.ai in a weekend. LLM agents investigate production incidents, store root causes and playbooks in persistent memory (SenseLab), and resolve recurring failures instantly.
┌─────────────────────────────────────────────────────┐
│ Incident 1 (novel) → Full investigation │
│ Incident 2 (recurring) → Pattern recognized, │
│ playbook applied │
│ │
│ The difference? Agents that remember. │
└─────────────────────────────────────────────────────┘
A multi-agent Council investigates incidents through coordinated phases:
flowchart TD
subgraph council [Agent Council]
T[Triage Agent] --> M[Memory Specialist]
M --> S1[MCP Specialist 1]
M --> S2[MCP Specialist 2]
S1 --> V[Verification Agent]
S2 --> V
V --> L[Lead Synthesizer]
end
subgraph memory [SenseLab Memory]
Patterns[Patterns]
Playbooks[Playbooks]
Traces[Decision Traces]
end
subgraph infra [Your Infrastructure — MCP]
GCP[Google Cloud]
DD[Datadog]
GH[GitHub]
Jira[Jira]
end
infra --> S1
infra --> S2
memory <--> M
L -->|writes findings| memory
L -->|suggests actions| Action[Action Executor]
Action --> infra
Novel incidents: Agents investigate from scratch, find the root cause, and store knowledge in SenseLab — patterns, playbooks, decision traces.
Recurring incidents: The Memory Specialist recognizes patterns from past investigations, specialists confirm with live data, resolution is fast and confident.
- Node.js 18+ (for MCP servers via npx)
- Python 3.11+ with uv (for backend)
git clone https://github.com/raia-live/sre-sample.git && cd sre-sample# Install + start (frontend + backend)
cd backend && uv sync && cd ..
cd frontend && npm install && cd ..
make devOr with Docker:
docker compose up --build3. Open http://localhost:3000
The onboarding wizard will appear on first launch and guide you through:
- Selecting your LLM provider (Anthropic, OpenAI) and entering your API key
- Entering your SenseLab API key (sign up here)
- Connecting MCP tools (your infrastructure data sources)
All configuration is saved automatically — no need to manually edit .env files.
For advanced users: You can still pre-configure via environment variables if you prefer. Copy
.env.exampleto.envand fill in your keys before starting. The onboarding wizard will detect existing config and skip completed steps.
Agents access your infrastructure via MCP (Model Context Protocol). You can add connectors through the Connectors page in the UI, or configure them manually in mcp-servers.json.
Supported out of the box:
| Tool | Package | What agents can do |
|---|---|---|
| Google Cloud | @google-cloud/gcloud-mcp |
Cloud Run, Logging, Monitoring, GKE, BigQuery |
| GitHub | @modelcontextprotocol/server-github |
Repos, PRs, Issues, Actions, deployments |
| Datadog | datadog-mcp |
Metrics, logs, traces, dashboards, monitors |
| Jira | mcp-jira-scoped |
Search issues, create tickets, manage sprints |
| Notion | @notionhq/notion-mcp-server |
Search pages, create docs, query databases |
| Sentry | @sentry/mcp-server |
Errors, performance, releases |
| PagerDuty | pagerduty-mcp |
Incidents, on-call schedules, escalations |
Any MCP-compatible server works — just add it via the UI or in mcp-servers.json.
Manual MCP configuration (advanced)
If you prefer config files over the UI, create mcp-servers.json in the project root:
cp mcp-servers.example.json mcp-servers.jsonExample configuration:
{
"gcp": {
"command": "npx",
"args": ["-y", "@google-cloud/gcloud-mcp"],
"env": { "CLOUDSDK_CORE_PROJECT": "your-project-id" }
},
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "ghp_..." }
},
"jira": {
"command": "npx",
"args": ["-y", "mcp-jira-scoped"],
"env": {
"JIRA_INSTANCE": "your-org",
"JIRA_USER_EMAIL": "you@company.com",
"JIRA_API_TOKEN": "ATATT...",
"JIRA_SCOPES": "read:jira-work,write:jira-work"
}
}
}Supports both stdio (local subprocess via npx) and HTTP (remote MCP servers).
- Multi-Agent Council: Triage, Memory, Specialists, Verification, Synthesis — working together
- Real-time Activity Feed: Watch agents reason, call tools, and build knowledge live
- Steering: Redirect agents mid-investigation with follow-up messages
- Stop & Resume: Halt investigations at any time; partial results are saved
- Parallel Threads: Fork multiple lines of investigation simultaneously
- Suggested Actions: After analysis, agents suggest concrete next steps based on connected tools
- One-click Execution: Create Jira tickets, Notion pages, GitHub issues directly from findings
- Safety Guards: Actions only use explicitly connected MCP tools — never the wrong tool
- Custom Skills (YAML): Define investigation behaviors, prompts, and tool access
- Self-Improvement: Agents propose skill updates after investigations
- Versioning & Rollback: Track skill evolution, revert if needed
- Effectiveness Rating: See how skills perform over time
- Pattern Recognition: Known failure patterns matched instantly
- Playbook Reuse: Validated remediation steps applied with confidence
- Cross-Investigation Learning: What one investigation discovers benefits all future ones
- Knowledge Graph: Incidents, Patterns, and Playbooks linked and traversable
- Outcome Propagation: Successful resolutions boost confidence; failures trigger re-evaluation
- MCP-First Architecture: Any MCP server works out of the box
- Managed Connectors: Pre-configured setups for GCP, GitHub, Datadog, Jira, Notion, Sentry, PagerDuty
- Connection Validation: Tools are tested before being marked as connected
| Component | Purpose |
|---|---|
| Backend (FastAPI + Python) | Council orchestration, agent execution, skills engine, WebSocket events |
| Frontend (Next.js + React) | Dashboard, incident workbench, skills UI, history, connectors, settings |
| SenseLab (cloud) | Persistent agent memory — patterns, playbooks, decision traces, knowledge graph |
| Anthropic API | LLM reasoning (Claude Sonnet) with tool_use |
| MCP Servers | Connect to any infrastructure tool via Model Context Protocol |
# Full local setup
make setup # Creates .env from template (optional — wizard handles this)
make dev # Starts backend + frontend
# Individual services
make dev-backend # Backend only (port 8000)
make dev-frontend # Frontend only (port 3000)
# Docker
make run # docker compose up --build
make stop # docker compose down
make clean # Remove all containers, volumes, and build artifactsThese are only needed if you want to pre-configure the app without using the onboarding wizard.
| Variable | Required | Description |
|---|---|---|
AMFS_API_KEY |
Yes | SenseLab API key from senselab.ai |
ANTHROPIC_API_KEY |
Yes | Anthropic API key |
MODEL |
No | claude-sonnet-4-6 (default) |
AMFS_HTTP_URL |
No | SenseLab endpoint (default: https://amfs-login.sense-lab.ai) |
LOG_LEVEL |
No | INFO (default), DEBUG for verbose |
Without persistent memory, every agent starts from zero. At scale with hundreds of services:
- Reliability: Known patterns are confirmed, not re-investigated
- Confidence: Playbooks validated by past outcomes get higher confidence
- Institutional Knowledge: What one agent learns, all agents benefit from
- Compounding Value: Each incident makes the system smarter
- Self-Improvement: Skills evolve automatically based on investigation outcomes
- Fork the repo
- Create a feature branch (
git checkout -b feature/my-feature) - Make your changes
- Run the backend tests:
cd backend && uv run pytest - Submit a PR
MIT