Versioned REST routes live under /api/v1/; MCP is mounted separately at
/mcp. A running deployment serves its authoritative OpenAPI schema at
/api/openapi.json and interactive Swagger docs at /api/docs. The tables below
are a curated map; use the schema for exact request and response shapes.
See also the public API stability contract and the API surface ownership charter.
Memory endpoints
| Endpoint | Method | Description |
|---|---|---|
/memories |
POST | Write a memory. LLM enrichment + embedding + entity extraction + contradiction detection. "persist": false for extract-only preview. Stored quarantined when the organization holds its agent's writes (quarantine in /settings), as are /memories/bulk items |
/memories/bulk |
POST | Write up to 100 memories. Batches embeddings, parallelizes enrichment, single transaction. Requires X-Bulk-Attempt-Id header (per-attempt idempotency); a retry with the same id resolves committed rows as duplicate_attempt instead of duplicating. Returns 200 (clean / all-error) or 207 Multi-Status (mixed) — read per-item status. From the broker's install credential, an item's metadata.rules_receipt (event_id, rule_set_hash: the rules delivery its session was under) is kept as system_metadata.rules_receipt; anyone else's is dropped. The broker's items may also come "status": "quarantined": a write its write gate refused, held for review, with the gate's account in metadata.write_gate (tool, paths, action, rule_ids) kept as system_metadata.hold (reason write_gate). From anyone else that status is an item error |
/memories |
GET | List memories (filter by type, status, agent; paginate) |
/memories/{id} |
GET | Full memory detail (embedding stats, entity links, RDF triple, temporal bounds) |
/memories/{id} |
PATCH | Update content or metadata. Re-embeds if content changes |
/memories/{id} |
DELETE | Soft delete (sets status to deleted). A held memory (quarantined) is a 404 here: it leaves quarantine only by a person's release or reject, and a reject deletes it |
/memories/{id}/status |
PATCH | Update lifecycle status. A memory held for review (quarantined) is moved only by a person, here or by a session rollback: active releases it, cancelled rejects it. A reject also soft-deletes it, with its auto-chunks, so no read returns it, search and recall included. A release then runs, in the background, what the held write skipped: its enrichment and governance verdict, its atomic-fact children, the near-duplicate merge it meant to make, entity extraction and contradiction detection. A held write's auto-chunks move with it, and are not released or rejected on their own (409) |
/memories/held |
GET | The memories held for review, newest first, with total; session_id narrows to one broker session. Each one's system_metadata.hold says why: below_trust (an agent below the organization's trust level) or write_gate (a file write the broker's gate refused, with the tool, paths, action and rule ids). A held write's auto-chunks aren't listed: they move with it. A person only |
/memories/rollback-session |
POST | Undo a broker session's writes ({"session_id"}): its live memories and the rows derived from them become outdated, what they had superseded or contradicted becomes active again (restored), and its held ones are rejected as a person rejects them (cancelled and soft-deleted), with their auto-chunks. A person only |
/memories/{id}/contradictions |
GET | View contradiction chain |
/memories |
DELETE | Bulk soft-delete. Skips held memories (quarantined), as every delete does, because a person decides on those. A fleet or tenant purge still removes them |
/memories/stats |
GET | Counts by type, agent, and status, plus pending: {embedding, enrichment, fanout} (live rows still owed background work), stranded: {embedding} (the pending-embedding rows written over an hour ago, which nothing will come back for; re-embed them with python -m core_storage_api.scripts.backfill_embeddings) and settled (all zero once the stranded rows are set aside). Benchmarks and other measure-after-ingest callers should poll until settled: true before measuring, and check stranded, since those rows are searched by keyword only — see BENCHMARKS.md. With no embedding provider configured, embedding_configured is false: rows are stored unembedded and searched by keyword, pending.embedding counts them, and settled ignores that count |
/search |
POST | Hybrid semantic + keyword search with graph-enhanced retrieval |
/recall |
POST | Search + LLM synthesis — summary is the answer to the query (the model reasons step by step internally; only its final answer is surfaced), alongside the source memories under memories (also mirrored to items for /search-shaped consumers — items is deprecated and scheduled for removal in v4.0.0; send items_alias: false to drop that copy now and halve the response, and read memories. The MCP recall brief already omits it by default). top_k is the result count — limit is accepted as an alias for it |
/ingest/preview |
POST | Extract 5-20 atomic facts from a URL or text (no writes) |
/ingest/commit |
POST | Write previewed facts as memories |
Knowledge graph endpoints
| Endpoint | Method | Description |
|---|---|---|
/entities |
GET | List entities (filter by type, search) |
/entities/upsert |
POST | Create or update entity |
/entities/{id} |
GET | Entity detail with relations and linked memories |
/relations/upsert |
POST | Create or update relation |
/graph |
GET | Full knowledge graph (entities + relations) |
Evolve, Insights, Agents, Crystallizer, Documents, Fleet, Admin
Karpathy Loop / Evolve
| Endpoint | Method | Description |
|---|---|---|
/evolve/report |
POST | Report an outcome (success/failure/partial) against recalled memories |
Insights
| Endpoint | Method | Description |
|---|---|---|
/insights/generate |
POST | LLM-powered analysis. Focus: contradictions, failures, stale, divergence, patterns, discover. When no LLM provider answers, the result carries skipped_reason: "llm_unavailable" and no findings, and prior insights are left as they were (MCP caura_insights says the same) |
Agents
| Endpoint | Method | Description |
|---|---|---|
/agents |
GET | List registered agents with trust levels |
/agents/{id} |
GET | Single agent detail |
/agents/{id}/trust |
PATCH | Set trust level (0-3) |
Memory Crystallizer
| Endpoint | Method | Description |
|---|---|---|
/crystallize |
POST | Trigger crystallization for a tenant |
/crystallize/all |
POST | Trigger for all tenants (admin key, nightly) |
/crystallize/reports |
GET | List crystallization reports |
/crystallize/latest |
GET | Most recent completed report |
Documents
| Endpoint | Method | Description |
|---|---|---|
/documents |
POST | Store or update a structured JSON document. Also mints a memory carrying the document's data, so the body is reachable by recall — the doc row embeds only data["summary"]. Independent of the summary: a doc without one is invisible to /documents/search and still mints. Not minted for collection="skills", _-prefixed collections, an empty data, or a payload over the memory size limit |
/documents/{id} |
GET | Retrieve document by ID |
/documents/query |
POST | Query by field equality filters |
/documents/search |
POST | Similarity search over documents written with a data["summary"]. With no embedding provider configured, documents are stored unindexed and this answers 501 EMBEDDING_NOT_CONFIGURED (MCP caura_doc op=search: the same code) |
/documents/{id} |
DELETE | Delete a document, and un-mint the memory its write minted |
Fleet
| Endpoint | Method | Description |
|---|---|---|
/fleet/heartbeat |
POST | Plugin heartbeat — upserts node status, returns pending commands. A node is bound to the credential that heartbeats it: an agent or install credential acts only as nodes bound to it, while a tenant credential acts as any node of its tenant and takes back one bound to a narrower credential |
/fleet/nodes |
GET | List fleet nodes with status (online/stale/offline) |
/fleet/nodes/{node_id}/release |
POST | Release a node from its credential, for example after rotating its key (tenant credential). With bind_agent_id or bind_install_uuid the node is bound to that credential at once; without either, its next heartbeat binds it to whichever credential sends it first. A node with no binding yet, including every node from before binding existed, is claimed the same way |
/fleet/commands |
POST | Queue a command for a node |
/fleet/commands |
GET | List command history; an agent or install credential sees only its own nodes' commands |
/fleet/commands/{command_id}/result |
POST | Report a command's result; an agent or install credential reports only on its own nodes' commands |
Admin + System
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Liveness check |
/version |
GET | Current version |
/tool-descriptions |
GET | Canonical MCP tool descriptions |
/admin/tenants |
GET | List all tenants (admin key) |
/admin/fleets |
GET | List fleets across all tenants (admin key) |
/admin/memories |
GET | List memories across all tenants with filters (admin key) |
/admin/memories/stats |
GET | Memory counts by tenant/type/status (admin key) |
/admin/memories/{id}/re-extract |
POST | Re-run one memory's entity extraction in the background, replacing the graph it has (admin key, tenant_id query). 202; 404 when the memory is gone or held, 409 when its organization has entity extraction off |
/admin/entity-extraction/rerun-lost |
POST | Re-run entity extraction for memories that lost it (a run that raised, was cancelled by a shutdown, or settled for the regex heuristic), oldest first, up to 50 per call and 3 per memory; answers with counts at once. core-operations calls it hourly (admin key) |
/settings |
GET / PUT | Per-tenant configuration. quarantine.below_trust (0 to 4; unset or 0 holds nothing) holds every write from an agent below that trust level as quarantined, for a person to release or reject; quarantine.below_trust_by_fleet overrides it per fleet ({fleet_id: level}). A raised level applies to every write sent after the change is saved, on every instance; a write caught between two quick changes gets a 503 to retry |
/audit-log |
GET | Audit log entries |
/mcp |
POST | MCP Streamable HTTP endpoint (mounted at app root, NOT under /api/v1) |
Auth: Most data endpoints require an X-API-Key; admin endpoints require
the admin key. Intentional public exceptions include the health/version/tool
description probes, /api/v1/whoami, and the plugin/skill bootstrap routes
(/api/v1/plugin-*, /api/v1/install-*, and /api/v1/skill/*). These public
routes expose generic software or identity-probe data, not tenant data.
Gateway-injected headers (trusted only behind the enterprise gateway):
| Header | Effect |
|---|---|
X-Agent-ID |
Scopes the request to this agent |
X-Org-Read-Only: true |
Plan-limit read-only mode — creates and other writes that grow the store return 403 PLAN_LIMIT_READ_ONLY. Deletes, memory status transitions, agent trust changes and PUT /settings stay allowed so an over-limit org can get back under its plan |
X-Tenant-ID |
Tenant identity when using the shared CAURA_API_KEY gate |
X-User-ID |
The person behind a dashboard session or JWT; the gateway sends none for an API key. Recorded in the audit trail only (see Audit attribution below), and read only when GATEWAY_SHARED_SECRET is set |
The identity headers are trusted on the gateway-header auth path. Set
GATEWAY_SHARED_SECRET so that path also requires a matching
X-Gateway-Secret. A network-exposed OSS deployment without a gateway should
set CAURA_API_KEY; that shared-key path authenticates first and prevents the
header-trust path from being reached.
This holds on both surfaces. /mcp is a separate ASGI mount with its own
auth middleware rather than a route behind get_auth_context, so "authenticates
first" is a property each surface has to implement for itself. When the key is
set, send it as X-API-Key (or a Bearer token) on MCP calls too; without it the
request is refused 401 before any identity header is consulted.
Audit attribution
A client says which Caura client it is with X-Caura-Surface. The value is the
client's own claim: it labels audit rows and metrics, and no authorization
decision reads it. The set is closed. Case and surrounding whitespace are
ignored. A value outside the set is dropped and recorded as null, and the
request goes ahead as normal (no 4xx).
X-Caura-Surface |
Client |
|---|---|
dashboard |
The enterprise web app, outside /prism |
prism |
The enterprise web app, on /prism |
broker |
caura-daemon, including the caura CLI and caura mcp-server, which reach core-api through it |
openclaw_plugin |
The OpenClaw plugin |
Calls to /mcp are recorded as mcp, from the transport. They take no header
for it, and a REST call cannot claim it.
The writes below record two keys in the audit row's detail:
user_id: the person the gateway vouched for (X-User-ID, above).surface: the allow-listedX-Caura-Surface, ormcp.
Both keys are always present on these rows, set to null when unknown. A
null user_id means nobody vouched for a person: an API key of any kind, a
call that did not come through the gateway, or a deployment without
GATEWAY_SHARED_SECRET. A row without the keys was written before they existed.
user_id is not author_user_id on keystone rows: that one is copied from the
request body.
action |
resource_type |
REST | MCP |
|---|---|---|---|
delete |
memory |
DELETE /memories/{id} |
caura_manage op=delete |
bulk_delete |
memory |
DELETE /memories, POST /memories/bulk-delete |
caura_manage op=bulk_delete |
conflict.review |
memory_conflict |
PATCH /conflicts/{id}/resolve |
— |
crystallize |
crystallization_report |
POST /crystallize |
— |
ingest_commit |
memory |
POST /ingest/commit |
— |
ingest_undo |
memory |
POST /ingest/undo/{run_id} |
— |
agent_tune |
agent |
PATCH /agents/{id}/tune |
caura_tune |
agent_trust_update |
agent |
PATCH /agents/{id}/trust |
— |
agent_fleet_update |
agent |
PATCH /agents/{id}/fleet |
— |
keystone.set |
keystone |
POST /keystones |
caura_keystones_set op=set |
keystone.delete |
keystone |
DELETE /keystones/{doc_id} |
caura_keystones_set op=delete |
quarantine.release |
memory |
PATCH /memories/{id}/status to active, on a held memory |
— |
quarantine.reject |
memory |
PATCH /memories/{id}/status to cancelled, on a held memory |
— |
session.rollback |
memory |
POST /memories/rollback-session, one row per memory it changed |
— |
Rate limiting (managed platform)
These limits apply to the managed platform at caura.ai. A self-hosted deployment enforces its own, looser per-route limits out of the box — see the self-hosted rate limiting section.
| Scope | Limit |
|---|---|
| Memory writes | 60 req/min per API key |
| Memory searches | 120 req/min per API key |
| General reads | 300 req/min per API key |
| Auth endpoints | 10 req/min per IP |
| Global DDoS floor | 1000 req/min per IP |
Exceeded limits return HTTP 429 with a Retry-After header. Rate-limited routes also carry X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset on successful responses, so a client can back off before it is throttled rather than after.
Those three headers describe the throttle only. The per-period plan quota is reported separately as X-Usage-Limit / X-Usage-Remaining on POST /memories, POST /memories/bulk and POST /search, and only where a usage meter is wired — a deployment without one (OSS standalone) omits them rather than reporting a placeholder.
Configuration is supplied through environment variables or .env. See
.env.example for the common OSS settings; the table below
also includes production-only safety controls.
The stock Compose file sets the storage service's DATABASE_URL to its bundled
PostgreSQL service. A custom storage deployment should set DATABASE_URL
directly. A complete ALLOYDB_HOST, ALLOYDB_USER, ALLOYDB_PASSWORD, and
ALLOYDB_DATABASE set (plus optional ALLOYDB_PORT) is also supported when
DATABASE_URL is absent; these are storage-service inputs, not aliases for the
POSTGRES_* fields.
| Variable | Default | Description |
|---|---|---|
POSTGRES_HOST, POSTGRES_PORT, POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB |
local PostgreSQL defaults | Inputs used by migration/dev helpers; the stock Compose file hardcodes its container connection values |
DATABASE_URL |
local PostgreSQL URL | Storage-service primary connection URL; set directly outside the stock Compose deployment |
READ_DATABASE_URL |
(empty) | Optional storage-service read-replica URL |
CORE_STORAGE_API_URL |
http://localhost:8002 |
Where core-api reaches the storage service (its writer, when the storage layer is split) |
CORE_STORAGE_SHARED_SECRET |
(empty) | Secret core-api and every storage caller send as X-Storage-Secret. Required: core-api refuses to start without it, and storage rejects a request without it. Docker Compose generates one; CORE_STORAGE_SHARED_SECRET_FILE reads it from a file instead |
ADMIN_API_KEY |
(empty) | Admin API key — bypasses tenant enforcement |
ADMIN_API_KEY_FILE |
(empty) | File holding the admin key, read only while ADMIN_API_KEY is blank. Docker Compose sets it to a key admin-key-init generates for the bundled scheduler |
CAURA_API_KEY |
(empty) | Shared perimeter key for a network-exposed OSS deployment |
GATEWAY_SHARED_SECRET |
(empty) | Secret required in X-Gateway-Secret before gateway identity headers are trusted |
JWT_SECRET |
change-me-in-production |
JWT signing secret; must be changed in production |
EMBEDDING_PROVIDER |
openai |
openai, local, or fake |
ENTITY_EXTRACTION_PROVIDER |
openai |
openai, gemini, openrouter, fake, or none (anthropic is refused at startup — no structured-output support) |
ENTITY_EXTRACTION_MODEL |
gpt-5.4-nano |
LLM model for enrichment and entity extraction; ignored (with a warning) for a provider whose model family it does not belong to, which then uses its own default |
OPENAI_API_KEY |
— | Required for OpenAI embeddings and enrichment |
USE_LLM_FOR_MEMORY_CREATION |
true |
LLM auto-classifies type, weight, title, summary on write |
ANTHROPIC_API_KEY |
— | Required for Anthropic |
OPENROUTER_API_KEY |
— | Required for OpenRouter |
GEMINI_API_KEY |
— | Required for Gemini (Developer API, from AI Studio) |
CORS_ORIGINS |
http://localhost:3000 |
Comma-separated allowed CORS origins |
ENVIRONMENT |
development |
development or production |
SETTINGS_ENCRYPTION_KEY |
— | Fernet key that encrypts tenant provider keys (api_keys.*) at rest. Required in production. Without it (dev, standalone) keys are stored as submitted; keys saved before encryption existed are encrypted on the tenant's next save |
INSTALLER_ALLOWED_API_URLS |
(empty) | Comma-separated extra origins that /install-plugin and /install-skill accept as api_url. The serving origin is always accepted; set this only when a proxy hides the public host from core-api |
PLATFORM_LLM_PROVIDER |
(empty) | Platform-default LLM: openai, vertex, or empty to disable |
PLATFORM_LLM_MODEL |
(empty) | Model override (e.g. gpt-5.4-nano, gemini-3.1-flash-lite-preview) |
PLATFORM_LLM_API_KEY |
— | OpenAI API key for the platform LLM singleton |
PLATFORM_LLM_GCP_PROJECT_ID |
— | GCP project for platform Vertex LLM |
PLATFORM_LLM_GCP_LOCATION |
us-central1 |
GCP region for platform Vertex LLM |
PLATFORM_EMBEDDING_PROVIDER |
(empty) | Platform-default embeddings: openai or empty to disable |
PLATFORM_EMBEDDING_MODEL |
(empty) | Embedding model override (e.g. text-embedding-3-small) |
PLATFORM_EMBEDDING_API_KEY |
— | OpenAI API key for platform embeddings |
With ENVIRONMENT=production, startup additionally requires
ADMIN_API_KEY, a non-default JWT_SECRET, SETTINGS_ENCRYPTION_KEY, and
either GATEWAY_SHARED_SECRET or CAURA_API_KEY. Standalone mode is rejected
in production.
caura/
├── core-api/ # Main FastAPI service
│ └── src/core_api/
│ ├── app.py # FastAPI app, lifespan, middleware
│ ├── mcp_server.py # MCP server (Streamable HTTP, 12 tools)
│ ├── constants.py # Limits and ranking parameters
│ ├── config.py # Settings (env vars)
│ ├── auth.py # API key + JWT auth, tenant enforcement
│ ├── routes/ # Route handlers
│ ├── services/ # Business logic
│ ├── providers/ # LLM/embedding abstraction + fallback
│ ├── pipeline/ # Composable write/search pipelines
│ └── tools/ # MCP tool implementations
│
├── core-storage-api/ # PostgreSQL CRUD microservice
│ └── src/core_storage_api/
│ ├── routers/ # Memory, entity, document, fleet CRUD
│ ├── services/ # ORM operations
│ └── database/ # Engine initialization and Alembic migrations
│
├── plugin/ # OpenClaw plugin (TypeScript)
│ └── src/
│ ├── tools.ts # Tool implementations
│ ├── agent-auth.ts # Per-agent credentials (agent-scoped mc_ keys)
│ ├── context-engine.ts # Auto-read/write lifecycle
│ ├── heartbeat.ts # 60s heartbeat → Caura API
│ └── educate.ts # Agent education delivery
│
├── common/ # Shared SQLAlchemy ORM models and constants
├── tests/ # Test suite
├── scripts/ # Smoke tests, benchmarks, export tools
├── docker-compose.yml # Production-like stack
├── docker-compose.dev.yml # Dev stack
└── .env.example # Common OSS configuration template
Typical results on a single-instance deployment (OpenAI embeddings + GPT-5.4 Nano):
| Operation | Mean | P50 | P95 |
|---|---|---|---|
caura_write |
~2000ms | ~2000ms | ~2300ms |
caura_recall |
~650ms | ~640ms | ~670ms |
caura_recall (with include_brief=true) |
~1300ms | ~1200ms | ~2100ms |
Write latency is dominated by LLM enrichment. Recall latency by the embedding API call.
See BENCHMARKS.md and the
performance guide for current methodology and reproducible
benchmarks.