External memory for AI assistants that remains useful after the context window ends.
Remembra stores facts, decisions, preferences, roles, constraints, relationships, and project history outside a model context window. It then retrieves only the authorized, relevant subset for the current session. The same memory service is available through MCP, HTTP, a TypeScript SDK, and a web dashboard.
Current release: @hilbras/remembra@5.7.1 Β· V5.7.1 release notes
Long-running assistants fail in predictable ways: they repeat questions, lose decisions, confuse project context, and treat untrusted transcript text as trusted instructions. Remembra separates memory storage from the conversation and gives the host a small, inspectable control plane.
conversation βββΊ memory_store / memory_digest βββΊ durable memory
β
new session βββΊ memory_context / memory_search βββββββββ
β
βββΊ bounded, relevant context only
V5 adds the production boundaries needed for shared deployments:
- Deterministic context assembly with explicit token and candidate budgets.
- Host-resolved tenant identity; public callers cannot select an organization.
- Strict and legacy modes with fail-closed migration/readiness checks.
- Bounded hybrid retrieval across keyword, vector, trust, scope, time, and relations.
- Verified recovery with signed snapshots, migration manifests, checkpoints, and rollback.
- Compatibility first: V4.9 clients, Markdown, the eleven memory types, and the original thirteen MCP tools remain available.
| Integration | Best for | Identity/auth model |
|---|---|---|
| MCP over stdio | OpenCode, Claude Code, Cline, Kimi Code | Local process; no network API key required |
| HTTP API | Web apps, ChatGPT actions, service-to-service calls | REMEMBRA_API_KEY plus a trusted host tenant resolver in strict mode |
| TypeScript SDK | Node and edge-compatible fetch clients | API key; SDK rejects caller-supplied identity fields |
| Host TypeScript APIs | Multi-tenant services and operators | Opaque, host-minted TenantContext objects |
| Web dashboard | Browsing, editing, auditing, graph, operations | Served by the HTTP process; data calls remain authenticated |
- Node.js
18.14.1or newer. - No API key is needed for keyword-only local MCP or HTTP-on-loopback use.
- SQLite is used by the CLI when
better-sqlite3is available; the Markdown backend remains supported for compatibility and library integrations.
npm install -g @hilbras/remembraMCP is the default mode. Configure your client to launch the remembra binary over stdio:
remembraFor Claude Code:
claude mcp add remembra -- remembraFor OpenCode, add the following to ~/.config/opencode/opencode.json:
{
"mcp": {
"remembra": {
"type": "local",
"command": ["remembra"]
}
}
}See client setup for Cline, Kimi Code, and other clients.
export REMEMBRA_API_KEY="$(openssl rand -hex 32)"
export REMEMBRA_HOME="$HOME/.remembra"
remembra --http --port 8787The server binds to loopback when no key is configured. A non-loopback host without a key is refused. Put a TLS reverse proxy in front of any public deployment.
curl http://127.0.0.1:8787/healthOpen http://127.0.0.1:8787/ for the dashboard. The shell is static and public on a keyed server; all data requests made by the page still require the API key.
curl -X POST http://127.0.0.1:8787/api/v1/memories \
-H "content-type: application/json" \
-H "x-api-key: $REMEMBRA_API_KEY" \
-d '{"type":"decision","content":"The project uses the /api/v1 namespace","scope":"project/demo","importance":4}'
curl "http://127.0.0.1:8787/api/v1/memories/search?query=api&scope=project/demo&limit=5" \
-H "x-api-key: $REMEMBRA_API_KEY"New integrations should use /api/v1/*. Existing unversioned routes remain supported for V4.9 compatibility.
The SDK is a side-effect-free fetch client; importing it does not start the CLI.
import { Remembra } from "@hilbras/remembra/sdk";
const memory = new Remembra({
endpoint: "http://127.0.0.1:8787",
apiKey: process.env.REMEMBRA_API_KEY,
});
await memory.store({
type: "fact",
content: "The project uses /api/v1",
scope: "project/demo",
});
const results = await memory.search({ query: "api", scope: "project/demo", limit: 5 });
const context = await memory.context({
query: "What architecture decisions should I remember?",
scope: "project/demo",
maxTokens: 4000,
limit: 50,
});
console.log(context.context);
console.log(context.tokenCount, context.retrievalMetadata.selectedCount);The SDK supports pagination, cancellation, structured errors, lifecycle operations, history, relations, batch operations, and the trusted tenant entity methods documented in docs/sdk.md.
A dependency-free Python client with synchronous and asynchronous surfaces is
available as hilbras-remembra; see docs/python.md.
Every memory has a type, scope, provenance, trust level, importance, retention policy, and optional typed relations. The model deliberately separates what was observed from whether it should be trusted.
| Type | Use it for | Example |
|---|---|---|
fact |
Stable knowledge | βThe project uses PostgreSQL 16.β |
preference |
User or team preferences | βPrefer concise answers.β |
decision |
A choice already made | βChose JWT over server sessions.β |
constraint |
A hard limit or prohibition | βNever commit secrets.β |
instruction |
Standing behavior | βRun tests before opening a PR.β |
role |
Persona or operating rule set | βAct as the staff engineer.β |
entity |
A named person, service, or repository | βbilling-service belongs to payments.β |
relationship |
A typed connection between entities | βbilling-service depends on ledger-db.β |
event |
A dated occurrence | βThe events table was migrated.β |
history |
Condensed chronology of work | βAuthentication was redesigned in March.β |
observation |
Raw signal awaiting validation | βp95 spiked after deployment.β |
system and verified content is treated as stronger evidence than trusted content; unverified content remains searchable but does not receive the standing-instruction boost. Conversation-derived roles and instructions land unverified until explicitly approved.
Treat retrieved memories as data with provenance, not automatically as commands. In particular, a global role or instruction is intentionally powerful and should be audited before it is trusted in a shared deployment. See security.
A scope narrows relevance; it never replaces tenant authorization. In V5, global means global within the authenticated organization, not global across all tenants. Retrieval combines:
- authorized tenant/project filters;
- standing role and instruction gates;
- keyword and optional embedding candidates;
- trust, provenance, pinned retention, importance, and recency;
- temporal filters and bounded one-hop relation expansion;
- deterministic tie-breaking and diversity selection.
Oversized context candidates are skipped and counted. Remembra never silently truncates a memory or returns an over-budget context.
memory_context and POST /api/v1/context use the same ranked, authorized retrieval path as search, then apply a deterministic token budget while walking the ranked results.
curl -X POST http://127.0.0.1:8787/api/v1/context \
-H "content-type: application/json" \
-H "x-api-key: $REMEMBRA_API_KEY" \
-d '{"query":"release and migration decisions","scope":"project/demo","maxTokens":4000,"limit":50,"explain":true}'The response contains:
{
"memories": [
{
"id": "memory-id",
"type": "decision",
"content": "The project uses the /api/v1 namespace",
"scope": "project/demo",
"trust": "trusted"
}
],
"context": "[memory-id] DECISION (scope: project/demo, trust: trusted)\nThe project uses the /api/v1 namespace",
"tokenCount": 812,
"retrievalMetadata": {
"query": "release and migration decisions",
"scope": "project/demo",
"maxTokens": 4000,
"tokenCounter": "conservative-estimate-v1",
"candidateCount": 50,
"selectedCount": 3,
"omittedCount": 2
}
}The default budget is 4000 tokens, the hard maximum is 100000, and the candidate cap is 100. Internal embedding vectors are never included in context responses. The context API is read-only and does not refresh recency.
Full contract: V5 context specification.
The HTTP process exposes the dashboard, health/readiness, metrics, memory operations, snapshots, administration, and V5 context/tenant routes.
| Surface | Examples | Purpose |
|---|---|---|
| Health and metrics | GET /health, GET /metrics |
Readiness, liveness, Prometheus metrics |
| Memory CRUD | POST /api/v1/memories, GET /api/v1/memories/:id, PUT, DELETE |
Store, inspect, patch, and forget memories |
| Retrieval | GET /api/v1/memories/search, POST /api/v1/context |
Ranked search and bounded context |
| Lifecycle | POST /api/v1/memories/:id/archive, .../revive, POST /api/v1/maintain |
Archive, restore, decay, and vector backfill |
| Graph/history | GET .../:id/history, POST .../:id/relate |
Diffs, typed relations, and backlinks |
| Batch/digest | POST /api/v1/memories/batch, POST /api/v1/memories/digest |
Bounded writes/read-only search and LLM extraction |
| Snapshots | GET /api/v1/snapshot, POST /api/v1/import |
Portable backup and idempotent restore |
| Administration | GET /api/v1/audit, GET /api/v1/quality, GET /api/v1/agents/:id |
Audit and operational visibility |
| Tenant entities | /api/v1/tenant/organization, /tenant/entities/..., /tenant/memberships/... |
Trusted organization, user, project, agent, and membership administration |
When REMEMBRA_API_KEY is set, data routes accept x-api-key or Authorization: Bearer. /health and the static dashboard shell are intentionally public; the shell contains no memory data. /api/v1 responses include X-Remembra-API-Version: v1.
The SDK preserves legacy response shapes and exposes server-managed identity rejection. For the complete route/error compatibility contract, see public API and stability.
The V4.9 manifest contains thirteen tools. V5 manifest version 2 adds memory_context without renaming or removing an existing tool.
| Tool | Purpose |
|---|---|
memory_store |
Persist one of eleven memory types |
memory_batch |
Bounded store, update, delete, selected export, or read-only search |
memory_update |
Patch fields with optional optimistic concurrency |
memory_archive |
Park a memory without deleting it |
memory_revive |
Return an archived memory to active storage |
memory_digest |
Extract and store memories from a transcript |
memory_search |
Retrieve relevant memories |
memory_context |
Build deterministic, token-bounded V5 context |
memory_list |
Browse stored memories with filters/pagination |
memory_get |
Fetch one memory, relations, and backlinks |
memory_relate |
Add, remove, or retype graph relationships |
memory_history |
View version history and line diffs |
memory_maintain |
Run decay, deletion, and embedding backfill |
memory_forget |
Permanently delete one memory |
Full argument schemas and session-flow guidance are in docs/tools.md.
V5 models an organization as the security boundary:
organization
βββ users
βββ projects
βββ agents
βββ memories
The host authenticates the caller, resolves current membership, and mints an opaque TenantContext. Remembra does not trust a tenant ID, organization ID, user ID, project ID, or agent ID supplied in an ordinary HTTP body, query parameter, MCP argument, or SDK payload.
| Mode | Behavior |
|---|---|
legacy |
Explicit V4.9 compatibility. Reads/writes only the legacy namespace and refuses mixed tenant data. |
strict |
Requires a current host-minted context for every data-plane operation; missing, stale, or mixed data fails closed. |
V4.9 memories do not silently acquire a tenant. A migration must explicitly assign them to an organization and produce a signed, checksummed manifest before strict rollout.
The CLI can bind a local process to one trusted tenant through environment variables:
export REMEMBRA_TENANT_MODE=strict
export REMEMBRA_TENANT_ID=org-demo
export REMEMBRA_TENANT_MEMBERSHIP_VERSION=membership-42
export REMEMBRA_TENANT_PROJECT_ID=project-demo
export REMEMBRA_SNAPSHOT_KEY="$(openssl rand -hex 32)"
remembra --httpREMEMBRA_TENANT_ID is the opaque organization selector. The membership version must match the host's current directory state. Raw HTTP headers, CLI arguments, and SDK fields cannot replace these bindings.
import { createTenantContext } from "@hilbras/remembra/tenant";
const tenant = createTenantContext({
organizationId: "org-demo",
membershipVersion: "membership-42",
projectId: "project-demo",
scopes: ["global", "project/project-demo"],
capabilities: ["tenant:read", "tenant:write"],
});
// Pass this opaque object only from trusted host code.
// A public transport resolver must return it after authentication.For organization administration, inject a TenantDirectory and TenantEntityService into the host application. Organization provisioning is default-deny unless an explicit authorizeBootstrap hook is supplied. Membership changes are versioned, bounded, and audited.
The complete isolation and migration contract is in V5 tenant specification. The security evidence requirements are in V5 threat model.
One process can still do everything β that is the default and it needs no configuration. Splitting the duties is a choice, not a requirement.
| Command | HTTP | Worker | Scheduler |
|---|---|---|---|
remembra (no subcommand) |
yes | yes | yes |
remembra serve [--port N] |
yes | no | no |
remembra worker |
no | yes | no |
remembra scheduler |
no | no | yes |
A role without an HTTP surface refuses --http and --port rather than
ignoring them: a worker that quietly bound a port would serve traffic from a
process whose entire premise is that it has no HTTP surface.
npm install redis # optional; not a dependency of this package
export REMEMBRA_REDIS_URL=redis://:password@host:6379
remembra serve # instance 1
remembra worker # and/or a separate worker processRedis is optional, and there is no fallback. With the variable unset, no
Redis code is imported and no shared store is opened β /health/ready is
byte-identical to a single-process install. With it set but unreachable,
startup fails rather than degrading to per-instance limits, because two
instances each allowing 60/min means a fleet allowing 120/min while every
instance reports a limit it is not enforcing. If the connection drops after
startup the limiter fails closed (503) and recovers on its own at the next
successful operation; readiness reports unready so a load balancer drains the
instance, while liveness stays up so it can be diagnosed.
Runtime dependencies are unchanged at four. redis is an optional peer, not a
dependency. The capability manifest advertises distributed whether or not the
variable is set β it describes the build, like webhooks, not the
deployment.
Not integration-tested against a live Redis. The quota scripts are verified by executing them through a Lua 5.3 VM with a shim of the Redis commands they use, and the failure matrix verifies our response to a store outage rather than Redis itself. See lock and jobs.
The CLI prefers SQLite for local durability and search:
$REMEMBRA_HOME/
βββ data.sqlite # default SQLite database
βββ data.sqlite-wal/-shm # only while SQLite is open
βββ .history/ # file-backend history, when applicable
βββ .remembra.lock # cross-process mutation lock
MemoryStore provides the readable Markdown backend and remains compatible with legacy data and explicit Markdown export/import. SQLite adds FTS5 when available, bounded candidate SQL, WAL mode, vector blobs, audit tables, and online backup support.
Memory IDs are globally unique. Relations, superseded references, snapshot references, and compressed references are validated against the same store before publication.
- Store: durable write with atomic file/database semantics.
- Update: optimistic concurrency through
expectedVersion; a stale writer receivesCONFLICTand writes nothing. - Archive/revive: reversible lifecycle transitions.
- Decay: unused memories may be archived; expired archived memories may be deleted.
- History: content changes retain bounded pre-images and line diffs.
- Maintenance:
remembra maintainperforms decay and embedding backfill.
# Portable snapshot; legacy mode accepts unsigned V4 snapshots.
remembra export backup.json
# Validate the entire snapshot without writing.
remembra import backup.json --dry-run
# Idempotent restore; existing IDs/duplicates are skipped.
remembra import backup.jsonIn strict mode, exports and imports use a canonical HMAC envelope and require REMEMBRA_SNAPSHOT_KEY. The complete snapshot/reference preflight happens before any write; per-record restore is idempotent but an operational failure after preflight can leave a partial application, so keep a verified backup. Tenant migration adds a signed manifest, checksum preflight, durable checkpoints, verified resume, failure records, and an explicit publication marker. For the V5.0.1 analyze/plan/apply commands, see the security and migration guide.
SQLite operators can use the verified recovery helpers from @hilbras/remembra/sqlite-recovery for online backup, integrity/schema checks, atomic restore, and retained-previous rollback. Close the live service before restoring and reject active SQLite sidecars.
See storage, migration, and public API.
Storage and keyword retrieval work without any provider key. Configure providers only when you need digest extraction or semantic search.
# Hosted providers
export OPENAI_API_KEY=...
export REMEMBRA_LLM=openai
export REMEMBRA_EMBEDDINGS=openai
# Fully local Ollama
export REMEMBRA_LLM=ollama
export REMEMBRA_LLM_MODEL=llama3.2
export REMEMBRA_EMBEDDINGS=ollama
export REMEMBRA_EMBEDDING_MODEL=nomic-embed-textProvider calls have per-attempt timeouts, bounded retries/backoff, a wall-clock budget, cancellation, and normalized errors. Embedding failure degrades to keyword search; LLM digest failure does not partially store a transcript. The provider adapter contract is public through @hilbras/remembra/providers.
See provider configuration for all variables and adapter examples.
Remembra is designed to fail closed at the boundaries that matter:
- API keys are compared with timing-safe equality; a non-loopback server without a key refuses to start.
- Request bodies, batches, queues, provider calls, retrieval candidates, context budgets, and page sizes are bounded.
- Tenant predicates are applied inside backend queries before limits, counts, ranking, and cursor creation.
- Caches and provider work use tenant-safe partitions.
- Public identity fields and tenant headers are rejected rather than trusted.
- Snapshots, migrations, directories, and SQLite restores reject symlinks, oversized inputs, tampering, and invalid references.
- Audit, structured logs, and Prometheus metrics are available for operational review.
- Optional PII redaction (
REMEMBRA_REDACT=1) and AES-256-GCM file-backend encryption (REMEMBRA_ENCRYPT_KEY) are available for higher-risk deployments; SQLite, snapshots, and transport need separate volume/backup/TLS controls.
- Use a long random
REMEMBRA_API_KEYfor every non-loopback HTTP deployment. - Terminate TLS at a reverse proxy; do not expose plain HTTP directly to the internet.
- Keep
REMEMBRA_HOME, snapshots, migration manifests, and encryption keys access-controlled. - Use strict mode with a host-resolved identity system for shared or multi-tenant deployments.
- Audit global roles and instructions regularly.
- Test encrypted/signed backups and restore procedures before relying on them.
- Monitor
/health,/metrics, structured logs, and queue/provider failures. - Use filesystem encryption and OS process isolation in addition to application-level controls.
Read security, self-hosting, V5.0.2 authorization, the final V5 threat model, and observability before exposing a deployment.
| Variable | Default | Purpose |
|---|---|---|
REMEMBRA_HOME |
~/.remembra |
Local data root |
REMEMBRA_API_KEY |
unset | HTTP API/UI data authentication |
REMEMBRA_HOST |
loopback-safe | HTTP bind address |
REMEMBRA_PORT |
8787 |
HTTP port; --port overrides it |
REMEMBRA_UI |
enabled | Set 0 to disable dashboard routes |
REMEMBRA_REDACT |
disabled | Set 1 for irreversible ingest-time PII redaction |
REMEMBRA_ENCRYPT_KEY |
disabled | 32-byte hex key for AES-256-GCM file-backend memory/history encryption (not SQLite/snapshots/transport) |
REMEMBRA_LLM |
openai |
Digest provider |
REMEMBRA_EMBEDDINGS |
none |
Embedding provider; none keeps keyword mode |
REMEMBRA_TENANT_MODE |
legacy |
legacy or fail-closed strict |
REMEMBRA_TENANT_ID |
β | Required organization selector in strict local mode |
REMEMBRA_TENANT_MEMBERSHIP_VERSION |
β | Required current membership version in strict local mode |
REMEMBRA_SNAPSHOT_KEY |
β | 64-character hex HMAC key for strict snapshots |
REMEMBRA_REDIS_URL |
unset | Shared state for locks and quota budgets across instances. Unset means single-host, and no Redis code is imported at all. See distributed deployments. |
REMEMBRA_WORKER_CONCURRENCY |
2 |
Simultaneous jobs for remembra worker |
REMEMBRA_WORKER_LEASE_MS |
60000 |
Claim lease a worker renews while a job runs |
REMEMBRA_SCHEDULER_INTERVAL_MS |
900000 |
How often remembra scheduler enqueues periodic work |
Additional limits and provider controls are documented in self-hosting, security, and providers.
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Transport adapters β
β MCP stdio Β· HTTP /api/v1 Β· TypeScript SDK Β· dashboard β
ββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MemoryService β
β authorization Β· ranking Β· context Β· lifecycle Β· snapshots β
β digest Β· relations Β· jobs Β· provider policy β
ββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MemoryBackend β
β SqliteBackend (default CLI) Β· MemoryStore (file backend) β
ββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββ΄ββββββββββββββββββββ
βΌ βΌ
βββββββββββββββββ ββββββββββββββββββββββ
β Provider β β Tenant directory β
β adapters β β in-memory / file β
βββββββββββββββββ ββββββββββββββββββββββ
β
βΌ (only when REMEMBRA_REDIS_URL is set)
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Shared state β optional β
β LockProvider (leases) Β· JobStore (durable ledger) Β· quota β
β reached only through a dynamic import of an optional peer β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key design properties:
- transports are thin; policy and authorization live in the service/backend boundary;
- V5.0.2 makes project/user/agent selectors conjunctive and requires explicit export authority;
- SQLite and file backends implement the same tenant-aware contract;
- atomic writes, advisory locking, and crash recovery protect local durability;
- relation/history/audit/job paths apply the same tenant filter as primary reads;
- provider failures are bounded and do not weaken storage correctness;
- cross-instance primitives are optional and additive: leases rather than held locks, atomic claims rather than check-then-act, and a shared store that fails closed rather than degrading to local state.
See architecture, memory model, and storage format for the detailed contracts.
git clone https://github.com/Hilbras/Remembra.git
cd Remembra
npm install
npm run build
npm test
npm run docs:checkUseful commands:
| Command | Purpose |
|---|---|
npm run build |
Compile TypeScript and copy the dashboard |
npm test |
Run the full test suite |
npm run docs:check |
Validate relative documentation links |
npm run security:check |
Run the tenant/security adversarial matrix |
npm run recovery:check |
Run migration, snapshot, and SQLite recovery tests |
npm run bench:scale |
Run deterministic 10K/50K scale benchmarks |
npm run bench:tenant |
Run isolated strict-tenant 10K/100K benchmarks |
npm run release:check |
Run the complete fail-closed release gate |
npm run dev |
Run TypeScript in watch mode |
The release gate includes build, tests, security/recovery matrices, documentation, audit, package contents, and benchmarks. See V5 release gates before publishing.
| Area | Documentation |
|---|---|
| First install | Getting started |
| Clients and MCP setup | Clients Β· Tools |
| HTTP and SDK | Public API Β· TypeScript SDK Β· Python SDK Β· Examples |
| V5 context | Context contract Β· Policy |
| Tenants and migration | Tenant contract Β· V5.0.2 authorization Β· V4.9 migration Β· V5.0.1 migration guide |
| Security | Security model Β· Threat model |
| Storage and recovery | Storage Β· Architecture |
| Providers | Providers |
| Operations | Self-hosting Β· Observability Β· UI |
| Distributed | Lock and jobs Β· V5.6.0 audit |
| Project process | Contributing Β· Changelog Β· V5 gates Β· V5.4.0 compatibility Β· V5.5.0 compatibility Β· V5.6.0 compatibility |
| Future architecture | V6 architecture specification |
- V4.9 remains supported: legacy HTTP routes, Markdown, existing clients, and all thirteen original MCP tools remain available.
- V5 is additive:
memory_context, tenant entities, and versioned APIs do not rename or remove the V4.9 surface. - V5.6 is additive and opt-in: locks, the job ledger, the durable worker, and the Redis adapter are new subpaths. A single-process deployment is unchanged β no new dependency, no new configuration, and an unchanged readiness payload. See V5.6.0 compatibility.
- Current focus: hardening the production memory platform, operational recovery, and measurable retrieval quality.
- V6 direction: see the V6 architecture specification for the security-first policy model, provider independence, offline-first core, migration lifecycle, and release roadmap.
- Schema boundary: tenantless V4 records use schema
3; tenant records use schema4.
MIT Β© Hilbras