Skip to content

feat(state): single-blob HTTP cache entry + drop cache_mutex (atomicity Phase 2) - #494

Merged
transfix merged 2 commits into
masterfrom
feat/state-atomicity-cache-payload
Sep 30, 2026
Merged

transfix merged 2 commits into
masterfrom
feat/state-atomicity-cache-payload

Conversation

@transfix

Copy link
Copy Markdown
Owner

Phase 2 / atomicity of docs/roadmap/STATE_LIFETIME_AND_ATOMICITY.md, building on the Phase 1 owning accessors (#492).

Before

An HTTP cache entry was 7 metadata child value-nodes (status/etag/last_modified/content_type/effective_url/fetched_at/freshness_ttl) plus the body on the entry node's data(). Reading or writing that spread across many node locks, so it needed a process-wide cache_mutex for two jobs: (1) atomicity — a reader must not see a half-written entry; (2) lifetime — a concurrent sweepExpired() must not free a node under a bare pointer.

After

The whole entry is one opaque blob on the entry node's data() (version byte + 8 length-prefixed fields). So:

  • Atomicity — a read or write is a single atomic node op under that node's own _mutex; a reader sees the whole old blob or the whole new blob, never a torn mix.
  • Lifetime — read_entry uses findDescendantShared() (Phase 1) to pin the node, so a concurrent sweepExpired() can only unlink it, never free it under the reader.
  • Writers of one key are already serialized by the per-key single-flight (flight_mutex), and the write is atomic.

With atomicity + lifetime + single-flight covered, there's nothing left for a global lock to protect — cache_mutex is removed. Cache reads take no process-wide lock (lock-free apart from the node's own brief _mutex); the only write coordination is the pre-existing single-flight.

Notes / scope

  • Entry child nodes are no longer individually addressable via state://…?children — they were internal-only and no consumer read them (verified by grep; only the test located the entry node, which still exists).
  • Test file unchanged — its assertions are entry presence/absence + expireAt, both preserved. All 12 HttpCacheTest pass, including ConcurrentDistinctKeysAreServedSafely now without cache_mutex; AriadneStateUri/StateLifetime/AriadneNetIntrinsics green (38/38 locally).
  • The blob is a simple portable length-prefixed format (LE uint64 lengths) with a version byte; a malformed/old-version blob decodes as a cache miss (safe re-fetch).

Not in scope: PR3 cache hardening (LRU budget, Vary/private, Expires) — still tracked separately.

…cache_mutex (Phase 2)

Phase 2 / atomicity of docs/roadmap/STATE_LIFETIME_AND_ATOMICITY.md, building on the
Phase 1 owning accessors (#492).

Before: an entry was 7 metadata child value-nodes (status/etag/last_modified/
content_type/effective_url/fetched_at/freshness_ttl) plus the body on the entry
node's data(). Reading/writing that spread across N node locks needed a process-wide
cache_mutex to avoid a torn (half-written) entry AND to stop a concurrent sweep
freeing a node under a bare pointer.

After: the whole entry is ONE opaque blob on the entry node's data() (version byte +
8 length-prefixed fields). A read or write is now a SINGLE atomic node op under that
node's own _mutex — a reader sees the whole old blob or the whole new blob, never a
mix. read_entry uses findDescendantShared() (Phase 1) to PIN the node, so a concurrent
sweepExpired() can only unlink it, never free it under the reader. Writers of one key
are already serialized by the single-flight flight_mutex, and the write is atomic, so
there is nothing left for a global lock to protect: cache_mutex is REMOVED. Reads no
longer take any process-wide lock (lock-free apart from the node's own brief _mutex).

Net: lock-free, torn-read-free cache reads; the cache's only write coordination is the
pre-existing per-key single-flight. Entry child nodes are no longer individually
addressable (they were internal-only; no consumer read them — verified). Test file
unchanged (its checks are entry presence/absence + expireAt, both preserved). All 12
HttpCache tests pass, incl. ConcurrentDistinctKeysAreServedSafely now without
cache_mutex; AriadneStateUri/StateLifetime/AriadneNetIntrinsics green (38/38).
…-vs-evict UAF)

Adversarial review of the cache_mutex removal found a HIGH UAF I introduced: with
cache_mutex gone, write_entry still used entries(key) (operator(), a bare unpinned
state&) and then mutated it via data()/expireAt, which lock the node and walk its
_parent. A concurrent evict/sweepExpired (a same-key store invalidation — an
independent URI handler no longer serialized against the fetch write) could free the
node between entries(key) returning and the data()/expireAt calls → use-after-free (or
a silently-lost write). read_entry and evict were already hardened with owning
accessors; write_entry was the one mutation site left on a bare ref.

Fix: write_entry now uses entries.sharedChild(key) (owning create-or-get), so the
entry node is PINNED across data()/expireAt — a concurrent evict can only unlink it,
never free it under the write. The entry's ancestors are the fixed, never-expiring
sys.net.http_cache.entries path, so the entry pin alone suffices (no chain needed).
Also refreshed the stale header comment to describe the single-blob storage. 12/12
HttpCache tests pass.
@transfix
transfix merged commit 50a3046 into master Sep 30, 2026
12 of 13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant