copier: parallel chunked copy step under the table lock (CO-4, LK-3) - #128
Conversation
cc2bec2 to
3149a5b
Compare
The chunker could cut ranges but nothing copied them; the applier needs a per-key uncut/in-flight/landed answer that only the copier can give. Promotes the in-transaction lock check to dbconn so copier and builder share one LK-1 confirmation.
A resumed copy discards changes for keys above the watermark that an earlier run may already have copied, so it now deletes that tail first in guarded batches. Claims are frontier-ordered under one lock, an uncommitted chunk stays in flight, and both relations are locked before the OID check.
3149a5b to
9119c2e
Compare
A zero watermark means nothing landed, not that the shadow is empty: a run whose first chunk never landed may have landed chunks above it. Also refuse a shadow that is the source itself and stop calling a lost connection LK-1.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
🤖 Review 1/2: adversarial correctness. Inspected: the whole diff at 0 blocking, 4 non-blocking. Invariants:
The code reads correct. All four findings are either guards the suite doesn't pin or an ordering edge in the resume premise. Non-blocking1. The per-chunk guard is unpinned. Every refusal test is actually refused by the resume clear. Test that kills it (mutant: drop
|
|
🤖 Review 2/2: OSS adoption and integration ease. Inspected at 0 blocking, 3 non-blocking. What works well for importers:
Non-blocking1. Copier refusals carry their cause only in message text (
Only lock loss can be told apart structurally, via 2. Three spellings of "an optional key". 3. This review was generated by Claude Code (claude-opus-5-5). |
aparajon
left a comment
There was a problem hiding this comment.
🤖 Stamped: 0 blocking, 7 non-blocking across the two review comments above.
This stamp was left by Claude Code (claude-opus-5-5).
morgo
left a comment
There was a problem hiding this comment.
🤖 Automated adversarial review, posted on Morgan Tocker's behalf.
The structure here is right, and the part I went looking hardest at — whether Position can ever lie to the applier — holds up. Cutting and registering under one lock is what makes "the frontier never runs ahead of an unregistered chunk" a property rather than a hope, and keeping the unlanded chunk in flight rather than rolling the frontier back is the correct shape: it fails toward deferral, which is the safe direction for CO-4.
The Chunker.Next mutex note I left on #127 was acted on, and more thoroughly than I suggested — the chunker now splits nextMu (held across the boundary query) from mu (taken only to read the cursor and publish the result), so Cut and Feedback no longer park behind a live query, and Copier.claim mirrors the same split with claimMu/mu. Lock order is claimMu → mu everywhere, never the reverse. Thank you.
Verified rather than assumed:
Run's clean exit is not turned into a cancellation by its own defer.return c.finish(context.Cause(ctx))evaluates the argument beforedefer stop(nil)fires, so a copy that covered the key space reports nil rather thancontext.Canceled. That ordering is load-bearing and invisible; it is worth a comment on thedefer, because reordering it intocause := context.Cause(ctx)after the defers would break every successful run.- A chunk whose
Commitfailed after the server committed is handled.copyChunkreturns the error, the chunk never lands, it stays in flight, and the watermark stops below it — so the committed rows sit above the watermark and the resume clear removes them before they can be read as landed. This is the case that justifiesclearAboveexisting at all, and it works. - Clearing under a live applier is safe.
Positionalready reportsCutas the resume watermark (or invalid) fromNewCopieronward, beforeRunis called, so every key above the watermark classifies asKeyUncutfor the entire duration of the clear and the applier discards rather than writing there. Nothing can land in the range being deleted while it is being deleted. - A resumed ledger cannot accept a chunk overlapping the landed prefix.
newLedgerseedscutfrom the watermark, sostartsAtFrontierdemandslower == watermark+1for the very first claim, and refuses everything oncecut == math.MaxInt64. - The clear loop terminates. It relies on
c.chunker.Rows()being positive — a zero would makeremoved < batchpermanently false and spin forever — andChunkerOptions.validaterejects a non-positive row count afterwithDefaults, soRows()is at leastMinRows. Sound, but it is an unstated dependency ofresume.goon the chunker's validation. guard's ordering does what its comment claims.LOCK TABLEresolves both names before the OID comparison, and ACCESS SHARE is sufficient because the things that would invalidate the proof — DROP, rename, most rewriting DDL — need ACCESS EXCLUSIVE.
One finding.
The resume clear re-scans from the same lower bound on every batch
batch := c.chunker.Rows()
for {
removed, err := c.clearBatch(ctx, pool, lower, batch)
...
if removed < batch {
return nil
}
}lower is computed once and never moves. Each batch runs
DELETE FROM shadow WHERE key IN (
SELECT key FROM shadow WHERE key >= $1 ORDER BY key LIMIT $2)with the same $1, so batch k must walk past everything the previous k-1 batches deleted before it reaches a live row. Deleted index entries are not removed by the delete itself — only by vacuum — so the scan traverses them. The work is quadratic in the number of rows cleared.
What makes this worth fixing rather than noting is that clearAboveSQL's own comment states the property that makes the fix trivially correct:
The subquery orders by the key so each batch removes a contiguous stretch off the primary-key index rather than an arbitrary sample
A contiguous stretch has a highest key. The cursor can advance to it. The chunker sitting next to this code does exactly that — advance(upper) sets next = upper + 1 — and the clear does not.
Failure scenario. A copy of a 50M-row table fails after the first chunk was pinned mid-insert and cancelled while later chunks committed (the situation TestCopierCancellationKeepsTheCancelledChunkInFlight constructs deliberately). Nothing landed contiguously, so the checkpointed watermark is zero, and per this PR's own reasoning the resume must clear the whole shadow. batch is Rows() on a freshly constructed chunker, i.e. DefaultInitialChunkRows = 1000, so that is 50,000 batches. With a bigint key at roughly 500 entries per btree leaf page, batch k traverses about 2k leaf pages of dead entries before finding a live one; summed, ~2.5×10⁹ buffer accesses instead of the ~100,000 an advancing cursor would need.
Estimate, not measurement: at sub-microsecond per cached page access that is on the order of tens of minutes of pure CPU, versus seconds. Two things soften it and neither bounds it — btree LP_DEAD hinting makes repeat traversals cheap per entry but does not remove them, and autovacuum may reclaim some of the tail concurrently, which makes the runtime depend on autovacuum timing rather than on the data. A resume that makes no visible progress for an unpredictable stretch, with every batch committing so there is no long transaction to alert on, is a bad failure mode for the path an operator reaches only after something already went wrong.
Fix, returning the batch's own high-water mark:
tag, err := tx.Exec(ctx, c.clearSQL, lower, limit) // → RETURNING the keyHave clearBatch scan DELETE … RETURNING key, report the maximum, and set lower = maxKey + 1 for the next batch — with a maxKey == math.MaxInt64 check to end the loop, since that is the one case where exactly batch rows were removed and there is genuinely nothing above them (incrementing would wrap).
Narrowing the range like that is safe here for the reason in the third bullet above: every key above the watermark reads as KeyUncut for the whole clear, so nothing can insert into a range a previous batch already passed. That is worth saying in the comment, because it is the non-obvious precondition the current fixed-$1 form does not need and the advancing form does.
No test would catch this. TestCopierResumesFromTheZeroWatermark clears 201 rows with InitialRows: 40, so about six batches — enough to prove correctness, far too few for the shape to show. A t.Log of the batch count, or an assertion that clearing N rows takes ⌈N/batch⌉ batches rather than more, would pin it without needing a large fixture.
Notes
land's invariant marker and its error tag disagree.
// INV: CO-4
if !c.ledger.land(chunk, inserted) {
return fmt.Errorf("%w (LK-3): landed chunk [%d, %d] was not in flight", ...)
}claim right above it has // INV: CO-4 over a (CO-4) error. These markers look greppable against docs/invariants.md, so one of the two spellings here is wrong — LK-3 is the defensible one for "was not in flight", which makes the comment the thing to change.
batch := c.chunker.Rows() couples the delete batch to the copy chunk size, inertly. It reads once, before any Feedback, so it is always InitialRows — the D12 sizing loop never influences it. That is fine, but the expression reads as if the clear adapts, and a future reader moving the read inside the loop would make delete batches follow insert timings, which is not a relationship anyone wants. A named constant, or opts.Chunker.InitialRows with a one-line reason, says what is actually meant.
checkShadow does not verify the precondition ON CONFLICT (key) needs. The doc says "what the shape cannot promise, checkShadow verifies before a connection is opened", and it covers the column list carrying the primary key — but ON CONFLICT (key) DO NOTHING additionally requires a unique index on that column in the shadow. A builder that defers index creation until after the copy (a normal optimisation, and one this design invites since the copy is the expensive part) produces a shadow that passes every check here and fails on the first chunk with SQLSTATE 42P10, after the lock session, the guard, and the clear have all run. The OID query in confirmRelations is already round-tripping to the catalog and could confirm the index in the same statement.
A transient error anywhere costs the whole copy back to the watermark. There is no retry: one lock_timeout on LOCK TABLE in guard, on any of the workers, returns from work, cancels the rest, and ends the run. That is a defensible first cut — resume exists — but combined with the zero-watermark case above it means a single transient failure before the first chunk lands discards every committed chunk and pays the clear to delete them. Worth stating in SAFETY.md as a known operational cost rather than leaving it to be discovered, since the obvious mitigation (checkpoint the watermark, retry the transient classes inside work) is a separate change.
The first resume-clear batch takes SHARE MODE on the shadow so a chunk transaction still open from the earlier run cannot commit a row behind the delete; Position.Cut becomes a Watermark; new tests pin the per-chunk OID guard, SET LOCAL ROLE, the ledger cut-through-MaxInt64 refusal, and the fence.
|
🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/128, follow-up commit Six of the seven findings are fixed in follow-up commit
The mutation run (27 of 32 killed) and the invariant checks on CO-4, LK-1, LK-3, ST-6, LK-2 and TM-2 confirm behaviour that is unchanged — no action. The two surviving defensive mutants in Source: block/pg-sprite#128, review comments 5869943436 and 5869943838 and review 5338688433 at head |
Adds
copier.Copier, the parallel chunked copy from a proven source into its built shadow, and promotes the in-transaction table-lock confirmation intopkg/dbconnso the copier and the shadow builder share it.Why
The chunker (previous PR in this stack) cuts consecutive key ranges but nothing copied them. The applier that follows needs more than a watermark: with several workers, chunks land out of order, so a captured change must be judged per key as uncut (discard), in flight (defer), or landed (apply). Only the copier knows which chunks are in flight, so it owns that answer.
The shadow builder already re-asserted the lock from inside its own transaction; the copier needs the identical check on every chunk connection. One implementation in
dbconnkeeps LK-1 enforced in one place.What
pkg/dbconn:(*TableLockSession).Confirm(ctx, conn)— nobody holds the lock → the newErrTableLockNotHeldnaming the table; another backend holds it →*TableLockHeldErrornaming it.Confirmreports what the server said; each writer wraps it once as its ownErrInvariantViolation(LK-1), so the message never states the same fact twice.schemachange.confirmTableLockmaps those onto its existingCauseLockUnconfirmed/CauseLockHeldElsewhererefusals (existing tests unchanged).pkg/copier:Shadowinterface (schema, source, shadow, both OIDs, copy columns) — the shapeschemachange.BuiltShadowsatisfies, sopkg/schemachange(builder and orchestrator) can import the copier without a cycle.NewCopierrefuses a shadow that is not the target's, or that is the source itself by name or by OID (ST-6) and a lock session that is missing, for another table, or already lost (LK-1).Copier.Run(ctx, pool): N workers (default 4) under the lock session'sBindcontext. Each chunk runs in its own transaction:SET LOCAL lock_timeout/statement_timeout,SET LOCAL ROLE owner,lock.Confirm(only "no holder" / "another backend" is LK-1; a lookup that got no answer is reported as the connection's error),LOCK TABLE source, shadow IN ACCESS SHARE MODE(so neither relation can be replaced for the rest of the transaction; a name that no longer resolves is ST-6), re-resolve both relation OIDs against the proof, then one frozen statementINSERT INTO shadow (cols) SELECT cols FROM source WHERE pk BETWEEN $1::bigint AND $2::bigint ON CONFLICT (pk) DO NOTHING. Elapsed time from an injectedprogress.ClockfeedsChunker.Feedback(D12).Position: cutting and registering a chunk happen under one lock, and the ledger refuses a chunk that does not start just above the cut frontier (CO-4 fail-closed), so the frontier never runs ahead of an unregistered chunk; a chunk is registered in flight before its transaction begins and leaves the in-flight set only when it commits — a chunk whose transaction did not commit stays inInFlight, so its keys never read as landed; the watermark advances over the contiguous landed prefix;Position.Classify(key)returnsKeyUncut/KeyInFlight/KeyLanded— the CO-4 rule the applier will call.Runreturns only after every worker has exited, so the copier holds no chunk transaction when a caller checkpoints the watermark (LK-3, client side; a cancelled statement ends on the server when it finishes or hitsstatement_timeout), then fails closed if the key space is not covered.DELETE … WHERE pk IN (SELECT pk … WHERE pk >= $1 ORDER BY pk LIMIT $2), $1 being the first key the resumed chunker will cut, each batch in the same guarded transaction shape), before its first chunk is cut. An earlier run may have landed chunks above W out of order, and the applier discards changes for keys above the resumed cut frontier, so a stale shadow row left there would surviveON CONFLICT DO NOTHING. A zero watermark means nothing landed, not that the shadow is empty — a run whose first chunk never landed may have landed chunks anywhere above it — so the clear then covers the whole shadow; a first run finds it empty and its one batch removes nothing. The first clear batch takesLOCK TABLE shadow IN SHARE MODEbefore its delete, so a straggling chunk transaction from the earlier run — one whoseINSERTis still open when the resume begins — commits (or aborts) before the clear reads the shadow, and a row it commits cannot slip in behind the delete. The fence also waits on any other open writer of the shadow and blocks new writes for the length of that one batch, bounded bylock_timeout; later batches do not fence.Position.Cutis aWatermark, likePosition.Watermarkand the ledger's frontier, so "nothing cut" and "cut through key K" have the one spelling the applier and the checkpoint will read (Watermark.Valid/Key);CutValidis gone.ErrInvariantViolationwrapping the cause until the orchestrator PR maps them onto the schema-change refusal classes;docs/refusal-classes.mdsays so).ALTER COLUMN TYPE … USINGchange needs a conversion-aware copy before the planner's copy-and-swap route (still refused today) is wired to the copier;copySQLand the design package map say so.SAFETY.mdcopier row,invariants.mdCO-4 / LK-1 / LK-3 Enforced today (LK-3 scoped to the client side), design package map (dbconn,copier) and D12,architecture.md.Tests (real PostgreSQL, PG 14/16/18): whole-table copy with 3 workers and 100-row chunks converges via
testutil.AssertConvergedwhile a sampler asserts everyPositionsnapshot is consistent (lowest in-flight chunk starts at watermark+1, nothing in flight ⇒ watermark = cut); a shadow row the applier commits while the copier's insert of that chunk is waiting on it is never overwritten; resume from a watermark removes stale shadow rows above it, keeps the row below it, and copies every key above it from the source; resume from the zero watermark removes stale rows down to the smallest int64 key and converges; cancellation with one chunk pinned mid-insert by an uncommitted shadow row returnscontext.Canceledwith exactly the pinned chunk in flight (Classifysays in-flight for its keys, landed for a chunk that landed above it), every key ≤ watermark present, and a resumed copier clears the tail, picks up a source change made in it meanwhile, and converges; lock loss mid-copy returnsErrInvariantViolationwrapping the session's loss with the interrupted chunk still in flight; a dropped-and-recreated shadow or source is refused (ST-6) with zero rows written, and a dropped source fails the write transaction's ownLOCK TABLE(SQLSTATE 42P01 → ST-6); a gone or rival-held lock is refused with zero rows written, and a session that already reported loss is refused byNewCopierbefore any connection opens; a lock confirmation that never got the server's answer is not an invariant violation; a shadow dropped and recreated between a chunk's claim and its transaction is refused per chunk (ST-6) with the impostor receiving no rows and that chunk left in flight; every row lands as the table owner (SET LOCAL ROLE, checked through a shadow column defaulting tocurrent_useron a table owned by a NOLOGIN role); a resume whose earlier run still has a chunk transaction open waits on it at the fence (an ungranted ShareLock inpg_locks, cut frontier unchanged, nothing in flight) and reads the committed row afterwards; the copy SQL, the resume-clear SQL, and the fence SQL are frozen as exact strings (TM-2); pure ledger tests for out-of-order landing, frontier-only claims, an unlanded chunk staying in flight, resume, and every claim refused once the key space is cut through the largest key.Before / after
References
docs/invariants.mdCO-4, LK-1, LK-3;docs/copy-and-swap-design.mdD12 and the package map.🤖 Drafted with Amp (Claude Opus 4.6); reviewed and edited by the author.