Skip to content

progress: copy operation and engine-measured work source (v4) - #131

Merged
Kiran01bm merged 4 commits into
mainfrom
kiran01bm/cs5c-progress-copy
Sep 30, 2026
Merged

Kiran01bm merged 4 commits into
mainfrom
kiran01bm/cs5c-progress-copy

Conversation

@Kiran01bm

@Kiran01bm Kiran01bm commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Adds the copy operation and a progress.WorkSource the tracker polls for engine-measured work, bumping the progress report to format_version 4. No consumer of the source lands here; the next PR in the stack wires copier.Copier to it.

Why

The copy-and-swap row copy (previous PR in this stack) is many short transactions across N workers, not one server statement, so PostgreSQL publishes no pg_stat_progress_* row for it. The tracker's only work path today is the concurrent-build one — a reserved session and a backend PID queried inside Progress. The copier already knows its own counters (Position.RowsInserted, the shadow's size in the catalog); what was missing is a way to hand them to the tracker under the same polling and lifetime rules the build path has, so an operator's Progress() during the copy reports rows and bytes instead of an empty step.

What

  • pkg/progress:
    • OperationCopy = "copy".
    • WorkSource interface — Work(ctx) (Work, error) — and Tracker.SetWorkSource / Tracker.StopWorkSource. Progress polls the source inside the existing pollMu critical section, on the observer's context, so a source sees one poll at a time and may read state that is not safe for concurrent observation. StopWorkSource is a fence like StopConcurrentBuild: it waits for an in-flight observation before clearing the source, so the engine can release the source's state right after it returns. The two setters are fences too: SetWorkSource and SetConcurrentBuild each replace a poll target, so each takes pollMu then mu (the one order every taker uses) and drains an in-flight poll before the old target's owner is free to release it.
    • The WorkSource contract puts Work on the engine's stop path: it must bound itself (memory reads, or catalog reads under a session statement_timeout) and honour ctx, since the fences wait for it and an observer's context may never end; it must take no lock the engine holds while calling Set*/StopWorkSource; and a catalog-reading Work is a CO-9 read site (pg_catalog. qualification, shadowing-search_path test).
    • A step's work comes from one place: SetWorkSource drops the build session and PID (so CancelBuild refuses with ErrNoActiveBuild and stopping the source does not revive the build), and SetConcurrentBuild drops the source. Start, StartStep and Finish reset the source like they reset the build fields; a source is polled only while the tracker is running.
    • On a source error the snapshot still carries the tracker state with the error alongside, matching the build path's query-error behavior.
    • FormatVersion 3 → 4. The Work doc now distinguishes server-observed counters (blocks, tuples, lockers) from engine-measured ones (rows, bytes); neither operation fabricates the other's.
  • docs/progress-report.md: version history (v4 splits work into two counter families selected by detail.operation), work presence rule ("measured work only" — and consumers pick the family from operation, never from work being present), the two counter families, the copy row in the operations table, polling semantics including the four handoff fences and the Work obligations, and a pinned copy-step example. SAFETY.md's pkg/progress dependency rationale names the same fences.

Consumers that reject unknown format_version values need a pin bump to 4; nothing else in the shape changed, and the existing build example is unchanged apart from the version.

Tests

pkg/progress/work_source_test.go (pure, -race): the copy-step JSON shape pinned as a literal (matching the doc example; rows and bytes carry distinct values so a swapped pair fails); the source polled once per observation; a source error returns the running snapshot with Work nil; no poll after Finish, after StartStep, after Start on an abandoned run, on a pending tracker, or after StopWorkSource; build and source mutually exclusive in both directions, including CancelBuild → ErrNoActiveBuild once a source owns the step; two concurrent pollers never overlap inside the source; StopWorkSource, SetWorkSource and SetConcurrentBuild each drain an in-flight observation (one table-driven test; removing pollMu from either setter fails exactly its subtest); SetAttempt/StartStep/Finish/Start do not wait behind an in-flight poll. Every test was checked against a mutant that removes the branch it covers. TestDocStatesCurrentFormatVersion pins the doc's stated version to FormatVersion.

Before / after

Before                                    After
┌─────────┐ Progress()                    ┌─────────┐ Progress()
│ Tracker │──▶ running? ──▶ build PID? ──▶│ Tracker │──▶ running? ──▶ source? ──▶ source.Work(ctx)
└─────────┘        │            │         └─────────┘        │            │   (engine-measured)
                   ▼            ▼                            ▼            ▼
             frozen/pending  pg_stat_progress_          frozen/pending  build PID? ──▶ pg_stat_progress_
                             create_index                                            create_index
                             (server-observed)                                       (server-observed)
copy step: work absent                    copy step: work = rows_copied/total, bytes_copied/total

References

🤖 Drafted with Amp (Claude Opus 4.6); reviewed and edited by the author.

The row copy has no pg_stat_progress view, so the tracker learns its
counters from the engine: a WorkSource polled inside Progress under the
same fence as the concurrent build, mutually exclusive with it.
The statement field, the pin claim, the docs index, and the core table
still described the report as server-observed only; the second pollMu
fence (StopWorkSource) was missing from the dependency-list rationale.
@Kiran01bm
Kiran01bm marked this pull request as ready for review September 29, 2026 07:22
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@aparajon

Copy link
Copy Markdown
Collaborator

🤖 Adversarial review (1/2): 0 blocking, 1 non-blocking.

I reviewed 1f50216, which adds the copy operation and the WorkSource seam and bumps format_version to 4. pkg/progress passes under -race.

The locking holds:

  • Progress snapshots the source under mu and calls it under pollMu.
  • StopWorkSource takes the same pollMu then mu order, so there is no inversion.
  • Start, StartStep and Finish reset the source under mu alone, so a poll in flight never gates execution.
  • The build and the source are mutually exclusive in both directions, and CancelBuild refuses once a source owns the step.

I made 12 deliberate breaks to the code and 9 were caught. The caught ones:

  • polling the source before the phase check
  • either setter keeping the other's fields
  • dropping the stop fence
  • Start or StartStep keeping the source
  • swallowing a source error, or attaching Work on one
  • reverting FormatVersion

The three survivors behave the same in every case the tracker can reach:

  • Finish keeping the source: the phase check already stops a finished tracker from polling, and only Start resumes it, which resets the source.
  • Checking the build before the source: the two are never set together.
  • Clearing the source before waiting on pollMu: the stop still drains the poll.

The PG 15 test job is red on TestExecuteCreateWithProgressReportsQualifiedStepStatementsInOrder in pkg/executor with cache lookup failed for function (SQLSTATE XX000). The diff touches no executor code, PG 14 and 16-18 pass, and main's last five runs are green. It looks like a race on a database-global event trigger in the long-lived database, not this change.

Non-blocking

1. A poll that never returns wedges the copy step, because StopWorkSource has only the observer's context to bound it.

SAFETY.md:96-100 states both handoffs are observer-gated. The build path's wait has two bounds, the poller's context and the reserved session's statement_timeout. The source path has only the first. StopWorkSource waits on pollMu for as long as source.Work(ctx) runs, and that call runs on whatever context the observer passed.

  • An observer that polls with a context carrying no deadline, while the source blocks, holds the engine's end of the copy step.
  • Nothing lands here that blocks: no source ships in this PR. So this is latent.
  • The docs say the copy's bytes_* come from "a size read from the catalog", so the next PR's source will likely make a database read inside Work.

The fix belongs in the contract. WorkSource should require Work to bound itself, with a memory-only read or a catalog read under the session's statement_timeout, and not rely on the caller's context. That gives the source path the same second bound the build path gets from statement_timeout.

Test case: fails at 1f50216
package progress_test

import (
	"context"
	"testing"
	"time"

	"github.com/block/pg-sprite/pkg/progress"
	"github.com/stretchr/testify/require"
)

// An observer that polls with a context carrying no deadline, against a
// source whose read has not returned, holds the step's StopWorkSource: the
// engine cannot finish the copy step until the observer gives up.
func TestStopWorkSourceIsBoundedOnlyByTheObserversContext(t *testing.T) {
	entered := make(chan struct{})
	source := fakeSource{work: func(ctx context.Context) (progress.Work, error) {
		close(entered)
		<-ctx.Done()
		return progress.Work{}, ctx.Err()
	}}
	tracker := runningTrackerWithSource(t, source)

	observer, cancelObserver := context.WithCancel(context.WithoutCancel(t.Context()))
	defer cancelObserver()
	go func() { _, _ = tracker.Progress(observer) }()
	<-entered

	stopped := make(chan struct{})
	go func() { tracker.StopWorkSource(); close(stopped) }()
	select {
	case <-stopped:
	case <-time.After(2 * time.Second):
		require.FailNow(t, "StopWorkSource has not returned 2s after the copy step ended; it is waiting on the observer's poll")
	}
}
--- FAIL: TestStopWorkSourceIsBoundedOnlyByTheObserversContext (2.00s)
        Error:      StopWorkSource has not returned 2s after the copy step ended; it is waiting on the observer's poll

The test passes once Work bounds itself independently of ctx.

This review was generated by Claude Code (claude-opus-5-5).

@aparajon

Copy link
Copy Markdown
Collaborator

🤖 Adoption and integration review (2/2): 0 blocking, 3 non-blocking.

OSS adoption

1. The package synopsis on pkg.go.dev still describes copy counters as future work. progress.go:1-3 says the package "deliberately contains copy counters that native operations leave empty so copy-and-swap can implement the same contract." With OperationCopy and WorkSource landing here, the one line an outside reader sees first undersells what the package now does. A sentence naming the two work families (server-observed builds, engine-measured copies) would match the new Work doc.

Integration ease for importers

2. The presence of work no longer means "a build row", and the version-history entry doesn't say so. Before v4, a non-nil Detail.Work meant PostgreSQL had published a pg_stat_progress_create_index row. An importer could key on presence alone. SchemaBot does exactly that (apply.go:1926 on main):

  • It exports the six block/tuple/locker counters and derives a percent from them.
  • For a copy step it would publish blocks_*/tuples_* zeros as if they were measured.
  • It would drop rows_*/bytes_* entirely.
  • The percent stays put only because an empty server_phase misses the phase table.

SchemaBot does not pin format_version either, so the bump alone won't flag this. The summary's "nothing else in the shape changed" is true of the keys but not of what their presence means. progress-report.md:21 could add one clause: consumers select the counter family by detail.operation, not by whether work is present. That is the one behavior change an importer has to act on.

3. A source that reads the catalog is a CO-9 read site. The docs name "a size read from the catalog" as a source of bytes_*, and CO-9 covers every progress probe. When the copier becomes a WorkSource in the next PR, its size read needs pg_catalog. qualification (or LocalSearchPath) plus the shadowing-search_path test that the other pkg/progress read sites carry. The WorkSource doc could say so, so that a third-party source written against this seam inherits the obligation too.

This review was generated by Claude Code (claude-opus-5-5).

@aparajon aparajon left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Stamped with 0 blocking findings. The non-blocking notes are in the two review comments above.

This stamp was left by Claude Code (claude-opus-5-5).

@morgo morgo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Automated adversarial review, posted on Morgan Tocker's behalf.

The shape is right. An operation the server publishes no progress row for needs a way to report what it measured, and modelling it as a source the tracker polls — rather than counters the engine pushes — keeps the tracker's "one read per observation, lifetime is the caller's context" property intact. Landing the interface and the fence before the consumer, with the format version bumped in the same commit as the new operation, is the right order.

Verified rather than assumed:

  • The new precedence in Progress is behaviour-preserving for the build path. The old guard was pid == 0 || session == nil || s.Phase != PhaseRunning; the new code hoists the phase test and then tests source before build. Because SetWorkSource clears session/buildPID and SetConcurrentBuild clears source, the two are mutually exclusive, so the reorder cannot shadow a build that would previously have been polled.
  • StopWorkSource takes the locks in the same order Progress does (pollMu then mu), matching StopConcurrentBuild and CancelBuild, so the new fence introduces no lock-order inversion among the tracker's own methods.
  • observeEngineWork cannot alias a snapshot. work is a fresh value per call and Snapshot is taken by value, so two observers never share a *Work.
  • A source is polled only while the tracker is running, because the phase test returns before the source is consulted, and Start/StartStep/Finish all null it.

SetWorkSource is StopConcurrentBuild without the fence, and its doc comment recommends it as a substitute

func (t *Tracker) SetWorkSource(source WorkSource) {
	t.mu.Lock()
	defer t.mu.Unlock()
	t.source = source
	t.session, t.buildPID = nil, 0
}

func (t *Tracker) StopConcurrentBuild() {
	t.pollMu.Lock()
	defer t.pollMu.Unlock()
	t.mu.Lock()
	defer t.mu.Unlock()
	t.session, t.buildPID = nil, 0
}

The clearing is character-for-character identical. The only difference is pollMu, and StopConcurrentBuild's own comment says what that difference buys:

the executor calls it before the build's own session can return to the pool, so no signal that read the build's PID completes after that backend could be running someone else's statement.

SetWorkSource's comment, meanwhile, advertises the clearing as a feature — "setting a source drops any concurrent build the step was polling" — and the type comment on Tracker states the rule as "Start, StartStep and Finish also clear those fields, but as resets under mu alone — by the time they run the step has already passed through its fence." SetWorkSource clears the same fields under mu alone but is called during the step, before any fence has run. It is the one clearing path that does not satisfy the precondition the comment gives for the others.

Failure scenario (observation). A step runs a concurrent index build; the executor has handed the tracker its reserved session. An operator calls Progress(ctx). The observer takes pollMu, reads session, pid, releases mu, and is inside dbconn.ConcurrentIndexProgress on that reserved pgx connection. The executor's step moves to the copy and calls SetWorkSource(copier), which returns immediately. Reading its doc, the executor concludes the tracker no longer holds the build and returns the reserved session to the pool. The pool hands that connection to another caller while the observer's query is still in flight on it — concurrent use of a single pgx connection, which is precisely the hazard the pollMu comment names.

Failure scenario (cancel), the worse one. Same start, but the operator calls CancelBuild(ctx). It takes pollMu, reads pid, and is inside the detached pg_cancel_backend send. SetWorkSource runs concurrently — it does not take pollMu, so it is not serialized against CancelBuild the way StopConcurrentBuild is — and the executor releases the build's backend. The signal lands on a recycled backend running an unrelated statement. This is the exact outcome StopConcurrentBuild's comment says the fence exists to prevent, reachable through a method whose documentation reads as the newer way to drop a build.

The mirror holds for SetConcurrentBuild. The new t.source = nil line drops the source under mu alone, so an observation already inside source.Work(ctx) keeps running. StopWorkSource's comment states the contract the engine needs — "the engine calls it before the source's state goes away" — but SetConcurrentBuild now offers a way to drop the source that does not honour it, and an engine transitioning copy → concurrent build will reach for it for the same reason.

TestWorkSourceAndBuildMutuallyExclusive (both directions) currently pins the un-fenced clearing as the intended behaviour, so the suite would not catch this being wrong.

Fix, in order of preference:

  1. Give both Set methods the fence. Four lines, and the asymmetry disappears:
    func (t *Tracker) SetWorkSource(source WorkSource) {
        t.pollMu.Lock()
        defer t.pollMu.Unlock()
        t.mu.Lock()
        defer t.mu.Unlock()
        t.source, t.session, t.buildPID = source, nil, 0
    }
    The cost is that a step transition waits behind an in-flight observation — the same cost StopConcurrentBuild already pays, at the same once-per-step frequency. The Tracker comment's objection ("a reset that waited behind an observation would make polling a gate on execution") was written about Start/StartStep/Finish, which run after the fence; it does not apply to a mid-step handoff that has had no fence at all.
  2. If the un-fenced behaviour is deliberate, say so at both call sites: that the clearing is bookkeeping only, that the corresponding Stop* is still required before the dropped state is released, and that an engine must never treat SetWorkSource as a replacement for StopConcurrentBuild. Then add a test that the fence is still needed, so the requirement is pinned somewhere other than a comment.

Option 1 is what I would do — the contract currently has three methods that clear the build's fields and only one of them is safe to clear them with, which is a distinction the next caller will not preserve.

Work runs arbitrary engine code under pollMu, and StopWorkSource waits for it

Progress holds pollMu across source.Work(ctx), and StopWorkSource acquires pollMu. So the engine's teardown blocks until an observer's callback into the engine returns. The WorkSource doc explicitly anticipates a source that takes locks — "a source sees one poll at a time and may read state that is not safe for concurrent observation" — which makes the inversion concrete rather than theoretical:

  • Engine goroutine holds the copier's internal mutex and calls StopWorkSource; it blocks on pollMu.
  • Observer holds pollMu and is inside Work, blocking on the copier's internal mutex.

Deadlock, and neither participant's context breaks it, because pollMu and a sync.Mutex inside the source are not context-aware. The observer's context bounds nothing here: cancelling it does not release pollMu.

The same shape gives a milder problem even without a cycle. StopConcurrentBuild waits behind a database round-trip, which the server and the connection bound; StopWorkSource waits behind whatever the source does, on whatever context the observer supplied — context.Background() from a long-poll handler is enough to make "waits for an in-flight observation" unbounded.

No consumer lands in this PR, so this is a contract gap rather than a live bug, and the next PR in the stack is exactly where it becomes one. Two sentences on WorkSource would close it: that Work must not acquire a lock the engine may hold when it calls StopWorkSource, and that it must return promptly and honour ctx, because both the tracker's other observers and the engine's own teardown wait on it.

Notes

The doc comment's "engine-measured" framing is the useful part of the version bump. Splitting the counter families explicitly — blocks/tuples/lockers as server-observed, rows/bytes as engine-measured, neither fabricating the other's — gives a consumer a rule it can apply to operations that do not exist yet, which is more durable than the operation-by-operation table. Worth making sure the format_version 4 entry in the doc's version history says the counter families were split, not only that an operation was added, since that is the part a strict consumer's validation would key on.

ErrNoActiveBuild after SetWorkSource is a good deliberate choice and is easy to misread as a bug. A caller that cancels a build and gets ErrNoActiveBuild has no way to tell "the build finished" from "the step moved to a copy". That is correct — there is nothing to cancel either way — but it is the kind of thing an operator-facing caller renders as "no build running" when the truthful answer is "the step is copying now". If the engine ever surfaces that error to a human, the snapshot's operation is the field that disambiguates it.

Kiran01bm and others added 2 commits September 30, 2026 18:30
Replacing a poll target under mu alone let the engine release the build
session or the source state while an observation was still reading it;
both setters now drain in-flight polls like the Stop* fences do, and the
WorkSource contract requires Work to bound itself and hold no engine lock.
@Kiran01bm

Copy link
Copy Markdown
Collaborator Author

🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/131, follow-up commit

Verdict: all eight findings across the three adversarial comments are fixed in the follow-up commit — one code change (both Set* setters now take the pollMu fence) plus contract wording in the WorkSource doc, Tracker comment, docs/progress-report.md, and SAFETY.md; nothing deferred or rejected.

# Finding Status Explanation
C3-F1 SetWorkSource / SetConcurrentBuild clear the other path's fields under mu alone, without the pollMu fence, so a caller may release the build session or source while an observation or CancelBuild is in flight fixed Both setters take pollMu then mu (same order as the Stop* fences); TestSourceHandoffsDrainInFlightObservation covers StopWorkSource, SetWorkSource, and SetConcurrentBuild, and dropping pollMu from either setter fails that setter's subtest.
C3-F2 Work runs under pollMu and StopWorkSource waits on it — a deadlock if the engine calls Stop*/Set* while holding a lock Work needs fixed WorkSource doc now requires that Work take no lock the engine holds while calling SetWorkSource, SetConcurrentBuild, or StopWorkSource, return promptly, and honour ctx; the stacked copier meets this (Copier.report releases c.mu before the tracker calls).
C1-F1 Work must bound itself rather than rely on the observer's context, otherwise StopWorkSource waits unbounded fixed Same doc: Work is memory-only or a catalog read on a session with statement_timeout set, because the fences wait for it and an observer's context may never end; SAFETY.md no longer says the wait is "bounded by the poller's context", and docs/progress-report.md carries the rule for external implementers.
C2-F3 A catalog-reading WorkSource is a CO-9 read site and the contract should say so fixed WorkSource doc and docs/progress-report.md: a Work that reads the catalog pg_catalog-qualifies every relation, function, and operator and is tested under a shadowing search_path; the stacked copier's unqualified pg_class/pg_table_size reads are fixed in that PR.
C2-F2 work presence no longer means a build is running; consumers must select the counter family by detail.operation fixed docs/progress-report.md "Work counters" states the rule and the failure it prevents (a presence-keyed consumer renders a copy step as a zero-block build); the work field row links to it and the v4 history entry says work is no longer a build signal.
C2-F1 Package synopsis still says copy counters are future work fixed pkg/progress/progress.go synopsis describes the shipped state: one snapshot shape, build counters from the server view, copy counters from the engine's WorkSource, each leaving the other's at zero.
C3-F3 v4 history should say the counter families were split, not only that an operation was added fixed Version-history entry rewritten: v4 added the copy operation and split work into server-observed and engine-measured families selected by detail.operation.
C3-F4 ErrNoActiveBuild after SetWorkSource is correct but deserves a clause in the CancelBuild doc fixed CancelBuild doc: a step whose work a WorkSource owns has no build to signal and reports ErrNoActiveBuild; the snapshot's operation tells the caller which case applies.

No action: the confirmed items — Progress check order (terminal → source → build), the mutual-exclusion tests, error-path behaviour, and the format_version bump — are unchanged. The test (PostgreSQL 15) failure at 1f50216 was the cross-package event-trigger harness race fixed by #133, now in this branch via the main merge.

Source: block/pg-sprite#131, review comments 5886048966 and 5886050903 and review 5355697487 at head 1f50216; fixes in the follow-up commit.

@Kiran01bm
Kiran01bm merged commit c98d704 into main Sep 30, 2026
16 checks passed
@Kiran01bm
Kiran01bm deleted the kiran01bm/cs5c-progress-copy branch September 30, 2026 22:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants