Skip to content

schemachange: pair indexes across a retyped column (ST-5, D8) - #138

Merged
Kiran01bm merged 2 commits into
mainfrom
kiran01bm/cs9-pair-type-change
Oct 4, 2026
Merged

Kiran01bm merged 2 commits into
mainfrom
kiran01bm/cs9-pair-type-change

Conversation

@Kiran01bm

@Kiran01bm Kiran01bm commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Closes the known gap carved out of #137: the cutover gate's index pairing now follows a column across a type change.

Why

ALTER COLUMN id TYPE bigint re-creates every index on id for the new type. On PostgreSQL 14–18 the shadow's copy of the primary key and of any secondary index on id carries int8_ops in pg_index.indclass while the source's carries int4_ops, even though pg_get_constraintdef still prints PRIMARY KEY (id). #137 pairs indexes by exact catalog definition, operator classes included, so those indexes landed in UnpairedSource / UnpairedShadow and the swap would have left _pgsprite_<hash>_new_pkey as the live table's primary-key name.

What

  • pkg/schemachange/dependents_retype.go (new):
    • retypedColumns(source, target) — the columns both models carry under one name with a different canonical type. Recorded both bare and pgx.Identifier{...}.Sanitize()-quoted, since pg_get_indexdef quotes a key column only when it needs to.
    • indexDefinition.relaxedAcross(retyped) — sets aside only what the server re-derives on a retyped key column: the type's default operator class (pg_opclass.opcdefault) and the column's own collation (indcollation = attcollation). An operator class or COLLATE written in the index still has to agree. Expressions and predicates are never relaxed (the server may render either differently for the new type).
    • pairIndexes(source, shadow, retyped) — exact pass first, then a relaxed pass over the exact pass's leftovers only, merged in source order. A relaxed definition that more than one index on either side renders pairs none of them — the relaxation cannot show those indexes are interchangeable, so they stay unpaired rather than pairing by position. An index on an untouched column pairs exactly or not at all; a predicate, DESC/NULLS ordering, uniqueness, or access-method difference still leaves the pair unpaired.
    • readIndexes reads the two per-key facts the relaxation needs (opcdefault, and whether the key's collation is the column's own); they ride on indexDefinition outside the pairing key.
  • GateCutover passes retypedColumns(sourceModel, targetModel) into the pairing; indexDependents renders the relaxed definition so the pairing key and the unpaired report agree.
  • Docs: D8 in docs/copy-and-swap-design.md and ST-5 in docs/invariants.md record the one set-aside difference and the ambiguity rule; the pkg/schemachange package-map row says the same. docs/refusal-classes.md notes that a server error from applying the statement to the empty shadow (e.g. an operator class that does not accept the new type, SQLSTATE 42804) is a wrapped *pgconn.PgError routed by SQLSTATE, not a RefusalCause.

Tests

  • Unit (dependents_retype_test.go): retyped-column derivation incl. a mixed-case "Qty"; pairing across a retyped column keeps source order; an opclass difference on an untouched column stays unpaired; only the default opclass and own collation are relaxed (a predicate or indoption difference still fails; an explicit opclass or explicit collation is kept across the retype); an expression over a retyped column is not relaxed; the relaxed pass runs only over exact leftovers; a relaxed key shared by two source or two shadow indexes leaves all of them unpaired.
  • Integration (cutover_gate_retype_integration_test.go, PG16 testcontainers):
    • TestGateCutoverPairsIndexesAcrossARetypedColumn — accounts(id integer PRIMARY KEY, balance integer, label text) with accounts_id_desc_idx (id DESC) and accounts_balance_idx; gates ALTER TABLE accounts ALTER COLUMN id TYPE bigint; asserts the int4_ops vs int8_ops precondition on the two primary keys and that all three indexes pair with nothing left over.
    • TestGateCutoverRetypePassLeavesADroppedColumnsIndexUnpaired — ALTER COLUMN id TYPE bigint, DROP COLUMN balance; the id indexes pair, accounts_balance_idx stays in UnpairedSource.
    • TestGateCutoverPairsAnIndexAcrossARetypeThatAddsACollation — ALTER COLUMN balance TYPE text; the index gains a collation ("" → default) and text_ops, and still pairs.
    • TestGateCutoverRetypeKeepsAnExplicitCollationApart — accounts_label_plain (label) and accounts_label_c (label COLLATE "C"), retype label to char(20); each pairs with the shadow index carrying its own collation.
    • TestGateCutoverRetypeKeepsAnExplicitOpclassApart — (label) and (label bpchar_pattern_ops) on a char(20) column, retype to text COLLATE "C"; the explicit opclass is kept and the two indexes pair with their own counterparts.
  • A mutant that returns the definition unrelaxed fails both the unit pairing test and the first integration test; a mutant that blanks every opclass and collation fails the explicit-opclass/collation tests; a mutant without the ambiguity guard fails the shared-key tests.
  • SKIP_INTEGRATION=1 go test ./..., go test -race ./pkg/schemachange/, make lint (0 issues) all pass locally.

Stack

Based on #137 (kiran01bm/cs9-fidelity-gate), which is based on #136 (kiran01bm/cs6-repair-policy). Retarget to the parent's base as each merges. No capability, verdict, or CLI surface changes; demo/tour.sh is unaffected.

ALTER COLUMN ... TYPE re-creates every index on that column for the new
type, so pg_index.indclass on the shadow's copy flips to the new type's
default operator class (int4_ops -> int8_ops) while the index still covers
the same columns the same way. The gate's exact-definition pairing left
such indexes unpaired, and a swap would have kept the LIKE-derived name
(_pgsprite_<hash>_new_pkey) on the live table.

GateCutover now derives the set of columns the statement retyped (same
name on both sides, different canonical type) and runs a second pairing
pass over the exact pass's leftovers only, with the operator class and
collation of every retyped key column set aside. Access method,
uniqueness, columns, ordering, predicate, and backing constraint must
still agree; an expression over a retyped column and every untouched
column stay exact. Pairs keep the source order.

Unit tests cover the retyped-column derivation (including a quoted name),
the relaxed pass over leftovers, and that only opclass/collation are
relaxed. Integration tests gate an int -> bigint change on a table with a
PK, a DESC index on the retyped column, and an index on an untouched
column, asserting the int4_ops/int8_ops precondition and full pairing;
and that a dropped column's index still stays unpaired alongside a
retyped one.

Design D8 and invariant ST-5 record the one set-aside difference.
@Kiran01bm
Kiran01bm force-pushed the kiran01bm/cs9-pair-type-change branch from de2c3a0 to 5d6f315 Compare October 1, 2026 11:10
@Kiran01bm
Kiran01bm marked this pull request as ready for review October 1, 2026 11:12
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@aparajon

aparajon commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

🤖 1/2: adversarial correctness review of 5d6f315. I read retypedColumns, relaxedAcross, the two-pass pairIndexes, and the one-line gateCutoverTx change, together with readIndexes, pairByDefinition and confirmNamesFree. I checked them against ST-5, CO-1 and LK-1, and against D8. Probes and mutations ran on real PostgreSQL 16, using the PR's accounts fixture.

1 blocking, 3 non-blocking.

The two-pass shape is right. The exact pass runs first, the relaxed pass sees only its leftovers, and an index on a column the statement did not retype pairs exactly or not at all. The gate stays fail-closed: every refusal before the pairing is unchanged, and extra pairs only widen the constraint-name check. The blocking finding is in what the relaxed pass sets aside. It blanks every opclass and collation on a retyped key column, including the ones the user wrote explicitly. The server keeps those on a rebuild, so they still tell indexes apart.

Blocking

1. The relaxed pass blanks explicit collations, so two indexes that differ only in collation pair by name order and swap names (ST-5, D8). dependents_retype.go:44-60, copy-and-swap-design.md:210-211

D8 relies on one premise: two indexes with the same key are interchangeable, so any pairing between them restores an equivalent catalog. The relaxed key breaks it. relaxedAcross clears the opclass and the collation of every retyped key column, whether the server derived them or the user chose them. A rebuild re-derives only the implicit ones. An explicit COLLATE "C" is printed by pg_get_indexdef and comes back unchanged.

Here is the example I ran. accounts.label text has accounts_label_plain (label), created first, and accounts_label_c (label COLLATE "C"). The statement is ALTER COLUMN label TYPE char(20). The opclass moves from text_ops to bpchar_ops, so neither index pairs exactly. The shadow ends with _label_idx (collation default) and _label_idx1 (collation C). Once both collations are blanked, all four indexes render the same key, and pairByDefinition hands them out in name order:

source (by relname)                    shadow it pairs with at 5d6f315         with the fix
accounts_label_c      COLLATE "C"      _label_idx    COLLATE default  WRONG    _label_idx1  COLLATE "C"
accounts_label_plain  COLLATE default  _label_idx1   COLLATE "C"      WRONG    _label_idx   COLLATE default

The gate mints CutoverReady with both names on the wrong index. The same run with two unique indexes, one under a nondeterministic ICU collation (und-u-ks-level2), puts accounts_label_ci_key on the case-sensitive unique index. After the swap, a later DROP INDEX accounts_label_key removes the case-insensitive uniqueness the user meant to keep. On 19ddb06 (the merge base) all four stay in UnpairedSource / UnpairedShadow. That loses the names, but visibly. So this is a regression from a visible gap to a silent wrong pairing.

The fix is to relax only what the server re-derives. readIndexes reads two more per-key arrays: pg_opclass.opcdefault, and whether indcollation[k] equals the key column's attcollation. relaxedAcross then blanks an opclass only when it is the type's default, and a collation only when it is the column's own. With that change (about 30 lines across dependents.go, dependents_retype.go, and the btreeOn test helper), equal relaxed keys again imply equal definitions on each side, so the D8 premise holds. go test ./pkg/schemachange/ passes in full with it, and so do the three tests in this comment. The ST-5 Enforced text (invariants.md:551-553) and D8 would then say "the default operator class and the column's own collation", which is what they already describe in prose.

Test that fails on 5d6f315 and passes with the fix
// Two indexes on one column that differ only in an explicit collation stay
// distinct across a retype: the server keeps the explicit collation and
// re-derives only the implicit one, so each source index pairs with the
// shadow index that carries its own collation (ST-5, D8).
func TestGateCutoverRetypeKeepsAnExplicitCollationApart(t *testing.T) {
	f := newShadowFixture(t)
	f.createAccounts(t)
	f.exec(t, `CREATE INDEX accounts_label_plain ON %s.accounts (label)`)
	f.exec(t, `CREATE INDEX accounts_label_c ON %s.accounts (label COLLATE "C")`)
	s := f.stage(t, "accounts", `ALTER TABLE %s.accounts ALTER COLUMN label TYPE char(20)`)
	shadow := s.built.ShadowTable()

	ready, err := f.gate(t, s)
	require.NoError(t, err)

	pairs := ready.Indexes().Pairs
	assert.Contains(t, pairs, schemachange.DependentPair{Kind: schemachange.DependentIndex, SourceName: "accounts_label_plain", ShadowName: shadow + "_label_idx"})
	assert.Contains(t, pairs, schemachange.DependentPair{Kind: schemachange.DependentIndex, SourceName: "accounts_label_c", ShadowName: shadow + "_label_idx1"})
}

--- FAIL: TestGateCutoverRetypeKeepsAnExplicitCollationApart (0.58s) on 5d6f315. Both assertions fail, and the pairs show accounts_label_c ↔ …_label_idx and accounts_label_plain ↔ …_label_idx1. It passes with the fix.

Non-blocking

1. The collation half of the relaxation is untested. Mutant M7 keeps every collation, and every test still passes. The PR's integration tests retype integer → bigint, which has no collation on either side. A retype that adds a collation (integer → text: "" becomes default, and int4_ops becomes text_ops) pairs only because the collation is blanked. It also pins the fix above, which must keep blanking the implicit collation.

Test that passes on 5d6f315 and kills M7
// Retyping an integer column to text gives its index a collation as well as
// a new operator class; both are set aside, so the index still pairs (D8).
func TestGateCutoverPairsAnIndexAcrossARetypeThatAddsACollation(t *testing.T) {
	f := newShadowFixture(t)
	f.createAccounts(t)
	s := f.stage(t, "accounts", `ALTER TABLE %s.accounts ALTER COLUMN balance TYPE text`)
	shadow := s.built.ShadowTable()

	ready, err := f.gate(t, s)
	require.NoError(t, err)

	assert.Contains(t, ready.Indexes().Pairs, schemachange.DependentPair{Kind: schemachange.DependentIndex, SourceName: "accounts_balance_idx", ShadowName: shadow + "_balance_idx"})
}

It passes on 5d6f315 and with the fix. --- FAIL: TestGateCutoverPairsAnIndexAcrossARetypeThatAddsACollation (0.11s) under M7.

2. TestPairIndexesRelaxedPassUsesOnlyExactLeftovers never reaches the relaxed pass. dependents_retype_test.go:155-170, dependents_retype.go:75

The exact pass pairs the only shadow index, so exact.UnpairedShadow is empty and pairIndexes returns at line 75. The test then holds whatever the second pass does. Running the relaxed pass over every index instead of the leftovers (M11), or letting indexesNamed keep everything (M17), both survive. Either mutant would give one shadow index two source partners. One unrelated shadow index makes the pass run:

Test that passes on 5d6f315 and kills M11 and M17
// The relaxed pass sees only what the exact pass left on both sides: with a
// leftover on each side it runs, and still cannot hand the shadow index the
// exact pass already took to a second source index.
func TestPairIndexesRelaxedPassRunsOnlyOverExactLeftovers(t *testing.T) {
	source := []indexEntry{
		{name: "orders_id_idx", definition: btreeOn("pg_catalog.int4_ops", "id")},
		{name: "orders_id_idx1", definition: btreeOn("pg_catalog.int8_ops", "id")},
	}
	shadow := []indexEntry{
		{name: "_new_id_idx", definition: btreeOn("pg_catalog.int8_ops", "id")},
		{name: "_new_qty_idx", definition: btreeOn("pg_catalog.int4_ops", "qty")},
	}

	got, err := pairIndexes(source, shadow, map[string]bool{"id": true})
	require.NoError(t, err)

	assert.Equal(t, DependentPairing{
		Pairs:          []DependentPair{{Kind: DependentIndex, SourceName: "orders_id_idx1", ShadowName: "_new_id_idx"}},
		UnpairedSource: []string{"orders_id_idx"},
		UnpairedShadow: []string{"_new_qty_idx"},
	}, got)
}

It passes on 5d6f315. --- FAIL: TestPairIndexesRelaxedPassRunsOnlyOverExactLeftovers (0.00s) under M11 and under M17.

3. A predicate over a retyped column stays unpaired, and the gate still passes. The design text names only expressions as held to an exact match, but the server re-renders predicates too. accounts (id) WHERE label <> '' reads back on the shadow as ((label)::text <> ''::text) after label becomes char(20). In the run that retyped id to bigint and label to char(20), every other index paired, but this one landed in UnpairedSource and its copy in UnpairedShadow as _pgsprite_…_new_id_idx3. That is the conservative outcome. But DependentPairing still calls it "not a fault", so the swap would keep the shadow-derived name without saying so. Two small steps would close this. Name predicates next to expressions in D8 and in the relaxedAcross comment. And consider having the gate refuse, or at least flag, an unpaired source index whose key and predicate columns all still exist on the shadow. That is the one case the user did not ask to drop.

Verified

  • ST-5 extends: GateCutover is still where the pairing is enforced, and the edited Enforced text names the right file. Blocking 1 is where the code sets aside more than that text says the server re-creates.
  • The gate stays fail-closed. The only gate change is the third argument at cutover_gate.go:152. Every refusal that runs before it is unchanged. That covers the lock, both OIDs, invalid indexes, fingerprints, fidelity and identities. The relaxed pass can only add pairs. The _old name set is Pairs + UnpairedSource, so it is the same either way, and constraintNamesTaken checks more pairs, never fewer.
  • CO-1 and LK-1 uphold: checkCutoverProofs, requireTableLock and the in-transaction confirmTableLock are untouched.
  • On PostgreSQL 16, across integer → bigint: a multi-column index (balance, id DESC NULLS LAST), a partial index WHERE id > 100, INCLUDE (id), a hash index, UNIQUE (id, balance), and a btree_gist EXCLUDE (balance WITH =) each pair with their shadow copy. An expression over the retyped column stays unpaired (the PR's unit test).
  • An index whose opclass the new type does not accept never reaches the gate. With USING brin (balance int4_minmax_multi_ops), ALTER COLUMN balance TYPE bigint fails BuildShadow with SQLSTATE 42804.
  • retypedColumns records each name both bare and Sanitize()-quoted. That matches quote_identifier, which pg_get_indexdef uses, including for keywords and for names with embedded quotes. A bare entry cannot collide with another key rendering. Dropped columns are excluded, since IntrospectTx filters attisdropped.
Mutant Caught by
M1 gate passes no retyped columns both new integration tests
M2 omit the quoted name TestRetypedColumnsNamesOnlyColumnsWhoseTypeChanged
M3 omit the bare name both new integration tests, TestRetypedColumns…
M4 a dropped column counts as retyped TestRetypedColumns…
M5 every shared column counts as retyped TestRetypedColumns…
M6 keep the opclass both new integration tests, TestPairIndexesPairsAcrossARetypedColumn
M7 keep the collation survives: killed by non-blocking 1's test
M8 relax in place, no slices.Clone equivalent today: nothing reads the entries' opclasses after pairing
M9 relax every key column 4 tests, incl. TestPairIndexesKeepsAnOpclassDifferenceOnAnUntouchedColumnUnpaired
M10 drop the len(retyped) == 0 early return equivalent: an empty relaxation is the exact pass
M11 relaxed pass over every index survives: killed by non-blocking 2's test
M12, M13 report the exact pass's unpaired lists both new integration tests, TestPairIndexesPairsAcrossARetypedColumn
M14 drop the exact pairs from the merge TestGateCutoverPairsIndexesAcrossARetypedColumn, TestPairIndexesPairsAcrossARetypedColumn
M15 merge in pass order, not source order TestPairIndexesPairsAcrossARetypedColumn
M16 indexDependents ignores the relaxation both new integration tests, TestPairIndexesPairsAcrossARetypedColumn
M17 indexesNamed keeps everything survives: killed by non-blocking 2's test

go build ./... and go vet ./pkg/schemachange/ pass on 5d6f315. So does go test ./pkg/schemachange/ (full package, PostgreSQL 16).

This review was generated by Claude Code (claude-opus-5-5).

@aparajon

aparajon commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

🤖 2/2: OSS adoption and integration ease, at 5d6f315. These are lenses, not correctness findings. 0 blocking, 3 non-blocking.

This closes the most visible gap #137 left. A widened key is the most common type change, and keeping accounts_pkey across it is what an operator expects without thinking about it. The change also adds no public API. DependentPairing keeps its shape, so an importer that already renders the pairing gets the improvement for free.

1. Say which pairs the relaxed pass made. An operator reviewing a cutover preview should see the difference between "this index is unchanged" and "this index is re-created for bigint and keeps its name". Today a DependentPair reads the same either way. A field on the pair would let SchemaBot render the second case as its own line, and it would make a wrong relaxed pairing (1/2, blocking 1) visible in review instead of after the swap. The field could be a Basis of exact / retyped, or the list of columns whose opclass or collation was set aside.

2. Give the remaining unpaired cases a reason. This narrows the unpaired set, so what is left is more likely to surprise someone. That covers an expression or a predicate over a retyped column (1/2, non-blocking 3), and a column the statement dropped. The suggestion from #137 still applies here: a reason on each unpaired entry would let an importer say "this index goes away" only when the statement dropped its column. Today the cases look the same as a pairing the engine could not make.

3. Document how a build fails when the new type rejects an index. A retype whose index opclass does not accept the new type, such as BRIN int4_minmax_multi_ops on a column widened to bigint, fails BuildShadow before any gate runs. The error is the wrapped server error (SQLSTATE 42804), not a *RefusalError, so RefusalCauseOf returns "". shadow.go:246 wraps it with %w, so an importer can reach the *pgconn.PgError with errors.As and route on its code. But refusal-classes.md does not say that build-time server errors pass through this way. One line there, or a cause for "the statement failed on the empty shadow", would save every importer from finding out from a failed run.

This review was generated by Claude Code (claude-opus-5-5).

@aparajon aparajon left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approving 5d6f315 with 1 blocking finding: the relaxed pass blanks explicit collations as well as derived ones. So two indexes on a retyped column that differ only in COLLATE pair by name order, and each ends up with the other's name (ST-5, D8). The fix is to set aside only the default opclass and the column's own collation. The 1/2 comment has the blocking finding with its test and a fix that keeps the package green, plus three non-blocking findings. Two of them come with tests: the collation half of the relaxation is unpinned, and the leftovers-only test never reaches the relaxed pass. The 2/2 comment has three non-blocking integration notes.

This stamp was left by Claude Code (claude-opus-5-5).

@morgo morgo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Automated adversarial review, posted on Morgan Tocker's behalf.

Approving. Running the relaxation as a second pass over only the first pass's leftovers is the right structure — it keeps the exact match authoritative, so an index that already has a perfect counterpart can never be stolen by a looser one, and it bounds the blast radius of the relaxation to the indexes that actually failed. Setting aside the operator class and collation rather than the whole key column is also the narrow version of the fix: everything that decides which rows the index covers still has to agree.

Verified rather than assumed:

  • The merge cannot drop an index out of both unpaired lists. merged.UnpairedSource/UnpairedShadow are taken wholesale from the relaxed pass, which looks like it would lose the exact pass's leftovers — but the relaxed pass is fed exactly exact.UnpairedSource/UnpairedShadow, so every exact leftover ends up either in partner or in the relaxed pass's own unpaired list. Worth checking because sourceNames() is Pairs plus UnpairedSource, and that is what confirmNamesFree checks the old-name space against; a name silently dropped here would be a derived name the gate never confirms is free, and the swap would fail under the cutover lock instead of refusing before it.
  • The early return is consistent with the merge. len(exact.UnpairedShadow) == 0 returns exact whole, which carries its own UnpairedSource; the merge path only replaces those lists when the relaxed pass actually ran over them.
  • relaxedAcross's shallow copy does not reach back into the caller's entries. relaxed := d shares the backing arrays of KeyColumns, Included and Options, but only the freshly cloned Opclasses and Collations are written. Since pairIndexes renders the same []indexEntry slice twice, an in-place clear would have rewritten the inputs between passes.
  • A constraint-backed index cannot cross-pair with a plain one. Constraint stays in the relaxed key, so accounts_pkey in the new integration test pairs through its own constraint definition rather than competing with accounts_balance_idx, even though both key on the retyped column.
  • The bare-and-sanitized keying matches the catalog form. pg_get_indexdef(…, k, true) prints the column name quoted only where quoting is required, and pgx.Identifier{}.Sanitize() always quotes, so entering both covers a case-sensitive name without a separate unquoting step.

Worth fixing before merge

The relaxed pass pairs by a key that is no longer unique, and pairByDefinition's "identical definitions are interchangeable" justification does not survive that. pairByDefinition takes candidates[0] when several dependents share a key. That is sound under exact matching — the doc's reasoning is that two source indexes with the same definition really are the same index twice, so an arbitrary pairing restores an equivalent catalog. After relaxedAcross the key is deliberately lossy, and two indexes that differ only in the thing it erased now collide.

Concretely, on a column the statement retypes:

CREATE INDEX accounts_code_idx         ON accounts (code);
CREATE INDEX accounts_code_pattern_idx ON accounts (code bpchar_pattern_ops);
ALTER TABLE accounts ALTER COLUMN code TYPE text;

Both source indexes relax to the same key, both shadow indexes relax to the same key, and the pairing between them is positional. If the two lists do not happen to agree in order — the source is ordered by ic.relname, the shadow by whatever names LIKE-built index creation produced — the swap puts accounts_code_idx on the pattern-ops index and accounts_code_pattern_idx on the default one. Every check the gate makes still passes, because the set of names is unchanged; the catalog afterwards has two of the operator's own indexes wearing each other's names. The same shape arises from the collation half (CREATE INDEX … (name) alongside CREATE INDEX … (name COLLATE "C")), which is if anything the more common pair, since both are the standard ways to make LIKE indexable.

TestPairIndexesRelaxedPassUsesOnlyExactLeftovers covers the adjacent hazard — a relaxed candidate stealing an index the exact pass already spoke for — but not this one, where both candidates are genuinely leftovers and the ambiguity is created by the relaxation itself. The cheap fix is to refuse to relax an ambiguous key: if more than one source leftover or more than one shadow leftover shares a relaxed key, leave all of them unpaired rather than guessing. That keeps the current outcome for every case the PR is aimed at and degrades to today's behaviour for the ambiguous ones, which is the conservative direction for a fidelity gate.

Notes

An index whose key columns are untouched can still be knocked out of pairing by a retype, and nothing in the relaxation can reach it. Predicate is pg_get_expr(indpred, …), and the deparser prints a constant with an explicit cast for every type other than the handful it treats as self-evident — int4 among them, int8 not. So CREATE INDEX ON accounts (created_at) WHERE balance > 0, with balance retyped int4 → int8, reads (balance > 0) on the source and (balance > '0'::bigint) on the shadow. The key column is created_at, which the statement did not touch, so relaxedAcross clears nothing and the index stays unpaired and keeps its generated shadow name after cutover.

The doc change sets aside the expression case explicitly — "an expression over such a column … is held to an exact match" — and I read that as deliberate. The predicate case is covered by the same sentence only implicitly ("everything else about the two definitions must still agree"), and it is the one that will surprise, because the index that loses its name does not look like an index on a retyped column at all. Either name predicates alongside expressions in that sentence, or say in pairIndexes' comment which parts of the definition the relaxation can and cannot reach.

The relaxation is keyed on schemadiff.Model type strings, which makes the gate's behaviour depend on how canonical those are. retypedColumns flags a column when targetType != c.Type as plain strings. A spelling difference that is not a type difference — int8 against bigint, varchar(64) against character varying(64) — relaxes a column that was never retyped, and relaxing a column that did not change can only loosen the gate: two indexes on that column differing genuinely in operator class would then pair. That is the same ambiguity as above arriving by a different route, so if the model's types are already canonicalised, a one-line assertion of that in retypedColumns' comment is worth more than it costs; if they are not, the ambiguity guard above handles both.

indexesNamed is quadratic, which is fine here and worth a word. slices.Contains over the leftover names, once per index, is nothing at catalog index counts — but it reads like an oversight next to the map the rest of the file uses. A one-line note that the leftover lists are small by construction (they are the residue of an exact pass) would stop the next reader from "fixing" it.

@Kiran01bm

Copy link
Copy Markdown
Collaborator Author

🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/138, follow-up commit

The relaxed pass now sets aside only what the server re-derives on a retyped key column (default operator class, column's own collation), a relaxed definition shared by more than one index on either side pairs none of them, and the docs name predicates, canonical types, and build-time server errors. The blocking finding and the "worth fixing before merge" item are fixed; the two integration-ease suggestions that need a new field on the pairing are tracked as internal follow-ups.

# Finding Status Explanation
C1-B1 Relaxed pass blanks explicit collations (and opclasses), so indexes that differ only in COLLATE pair by name order and swap names fixed readIndexes now reads pg_opclass.opcdefault per key and whether indcollation[k] equals the key column's attcollation; relaxedAcross blanks an opclass only when it is the type's default and a collation only when it is the column's own. A written opclass or COLLATE stays in the relaxed key. New unit tests TestPairIndexesKeepsAnExplicitOpclassAcrossARetype / …ExplicitCollationAcrossARetype, and PG tests TestGateCutoverRetypeKeepsAnExplicitCollationApart (your label / label COLLATE "C" → char(20) run) and TestGateCutoverRetypeKeepsAnExplicitOpclassApart (label / label bpchar_pattern_ops → text COLLATE "C"); each source index pairs with the shadow index carrying its own collation or opclass. D8 and ST-5 now say "the default operator class and the column's own collation".
C1-N1 Collation half of the relaxation untested fixed TestGateCutoverPairsAnIndexAcrossARetypeThatAddsACollation retypes balance integer → text, so the index gains both a collation and text_ops and pairs only because the own collation is set aside. A mutant that keeps every collation fails it.
C1-N2 TestPairIndexesRelaxedPassUsesOnlyExactLeftovers never reaches the relaxed pass fixed Replaced by TestPairIndexesRelaxedPassRunsOnlyOverExactLeftovers: two source indexes on the retyped id (int4_ops and int8_ops) against one shadow id index (int8_ops) and an unrelated qty index. The exact pass takes the int8_ops pair; the int4_ops source index would also match that shadow index once relaxed, but the relaxed pass sees only the leftovers, so it stays in UnpairedSource and the qty index in UnpairedShadow. A mutant that runs the relaxed pass over the full lists fails it.
C1-N3 Predicate over a retyped column stays unpaired silently; consider refusing or flagging an unpaired source index whose columns all still exist fixed / deferred Fixed: D8 and the relaxedAcross comment now name predicates alongside expressions as held to an exact match. Deferred: refusing or flagging an unpaired source index whose key and predicate columns all still exist on the shadow is a change to what DependentPairing reports, not to the pairing itself; tracked as an internal follow-up together with the per-entry reasons in C2-2.
C2-1 Say which pairs the relaxed pass made (Basis on DependentPair) deferred Agreed. DependentPair keeps its shape in this PR; a Basis (exact / retyped, or the columns whose opclass or collation was set aside) is tracked as an internal follow-up with the unpaired-reasons work so the preview renders both in one change.
C2-2 Give the remaining unpaired cases a reason deferred Already tracked from the previous PR's round as an internal follow-up (dropped by the statement vs no counterpart vs left unpaired by the relaxed rule); this round adds the "all columns still exist" case from C1-N3 to it.
C2-3 Document that build-time server errors pass through as wrapped *pgconn.PgError, not a RefusalCause fixed docs/refusal-classes.md gains a paragraph after the cutover rows: a statement the server rejects on the empty shadow (an opclass that does not accept the new type, SQLSTATE 42804) surfaces as the server's *pgconn.PgError wrapped with the shadow's name, reachable with errors.As, with RefusalCauseOf returning the empty cause; importers route it by SQLSTATE. No new cause is minted, since the server would reject the same statement on the source.
R1 Restates C1-B1 and C1-N1/N2 no action Covered by the C1 rows above.
R2-1 Relaxed key is no longer unique, so candidates[0] in pairByDefinition is no longer justified fixed pairIndexes computes the relaxed keys that more than one leftover index on either side renders and excludes every index rendering one from the relaxed pass; those stay in UnpairedSource / UnpairedShadow in input order. New tests TestPairIndexesLeavesSourceIndexesThatRelaxAlikeUnpaired and …ShadowIndexesThatRelaxAlikeUnpaired. With C1-B1 narrowing the relaxation, the guard is the backstop for a written opclass that happens to be the default for another type. D8 records the rule.
R2-N1 A predicate over a retyped column knocks an index on untouched key columns out of pairing fixed Documented in D8 and the relaxedAcross comment (same change as C1-N3); the index stays unpaired rather than pairing loosely.
R2-N2 Relaxation keyed on schemadiff.Model type strings depends on their canonicality fixed Verified: the model renders every column type with format_type, so int8 and bigint cannot differ as strings; retypedColumns' comment now asserts this. The R2-1 guard covers the other route.
R2-N3 indexesNamed is quadratic fixed Comment added: the leftover lists are the residue of the exact pass and small by construction, so the linear scan is deliberate.

Verified sections of both reviews (merge cannot drop an index from both unpaired lists, early return consistent with the merge, shallow copy does not reach the caller's entries, constraint-backed indexes cannot cross-pair, bare-and-sanitized keying) need no action; the merge path is unchanged and the new ambiguity exclusion feeds the same UnpairedSource / UnpairedShadow lists.

Source: block/pg-sprite#138, review comments 5931555831 and 5931557266, reviews 5379260375 and 5383581722 at head 5d6f315; fixes in the follow-up commit.

The relaxed pairing rule for a key column the gated statement retyped
set aside the operator class and collation outright, so an index with
a written operator class or COLLATE clause paired with a shadow index
that lost it. Narrow the rule to what the server re-derives: the
default operator class (pg_opclass.opcdefault) and the column's own
collation (indcollation = attcollation). A written opclass or collation
still has to agree.

Leave a relaxed definition that two indexes on either side share
unpaired rather than pairing by position: identical indexes are
interchangeable only when their full definitions match, and the
relaxation cannot show that.

Document that expressions and predicates over a retyped column stay
exact, that schemadiff types are canonical (format_type), and that a
server error from applying the statement to the empty shadow is a
wrapped *pgconn.PgError routed by SQLSTATE, not a RefusalCause.

Tests: explicit opclass and explicit collation kept apart across a
retype (unit and PG), an added collation still pairs, the relaxed pass
runs only over exact leftovers, and shared relaxed keys on either side
stay unpaired.
@Kiran01bm
Kiran01bm enabled auto-merge (squash) October 4, 2026 22:16
@Kiran01bm
Kiran01bm merged commit 7ffb964 into main Oct 4, 2026
16 checks passed
@Kiran01bm
Kiran01bm deleted the kiran01bm/cs9-pair-type-change branch October 4, 2026 22:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants