Skip to content

checksum: digest chunks with SHA-256 instead of md5 so a FIPS-mode OpenSSL does not fail the pass - #143

Merged
Kiran01bm merged 4 commits into
mainfrom
kiran01bm/cs6-digest-sha256
Oct 5, 2026
Merged

Kiran01bm merged 4 commits into
mainfrom
kiran01bm/cs6-digest-sha256

Conversation

@Kiran01bm

@Kiran01bm Kiran01bm commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Switches the checksum verifier's chunk digest from md5 to SHA-256, so a verification pass does not fail on a host whose OpenSSL runs in FIPS mode.

Why

PostgreSQL built against OpenSSL routes md5() through it, and an OpenSSL in FIPS mode refuses MD5. On such a host every chunk digest errors, so the verifier fails closed on every pass — correct behaviour for a fault, but the fault is the hash choice, not the data. The #134 review flagged it (lens 2/2, finding 3).

The digest never leaves the process: a Report is in memory and nothing persists it, so changing the hash has no compatibility surface — no checkpoint, no stored fingerprint, no JSON contract carries it.

What

  • pkg/checksum digestSQL: each row's record text is turned into bytes with convert_to(…, pg_catalog.getdatabaseencoding()) — the database's own encoding as the target, so no conversion happens and the bytes hashed are the ones the server stores, as md5(text) hashed them — and hashed with pg_catalog.sha256; the per-row hashes are aggregated as bytes (string_agg(bytea, ''::bytea ORDER BY pk)), hashed again, and only the chunk's hash is encoded as hex. Every function stays pg_catalog-qualified (CO-9). Both sides still run the identical frozen statement; only the table differs.

    SELECT pg_catalog.count(*),
           pg_catalog.encode(pg_catalog.sha256(COALESCE(pg_catalog.string_agg(
             pg_catalog.sha256(pg_catalog.convert_to(ROW("id"::bigint, "qty"::numeric(10,2))::text, pg_catalog.getdatabaseencoding())),
             ''::bytea ORDER BY "id"), ''::bytea)), 'hex')
    FROM "schema"."table"
    WHERE "id" BETWEEN $1::bigint AND $2::bigint
  • Digest.Hash doc: hex SHA-256; the SHA-256 of the empty input for an empty range.

  • Tests: TestDigestSQLIsFrozen re-pinned (TM-2); TestVerifierIgnoresTheSessionSearchPath decoys are sha256(bytea), convert_to(text, name), getdatabaseencoding() and encode(bytea, text) in place of md5(text); TestDigestHashesARowThatIsNotValidUTF8 digests a non-UTF-8 byte on a SQL_ASCII database (new testutil.NewDatabaseWithEncoding); TestVerifierReportsAChunkTheShadowIsMissingEntirely pins the empty side's zero rows and empty-input hash.

  • Test fixtures that generated filler with md5() (testutil.WorkloadTable, the preflight index-bytes test, the concurrent-index progress test, the Supabase realtime probe) now use sha256(), with every fixture's byte sizes unchanged, so the suite itself runs on a FIPS-mode server via PG_DSN. No md5() call remains under the repo.

  • Docs: D7 in copy-and-swap-design.md states the SHA-256 digest and the FIPS reason; the pkg/checksum package-map row, the CO-9 decoy list in invariants.md, and the checksum idiom row in mysql-vs-postgresql.md follow.

Decisions to veto

  • Bytes, not hex, between the two hash levels. The per-row hashes are aggregated as bytea and only the final digest is hex. Hex at both levels would also work; bytes halve the aggregate's input and keep one encode at the end.
  • convert_to(…, getdatabaseencoding()) rather than 'UTF8' or a text→bytea cast. There is no such cast; convert_to is the catalog's way. Targeting the database's own encoding performs no conversion, so a SQL_ASCII database holding bytes that are valid in no encoding digests like any other instead of failing every pass with 22021. On a UTF-8 database the two spellings are the same statement.
  • No FIPS-specific test. CI cannot run a FIPS-mode OpenSSL; what it proves is that the statement is frozen, both sides agree, differences are still found, and the functions resolve in pg_catalog. With the fixtures off md5(), an adopter can run make test-db against a FIPS host to check the claim. sha256() is available from PostgreSQL 11 and runs on every supported major (14 → 18).

Release note

Checksum digests are now SHA-256 hex (64 characters); Mismatch.Source.Hash and Mismatch.Shadow.Hash were md5 hex (32 characters). Nothing persists a digest, so no stored value changes meaning.

Review round 1

Addressed in the follow-up commit: getdatabaseencoding() as the convert_to target with the SQL_ASCII test and the CO-9 decoy; the empty-chunk test; the two doc spellings (sha256 over convert_to(…), ''::bytea separator); fixtures off md5(); the release-note line above.

Verification

go test ./pkg/checksum/ ./internal/testutil/ (PG16), the SQL_ASCII and empty-chunk tests on PG18, each new test shown to fail against its mutant ('UTF8' target → 22021; COALESCE dropped → NULL scan), the preflight index-bytes and concurrent-index progress tests on the new fixtures, SKIP_INTEGRATION=1 go test ./..., make lint (0 issues). The Supabase realtime probe runs in CI's Supabase job.

Kiran01bm and others added 3 commits October 3, 2026 18:25
…enSSL does not fail the pass

PostgreSQL built against OpenSSL routes md5() through it, and an OpenSSL
in FIPS mode refuses MD5, so every chunk digest — and with it every
verification pass — failed on such a host. The digest statement now hashes
each row's record text as UTF-8 bytes with pg_catalog.sha256, aggregates
the per-row hashes as bytes in key order, hashes the aggregate again, and
renders only the chunk's hash as hex. The digest never leaves the process
(a Report is in memory; nothing persists it), so the hash has no
compatibility surface. Both sides still run the identical frozen statement
and every function stays pg_catalog-qualified (CO-9).

Tests: the frozen digest statement is re-pinned; the shadowing-search_path
test's decoys are sha256, convert_to and encode in place of md5.

Docs: D7 states the SHA-256 digest and why; the package map, the CO-9
decoy list and the MySQL-vs-PostgreSQL idiom table follow.
@Kiran01bm
Kiran01bm marked this pull request as ready for review October 5, 2026 01:59
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@aparajon

aparajon commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

🤖 1/2: adversarial correctness review of 9f4752f. I read digestSQL, the Digest contract, every caller in pkg/checksum (verify, check, repair reread), and the two changed tests, against CO-1, CO-2, CO-9 and TM-2. I ran the package on real PostgreSQL 16, mutated each part of the new statement, and probed the encoding path on databases created with other server encodings.

0 blocking, 3 non-blocking.

The swap is sound. Both sides still run one statement that differs only in the table it reads, so the hash function, the UTF-8 conversion, the ORDER BY key and the empty-range COALESCE are applied identically to source and shadow, and nothing but the data can make them differ. The aggregate is at least as strong as before. Each row contributes a fixed-width 32-byte hash in key order, the row text includes the key, and the row count is compared separately. A missing, extra, changed or moved row still changes the digest, and nothing can be read two ways at a concatenation boundary. sha256(bytea) is in core from PostgreSQL 11, and the 14 to 18 CI matrix is green at this head. No production path calls md5() any more. The remaining hashtext calls key advisory locks and use PostgreSQL's own hash, not OpenSSL. Nothing persists a Digest: CleanWatermark and VerifiedShadow carry watermarks and OIDs, and the checkpoint row stores no digest. So a resume across versions never compares an md5 digest with a SHA-256 one.

Non-blocking

1. Converting to UTF-8 adds a failure the md5 digest did not have: on a SQL_ASCII database, a row holding a non-UTF-8 byte fails every pass over its chunk. digest.go:107, digest.go:86-88

md5(text) hashed the record text's bytes in the server encoding as they were. convert_to(…, 'UTF8') converts them first, and that conversion validates. A SQL_ASCII database stores whatever bytes a client sent, so a row containing 0xff digests fine on main and errors at this head with invalid byte sequence for encoding "UTF8": 0xff (SQLSTATE 22021). The pass fails closed, so CO-1 holds and no data is at risk. But the change trades one "every pass fails on this kind of database" for another, which is the very failure the PR sets out to remove. The user also sees an error that looks like corrupt data in their table, not a limit of the verifier.

The stated reason for the conversion does not hold either. client_encoding applies only on the wire. It never reaches a server-side function's input, so the md5 digest never depended on it.

Converting to the database's own encoding keeps everything the PR wants and drops the failure. convert_to(x, pg_catalog.getdatabaseencoding()) performs no conversion. It hashes the bytes md5(text) hashed and is identical to 'UTF8' on a UTF-8 database. I made that one-word change: the whole package passes, and so does the test below. CO-9's decoy set would gain a getdatabaseencoding() impostor. The comment at L86-L88 would say the record is hashed as the server stores it.

Test that fails on 9f4752f, and passes on main and with the fix
// A SQL_ASCII database stores whatever bytes a client sends, including
// bytes that are not valid UTF-8. The digest hashes each row's rendering in
// the server's own encoding, so such a row digests like any other instead
// of failing every pass over its chunk.
func TestDigestHashesARowThatIsNotValidUTF8(t *testing.T) {
	server := testutil.StartPostgres(t)
	admin, err := pgxpool.New(t.Context(), server)
	require.NoError(t, err)
	t.Cleanup(admin.Close)
	name := fmt.Sprintf("sql_ascii_%d", os.Getpid())
	_, err = admin.Exec(t.Context(), "CREATE DATABASE "+name+" ENCODING 'SQL_ASCII' LC_COLLATE 'C' LC_CTYPE 'C' TEMPLATE template0")
	require.NoError(t, err)
	t.Cleanup(func() {
		_, err := admin.Exec(context.WithoutCancel(t.Context()), "DROP DATABASE IF EXISTS "+name+" WITH (FORCE)")
		assert.NoError(t, err)
	})
	u, err := url.Parse(server)
	require.NoError(t, err)
	u.Path = "/" + name
	pool, err := dbconn.NewPool(t.Context(), dbconn.Config{URL: u.String()})
	require.NoError(t, err)
	t.Cleanup(pool.Close)

	f := proofFixture{pool: pool, schema: testutil.NewSchema(t, pool)}
	f.exec(t, `
		CREATE TABLE %s.orders (
			id bigint PRIMARY KEY,
			note text NOT NULL
		)`)
	f.exec(t, `INSERT INTO %s.orders VALUES (1, pg_catalog.convert_from('\xff'::bytea, 'SQL_ASCII'))`)
	target := f.prove(t, "orders")
	sql := digestSQL(target, f.schema, "orders", []columnType{
		{name: "id", typeName: "bigint"},
		{name: "note", typeName: "text"},
	})

	tx, err := pool.Begin(t.Context())
	require.NoError(t, err)
	t.Cleanup(func() { _ = tx.Rollback(context.WithoutCancel(t.Context())) })
	d, err := digest(t.Context(), tx, sql, 1, 1)
	require.NoError(t, err)
	assert.Equal(t, int64(1), d.Rows)
}

--- FAIL: TestDigestHashesARowThatIsNotValidUTF8 (0.07s) on 9f4752f with ERROR: invalid byte sequence for encoding "UTF8": 0xff (SQLSTATE 22021). It passes on main (1f46fb9) and with 'UTF8' replaced by pg_catalog.getdatabaseencoding().

I also tried a database in the one server encoding that has no conversion to UTF-8 at all, where this head would fail on every row. A UTF-8 client cannot connect to such a database in the first place, so SQL_ASCII is the only case that is reachable.

2. The empty side of a chunk is documented but untested: dropping the COALESCE survives the package. digest.go:19-21, digest.go:109

Digest.Hash promises "the SHA-256 of the empty input for an empty range". Without the COALESCE, a side with no rows in the chunk yields NULL, the scan fails, and the pass errors where it should report a mismatch. No test reaches a chunk that one side holds no rows of, so the mutant survives at this head and on main alike. The gap is older than this PR, but the PR rewrote both that expression and its doc line, and one test pins them:

Test that passes on 9f4752f and fails with the COALESCE removed
// A chunk the shadow holds no rows of is a mismatch, not a failed pass: the
// empty side digests as zero rows and the SHA-256 of the empty input, so the
// report names the chunk and the policy decides what happens next.
func TestVerifierReportsAChunkTheShadowIsMissingEntirely(t *testing.T) {
	f := newVerifierFixture(t)
	target, lock, shadow := f.prepare(t)
	f.exec(t, "DELETE FROM "+f.shadowName(shadow)+" WHERE id > 2000")

	report, err := f.verify(t, f.pool, target, shadow, lock, copier.NewWatermark(math.MaxInt64))
	require.NoError(t, err)
	require.Len(t, report.Mismatches, 1)
	empty := report.Mismatches[0]
	assert.Equal(t, chunk(t, 2001, math.MaxInt64), empty.Chunk)
	assert.Equal(t, int64(500), empty.Source.Rows)
	assert.Equal(t, int64(0), empty.Shadow.Rows)
	assert.Equal(t, "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", empty.Shadow.Hash)
}

It passes on 9f4752f. With the COALESCE removed: --- FAIL: TestVerifierReportsAChunkTheShadowIsMissingEntirely (0.52s), digest chunk [2001, 9223372036854775807] of shadow …: can't scan into dest[1] (col: encode): cannot scan NULL into *string.

3. Two doc spellings of the new digest do not run as written. mysql-vs-postgresql.md:175 gives sha256(row::text), but sha256 takes bytea and there is no implicit cast from text: ERROR: function sha256(text) does not exist. The idiom needs the convert_to(row::text, …) the code uses. copy-and-swap-design.md:191 keeps the text separator '' in string_agg(row_hash, '' ORDER BY pk). The rows are bytea now, so it is ''::bytea.

Verified

  • Head passes. All 37 tests in pkg/checksum pass against PostgreSQL 16. CI's 14, 15, 16, 17 and 18 jobs are green at this SHA.
  • CO-1 is upheld. The gate still compares every chunk and fails closed. A mutant that hashes only the key column, and a mutant that orders the shadow side differently from the source, are each caught by 13 or more behavioral tests.
  • CO-9 is upheld, and its Enforced line names the new decoys. As before with md5, the decoys sit behind the transaction's LocalSearchPath("pg_catalog"), so they cannot catch a dropped pg_catalog. qualifier. Only TestDigestSQLIsFrozen does, which is the two-layer design the entry describes.
  • TM-2 is upheld. The frozen string is re-pinned, and the behavioral tests are unchanged and pass.
  • CO-2 and CO-3 are not touched. No persisted state carries a digest.
  • Performance. A digest of one million 4-column rows on PostgreSQL 16 (arm64) took about 805 ms with md5 and about 723 ms with this head's statement. The getdatabaseencoding() variant took about 735 ms. On a CPU without SHA instructions, SHA-256 would cost more per byte than md5. I did not measure that case, and the row-to-text rendering dominates either way.
  • Mutations at digest.go, each restored with git checkout. Every one fails TestDigestSQLIsFrozen. The column below is the behavioral tests, with the frozen test excluded:
Mutant Caught by (behavioral)
Row hash over the key column only 13 tests, e.g. TestVerifierLocatesAChangedShadowRow, TestCheckAbortsOnDivergenceAndLeavesTheShadowAlone
Shadow side aggregated in the opposite key order 16 tests, e.g. TestVerifierReportsACompleteCopyClean
pg_catalog. dropped from sha256 / convert_to / encode survives (pinned by the local search_path; frozen test only)
COALESCE dropped survives, here and on main (finding 2)
ORDER BY dropped survives, here and on main (heap order equals key order in the fixtures; fails closed if it diverges)
'UTF8' → 'LATIN1' survives (no fixture holds non-ASCII text; finding 1's test is the first)
Inner per-row hash dropped survives (equivalent: record text is self-delimiting)
Outer hash dropped survives (equivalent: the bytes are compared directly)

This review was generated by Claude Code (claude-opus-5-5).

@aparajon

aparajon commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

🤖 2/2: OSS adoption and integration ease, at 9f4752f. These are lenses, not correctness findings. 0 blocking, 2 non-blocking.

This is a good adoption change. Regulated users often run PostgreSQL on hosts with OpenSSL in FIPS mode. Before this PR, every verification pass failed for them, with an error that pointed at the database rather than at the engine's choice of hash. The change costs importers nothing: the digest stays inside pkg/checksum, the verdict and plan JSON are untouched, and on the host I measured the new statement is no slower. 1/2's finding 1 is the one thing I would fix before release. Without it, the PR swaps FIPS hosts for SQL_ASCII databases holding non-UTF-8 bytes, and those users get an error that reads like corruption in their own table.

1. The test suite still generates data with md5(), so nobody can run it on a FIPS-mode server to check the claim. testutil/workload.go, lines 67 and 242

The PR is right that CI cannot run a FIPS-mode OpenSSL. An adopter who wants to confirm FIPS support on their own host would point PG_DSN at it and run make test-db, though. The shared workload fixture fills blob with repeat(md5(g::text), 300) and rewrites it with md5(random()::text), so the fixture fails before any product code runs. A few other fixtures do the same (preflight_integration_test.go:90, cic_build_integration_test.go:130, the Supabase realtime test). Switching them to encode(sha256(convert_to(g::text, 'UTF8')), 'hex'), or to repeat/lpad filler, makes the suite itself FIPS-clean, and then "runs on a FIPS host" is something an adopter can check rather than take on trust. This could be a follow-up, but it fits in this PR.

2. Digest.Hash is exported, and its format changes from 32 to 64 hex characters. digest.go:19-22

"No compatibility surface" holds inside the repo. For importers, though, Mismatch.Source.Hash and Mismatch.Shadow.Hash are the public fields an orchestrator would render when it shows an operator which chunk differed, and both are carried in DivergenceError. SchemaBot pins a release that predates this change and does not import pkg/checksum yet, so nothing breaks today. One line in the release notes ("checksum digests are now SHA-256 hex, 64 characters") is enough, so that an importer storing or column-sizing the value is not surprised.

This review was generated by Claude Code (claude-opus-5-5).

@aparajon aparajon left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approving 9f4752f with 0 blocking findings. Source and shadow still run one identical statement, so the hash swap cannot hide a difference or invent one. The gate still fails closed (CO-1 upheld), every function stays pg_catalog-qualified with the decoys updated (CO-9 upheld), and no persisted state carries a digest, so no md5 digest is ever compared with a SHA-256 one. I checked this on real PostgreSQL and with mutations.

The 1/2 comment has 3 non-blocking findings:

  • convert_to(…, 'UTF8') makes every pass fail on a SQL_ASCII database holding a non-UTF-8 byte, where md5 worked. Using pg_catalog.getdatabaseencoding() as the target fixes it, and a test that ran is attached.
  • The empty-range COALESCE is documented but untested. A test that ran is attached.
  • Two doc spellings of the digest (sha256(row::text), and a text '' separator) do not run as written.

The 2/2 comment has 2 non-blocking adoption notes: test fixtures still call md5(), and the exported Digest.Hash width change deserves a release-note line.

This stamp was left by Claude Code (claude-opus-5-5).

sha256 takes bytes, and convert_to with 'UTF8' as the target validated the
record text as UTF-8 on the way to them. A SQL_ASCII database stores
whatever bytes a client sends, so a row holding a byte that is valid in
no encoding failed every pass over its chunk with 22021. Converting to
pg_catalog.getdatabaseencoding() instead performs no conversion and hashes
the bytes as the server stores them, which is what md5(text) hashed; on a
UTF-8 database the statement is unchanged. The CO-9 decoy set gains an
impostor getdatabaseencoding, and a test on a SQL_ASCII database pins the
non-UTF-8 row.

Pin the empty side of a chunk: a chunk the shadow holds no rows of reports
zero rows and the SHA-256 of the empty input, not a scan of NULL.

Correct the two doc spellings of the digest that did not run as written:
sha256 takes bytea, so the idiom needs convert_to, and the aggregate's
separator is ''::bytea.

Move the test fixtures that generated filler with md5() to sha256(), with
every fixture's byte sizes unchanged, so the suite itself runs on a
FIPS-mode server and an adopter can check the claim with PG_DSN.

Release note: checksum digests are SHA-256 hex, 64 characters, where
Mismatch.Source.Hash and Mismatch.Shadow.Hash were 32.
@Kiran01bm

Copy link
Copy Markdown
Collaborator Author

🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/143, follow-up commit

All five non-blocking findings are fixed in this PR; nothing is deferred. Both attached tests were taken as written (the SQL_ASCII one now goes through a testutil.NewDatabaseWithEncoding helper), each was shown to fail against its mutant, and go test ./pkg/checksum/ ./internal/testutil/ is green on PG 16 with the two new tests also run on PG 18.

# Finding Status Explanation
C1-1 convert_to(…, 'UTF8') fails every pass on a SQL_ASCII database holding a non-UTF-8 byte (22021), where md5(text) worked; the client_encoding rationale does not hold fixed Taken as proposed: the target is pg_catalog.getdatabaseencoding(), which performs no conversion and hashes the bytes as the server stores them. digestSQL's comment and the D7 / mysql-vs-postgresql.md idiom say so; the client_encoding claim is gone from the comment and the PR body. TestDigestHashesARowThatIsNotValidUTF8 is added (fails at the reviewed head with 22021, passes now) and TestVerifierIgnoresTheSessionSearchPath gains an impostor getdatabaseencoding() that names no encoding, listed in CO-9's Enforced line.
C1-2 The empty side of a chunk is documented but untested; dropping the COALESCE survives the package fixed TestVerifierReportsAChunkTheShadowIsMissingEntirely added as written: 500 source rows, 0 shadow rows, shadow hash e3b0c4…b855. Verified the mutant: with the COALESCE removed it fails on cannot scan NULL into *string.
C1-3 Two doc spellings of the digest do not run as written (sha256(row::text); text '' separator) fixed mysql-vs-postgresql.md now reads sha256(convert_to(row::text, getdatabaseencoding())); D7 in copy-and-swap-design.md uses ''::bytea and the same row_hash spelling.
C2-1 The test suite still generates filler with md5(), so an adopter cannot run it on a FIPS-mode server to check the claim fixed All five fixtures moved to encode(sha256(convert_to(…, 'UTF8')), 'hex') with byte sizes unchanged (repeat(…, 150) for the 9600-byte workload blob, repeat(…, 2) for the 128-byte CIC payload, left(…, 32) where a 32-character value was used). No md5( call remains under the repo; make test-db with PG_DSN at a FIPS host now exercises only product code against it.
C2-2 Digest.Hash is exported and its width changes from 32 to 64 hex characters fixed A "Release note" section in the PR body states it (SHA-256 hex, 64 characters; Mismatch.Source.Hash / Mismatch.Shadow.Hash were 32; nothing persists a digest), and the follow-up commit's body carries the same line so it travels with the squash.

Verified section (CO-1, CO-9, TM-2, CO-2/CO-3 upheld; mutation table; performance): no action. Of the mutants that survived for want of a fixture, COALESCE dropped is now caught by the empty-chunk test, and a convert_to target that validates (the reviewed head's 'UTF8') by the SQL_ASCII test.

Source: block/pg-sprite#143, review comments 5987177880 and 5987178438 and review 5409528259 at head 9f4752f; fixes in the follow-up commit.

@Kiran01bm
Kiran01bm enabled auto-merge (squash) October 5, 2026 03:38
@Kiran01bm
Kiran01bm merged commit d5309b9 into main Oct 5, 2026
16 checks passed
@Kiran01bm
Kiran01bm deleted the kiran01bm/cs6-digest-sha256 branch October 5, 2026 03:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants