Skip to content

Release V2.0 - #532

Open
eigmax wants to merge 113 commits into
mainfrom
feat/upgrade-plonky3
Open

eigmax wants to merge 113 commits into
mainfrom
feat/upgrade-plonky3

Conversation

@eigmax

@eigmax eigmax commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Core features:

  1. JIT MIPS emulator
  2. CPU-less chip arch
  3. Jaggled reduction for chip area
  4. WHIR PCS

Proximity Test
UDR >=100-bit security by default

Proof size: < 600KiB

@eigmax
eigmax force-pushed the feat/upgrade-plonky3 branch from b1de559 to c6c37f8 Compare September 21, 2026 00:56
Squash of feat/upgrade-plonky3 (1690 commits) merged with main at 89a2b74.
@eigmax
eigmax force-pushed the feat/upgrade-plonky3 branch from c6c37f8 to a773261 Compare September 21, 2026 00:57
Main removed it in 671e4ce; this branch keeps parity. The syscall
BOOLEAN_CIRCUIT_GARBLE (0x00_01_00_31), its executor event and syscall
implementation, the BooleanCircuitGarble and BooleanCircuitGarbleControl
chips (MipsAirId 50 and 52) with their split threshold and shape-enumeration
cluster, the guest-side zkm-zkvm/zkm-lib bindings, the Lean models and the
docs entries are gone; the AIR cost artifact loses the two keys.

The remaining discriminants are unchanged, so no other chip's id moves.
Gated on the GPU pair against block 25907955 with the regenerated
zerocheck kernels.
zkm-primitives gained blake3, tracing and the serial_test dev-dependency
with the imm-wrap-vk mode (#523), and zkm-sdk dropped alloy-signer. The
release squash carried the pre-merge lock; this is the one every build on
the box resolved to, so `--locked` builds agree with it.
…public-value copies to agree

The State bus carries the two-pc state `(pc, next_pc)`; its entry endpoint is
the public-value pair `(start_pc, start_next_pc)` and its exit endpoint
`(next_pc, next_next_pc)`. The verifiers chained `start_pc` across shards and
nothing else, so row 0 of every shard -- which takes its `next_pc` from that
endpoint -- continued at an address of the prover's choosing: execute the real
instruction at the chained `start_pc`, then jump anywhere. The executor never
splits a branch from its delay slot, which is what made the old Cpu chip's
local `next_pc = pc + 4` rule sufficient; the frame redesign moved the state
onto the bus and lost that rule.

No shard boundary falls inside a delay slot, so an execution shard enters and
exits with the sequential lookahead. The normalize program and the host
verifier now require `start_next_pc = start_pc + 4` and `next_next_pc =
next_pc + 4` on execution shards -- the halting row included, whose `next_pc`
is 0 and whose lookahead the frame sends as `next_pc + 4` -- and equal-to-pc
on non-execution shards, whose endpoints cancel on the bus. With `start_pc`
chained this gives `start_next_pc = prev.next_next_pc` without a new public
value.

Writing the test exposed a second gap: a shard proof carries its public values
twice, on `ShardProof` (read by the machine-level checks) and inside
`JaggedShardProof` (what the constraints are evaluated against and what the
recursion witness is built from). Both are observed into the transcript, which
binds each to the proof but not to the other, so a proof built with a
divergent pair passed the pc chain on one copy and the constraints on the
other. `verify_shard` now requires them equal.

`two_pc_lookahead_is_pinned_per_shard` forges each lookahead against the host
verifier and the normalize program, forges the machine-level copy alone, and
runs the honest fibonacci proof as the positive control for the halting row.

This changes the normalize program, so every leaf vk moves: the vk map,
vk_root and the outer artifacts are regenerated by the campaign that follows.
…s endpoints cancel

The per-shard lookahead pin (c01952b) asserted `start_next_pc = start_pc`
and `next_next_pc = next_pc` on shards without instruction chips. That is
not an invariant of the honest prover: a deferred precompile shard takes its
public values from the running prover state, `start_pc = next_pc = P` but
`start_next_pc = next_next_pc = P + 4` (the previous execution shard's
lookahead), whereas the executor's own finalization gives an empty shard
`(P, P)`. Every reth block has such shards; fibonacci has none, which is why
the positive control passed and production did not: 288 leaf failures, all
`DivFAssert 4/0`.

On a shard with no row on the State bus the two endpoints only have to
cancel against each other, so the rule there is `start_next_pc =
next_next_pc` and nothing more; the next execution shard re-derives its
lookahead from the chained `start_pc`. The execution-shard pins are
unchanged -- no production shard starts in a delay slot, because the native
producer defers a branch's fence check to its delay slot exactly as the
interpreter does.
The release squash carried hunks that were never run through rustfmt, and
the Rustfmt job on PR #532 fails on 46 files. Formatting only; no guest
crate is touched, so the production ELF is unchanged.
The normalize and compose programs seeded their running public values with
`MaybeUninit::zeroed()` Felts. A Felt is a variable handle, so a zeroed one is
variable 0, and every placeholder aliased it until its first real assignment;
the DSL cannot see the alias. Each placeholder is now a fresh `builder.uninit()`,
as the few already were. This allocates variable ids, so the recursion programs
and their vks move.
The memory argument orders accesses by (shard, timestamp) and assumes both
comparands are bounded -- shards below 2^16, timestamps below 2^TIMESTAMP_BITS
-- and that every value limb is a byte. A local entry is where a shard's view
of an address begins and ends, and its initial side was a free witness: nothing
bounded it, so a prover could open an address's chain at a shard, timestamp or
value the ordering proof was never made for.

Each entry now carries a 16-bit limb per shard and a (16, TIMESTAMP_HIGH_LIMB_BITS)
split per clk, range-checked through the byte and range tables with the same
primitive the instruction frames use, and both values are byte-checked. The
CUDA trace kernel fills the same columns through the shared header, so device
and host traces stay in parity; the compiled zerocheck kernels are regenerated
with the pair. This changes the MemoryLocal AIR, so the core vk moves.
Three tests still assumed `ZKMProofShape::generate` emits normalize
shapes. It does not: a normalize key is a function of the shard's chip
name set, which the block decides, so those keys are collected from real
proofs and only compose, deferred and shrink are enumerated
(`generate_emits_no_normalize_shapes` pins that).

- `generate_uses_stacked_shapes_for_recursion` and
  `arity1_classes_cover_full_sweep` measured the dropped enumeration
  against itself; removed.
- `normalize_dummy_core_vk_chip_information_is_faithful` hardcoded the
  preprocessed set as {Program, Byte}; the machine now carries Range as
  well. The test derives the set from the machine and asserts the two
  facts the dummy's faithfulness rests on: every preprocessed chip is
  included in every shard, and the dummy VK built from a shape carrying
  them reports exactly that set.
…ize program

Three changes moved the leaf vks: the per-shard 2-pc lookahead pin (a97ad83),
fresh variables for the public-value placeholders (6e86644), and the
local-memory range checks that change the MemoryLocal AIR (7d1c985). The map
is the 29 enumerated compose/deferred/shrink keys plus 325 leaf keys collected
from production and the two cached reth blocks under the new programs; 354
entries in all. The root is written from it. Gated with VERIFY_VK=true on
blocks 25907955 and 25955640 before deployment.
…script

prove_shard_with_data, commit_traces, prove_shard_with_data_boxed and the
normalize test witness: the doc comments state what is observed and drawn,
in order, and what each commitment binds (C_main = compress(root, H(n,
h, w)) on the inner ring, the root on the outer); no comments sit between
arguments. Comments only: the code is token-identical.
…osing compose; one allowlist tree height

The stacked-WHIR open combines every stripe of both committed rounds with
the powers of one challenge, so its batching error is (t - 1) * L / |F|,
94 bits at t = 256 and the binding term of the schedule. The prover now
grinds `WhirConfig::batch_pow_bits` = 8 bits between the absorbed stripe
claims and the challenge, the host and recursive verifiers check the
witness, and the term sits at 102 bits: the schedule is at 100 bits in the
unique-decoding regime. The transcript profile moves with it; the map is
regenerated in the next commit.

The compose that closes the tree (is_complete = 1) is snapped to its own
rows below arity REDUCE_BATCH_SIZE and so has its own key, which the
enumeration never produced: `ZKMProofShape::CompressRoot` is emitted next
to every compose tuple (43 shapes, 35 distinct keys; the arity-3 pair
coincides). `wrap_committed_stripes_are_pinned` pins the wrap shard's
committed geometry, 8 preprocessed and 16 main stripes.

The standalone verifier carried its own copy of the allowlist tree height,
12, while the prover pads the key list to 2^14, so the root it recomputed
matched no proof (vk_root mismatch on every input). `VK_MERKLE_TREE_HEIGHT`
is now one definition in zkm-recursion-core, re-exported by both.
Every `ZR-nn` tag and milestone label in a comment is replaced by the
property it stood for: the chord denominator (x2 - x1)^{-1}, the
preceding-root bind (the preprocessed round's root is the key's), the
opening cross-bind, coverage of the flat layout, the stacking height as
part of the protocol. Two comments that sat between arguments move out;
one orphaned test doc stacked on another test is dropped.

Comments only, except one assert message in the global accumulation test
helper, which now states the constraint instead of an audit id.
…e binds

Argument notes move into the function docs, line-number citations become
symbol names, the prefix-sum and bit-width reasoning is written as the
identities it rests on, and the history of how each piece came to be is
dropped. Comments only.
Argument notes move into the function docs and call-site notes above the
calls; the cross-bind and the padding count are written as the identities
they enforce; wire-format history and line citations are dropped. Comments
only.
Argument notes move into the function docs and above the calls; the
public-values balance, the cross-bind, the row-count walk and the GKR
packing are written as identities; line-number citations and the history
of each check are dropped. Comments only.
The module and function docs describe the stage as it is; the orientation
of the branching-program inputs, the embedding factor and the chunked sum
are written as identities; porting notes, line citations and the argument
comments are removed. Comments only.
The seeding, the evaluator's column count, the fixed stacking height and
the program-wide key are stated as the invariants they are; the history of
the clamp and of the per-shard key, porting notes, line citations and the
argument comments are removed. Comments only.
The raw-height, fixed-shape and height-forgery tests are described by the
property each asserts; the history of the restructure they guarded, commit
citations and the retired flags are removed. Comments only.
Argument notes move into the function docs; the interaction-axis size, the
Merkle path length and the opening shapes are stated as the formulas the
program is built from; line citations and measured traces are removed.
Comments only.
…ipt and checks

The per-round operations, the reconstruction and the position of the
opening observation are stated as formulas and orderings; line citations,
the pv_challenge history and legacy/FIX-off wording are removed. Comments
only.
Every eprintln! in host code becomes tracing::info! or, for a failure,
rejection or mismatch, tracing::warn!; the two eprint! uses become single
lines. Command-line tools, examples and benches that printed this way now
install a subscriber: `setup_cli_logger` (new, in zkm-core-machine) logs
to stderr at info unless RUST_LOG says otherwise, so their reports still
show by default and stdout stays free for data.

Left as they are: the guest crates (a change there rebuilds the guest ELF)
and zkm-build's relay of a child compiler's stderr, which is output of
that process, not a report of ours.
…ng grind

35 enumerated keys plus 355 leaf keys collected from production and the two
cached reth blocks under the new transcript, 390 entries; the root and the
recorded profile digest are rewritten. Gated on blocks 25907955 and
25955640 with VERIFY_VK=true on the server AND an independent host
verification of the dumped compressed proof against this map; a proof
made before the map is rejected, as is a one-byte tamper.
The machine has 10 precompile families, so 6 + 2*10 = 26 clusters and
26 * 12 * 12 * 5 * 5 = 93,600 shapes; the pinned 28 and 100,800 counted a
family the machine no longer has.
Every non-doc comment inside a function signature (parameters, generics,
where clause), a function body, or between a function's doc and `fn` is
removed.  Where one carried a soundness or safety reason, that reason is a
sentence (or a `# Safety` / `# Panics` section) in the function's doc.
Doc and item comments drop narrative history, file:line citations and
stacked summaries in favour of the claim each item proves; orphan docs of
removed fields are dropped.

Code is token-identical to the parent up to rustfmt reflow, except where a
branch held only a comment: empty `else {}` blocks are removed, and the
jagged verifier's main-round and padding branches (both `row_count = h`)
are one `else`.
`validate_jagged_rounds` checks, before any transcript step, that each
round's commit, prover data and WHIR data agree on area and stacking height,
that claims and row points match the packed chips and their widths, that
every area is a positive whole number of stripes covering its packed cells,
and that the opening area is next_pow2 of the summed areas.  The device
prover calls it so both backends reject the same malformed state early.

The prove_trusted_evaluations doc now states the implemented protocol:
one fresh z_col shared by both rounds, weights chi_k(z_col)*chi_j(z_row),
zero claims for zero-row chips, and h_i = 1 for a missing height.
make_basefold_merkle_proofs returns VkNotAllowed (digest and map size, same
message text as before) when vk_verification is on and a child vk is not in
the map; the deferred-input and first-layer builders, compress, shrink and
wrap_bn254 propagate it.  compress drains every proof still in flight before
returning the error, so no worker blocks on a bounded channel.  One disallowed
vk now fails the proof being built instead of killing the prover process.
`actions/checkout` cleans the workspace, and `git clean -ffdx` takes `target/`
with it, so each run on the self-hosted runner was a cold build of everything.
Every file under the runner's target directory carried the job's own start
time. Both self-hosted jobs now keep the cache outside the checkout, where
cleaning cannot reach it.

The time-consuming job also ran its first command with no RUSTFLAGS and the
rest with `-C target-cpu=native`, which invalidates every artifact the first
command produced. Both jobs now set one value for the whole job, which also
lets them share the cache with each other instead of rebuilding it in turn.

The time-consuming workflow had no concurrency group, so a superseded push
queued behind the run it replaced on a runner with one slot. One recent pull
request job waited two hours and sixteen minutes to start.

Measured before this change: the pull-request test job 37 min, the
time-consuming job 1 h 54 m, run one after the other. Build dominates both;
once warm, most packages finish in seconds.
@eigmax
eigmax force-pushed the feat/upgrade-plonky3 branch from f68bde0 to ddc1ba6 Compare September 24, 2026 07:05
Job 107495508244 failed with `process didn't exit successfully ... (signal:
15, SIGTERM)` and no failing test. The runner's out-of-memory daemon sent it:

  earlyoom: mem avail: 12171 of 126759 MiB (9.60%), swap free: 0
  earlyoom: sending SIGTERM to "zkm_sdk-d9aeff3": badness 1330, VmRSS 74964 MiB

The regenerated v2.0.0 circuit is a third larger than the one it replaced,
31,896,387 constraints against 23,630,444, with an 11.26 GB proving key and a
4.64 GB circuit. Two Groth16 cases overlapped at two test threads and together
crossed the limit.

Those cases are single-process memory hogs and gain nothing from parallelism:
each spends about three minutes reading its R1CS and proving key before it
proves anything.
The previous attempt, running the suite single-threaded, was wrong. The peak is
cumulative across cases inside one process rather than concurrent: every case
that drives gnark loads a multi-gigabyte R1CS and proving key, and the process
does not return that memory when the case ends. Measured on the same binary:

  whole suite, one process, single-threaded   102 GiB
  one heavy case in its own process            52 GiB

No thread count makes the first fit a 126 GB runner whose out-of-memory daemon
fires at 10% free, which is why two runs died with no failing test. Running the
five circuit-loading cases one process each bounds the peak at one case, and
the remaining cases still share a process.

Verified before pushing: each of the five filters reports "running 1 test" and
passes, so none is silently matching nothing, and the remainder runs its 22.
Serialising the remainder was left over from the previous attempt and is not
needed once the five circuit-loading cases run in their own processes. Measured
on the same binary: the remainder passes at two threads with a 30 GiB peak in
102 s, against 150 s single-threaded, so the flag only cost time.

The five heavy cases keep one process each, which is what bounds the peak; each
matches a single test, so their thread count is immaterial.
test_e2e_prove_plonk peaks at 83.4 GiB, and it peaks there with
GOMEMLIMIT=40GiB and GOGC=50 set: the PLONK constraint system has 43.2M
constraints, so the solver's witness is live data, and a soft heap limit
cannot bound a live set. The stale "52 GiB per case" figure is replaced
by the measured one.

The comment now also names the mitigation that does work on this runner.
earlyoom sends SIGTERM only when memory and swap are both under 10%, and
this box's 24 GiB swap sat permanently full, so half the trigger was
always satisfied; the box now carries a second 64 GiB swapfile. Two runs
that would previously have been killed, at 67 GiB and 83.4 GiB, both
completed with no earlyoom event.

Comment only: the commands are unchanged.
The schedule targeted 100 bits per transcript component, which is the
minimum-component convention: with roughly two dozen components, a union bound
over them lands near 97.7, so a configuration that read "100 bits" per round
was a sub-100-bit argument once its components were summed.  Every round's
query count is now solved for a per-component target that makes the union
clear 100, the target and every grinding difficulty are env-backed accessors
rather than literals, and the digest pins the defaults so a proof records the
schedule it was produced under.

The profiles now sum to 100.70 bits for a core shard proof, 100.76 for a
recursion node and 100.83 for the outer wrap.  Two bounds came out of doing
the accounting properly rather than by assumption:

  * the wrap grind stops at 26 rather than going higher, because its
    proof-of-work witness is one KoalaBear element: only ORDER ~ 2^31 indices
    exist, so a 30-bit grind would leave two expected witnesses and find none
    about one open in seven;
  * no target lifts the inner rings much past 101 bits at all, because the
    fifteen folding steps of a WHIR round are field-limited.  That ceiling is
    now written down where the target is set.

WHAT IT COST, AND WHY IT NO LONGER DOES.  The retarget made one reth block
take 166 s instead of 34 s, all of it proof of work, and 123 s of that was the
LogUp-GKR grind alone.  Two defects, both here: `deterministic_grind` handed
rayon's `find_first` the whole prime field when a witness is expected every
2^bits indices, and the WHIR query, folding and stripe-batching grinds called
p3's `challenger.grind` directly, which consults no accelerator registry.  The
search now runs over growing windows, returning the same smallest witness, and
those grinds go through `accelerated_grind`, which uses a registered
accelerator when there is one and is otherwise exactly what it replaced.  With
the paired ziren-gpu change the same block is 33.4 s against 33.9 s for the
old schedule.

The transcript profile digest moves to 6a40de2d, so every recursion verifying
key moves with it.  The map is regenerated by enumeration over the declared
shard shapes, 5 h 44 m, 127 of about 7,000 shapes skipped for fitting no pin
class; the key count is unchanged at 6,382, which is the check that matters,
and the allowlist root moves 004cb00d -> 0082e14c.

Also here: the soundness census becomes a repository binary that emits the
schedule beside the machine figures, so the published soundness model can be
diffed against the code instead of transcribed from it, and the paper is
carried in the repository with its figures and built PDF.
Pull-request feedback took four hours, of which about eighty-five percent was
queueing.  There is one self-hosted slot, and the time-consuming workflow held
it for 3 h 39 m per push, so the pull-request test job started three and a half
hours after the clippy job it was queued behind:

  Time-consuming / Cargo Test   15:13:24 -> 18:52:50   3 h 39 m
  ci.yml / Cargo Test           18:52:53 -> 19:24:39      32 m

The two cannot share the runner: a Groth16 case peaks near 73 GB of its 123 GB.
They also cannot be merged, for the reason already recorded in the workflow --
the gnark cases never return their R1CS and proving key, so one process per case
is what caps the peak.

So the split is by cost rather than by branch.  The end-to-end tests that
rotted unnoticed, which is why this workflow runs on the feature branch at all,
still run on every push.  The five Groth16 and PLONK proofs move to a nightly
schedule, main, and manual dispatch; at one completed run in the last twenty,
per-push runs of them bought nothing to lose.

The pull-request test job now makes one nextest pass over the twelve packages
instead of twelve sequential cargo invocations.  `cargo test` runs one test
binary at a time, so a single slow test left most of the machine idle; measured
on the four heavy packages, 830 s sequentially against 495 s in one pass, 346
tests, same result.  nextest installs itself if absent so a fresh runner does
not fail on a missing tool.

Both jobs now print each phase's duration and the memory left behind it, so the
next slow run says where the time went rather than only that it took it.
…al path

One test walked all 770 conformance vectors at about 0.59 s each, so it took
453 s, and libtest cannot parallelise inside a test: it was the floor for the
whole suite no matter how the packages were scheduled. Eight tests now each
check the vectors whose index is congruent to their shard, which both runners
can spread across threads:

  nextest, 8 threads     453 s -> 59.3 s   shards within 1.6 s of each other
  plain cargo test       453 s -> 121 s
  whole package          191 s -> 114 s

Sharding by index rather than by mnemonic keeps the split independent of what
the file contains: adding vectors redistributes them, where a hand-written list
of mnemonics would leave a new one unchecked.

Two properties the single test carried implicitly become their own tests,
because a shard cannot assert them from inside: that every vector is checked by
exactly one shard, and that the file's mnemonics are all reached. A vector no
test touches is worse than a failing one, so that check is worth its 11 ms.

A failure still names the vector, its mnemonic and now its shard.
The test defaulted to /data/stephen/cannon-mips/..., a path on one developer's
machine, and returned early with an info log when that path was absent. So it
reported success in CI while executing nothing, and it does so for a result the
paper states: that the 49 Cannon programs applicable to a little-endian guest
all pass. A green test backing a published claim that never ran is worse than
no test.

It now requires CANNON_MIPS_TESTS, fails rather than skips when the directory is
missing or unreadable, and is #[ignore]d by default because the programs are not
vendored here. CI runs it explicitly, on the runner where the suite is present,
so the claim is exercised where it can be seen. It takes 1.6 s.

The suite stays because it is the only oracle in the tree that we did not write:
the spec vectors come from our own ISA document and generator, while these are
another team's reading of the architecture manual.
The programs are not ours and are not vendored here, so the test could only run
against a checkout someone happened to have: it defaulted to a path on one
developer's machine and returned green when that path was absent, which is why
it has never run in CI. Rather than carry a test that depends on an external
working copy, the tree keeps the one vector source it owns -- the 770
specification vectors, assembled from the ISA document and checked against
Unicorn as the independent oracle.

The Cannon conformance result stands as a measurement that was taken, and the
paper reports it as such; nothing in this commit changes what was observed, only
where the check lives.
…oved

The grinding schedule this branch settles moves the wrap verifying key, so the
Groth16 circuit that pins it, the PLONK circuit and the imm-wrap-vk variant were
all regenerated, each with three fresh phase-2 contributions. The embedded keys
were still the previous ceremony's.

In standard mode nothing noticed: the vk hash committed into the Groth16 public
inputs is vk.hash_bn254() alone, which does not involve these bytes. In
imm-wrap-vk mode it is hash_vkey_with_part_vk over the embedded part_stark_vk,
so a stale copy made the host compute a hash the proof does not carry, and
test_groth16_public_values_imm_wrap_vk failed with "the verifying key does not
match the inner groth16 bn254 proof's committed verifying key". That test is the
only thing in the tree that reads part_stark_vk, which is why the mismatch
survived a passing standard suite.

The history entry is updated alongside, since release.sh keeps one per version
and test_get_part_stark_vk asserts the v2.0.0 entry equals what is embedded.

Verified on the GPU box against the installed v2.0.0 artifacts: the three
imm-wrap-vk tests pass where one failed before, test_groth16_public_values still
passes, and test_get_part_stark_vk resolves v2.0.0 to the same bytes.
The Contributions paragraph opened by saying what the paper is not, and the
related-work paragraph on Ceno repeated the move. Both now say what is claimed:
a machine, and of the commitment layer the composition, its parameter schedule
and its recursive verifier.

Also carries the section, figure and bibliography edits from this round of
review, and the rebuilt PDF (55 pages, no undefined references).
The self-installing step wrote into `~/.cargo/bin`, which does not exist on the
self-hosted runner: its CARGO_HOME is /data/.cargo. tar therefore failed with
"Cannot open: No such file or directory" and took the whole job down with exit 2
before a single test ran, so the pass I measured locally never happened in CI.

It now resolves CARGO_HOME with a $HOME/.cargo fallback, creates the directory so
a fresh runner is covered too, puts it on PATH in case that bin is not already
there, and prints the version -- a step that silently installs nothing and then
fails later is harder to read than one that says which nextest it got.

Verified by running the same sequence against a CARGO_HOME that did not exist:
the directory is created and cargo-nextest 0.9.146 installs and runs.
The imm path had no CI coverage at all. In that mode the hash committed to the
proof is hash_vkey_with_part_vk over the embedded PART_STARK_VK_BYTES, so when a
ceremony moves the wrap vk and the tree does not follow, this is the only test
that fails -- the standard suite stayed green through exactly that drift, and it
was caught by hand. It runs in the `snark` job, in its own process per feature
set, because it loads the imm proving key and because it sets ZKM_IMM_WRAP_VK at
runtime. That variable is deliberately not exported: zkm-build would then pass
--features imm-wrap-vk to the plain `guests` workspace, which has no such
feature, and the guest build fails outright.

Verified on the box before pushing, asserting on the test's own result line
rather than the last one in the log: 1 passed in the default features (188s) and
1 passed with `ark` (145s).

Also drops the measurement anecdotes and commit archaeology from the comments in
both workflows, 200 comment lines down to 56, keeping the reasons a reader acts
on and removing the numbers that only dated them.
I replaced the twelve-package loop with one nextest pass on the strength of
830s -> 495s. That measurement was taken at 16 test threads on the 124-core GPU
box, and this runner has 16 cores and 123 GB, so the workflow runs at the job's
RUST_TEST_THREADS of 2. Measured again at that concurrency, on one tree, with the
same 636 tests: the loop takes 815s and nextest 1053s. nextest gives every test
its own process, and at 2-way concurrency that spawn cost is larger than the
cross-binary scheduling it buys. It also does not run doctests, so it needed a
second pass to keep them, and it needed installing on the runner -- where it went
to $HOME/.cargo/bin, which does not exist here, and took the job down before a
single test ran.

So the loop comes back. The one thing worth keeping from the detour is that a
failing package no longer hides the rest: each rc is collected and the step fails
at the end naming every failure, instead of aborting at the first under `bash -e`.

Verified: the loop run is 636 passed, 0 failures; and the collection logic
reports both of two injected failures and exits 1 rather than stopping at the
first.
A schedule runs on the default branch only, so gating these proofs on
schedule/dispatch/main left the feature branch -- where all the work happens --
covered by manual dispatch alone. That is how a moved wrap vk reached the
published v2.0.0 artifacts this week with a green suite behind it: the only tests
that notice are the ones that were not running.

A push that touches crates/recursion, crates/pcs or crates/verifier/bn254-vk now
runs them. The decision is a `git diff --name-only` over the whole push in a job
on a hosted runner, so it never costs the one self-hosted slot and adds no
third-party action to a runner that holds credentials. A push with no usable base
(new branch, force-push) answers yes, because an extra run costs slot time while a
missed wrap-vk move ships artifacts no released verifier accepts.

`!cancelled()` keeps a failure of that job from silently suppressing the
schedule, dispatch and main paths, which a bare `needs:` would.

Verified against six real commits, two that move the vk and four that do not:
60dd5dd (bn254-vk) and 4ca047b (pcs) classify true; b158194, ada0a46,
259f146 and 42223ce classify false. The no-base path returns true for an empty,
all-zero and unknown base, and diffs a real one.
…explicit

deps: the 22 p3-* crates were pinned to `branch = "zkm/whir-pcs"`. A branch is a
mutable pointer, and these crates are the field, the challenger and the Merkle tree
-- the transcript -- so `cargo update` could change every challenge, hence every vk,
with no diff to review. It already had: Ziren locked bca38384 while ziren-gpu, which
production is built from, locked 4dd0d47a, whose four extra commits add the
`grind_hook` the device grind installs through. A CPU-side build here resolved a
Plonky3 production does not use, and an unregistered grind falling back to the host
cost +123s on the LogUp grind alone when that happened for real. Both now resolve
4dd0d47a. Cargo rejects `branch` and `rev` together, so the branch name survives as
a comment.

Verified inert rather than assumed: the net bca38384..4dd0d47a diff is +90 lines in
three files with no deletions (the merkle-tree commit between them was reverted),
all behind `if let Some(hook) = grind_hook::get()`, which nothing in this repo
installs. `cargo check -r --workspace --all-targets` passes, Cargo.lock moves by
exactly one line, the recomputed transcript profile digest still matches the pin at
6a40de2d0355a10a, and vk_map.bin and vk_root.bin are byte-identical. No vk
regeneration, no ceremony.

deps: dropped `p3-mds` and `p3-multilinear-util`, which no .rs file references. This
is manifest hygiene only, not a build saving: both arrive transitively (p3-mds via
p3-koala-bear, p3-mersenne-31, p3-monty-31 and p3-poseidon2; p3-multilinear-util via
p3-commit) and still compile.

ci: core-machine gets a third test thread instead of the job's 2. Its cases prove at
a quarter of the area fence, ~26 GB each, so three fit in roughly 78 GB of the
runner's 123 GB, while the sdk and end-to-end cases that peak near 77-102 GB stay at
2. The estimate cannot be checked from here, so the step measures itself: a sampler
records the lowest MemAvailable and reports it as a notice, and if three was too
many the next run says so with a number rather than with an earlyoom kill. Context:
that job is 166 minutes and 99% of it is test execution, not compilation.

paper: `guest` appeared 61 times and `host` 41, neither ever defined, with the first
use of `guest` ahead of the section that explains it. Defined once in Preliminaries,
where it also records that the verifier trusts nothing the host computes -- load
bearing in the soundness argument and previously implicit. "Guest cycles count
executed instructions" is checked against the executor, which has exactly one
`global_clk += 1`.

paper: stopped tracking the PDF. A committed build output drifts from the sources it
represents and nothing was checking; 15 blobs and 3.1 MB of history for a file
regenerated on every edit. Build it with `make` in docs/paper. Removed from the
index, not from history.
@eigmax
eigmax force-pushed the feat/upgrade-plonky3 branch from 5cf2297 to 2c93774 Compare September 26, 2026 10:43
Section 4 was the least abstract part of the main body: 26 \texttt{} identifiers in
173 lines, and most were the implementation's own struct names -- MemoryBump,
MemoryLocal, MemoryGlobalInit, MemoryGlobalFinalize, Global, Byte, Range, Program --
already catalogued in Appendix~\ref{sec:app-chips}. Each paragraph's header states
the role and the prose then named our Rust type, so the section read as a
description of one codebase rather than of a construction.

The prose now names the role: a dedicated unit for the register shadow read, a
shard-boundary unit for the local memory rows, two further units for initial and
final memory, the accumulation unit, the byte table, a range table, the program
table. The mathematics and the arguments are unchanged; only the names are.

What deliberately stays concrete: the MIPS mnemonics, which are specification terms;
\texttt{reth}, a real program; and the sentence that points at the appendix and names
the three families dominating trace area, which is the one place the abstraction
should meet the realisation.

Caught in review before building: stripping \texttt{ from "On a \texttt{reth}" left a
dangling brace and the file was unbalanced at 303/304. Restored, and the count is
even again. Rebuilt: 55 pages, no errors, no undefined references.
…#534)

WeierstrassDoubleAssign divided by 2y without constraining it to be
nonzero, so a row doubling a point with y = 0 could claim any slope.
The chip now carries an inverse check (2y * inv = is_real) through a
FieldOpCols Div, the executor refuses to double a point with y = 0, and
the tests cover the honest and forged zero-y rows.  The new columns move
the core vk; mips_costs.json is regenerated for the double chips.
mark_gadget / annotate hooks for leading_one, gt_bytes, field_op,
field_den, field_inner_product, field_sqrt, field_lt, divrem, the
keccak round and the decompress chips.  The hooks are no-ops for the
prover builders, so no constraint, trace or vk changes.
)

The Global chip's is_send / is_receive choose the sign of the lifted point's
y through the gated ranges on y.0[6], but nothing made them bits or tied them
to is_real: a real row with both flags 0 accepted both (x, y) and (x, -y),
so its accumulated digest was the prover's choice.  Every sender on the local
Global bus puts exactly one constant flag, so this was not reachable through
the bus, but the chip did not state the invariant itself.

eval_single_digest now asserts both flags boolean and is_send + is_receive =
is_real, as SP1 does since v6.0.0.  No columns change; the host and device
trace already set exactly one flag on real rows and neither on padding rows.
…#536)

KeccakSpongeControl bound the syscall's input address, output address and
length only on the first block: bus B chained (clk, block, state) and nothing
else, so on a multi-block call a prover could absorb later blocks from any
address, stop after any number of blocks, and write the digest anywhere.  The
first and final flags were also not tied to is_real, so a row with is_real = 0
could receive the syscall and write an unpermuted state as the digest.

- bus B carries (input_address, output_address, input_length), the send with
  input_address + 4 * rate;
- the first block has block = 0 and binds the new input_length column to the
  length word it reads, whose top byte is below 64 (length < 2^30 words);
- the final block has (block + 1) * rate = input_length;
- is_first_block and is_final_block each imply is_real.

The new column is last, so no existing offset moves; the chip's trace is
generated on the host.
The WeierstrassDouble inverse column (#534), the Global flag constraints
(#535) and the keccak sponge's bus-B context and new column (#536) move every
leaf key, so the enumerated map is rebuilt: 6,517 shapes enumerated, 6,366
keys, root 0035eacfa17d72ff (was 004cb00d5f734d88).

The SNARK circuits do not move: the Groth16 constraint system and the wrap vk
(part_stark_vk.bin) are byte-identical to v2.0.0's, and the vk root is a
public input.  Checked with VERIFY_VK=true against the released v2.0.0 keys:
zkm-sdk test_groth16_public_values (a 31.9M-constraint Groth16 proof, proved
and verified) and the three imm-wrap-vk tests pass, as does
test_get_part_stark_vk.
trusted_setup.sh and trusted_setup_imm_wrap_vk.sh clone ProjectZKM's
semaphore-gnark-11 zkm2-par branch, which takes gnark from ProjectZKM/gnark
mpc-setup-par: InitPhase2 accumulates per wire in parallel and Contribute
scales Z and L in parallel.  The output is byte-identical to the sequential
tool apart from the random proof of knowledge; p2n drops from ~8 h to ~3.4 h
and each p2c from ~50 min to ~2-3 min on the wrap circuit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants