Skip to content

test(cow): drive the engine against a forked mainnet, and fix what it found - #691

Open
mfw78 wants to merge 7 commits into
mainfrom
cow/655-anvil-e2e
Open

mfw78 wants to merge 7 commits into
mainfrom
cow/655-anvil-e2e

Conversation

@mfw78

@mfw78 mfw78 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Part of #655. Closes #692. Closes #693.

Drives a real TWAP from registration to retirement against the deployed contracts, and fixes the three defects that run surfaced.

The end-to-end harness, the chain leg through to Post and the submit that follows it, plus the poll-classification fix the harness found. #694 is folded in here rather than shipped alone, since it exists because of this run.

What it does

Boots the shipped shepherd binary against an anvil fork of a real mainnet node, registers a conditional order on the deployed fork registry, and asserts the module indexes it and polls it through the structured generator wire.

Skipped unless SHEPHERD_FORK_RPC names an endpoint to fork, and skipped when anvil or cast is absent, so CI is unaffected:

just test                                                  # skips
SHEPHERD_FORK_RPC=http://<node>:8545 cargo test -p shepherd-engine --test anvil_fork_e2e

Why the binary and not BootScenario

The runtime's test harness wires a FakeNode and sets config.chains from test_chain_configs().
There is no seam for a real endpoint, and only the binary reads [chains.1] rpc_url.
Adding one would be a nexum-runtime change plus the lock-step pin bump through videre and shepherd, to serve a test.

Nothing is deployed

ComposableCow at 0xf9ba6F64... and OwnedTWAP at 0x4e17a65d... went out in the same CREATE2 broadcast at block 25674440, and CowAccount7702 at 0x15236F06... is live too, so an anvil fork already has all three.
Owned gates only setDescriptor and setModule, so registering against the handler needs no owner.

Neither the registry nor the handler has emitted a log in the 258000 blocks since, so every log the run sees is one it created.

The lifecycle it covers

The registered fixture is two parts 600 seconds apart, and the run follows all of it:

stage asserted
registration indexed commitment: from a real ConditionalOrderCreated
part 0 poll ... -> Post over a live eth_call, then submitted
between parts the commitment is gated to part 1's start and nothing polls
part 1 its own order, its own journal key, submitted again
end of series NextPoll::Never, and the commitment retires

The gap between the parts is load-bearing rather than incidental. Part 0's post carries nextPollTimestamp at part 1's start, so the run has to push the fork's clock forward to reach it. Removing that jump makes the second submission never arrive, which is how the gating is proven rather than assumed.

Part 1 mints a different order, with a later validTo and so a different uid, and the run asserts the two journal keys differ rather than that a second line turned up.

The retirement after the final part is the teardown from #685 and #686. It was only ever unit-tested; this is the first time it runs against the deployed registry and handler.

The 7702 delegation is load-bearing

The owner is an anvil EOA delegated to CowAccount7702, the delegate #658 names for the acceptance smoke, set with anvil_setCode writing the 0xef0100 ++ implementation designator.

It is not a convenience. The registry builds an ERC-1271 signature only for a POST verdict, and that path asks the owner for supportsInterface(ISignatureVerifierMuxer):

  • A bare EOA has no code, so the call succeeds with empty returndata, the bool decode fails, and the whole poll reverts with no payload.
  • CowAccount7702 has no supportsInterface either, but it reverts FnSelectorNotRecognized, which is a real revert, so _buildSignature catches it and takes its documented "assume a non-Safe wallet" branch.

So the delegate is what makes a post reachable at all. Without it the run stops at TryNextBlock and never exercises the wire this issue is about.

Two defects it surfaced

#692, an empty-data revert re-polls forever. Fixed here.
classify_revert read rpc.data, found None, and fell through to Verdict::TryNextBlock silently, so such a commitment was re-polled every block indefinitely with nothing dropping, backing off or reporting it.

The poll interface is closed: a generator answers through GeneratorResult codes and the registry's own refusals carry selectors. An empty payload is outside that, so it names a defect. It does not say whose, so the classifier no longer guesses: classify_revert returns a Refusal and the keeper attributes it with one eth_getCode.

A codeless owner cannot answer the registry's ERC-1271 probe, which makes that call revert in the caller's own frame with nothing attached. That is the registration answering for itself, and it drops, loudly. Every other cause points outward at a gas cap too low for the handler or a registry address that is not the fork; those are identical for every commitment, so a drop would delete a whole watch set to report an operator's typo. They back off for an hour and name the likely cause.

#693, a venue receipt mismatch destroyed the commitment. Fixed here.
The run reached Post, submitted, and a receipt that did not match made retry_action drop the commitment permanently.
#682's scope included "decide the retry_action fallback for non-Denied faults, which bypasses the table and drops", and #688 did not address it, because retry_action lives in videre-sdk.

The fix is a CoW fault policy, over the FaultPolicy seam added in nullislabs/videre-nexum-module#88, and the videre pin moves from fd8af02 to 9ab1515 across all eight declarations.

A receipt this keeper cannot correlate ends the submission but not the commitment. Those are different questions and this is the one place they part:

  • The uid is keccak(order) ++ owner ++ validTo, a pure function of the body that was sent, so re-posting it yields the same disagreement every time, and the order cannot be asked after either, since the only handle on it is a uid this keeper does not trust. The reservation is released. Leaving it would have reconcile re-post the same body on every tick forever, which is worse than the drop it replaced.
  • The commitment survives, because its next poll mints a later part with a later validTo and a different uid. A repeat at a later block says the disagreement is systematic, and then it drops.

A second bug went with it. Denials now route through the shipped table for the reconcile pass as well as the submit path. They did not before: run.rs classified Denied with classify_denied, but reconcile used the platform default internally, which knows nothing of classification.toml. So a stranded reservation for a clearable refusal like InsufficientBalance was released while the same refusal on the submit path backed off.

An open question I could not close

The chain behaves exactly as the fix assumes. Registering from a genuinely codeless address and calling getTradeableOrderWithSignature with the module's own call shape reverts with empty data, and OwnedTWAP.poll returns POST for that same owner.

Under the engine, though, the same registration polls to TryNextBlock with no warning logged at all, which means the eth_call succeeded and decoded a TRY_NEXT_BLOCK generator code rather than reverting. I could not reconcile that with the direct call, so the end-to-end case for this path is not in this PR. The behaviour is covered by unit tests over classify_revert and the attribution instead.

Worth knowing for whoever picks it up: anvil's account zero already carries an EIP-7702 delegation on mainnet, which a fork inherits. An earlier version of this investigation was wrong for that reason, and the harness now registers from a fresh address so an owner's code is only ever what the test set.

The mock had to become faithful

The second commit is why the run gets past Post.

orderbook-mock answered with a counter packed into 56 bytes and ignored the posted body. A submitting venue re-derives the UID locally and refuses a receipt that disagrees, so no submit through this mock could ever complete: the run reached Post and then dropped the commitment on receipt mismatch.

It now parses the posted OrderCreation and returns the UID that order actually has, through order_data() and the chain's settlement domain, so the derivation is the one the venue checks with rather than a second implementation of it. A body it cannot read is refused rather than answered with something the venue would reject anyway, and --chain-id selects the domain because a UID from the wrong domain is refused exactly like a synthetic one.

That is what lets this PR assert the submit lands, not just the poll.

One trap worth naming: cargo test -p shepherd-engine does not rebuild another package's binary, so a stale orderbook-mock fails as a receipt mismatch rather than as a stale build. The test now compares mtimes and says so.

What is deliberately not in this PR

The Sepolia script migration. scripts/lib.sh, e2e-onchain.sh and baseline_latency.py still target Sepolia. This PR adds the fork harness beside them rather than rewriting them in the same change.

#660's restart-gap test. It belongs on this harness and lands next.

Verification

243 tests pass with the fork test skipped, which is the CI path.
With SHEPHERD_FORK_RPC set against a real node the fork test passes in about 25 seconds, asserting indexed commitment:, then poll commitment:... -> Post, then the submit landing, with no decode failure and no eth_call failed on the way.
cargo fmt --check, cargo clippy --workspace --all-targets --all-features -- -D warnings, RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps and just build-modules all exit 0.

Clippy caught a genuine leak while writing this: start_anvil dropped its child on the panic path without reaping it. The child is owned before the readiness loop now.

AI Assistance: Claude Code used for the harness and the environment survey.

Boots the shipped `shepherd` binary against an anvil fork of a real
mainnet node, registers a conditional order on the deployed fork
registry, and asserts the module indexes it and polls it to `Post`
through the structured generator wire.

Drives the binary rather than `BootScenario`: that harness wires a
`FakeNode` and hardcodes its chain config, so it has no seam for a real
endpoint, and only the binary reads `[chains.1] rpc_url`. Adding one
would be a nexum-runtime change plus the pin cascade, to serve a test.

Nothing is deployed. The registry, `OwnedTWAP` and `CowAccount7702` are
all live on mainnet, so an anvil fork already has them.

The owner is an EOA delegated to `CowAccount7702`, which is what #658
names for the acceptance smoke. It is load-bearing rather than
incidental: the registry builds an ERC-1271 signature only for a POST,
and that path asks the owner for `supportsInterface`. A bare EOA
answers with empty returndata and the whole poll reverts with no
payload; the delegate answers `FnSelectorNotRecognized`, a real revert,
which `_buildSignature` catches as a non-Safe wallet.

Skipped unless `SHEPHERD_FORK_RPC` names an endpoint to fork, since CI
reaches no such node, and skipped when `anvil` or `cast` is absent.

AI Assistance: Claude Code used for the harness and the environment survey.
The mock answered with a counter in 56 bytes, ignoring the body. A
submitting venue re-derives the UID locally and refuses a receipt that
disagrees, so no submit through this mock could ever complete: the fork
run reached `Post` and then dropped the commitment on `receipt
mismatch`.

It now parses the posted `OrderCreation` and returns the UID that order
actually has, via `order_data()` and the chain's settlement domain, so
the derivation is the one the venue checks with rather than a second
implementation of it. A body it cannot read is refused rather than
answered with something the venue would reject anyway.

`--chain-id` selects the settlement domain, because a UID from the
wrong domain is refused exactly like a synthetic one.

This closes the venue leg of the fork run, which now asserts the
submit lands.

AI Assistance: Claude Code used for the mock and the tests.
@mfw78
mfw78 force-pushed the cow/655-anvil-e2e branch from 4745789 to 303f3aa Compare September 8, 2026 23:31
`classify_revert` fell through to `TryNextBlock` whenever it could not
find a selector, so a refusal with no readable payload re-polled on
every block indefinitely, silently, with nothing to distinguish it from
a healthy scheduling gap.

The poll interface is closed: a generator answers through
`GeneratorResult` codes and the registry's own refusals carry selectors.
An empty payload is outside that, so it names a defect. It does not say
whose, and the candidates are not alike, so the classifier no longer
guesses. `classify_revert` returns a `Refusal` and the keeper attributes
it with one `eth_getCode`.

A codeless owner cannot answer the registry's ERC-1271 probe, which
makes that call revert in the caller's own frame with nothing attached.
That is the registration answering for itself and it drops, loudly.

Every other cause points outward: a gas cap too low for the handler, or
a registry address that is not the fork. Those are the same for every
commitment, so a drop would delete a whole watch set to report an
operator's typo. They back off for an hour and name the likely cause.
A failed `eth_getCode` reads as unattributable and takes that path too.

Reachable in production by an ordinary mistake, and the codeless case
is recoverable rather than permanent: an EOA can gain code through an
EIP-7702 delegation, which is exactly the flow #658 describes.

Closes #692.

AI Assistance: Claude Code used for the fix and the tests.
Extracts the anvil fork, the orderbook mock and the engine into a
`Harness` so a second case can reuse them.

The run also stops registering from anvil's account zero. That is the
well-known test key, and on mainnet it already carries an EIP-7702
delegation, which a fork inherits: an owner meant to have known code
would silently have whatever mainnet gave it. The run funds and
impersonates a fresh address instead, so the owner's code is exactly
what the test put there.

AI Assistance: Claude Code used for the refactor.
@mfw78 mfw78 changed the title test(cow): drive the engine against a forked mainnet test(cow): drive the engine against a forked mainnet, and fix what it found Sep 9, 2026
Bumps the videre pin to 9ab1515 and supplies the CoW fault policy the
seam there exists for.

A receipt this keeper cannot correlate used to remove the commitment.
The fork run reached `Post`, submitted, got a uid that did not match the
one derived from the order it sent, and destroyed a commitment whose
next poll would have minted a perfectly good order.

The uid is `keccak(order) ++ owner ++ validTo`, a pure function of the
body that was sent, so re-posting it yields the same disagreement every
time, and the order cannot be asked after either: the only handle on it
is a uid this keeper does not trust. So the submission ends and its
reservation is released. Leaving it would have reconcile re-post the
same body on every tick forever.

The commitment is not the submission, and it survives: the next poll
mints a later part with a later `validTo` and a different uid. A repeat
at a later block says the disagreement is systematic rather than one
order's, and then it drops.

That parts from `is_terminal` deriving from `action`, deliberately and
in the one place the two questions differ.

Denials now route through the shipped table for the reconcile pass as
well as the submit path. They did not before: reconcile took the
platform default, which knows nothing of `classification.toml`, so a
stranded reservation for a clearable refusal was released while the same
refusal on the submit path backed off.

Closes #693.

AI Assistance: Claude Code used for the policy and the tests.
@mfw78
mfw78 force-pushed the cow/655-anvil-e2e branch from 5880966 to b200f14 Compare September 9, 2026 06:11
The run stopped after part 0 posted, so it proved the chain and venue
legs connect and nothing else. A TWAP is a series, and the parts of it
that had never been exercised end to end are the ones the last few
changes were about.

It now advances the fork's clock past the first part's window. That is
load-bearing: part 0's post carries `nextPollTimestamp` at part 1's
start, so the commitment is gated until then and nothing polls in
between. Without the jump the second submission never arrives.

Part 1 mints its own order, with a later `validTo` and so a different
uid, and submits under its own journal key: the run asserts the two keys
differ rather than that a second line appeared.

After the final part the generator reports no successor, which is
`NextPoll::Never`, and the commitment retires. That teardown was only
ever unit-tested; this is the first time it runs against the deployed
registry and handler.

AI Assistance: Claude Code used for the lifecycle run.
Two parts could pass for a fluke and the run only ever asserted on log
lines, so a teardown that logged but left rows behind would have gone
unnoticed. It is five parts now, and the run reads the engine's own
redb store back at the end: after retirement no `commitment:`, `due-`,
gate, refusal, `watch-sub:`, `exp-t:`, `submitted:`, `parked:`, `root:`
or `context:` key survives.

The engine is asked to exit rather than killed. redb loses whatever it
had not flushed, and what is left is an older consistent snapshot, so a
killed engine reads as a run that stopped mid-series: the first version
of this failed that way, reporting a teardown that had in fact
happened.

Same backend the runtime writes, and the same keccak namespace, so the
read cannot drift from the write.

AI Assistance: Claude Code used for the store readback and the longer run.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cow: a venue receipt mismatch destroys the commitment permanently cow: an empty-data revert re-polls the commitment on every block forever

1 participant