ci: Fix sticky caches - #649
Merged
Merged
Conversation
RunsOn files a sticky-disk snapshot under the run's git ref and falls back to the default branch's snapshot for a ref that has none. Nothing ran these jobs on main, so the ci-lean-test and ci-rust-test lineages never had a main snapshot: every new PR branch and every merge-queue entry started from an empty disk and rebuilt about 1,300 Lean modules, ten minutes more than a warm run. A push run after each merge keeps a main snapshot at most one merge old.
…the PR branch The five matrix partitions used their own sticky lineage, which had no main snapshot either, so every merge-queue run was five cold builds. They now use ci.yml's lean-test lineage, which the push-to-main runs keep warm; the toolchain, codegen target and cached paths are the same. A comment event runs under the default branch's ref, so an issue_comment-triggered run would snapshot a PR's build as main's. The comment now only relays: a small job dispatches this workflow on the PR's head branch with the PR number as input, and the dispatched run does what the comment run did, under the branch's own ref. Fork branches and branches without the dispatch trigger get a comment explaining why nothing ran.
Same reasoning as ci.yml: the nix-x86-64-v4 lineage never had a main snapshot, so every new ref started from an empty disk.
The push run exists to snapshot build products the merge queue already tested. lean-test keeps the all-targets build and IxTcVerify and skips the codegen check, the toolchain diff and the test tiers; rust-test keeps clippy and check and builds the test binaries without running them, skipping rustfmt and cargo-deny. cuda-compile is all build steps and its main-keyed cache entry is what a new branch's restore-keys prefix can see, so it runs unchanged.
Cargo judges a workspace crate fresh by source mtimes, and a fresh checkout renews every mtime, so the restored target directory never spares the repo's own crates: a warm rust-test run still spent about five and a half minutes recompiling them, and lean-test rebuilds ix_rs for two minutes every run. sccache keys rustc outputs by content, so unchanged crates return as cache hits; linking and build scripts still run, so thin-LTO link time remains. cuda-compile is left out with a note on wiring both rustc and nvcc when the GPU benchmark lands after #644; the container's S3 credentials and multi-stark's single --lib nvcc call both need attention first.
Only the `nix develop` step can use it: the sandboxed `nix build` and `nix flake check` see neither the wrapper nor the network and stay on Cachix. The dev shell inherits RUSTC_WRAPPER and keeps the outer PATH, so nothing in the flake changes.
…enchmark on the PR branch The build jobs of bench-main and bench-pr share one native-r8i sticky lineage in place of the tarballed Lake cache and rust-cache: bench-main's push-to-main run leaves the lineage's main snapshot, and a !benchmark run, dispatched on its PR branch, restores that snapshot and writes its own to the branch. Each compile env gets its own lineage holding the compile workspace's .lake, since concurrent jobs on one lineage keep only the last clean completion; the FLT package cache and its matrix key go with it. sccache covers the workspace-crate rebuilds in both build jobs and in bench-pr's from-scratch base builds. The relay dispatched on the default branch, which made every !benchmark run a main job to RunsOn; it now dispatches on the PR's head branch and refuses forks with a comment, since a fork branch cannot be dispatched and would otherwise snapshot into main's lineage.
cargo-deny-action runs in Docker and inherits the job's RUSTC_WRAPPER=sccache, but the container has no sccache, so cargo failed on `sccache rustc -vV` before deny could fetch the crate graph. Deny never compiles; an empty wrapper is one cargo ignores.
samuelburnham
marked this pull request as ready for review
September 29, 2026 19:51
samuelburnham
enabled auto-merge
September 29, 2026 19:53
arthurpaulino
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RunsOn sticky disk cache snapshots are keyed on the branch in which they originated, with a fallback to main if none is found for a given branch. However, since we only run CI on
merge_group, there is no main cache for a fresh PR ormerge_grouprun to use, which means we have been only restoring the cache on subsequent PR runs.This PR runs the build-only steps for
lean-testandrust-teston push to main so that the build artifacts persist in the cache properly. It also adds a https://github.com/mozilla/sccache/ S3 backend so that the Rust crates can cache their compilation in many cases, since they are cached bymtimerather than file edits like inlake.Lastly, the
!benchmarkand!merge-testscomment workflows currently run on the default branch, which means that they would also write to the cache for main despite creating PR build artifacts. Instead we want them to write to their PR cache so that the main cache is kept clean, so this PR changes them to run onworkflow_dispatchwith the PR branch correctly checked out.Note
sccache should be integrated with CUDA compilation after #644