Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
ac980a8
cuda: hash single-chunk BLAKE3 rows with one thread per message
samuelburnham Sep 11, 2026
38f7cc0
cuda: fuse wide NTTs at any height and trim LDE memory passes
samuelburnham Sep 11, 2026
3c87a59
cuda: per-device prover binding, host pools per device, GPU trace wit…
samuelburnham Sep 15, 2026
9869234
cuda: remove the Merkle tree cache, keep per-construction device admi…
samuelburnham Sep 15, 2026
9b172ed
cuda: hash height groups wider than 32 KiB per row on the host
samuelburnham Sep 15, 2026
39d59cb
cuda: share Goldilocks arithmetic with generated trace kernels
samuelburnham Sep 16, 2026
d4d09c7
cuda: admit lookup graphs only for resident LDEs that kept their trace
samuelburnham Sep 17, 2026
89747d3
cuda: let trace generators keep device state, released after the look…
samuelburnham Sep 17, 2026
2ec82fa
cuda: spans for LDE shapes, commit sources and the FRI opening phases
samuelburnham Sep 17, 2026
a1fcbf5
cuda: cache the FRI coset per process; double-buffer staged uploads
samuelburnham Sep 17, 2026
82c6c9b
cuda: collect lightweight prover operation metrics
samuelburnham Sep 17, 2026
1691727
cuda: sppark's Goldilocks NTT as an optional comparison backend
samuelburnham Sep 17, 2026
fa64c28
cuda: build sppark's runtime without exceptions or threads
samuelburnham Sep 17, 2026
aa714fd
cuda: the sppark fork's dev branch carries the runtime mode
samuelburnham Sep 17, 2026
d964550
cuda: pin sppark at the fork's dev branch, 6c5d826
samuelburnham Sep 17, 2026
44ea38c
cuda: sppark adapter resolves the device by ordinal and fences before…
samuelburnham Sep 17, 2026
9fef499
cuda: resident coset LDE through sppark behind a runtime switch
samuelburnham Sep 17, 2026
62cee20
cuda: admit the sppark panel scratch, dispatch by height, count shape…
samuelburnham Sep 17, 2026
24ef583
cuda: every transform of a proof through sppark above the height thre…
samuelburnham Sep 17, 2026
a15a979
cuda: the sppark adapter parses its settings without strtoul
samuelburnham Sep 17, 2026
4739cb7
cuda: sppark dispatch by whole shape, admitted everywhere, labelled i…
samuelburnham Sep 17, 2026
9c51794
cuda: sppark panels transform their columns in batched launch sequences
samuelburnham Sep 17, 2026
0ac94c1
cuda: sppark height threshold at 2^18, short shapes in the resident b…
samuelburnham Sep 17, 2026
e7be72c
cuda: sppark panel glue as tiled transposes, the fused expansion meas…
samuelburnham Sep 17, 2026
1326c65
README: the sppark transform backend and its settings
samuelburnham Sep 17, 2026
fde18f8
README: AIUR_METRICS is a path for Ix, any value for the counters
samuelburnham Sep 17, 2026
41fd44c
build: pin sppark at the fork's batched entry and patch record
samuelburnham Sep 17, 2026
ce77de2
cuda: one dispatch decision per transform, tables built for that back…
samuelburnham Sep 17, 2026
b6104b6
cuda: the sppark panels' scratch comes from the stream's memory pool
samuelburnham Sep 17, 2026
9c62a10
Use sppark for all CUDA transforms
samuelburnham Sep 17, 2026
bc864c1
Enable device snapshots with event-only tracing filters
samuelburnham Sep 17, 2026
7eae079
Validate transform plans against CUDA buffer dimensions
samuelburnham Sep 17, 2026
6516490
Format src/cuda/mmcs.rs with rustfmt
samuelburnham Sep 28, 2026
c9532e3
Drop the unused metrics layer, debug knobs and experiment docs
samuelburnham Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
139 changes: 131 additions & 8 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

16 changes: 12 additions & 4 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ edition = "2024"
authors = ["Argument Engineering <engineering@argument.xyz>"]
license = "MIT OR Apache-2.0"
rust-version = "1.98"
links = "multi_stark_cuda"

# See more keys and their definitions at https://doc.rust-lang.org/cargo/reference/manifest.html

Expand Down Expand Up @@ -32,6 +33,14 @@ p3-merkle-tree = { git = "https://github.com/Plonky3/Plonky3", rev = "3152b14a89
p3-symmetric = { git = "https://github.com/Plonky3/Plonky3", rev = "3152b14a89067c83775a8076cc262ffc48a1fd7c" }
p3-util = { git = "https://github.com/Plonky3/Plonky3", rev = "3152b14a89067c83775a8076cc262ffc48a1fd7c" }

# The CUDA transform engine. The fork supplies column batching, borrowed
# streams and a runtime compatible with the Lean executable's C++ linkage.
[dependencies.sppark]
git = "https://github.com/argumentcomputer/sppark"
rev = "e10e107673aa22861f0f8b9758fc62169ab919ae"
optional = true
features = ["cuda"]

[dev-dependencies]
criterion = "0.5"
p3-baby-bear = { git = "https://github.com/Plonky3/Plonky3", rev = "3152b14a89067c83775a8076cc262ffc48a1fd7c" }
Expand All @@ -45,10 +54,9 @@ harness = false

[features]
parallel = ["p3-maybe-rayon/parallel"]
# Use the first-party CUDA Goldilocks DFT/LDE backend. Enabling this feature
# requires a CUDA toolkit at build time and an NVIDIA GPU at runtime; the
# default CPU build never invokes nvcc or links the CUDA runtime.
cuda = ["dep:itertools", "dep:rayon"]
# CUDA requires the toolkit at build time and an NVIDIA GPU at runtime.
# CPU builds neither compile sppark nor link CUDA.
cuda = ["dep:itertools", "dep:rayon", "dep:sppark"]

# Similar to `release`, but preserves debug info
[profile.dev-ci]
Expand Down
27 changes: 25 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ lookup arguments for shared state.
a batteries-included Goldilocks/Blake3 instantiation is provided
- **Serialization** — `Proof::to_bytes` / `Proof::from_bytes` via bincode
- **Parallel proving** — opt-in via the `parallel` feature flag
- **CUDA transforms** — opt-in first-party Goldilocks DFT/LDE backend via the
- **CUDA transforms** — opt-in sppark Goldilocks DFT/LDE backend via the
`cuda` feature; normal builds remain independent of CUDA

## Reference configuration
Expand Down Expand Up @@ -80,7 +80,7 @@ are enabled by default via `.cargo/config.toml`.
## CUDA acceleration

The optional `cuda` feature routes the production Goldilocks configuration's
PCS and quotient transforms through first-party CUDA kernels. CUDA and CPU
PCS and quotient transforms through sppark. CUDA and CPU
proofs generated by the same revision are byte-identical and use the same CPU
verifier. Canonical field serialization changes newly generated CPU proof bytes
relative to revisions before this backend; previously generated proofs remain
Expand All @@ -104,6 +104,29 @@ hardware and downstream Ix measurements live in
See [cuda/README.md](cuda/README.md) for architecture, build settings, the NVIDIA
correctness harness, platform limitations, and benchmark commands.

### sppark transforms

Every CUDA build uses [sppark](https://github.com/argumentcomputer/sppark)'s
Goldilocks NTT, pinned to the fork's `dev` branch. Its borrowed streams keep
panel allocation, transforms and freeing on the caller's stream. Tiny generic
host matrices retain the CPU DFT. CPU builds do not depend on sppark or CUDA.

An immutable per-device plan fixes each transform's panel size, launch groups
and coset powers. Admission and execution share the plan. Dimensions must fit
the compiled domain (currently 2^28 rows), checked byte arithmetic and at least
one column within the panel budget; invalid shapes fail before allocation.

`MULTI_STARK_SPPARK_PANEL_BYTES` caps the transform scratch, 4 GiB by default;
a budget smaller than one column is a configuration error. Columns are
transformed in launch groups sized to the device's L2 cache. Both are captured
when the device DFT is constructed. The fork's `SPPARK_NO_CXX_RUNTIME` mode
aborts on CUDA errors with diagnostics.

```sh
cargo test --release --features parallel,cuda --lib cuda::sppark::tests::
MULTI_STARK_CUDA_BENCH_SHAPES="20,533,2" cargo run --release --features parallel,cuda --example cuda_resident_lde_bench
```

## License

MIT or Apache-2.0
Loading
Loading