Repository navigation
Stencil 3D device kernels: 6D matrix views instead of the 2p+1 assumption - #97
Merged
Merged
Conversation
The 3D CUDA kernels stencil_dot_3d and stencil_transpose_3d take the matrix data as Array6D<double> views (cunumpy >= 0.6.1) instead of raw pointers whose shape was derived from 2 * p + 1 diagonals. All six dot/transpose kernels (pyccel and CUDA, 1D to 3D) read the number of diagonals from the matrix data: the pads of the matrix are q = (n - 1) // 2, diagonal d of row i is the column i - q + d, the last owned row uses n - 1 + add. With q = p the results are bitwise those of the previous kernels. - stencil_transpose_<n>d: drop the e_in argument (only needed for the raw pointer in 3D). - StencilMatrix: dot, vdot and transpose no longer raise for matrices with fewer diagonals than 2 * p + 1; spaces with shifts > 1 still raise. transpose(out=...) asserts that out has the pads of the matrix. - Parity/emulation cases: matrices with fewer diagonals, non-periodic rectangular blocks, pads 0; dense references in test_device_matvec.py, a distributed check in test_mpi_device.py. - CUDA_STRATEGY.md: notes for this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 7, 2026
…6d-views Stack #703 -> #706: feectools submodule points to the head of struphy-hub/feectools#97, which now contains #96. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
max-models
added this pull request to stack #99
October 8, 2026 05:42
spossann
reviewed
Oct 8, 2026
spossann
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 9, 2026
**Stack:** part 3 of 3 — based on #706 (`feectools-stencil-6d-views`, merge that first, after #703), next: none (top of the stack). Corresponding update in feectools: struphy-hub/feectools#98. Mirrors the feectools stack struphy-hub/feectools#96 → struphy-hub/feectools#97 → struphy-hub/feectools#98. `feectools-stencil-6d-views` is merged into this branch (no rebase), and the `feectools` submodule now points at the stacked head of struphy-hub/feectools#98, `6d88806`, which contains feectools#96 and #97. Against its base (#706) this PR only moves the submodule `d16a6ad` → `6d88806`. After each feectools PR of the stack is merged into `devel-tiny`, the `feectools` submodule of the struphy stack must be moved to the corresponding merged `devel-tiny` commit (the `pr-feectools-submodule` check fails until then). --- **Solves the following issue(s):** Part of #689 (`VlasovAmpereOneSpecies` end to end on the GPU, no host/device transfers in the time loop), tracked in #650. This PR moves the `feectools` submodule to struphy-hub/feectools#98 ("Inner products stay on the device"). On the CuPy backend, every CG iteration of a struphy solve copied two scalars from the device to the host; after this change it copies one (the convergence test). **The `pr-feectools-submodule` check fails until struphy-hub/feectools#98 is merged** into `devel-tiny`. After that, the submodule should point at the merge commit. **Core changes:** - `feectools` submodule: `d16a6ad` (base #706) → `6d88806` (branch `inner-on-device` of struphy-hub/feectools#98, stacked on feectools#96 and #97; originally `a15a8e8`). No struphy code changes. - What changes in feectools, and only on the CuPy backend: - `StencilVector.inner`/`BlockVector.inner`/`dot_inner` return a 0-d device array, in serial and with MPI. - `axpy` accepts such a scalar without copying it to the host. - CG, PCG, BiCG, BiCGStab and PBiCGStab keep alpha/beta on the device and copy only the residual norm, once per iteration. On NumPy, results and types are unchanged. Copies to the host per iteration, counted under cunumpy's fake CuPy: | Solver | before | after | | --- | --- | --- | | CG | 2 | 1 | | PCG | 3 | 1 | | BiCG | 4 | 1 | | BiCGStab | 6 | 1 | | PBiCGStab | 5 | 1 | - struphy call sites of `.inner` that run on CuPy now get a 0-d device array. I checked them: - the scalars in `models/scalars.py` write it into `xp` buffers (`local_value[0] = ...`), which works on the device; - the model energy methods (`linear_mhd.py`, `shear_alfven.py`, ...) return it, and the scalar machinery handles it like the existing `dot_inner` results; - `PolarVector.dot` adds a NumPy scalar to it, which works, but polar splines are not on CuPy yet (#695); - the multigrid smoothers run their own PCG with `.inner`. It works with device scalars: the arithmetic stays on the device, and the comparisons with 0 copy to the host implicitly. - This removes the first blocker listed in #705 ("CG inner products ... The Schur solve is outside the transfer guard because of this"). The remaining copy per CG iteration still counts as a transfer for `assert_no_transfers`. The transfer guard could then cover the Schur solve with an allowance of one 8-byte copy per iteration. **Model-specific changes:** None. **Documentation changes:** None in struphy. The feectools PR documents the change in feectools' `CUDA_STRATEGY.md` (section "Inner products on the device"). **Testing** (macOS, no GPU; struphy kernels compiled with GNU/Fortran; **GPU tests not run**): - feectools (see struphy-hub/feectools#98): - serial suite: 9456 → 9480 passed, with the same 6 unrelated failures in `ddm/tests/test_cart_*d.py`; - `mpirun -n 2` linalg MPI suite: 738 → 739 passed; - the new fake-CuPy tests check the copy counts and that the results equal NumPy's. - struphy, solver tests on the NumPy backend, with feectools `devel-tiny` (before) and this branch (after): - `propagators/tests/test_poisson.py -m "not mpi" -k "not multigrid"`: 40 passed before, 40 passed after; - `linear_algebra/tests/test_saddlepoint_massmatrices.py`: the first case (`SaddlePointSolverUzawaNumpy`) passed before and after. I stopped the remaining saddle-point and multigrid tests: with the feectools kernels uncompiled in the checkouts, each test took more than 10 minutes. On NumPy, feectools' change is limited to identity helpers and an equivalent loop structure in BiCGStab, and the feectools solver tests check the same iteration counts and results. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 2 of 3 — based on #96 (
device-kronecker-solve, merge that first), next: #98 (inner-on-device).device-kronecker-solveis merged into this branch (merge commitd16a6ad, no rebase, so review comments stay); the only conflict wasCUDA_STRATEGY.md, where both new sections are kept ("Kronecker solver on the device (#96)" first, then "6D matrix views"). The diff against the base shows only this PR's change. When #96 is merged, GitHub retargets this PR todevel-tiny.Companion struphy PR: struphy-hub/struphy#706 (part 2 of the struphy stack struphy-hub/struphy#703 → struphy-hub/struphy#706 → struphy-hub/struphy#716); its
feectoolssubmodule points at this branch's new headd16a6ad, which contains #96.Corresponding PR in struphy: struphy-hub/struphy#706
Summary
The 3D device stencil kernels from #88 (
stencil_dot_3d,stencil_transpose_3d) take the matrix data as 6D views (Array6D<double>, cunumpy >= 0.6.1) instead of raw pointers, and all dot/transpose kernels read the number of diagonals from the matrix data instead of assuming2 * p + 1.StencilMatrix.dot,vdotandtransposetherefore no longer raiseNotImplementedErrorfor matrices with fewer diagonals than the pads of their spaces allow (blocks between spaces of different degree, derivative-type stencils). Related to struphy-hub/struphy#650 and struphy-hub/struphy#688 (6D array views in cunumpy).Changes
stencil_dot_3d(Array6D<double> mat, Array3D<double> x, Array3D<double> out, s_in, p_in, add, s_out, e_out, p_out)andstencil_transpose_3d(Array6D<double> mat, Array6D<double> matT, s_in, p_in, add, s_out, e_out, p_out). Strided views as in 1D/2D (Array2D/Array4D), so the hand-computed strides are gone. The pyccel versions takefloat[:, :, :, :, :, :]with the same arguments in the same order.n = mat.shape[ndim + k]diagonals, its pads areq = (n - 1) // 2, diagonaldof rowiis the columni - q + d(atx[i - q + d - s_in + p_in]), and the last owned row usesn - 1 + add. The transpose maps diagonaldofmatTto diagonalq + i - jofmat. Withq = pthis is the old loop, in the same order. The pyccel kernels are one loop nest that picks the number of diagonals per row, instead of the 2/4/8 spelled-out (interior, last row) combinations.e_inremoved fromstencil_transpose_1d/2d/3d. It was only needed for the row extents of the raw pointer in 3D, and the argument list is the same for all dimensions.StencilMatrix: the2 * p + 1check is gone fromdot,vdotandtranspose.transpose(out=...)asserts thatouthas the pads of the matrix. Spaces with shifts > 1 still raiseNotImplementedError(see below).MATRIX_CASES(GPU parity and CPU emulation) gets 2 new 1D cases, 2 in 2D and 4 in 3D. They cover fewer diagonals in some directions, non-periodic rectangular blocks between spaces of different size per direction, a derivative-type rectangular block with fewer diagonals in every direction, and pads 0 (one diagonal).test_device_matvec.py: the old "reject" test is replaced bytest_fewer_diagonals_match_dense_reference, which checksdot,vdot,transposeandtranspose(out=...)againsttoarray()for 8 such matrices. A new test checks that shifts > 1 raise.test_mpi_device.py: a matrix with fewer diagonals against the global field, and the adjoint identity of its transpose for a non-symmetric (one-sided) stencil.CUDA_STRATEGY.md: new section "6D matrix views". The2 * p + 1/ raw-pointer limitation and the open question about 6D views are removed, and shifts > 1 are listed as an open question.cunumpy >= 0.6.1was already required ondevel-tiny(Update to cunumpy 0.6.1 #94), so it is unchanged.Behaviour changes
2 * p + 1(StencilMatrix(V, W, pads=q)withq < p) work indot,vdotandtransposeon both backends, where they used to raise.stencil_transpose_<n>dtakes one argument fewer (e_in). OnlyStencilMatrixcalls these kernels; struphy does not call them directly.matvec_<n>d,transpose_<n>d) disagree withtoarray()for shifts > 1. Their transpose is also not the adjoint of their product, and psydac never tested them ("TODO: verify for s>1"). Supporting shifts would first need a verified data layout.dot5.9 → 5.1 ms,transpose6.8 → 6.8 ms.Testing
Local runs on macOS without a GPU (cunumpy 0.6.1, pyccel 2.2.1 with C, Open MPI 5). The GPU tests were not run. Before → after:
test_cuda_parity.py,test_cuda_emulation.py,test_device_matvec.py: 29 passed / 31 skipped → 37 passed / 49 skipped. More GPU parity cases are skipped now, and the CPU emulation covers all cases, the new ones included.feectools/linalg -m "not mpi and not petsc": 7942 passed → 7950 passed.feectools -m "not mpi and not petsc": 9456 passed, 6 failed → 9464 passed, 6 failed. The 6 failures are the pre-existing ones inddm/tests/test_cart_2d.py/test_cart_3d.py.mpirun -n 2 … feectools/linalg -m "mpi and not petsc" --with-mpi: 738 passed → 739 passed.test_mpi_device.pyalso passes with 1, 3 and 4 ranks.-DCUNUMPY_BOUNDS_CHECK: no out-of-bounds index.dot,vdotandtransposeagainsttoarray()for every case with a dense meaning.🤖 Generated with Claude Code