Skip to content

feectools: stencil 3D device kernels with 6D views - #706

Closed
max-models wants to merge 2 commits into
feectools-device-kronecker-solvefrom
feectools-stencil-6d-views
Closed

max-models wants to merge 2 commits into
feectools-device-kronecker-solvefrom
feectools-stencil-6d-views

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 2 of 3 — based on #703 (feectools-device-kronecker-solve, merge that first), next: #716 (feectools-inner-on-device).

Mirrors the feectools stack struphy-hub/feectools#96 → struphy-hub/feectools#97 → struphy-hub/feectools#98. feectools-device-kronecker-solve is merged into this branch (no rebase), and the feectools submodule now points at the stacked head of struphy-hub/feectools#97, d16a6ad, which contains feectools#96. Against its base (#703) this PR only moves the submodule 64b10fa → d16a6ad. After each feectools PR of the stack is merged into devel-tiny, the feectools submodule of the struphy stack must be moved to the corresponding merged devel-tiny commit (the pr-feectools-submodule check fails until then).


Solves the following issue(s):

Corresponding update in feectools: struphy-hub/feectools#97

Related to #650 (CUDA) and #688 (6D array views in cunumpy). With cunumpy 0.6.1's Array6D, the 3D stencil kernels of feectools take the matrix data as views instead of raw pointers. They no longer assume 2 * p + 1 diagonals, and StencilMatrix.dot/vdot/transpose work for matrices with fewer diagonals (blocks between spaces of different degree, derivative-type stencils) where they used to raise NotImplementedError.

Core changes:

This PR only moves the feectools submodule from f97ea5a (devel-tiny, feectools#95) to d16a6ad, the head of struphy-hub/feectools#97 (stencil-6d-views, now stacked on feectools#96; against the base #703 the move is 64b10fa → d16a6ad). No struphy code changes: struphy does not call the stencil kernels directly, and the e_in argument that feectools#97 removes is only used inside StencilMatrix.

The pr-feectools-submodule check fails until struphy-hub/feectools#97 is merged into devel-tiny and the submodule points at devel-tiny again.

Model-specific changes:

None.

Documentation changes:

None in struphy. feectools' CUDA_STRATEGY.md has a new "6D matrix views" section.

🤖 Generated with Claude Code

Point the feectools submodule at struphy-hub/feectools#97 (branch
stencil-6d-views): the 3D stencil dot/transpose kernels take the matrix
data as 6D views, and StencilMatrix.dot/vdot/transpose no longer raise for
matrices with fewer than 2p+1 diagonals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…6d-views

Stack #703 -> #706: feectools submodule points to the head of
struphy-hub/feectools#97, which now contains #96.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack #703 -> #706 -> #716: feectools submodule points to the head of
struphy-hub/feectools#98, which now contains #96 and #97.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel to feectools-device-kronecker-solve October 7, 2026 21:59
@max-models max-models closed this Oct 8, 2026
spossann added a commit that referenced this pull request Oct 9, 2026
**Stack:** part 3 of 3 — based on #706 (`feectools-stencil-6d-views`,
merge that first, after #703), next: none (top of the stack).

Corresponding update in feectools:
struphy-hub/feectools#98. Mirrors the feectools
stack struphy-hub/feectools#96 →
struphy-hub/feectools#97 →
struphy-hub/feectools#98.
`feectools-stencil-6d-views` is merged into this branch (no rebase), and
the `feectools` submodule now points at the stacked head of
struphy-hub/feectools#98, `6d88806`, which
contains feectools#96 and #97. Against its base (#706) this PR only
moves the submodule `d16a6ad` → `6d88806`. After each feectools PR of
the stack is merged into `devel-tiny`, the `feectools` submodule of the
struphy stack must be moved to the corresponding merged `devel-tiny`
commit (the `pr-feectools-submodule` check fails until then).

---

**Solves the following issue(s):**

Part of #689 (`VlasovAmpereOneSpecies` end to end on the GPU, no
host/device transfers in the time loop), tracked in #650. This PR moves
the `feectools` submodule to struphy-hub/feectools#98 ("Inner products
stay on the device"). On the CuPy backend, every CG iteration of a
struphy solve copied two scalars from the device to the host; after this
change it copies one (the convergence test).

**The `pr-feectools-submodule` check fails until
struphy-hub/feectools#98 is merged** into `devel-tiny`. After that, the
submodule should point at the merge commit.

**Core changes:**

- `feectools` submodule: `d16a6ad` (base #706) → `6d88806` (branch
`inner-on-device` of struphy-hub/feectools#98, stacked on feectools#96
and #97; originally `a15a8e8`). No struphy code changes.
- What changes in feectools, and only on the CuPy backend:
- `StencilVector.inner`/`BlockVector.inner`/`dot_inner` return a 0-d
device array, in serial and with MPI.
  - `axpy` accepts such a scalar without copying it to the host.
- CG, PCG, BiCG, BiCGStab and PBiCGStab keep alpha/beta on the device
and copy only the residual norm, once per iteration.

  On NumPy, results and types are unchanged.

  Copies to the host per iteration, counted under cunumpy's fake CuPy:

  | Solver | before | after |
  | --- | --- | --- |
  | CG | 2 | 1 |
  | PCG | 3 | 1 |
  | BiCG | 4 | 1 |
  | BiCGStab | 6 | 1 |
  | PBiCGStab | 5 | 1 |

- struphy call sites of `.inner` that run on CuPy now get a 0-d device
array. I checked them:
- the scalars in `models/scalars.py` write it into `xp` buffers
(`local_value[0] = ...`), which works on the device;
- the model energy methods (`linear_mhd.py`, `shear_alfven.py`, ...)
return it, and the scalar machinery handles it like the existing
`dot_inner` results;
- `PolarVector.dot` adds a NumPy scalar to it, which works, but polar
splines are not on CuPy yet (#695);
- the multigrid smoothers run their own PCG with `.inner`. It works with
device scalars: the arithmetic stays on the device, and the comparisons
with 0 copy to the host implicitly.
- This removes the first blocker listed in #705 ("CG inner products ...
The Schur solve is outside the transfer guard because of this"). The
remaining copy per CG iteration still counts as a transfer for
`assert_no_transfers`. The transfer guard could then cover the Schur
solve with an allowance of one 8-byte copy per iteration.

**Model-specific changes:**

None.

**Documentation changes:**

None in struphy. The feectools PR documents the change in feectools'
`CUDA_STRATEGY.md` (section "Inner products on the device").

**Testing** (macOS, no GPU; struphy kernels compiled with GNU/Fortran;
**GPU tests not run**):

- feectools (see struphy-hub/feectools#98):
- serial suite: 9456 → 9480 passed, with the same 6 unrelated failures
in `ddm/tests/test_cart_*d.py`;
  - `mpirun -n 2` linalg MPI suite: 738 → 739 passed;
- the new fake-CuPy tests check the copy counts and that the results
equal NumPy's.
- struphy, solver tests on the NumPy backend, with feectools
`devel-tiny` (before) and this branch (after):
- `propagators/tests/test_poisson.py -m "not mpi" -k "not multigrid"`:
40 passed before, 40 passed after;
- `linear_algebra/tests/test_saddlepoint_massmatrices.py`: the first
case (`SaddlePointSolverUzawaNumpy`) passed before and after.

I stopped the remaining saddle-point and multigrid tests: with the
feectools kernels uncompiled in the checkouts, each test took more than
10 minutes. On NumPy, feectools' change is limited to identity helpers
and an equivalent loop structure in BiCGStab, and the feectools solver
tests check the same iteration counts and results.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant