Skip to content

feectools: Kronecker solver on the device - #703

Closed
max-models wants to merge 2 commits into
develfrom
feectools-device-kronecker-solve
Closed

max-models wants to merge 2 commits into
develfrom
feectools-device-kronecker-solve

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 1 of 3 — based on devel, next: #706 (feectools-stencil-6d-views), then #716. Merge this one first.

Mirrors the feectools stack struphy-hub/feectools#96 → struphy-hub/feectools#97 → struphy-hub/feectools#98. The feectools submodule here points at the head of struphy-hub/feectools#96 (64b10fa, part 1 of that stack). After each feectools PR of the stack is merged into devel-tiny, the feectools submodule of the struphy stack must be moved to the corresponding merged devel-tiny commit (the pr-feectools-submodule check fails until then).


Solves the following issue(s):

Corresponding update in feectools: struphy-hub/feectools#96

Part of #650 and #689. The MassMatrixPreconditioner (for example the preconditioner of EfieldWeightsCoupling in LinearVlasovAmpereOneSpecies) applies its inverse with feectools' KroneckerLinearSolver. On the CuPy backend, every one of those solves, in every CG iteration, copied the data to the host and back, because the 1D solvers use LAPACK/SuperLU. struphy-hub/feectools#96 removes those copies: the 1D matrices are inverted once on the host, and device data is solved with one matrix product per direction on the device.

Core changes:

Model-specific changes:

None.

Documentation changes:

None in struphy. feectools' CUDA_STRATEGY.md has the notes for the change.

Testing: see struphy-hub/feectools#96 for the feectools tests. The GPU tests there are not run yet (no GPU here). struphy's test suite runs in CI.

Not covered here: from reading the code, MassMatrixPreconditioner may not get as far as the solve on the CuPy backend. is_circulant and FFTSolver.__init__ assert isinstance(M_arr, xp.ndarray), but M_arr = M.toarray() is a NumPy array. I have not run this on a GPU. If it fails, it needs a struphy-side fix.

🤖 Generated with Claude Code

Moves the feectools submodule to struphy-hub/feectools#96 (branch
device-kronecker-solve): KroneckerLinearSolver solves device data with
dense inverses of the 1D matrices (one GEMM per direction) instead of
copying it to the host for LAPACK/SuperLU, so MassMatrixPreconditioner
solves on the CuPy backend make no host/device transfers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
…6d-views

Stack #703 -> #706: feectools submodule points to the head of
struphy-hub/feectools#97, which now contains #96.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack #703 -> #706 -> #716: feectools submodule points to the head of
struphy-hub/feectools#98, which now contains #96 and #97.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models max-models closed this Oct 8, 2026
spossann added a commit that referenced this pull request Oct 9, 2026
**Stack:** part 3 of 3 — based on #706 (`feectools-stencil-6d-views`,
merge that first, after #703), next: none (top of the stack).

Corresponding update in feectools:
struphy-hub/feectools#98. Mirrors the feectools
stack struphy-hub/feectools#96 →
struphy-hub/feectools#97 →
struphy-hub/feectools#98.
`feectools-stencil-6d-views` is merged into this branch (no rebase), and
the `feectools` submodule now points at the stacked head of
struphy-hub/feectools#98, `6d88806`, which
contains feectools#96 and #97. Against its base (#706) this PR only
moves the submodule `d16a6ad` → `6d88806`. After each feectools PR of
the stack is merged into `devel-tiny`, the `feectools` submodule of the
struphy stack must be moved to the corresponding merged `devel-tiny`
commit (the `pr-feectools-submodule` check fails until then).

---

**Solves the following issue(s):**

Part of #689 (`VlasovAmpereOneSpecies` end to end on the GPU, no
host/device transfers in the time loop), tracked in #650. This PR moves
the `feectools` submodule to struphy-hub/feectools#98 ("Inner products
stay on the device"). On the CuPy backend, every CG iteration of a
struphy solve copied two scalars from the device to the host; after this
change it copies one (the convergence test).

**The `pr-feectools-submodule` check fails until
struphy-hub/feectools#98 is merged** into `devel-tiny`. After that, the
submodule should point at the merge commit.

**Core changes:**

- `feectools` submodule: `d16a6ad` (base #706) → `6d88806` (branch
`inner-on-device` of struphy-hub/feectools#98, stacked on feectools#96
and #97; originally `a15a8e8`). No struphy code changes.
- What changes in feectools, and only on the CuPy backend:
- `StencilVector.inner`/`BlockVector.inner`/`dot_inner` return a 0-d
device array, in serial and with MPI.
  - `axpy` accepts such a scalar without copying it to the host.
- CG, PCG, BiCG, BiCGStab and PBiCGStab keep alpha/beta on the device
and copy only the residual norm, once per iteration.

  On NumPy, results and types are unchanged.

  Copies to the host per iteration, counted under cunumpy's fake CuPy:

  | Solver | before | after |
  | --- | --- | --- |
  | CG | 2 | 1 |
  | PCG | 3 | 1 |
  | BiCG | 4 | 1 |
  | BiCGStab | 6 | 1 |
  | PBiCGStab | 5 | 1 |

- struphy call sites of `.inner` that run on CuPy now get a 0-d device
array. I checked them:
- the scalars in `models/scalars.py` write it into `xp` buffers
(`local_value[0] = ...`), which works on the device;
- the model energy methods (`linear_mhd.py`, `shear_alfven.py`, ...)
return it, and the scalar machinery handles it like the existing
`dot_inner` results;
- `PolarVector.dot` adds a NumPy scalar to it, which works, but polar
splines are not on CuPy yet (#695);
- the multigrid smoothers run their own PCG with `.inner`. It works with
device scalars: the arithmetic stays on the device, and the comparisons
with 0 copy to the host implicitly.
- This removes the first blocker listed in #705 ("CG inner products ...
The Schur solve is outside the transfer guard because of this"). The
remaining copy per CG iteration still counts as a transfer for
`assert_no_transfers`. The transfer guard could then cover the Schur
solve with an allowance of one 8-byte copy per iteration.

**Model-specific changes:**

None.

**Documentation changes:**

None in struphy. The feectools PR documents the change in feectools'
`CUDA_STRATEGY.md` (section "Inner products on the device").

**Testing** (macOS, no GPU; struphy kernels compiled with GNU/Fortran;
**GPU tests not run**):

- feectools (see struphy-hub/feectools#98):
- serial suite: 9456 → 9480 passed, with the same 6 unrelated failures
in `ddm/tests/test_cart_*d.py`;
  - `mpirun -n 2` linalg MPI suite: 738 → 739 passed;
- the new fake-CuPy tests check the copy counts and that the results
equal NumPy's.
- struphy, solver tests on the NumPy backend, with feectools
`devel-tiny` (before) and this branch (after):
- `propagators/tests/test_poisson.py -m "not mpi" -k "not multigrid"`:
40 passed before, 40 passed after;
- `linear_algebra/tests/test_saddlepoint_massmatrices.py`: the first
case (`SaddlePointSolverUzawaNumpy`) passed before and after.

I stopped the remaining saddle-point and multigrid tests: with the
feectools kernels uncompiled in the checkouts, each test took more than
10 minutes. On NumPy, feectools' change is limited to identity helpers
and an equivalent loop structure in BiCGStab, and the feectools solver
tests check the same iteration counts and results.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant