Repository navigation
KroneckerStencilMatrix: grouped factors, ComposedKroneckerStencilMatrix, KroneckerSumSolver - #100
Merged
Merged
Conversation
- KroneckerStencilMatrix factors may act on several consecutive axes (e.g. 2d x 1d on a 3d space); new `axes` property, stricter checks (StencilMatrix factors, codomain npts, number of axes). - `A @ B` of two Kronecker matrices with the same axis groups returns a KroneckerStencilMatrix with factors C_k = A_k @ B_k (computed in sparse format on process-local spaces with wider pads, rows of B gathered across processes). Domain/codomain stay the original spaces; copies of the operands are kept in `factors` and used by `dot`. - dot/tostencil use the factor's own pads; tostencil raises ValueError if the band does not fit into the domain pads. - KroneckerLinearSolver and kronecker_solve take `factor_ndims` for solvers of factors with several axes (serial along grouped axes). - Fix __imul__ on the tuple of factors; docstrings and type annotations. - Bump version to 0.6.0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A @ B of Kronecker matrices with the same axis groups now returns a ComposedKroneckerStencilMatrix (subclass of ComposedLinearOperator): `multiplicands` are the operands (dot goes through them), `mats` the exact factors C_k of the product (process-local, wide band; used by tosparse and for KroneckerLinearSolver). Scaling, copy, transpose and further products keep the type. - KroneckerStencilMatrix loses the `factors` argument: its factor pads always fit into the domain pads, so dot and tostencil always work. - ComposedLinearOperator.multiplicants -> multiplicands (correct spelling); `multiplicants` remains as a deprecated alias. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Exact inverse of sum_d M_1 x ... x S_d x ... x M_n + sigma M_1 x ... x M_n
(e.g. a Laplacian on a tensor-product grid): per direction the generalized
eigenproblem S_d U_d = M_d U_d Lambda_d is solved, and A^{-1} = U Lambda^{-1} U^T
with U = U_1 x ... x U_n. U^T and U are applied with KroneckerLinearSolver
(also along distributed axes), Lambda^{-1} locally; vanishing eigenvalues are
skipped (pseudo-inverse). A direction without stiffness term is given as None.
Tests against dense solves, regular and singular, serial and with MPI.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rties Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
🟡 Changes recommended
Parallel products can discard valid periodic entries, and incompatible process-local factor layouts are not rejected.
2 open findings
What changed in this PR
Extends Kronecker linear algebra with grouped factors, composed products, and fast diagonalization.
Changes:
- Adds grouped-axis Kronecker matrices and composed products.
- Adds grouped-factor and Kronecker-sum solvers.
- Renames
multiplicantsand exposes derivative metadata.
| File | Description |
|---|---|
pyproject.toml |
Bumps version to 0.6.0. |
feectools/linalg/kron.py |
Implements the new Kronecker functionality. |
feectools/linalg/basic.py |
Renames composed-operator multiplicands. |
feectools/feec/derivatives.py |
Adds derivative properties. |
feectools/api/fem_common.py |
Uses the renamed property. |
feectools/api/fem_bilinear_form.py |
Uses the renamed property. |
feectools/linalg/tests/test_kron_stencil_matrix.py |
Adds extensive serial and MPI coverage. |
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
- KroneckerStencilMatrix: the rows owned by a factor must contain the rows of the codomain on this process (ValueError otherwise). dot and tostencil index the factor rows by global row, so both process-local factors and factors owning all rows (e.g. from tokronstencil) work in parallel; before, other rows were silently used. - Products: sum duplicate COO entries of the local factor (periodic factors with 2p + 1 > n have two diagonals in the same column) before gathering the rows of other processes; before, np.unique dropped one of them in parallel. - Tests for full-row factors, the row check and periodic duplicates (serial and MPI). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Member
Author
|
@max-models this is ready for review! |
spossann
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 9, 2026
…ass preconditioners; bump feectools (#731) ## Summary Points the `feectools` submodule to the branch of struphy-hub/feectools#100, so that the struphy test suite runs against it before that PR is merged. struphy-hub/feectools#100 changes `KroneckerStencilMatrix` and `KroneckerLinearSolver`: - factors on groups of axes (e.g. 2d x 1d); - `A @ B` returns a new `ComposedKroneckerStencilMatrix` (operands in `multiplicands`, exact Kronecker factors of the product in `mats`); - `ComposedLinearOperator.multiplicants` → `multiplicands`, with a deprecated alias; - `KroneckerSumSolver`, the exact inverse of sums of Kronecker products by fast diagonalization; - `factor_ndims` in `KroneckerLinearSolver`; - docstrings and type annotations; feectools 0.7.0. Existing calls are unchanged, so struphy needs no code changes. struphy builds these objects in `feec/preconditioner.py` (`MassMatrixPreconditioner`) and `feec/projectors.py`. ### `multiplicants` → `multiplicands` - All struphy uses (preconditioner, projectors, multigrid coarsening) now use `multiplicands`. - `linop_nbytes` in `feec/memory.py` adds up all child attributes, so a `ComposedKroneckerStencilMatrix` counts both its operands and the factors of the product. ### Refactor of `MassMatrixPreconditioner` - The construction of the Kronecker approximation is split into module-level helpers in `feec/preconditioner.py`. They cover the boundary conditions, gathering array weights over MPI, the 1d weight, the 1d mass matrix and its BCs, the 1d solver, the process-local 1d factor, the block-diagonal assembly and the operator to invert. - `MassMatrixDiagonalPreconditioner` now calls these helpers instead of repeating the same ~200 lines. - Values and the order of MPI calls are unchanged. The copy into the process-local 1d stencil matrix is now vectorized, and the block-diagonal operators are no longer hard-coded to 3 components. - Docstrings, comments and type annotations. ### `MassMatrixPreconditioner`: new option and fixes - New option `weight_reduction="midpoint" | "average"` sets how the weight is reduced to 1d in the directions other than `dim_reduce`. The default `"midpoint"` gives the same results as before, verified for M0, M1, M2, Mv, M1n and array weights, serially and on 2 ranks. `"average"` uses the mean over those directions: Gauss quadrature of the Derham for array weights, 16-point Gauss–Legendre for callables. - Array weights are reduced with a single `Allreduce`, instead of a subcommunicator plus `Bcast`. - The mass matrix is located in the composed operator by identity, and the approximate inverse is applied exactly there. Both preconditioners share this code (`_apply_composed`). Before, a mass matrix in the left-most position was applied instead of inverted. - `FFTSolver` copies its column, which was a view of the dense matrix also used for the stencil factor. It now stabilises a singular matrix once, at construction. `is_circulant` is vectorized. PCG iterations (tolerance 1e-8) on HollowTorus, grid 12×16×6, degree (2,3,2), as a reference for choosing `weight_reduction`: | | M0 | M1 | M2 | M3 | Mv | M1n | M2n | Mvn | |---|---|---|---|---|---|---|---|---| | midpoint | 13 | 22 | 22 | 13 | 31 | 673 | 600 | 1000 (not converged) | | average | 13 | 15 | 15 | 13 | 30 | 798 | 727 | 1000 (not converged) | On HollowCylinder both need 2 iterations for every matrix. ### Kronecker preconditioners: base class, stiffness preconditioner **`KroneckerPreconditioner`** is a new base class for preconditioners of the form `P = B E · S Ã⁻¹ S · Eᵀ Bᵀ`: - `Ã` (`matrix`) approximates the core operator, and its exact inverse is `solver`. Subclasses build the approximation. - `S` is an optional diagonal scaling `D̂^{1/2} D^{-1/2}` (Loli–Sangalli–Tani), with `D` the diagonal of the core operator and `D̂` that of the approximation. - The base class holds the common interface (`solve`, `dot`, the composition with `B E`, the properties). **Mass preconditioners** - `MassMatrixPreconditioner` is a subclass. It gets the options `diagonal_scaling` (default `False`) and `dim_reduce=None` (unit weights in all directions), and the method `update_mass_operator`. With `diagonal_scaling=True`, the PCG iterations for M1 on a HollowTorus drop from 20 to 8. - `MassMatrixDiagonalPreconditioner` is now `MassMatrixPreconditioner(dim_reduce=None, diagonal_scaling=True)`, kept for its name in the solver options. It no longer assembles a separate 3d logical mass matrix; `D̂` comes from the Kronecker approximation. Results are unchanged up to rounding (1e-13), checked on HollowTorus and IGAPolarCylinder, including `update_mass_operator`, serially and on 2 ranks. **`StiffnessPreconditioner`** (new) preconditions the stabilized stiffness operators `Gᵀ M1 G + σ M0`, `Cᵀ M2 C + σ M1` and `Dᵀ M3 D + σ M2`: - With Kronecker mass matrices, each diagonal block is a sum of Kronecker products of 1d stiffness and mass matrices. It is inverted exactly with `KroneckerSumSolver`, one solver per component (the components have different spline types, e.g. DNN/NDN/NND for M1). For grad, this is exact on the logical cube. - `curl` and `div` neglect the off-diagonal blocks (block Jacobi). Block Jacobi alone overestimates the operator on the kernel of the derivative, so a kernel correction `σ⁻¹ d₋ P₋ d₋ᵀ` is added (default, needs `σ > 0`): with `d₋ = G` and the grad preconditioner for curl, and `d₋ = C` and the curl block Jacobi for div. - Weights: `weights="average"` (default) approximates the mass-matrix weights by one common separable shape per direction plus a mean per term, which keeps the fast diagonalization exact. `weights="unit"` uses the logical cube. Diagonal scaling (default on) adds the geometry pointwise; for it, the diagonal of the stiffness operator is computed by probing. - Polar splines are not supported yet. PCG iterations (tolerance 1e-8), HollowTorus, grid 12×16×6, degree (2,3,2): | | no preconditioner | mass preconditioner | `StiffnessPreconditioner` (defaults) | |---|---|---|---| | grad, σ = 0 | 187 | – | 41 | | curl, σ = 1 | > 3000 | 716 | 46 | | div, σ = 1 | > 3000 | 636 | 48 | On a stretched cuboid (constant, anisotropic weights), grad needs 2 iterations. For large σ (1e4), the mass preconditioner is as good or slightly better (about 13 vs 25 iterations). ## Tests Run locally with the submodule branch: all preconditioner tests (`test_preconditioner_transpose.py`, the new `test_stiffness_preconditioner.py`, `test_mass_matrices.py::test_mass_preconditioner*`, `test_reduced_weight_1d`) pass serially and on 2 ranks (52 passed each). The new stiffness tests cover the approximation against the operator on the unit cube, PCG iteration counts for grad/curl/div, the kernel correction, `weights="average"`, and independence of the MPI decomposition. The rest is left to CI. Don't raise the `feectools` pin in `pyproject.toml` until 0.7.0 is on PyPI. Before merging, the submodule pointer should move to the merged `devel-tiny` commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Improves
KroneckerStencilMatrixandKroneckerLinearSolverinfeectools/linalg/kron.py.Factors on groups of axes. A factor of a
KroneckerStencilMatrixmay now act on several consecutive axes, e.g. a 2d x 1d product on a 3d space:axesproperty.ndimis now the number of axes of the domain; before, it was the number of factors, which is the same for 1d factors.StencilMatrix, the codomainnptsare checked as well, and the factors must cover exactlyV.ndimaxes.dot,tostenciland__getitem__work with grouped factors.dotandtostencilnow use each factor's own pads and row offsets, so a factor's pads may be smaller than the pads of the space.Matrix product
A @ B→ComposedKroneckerStencilMatrix. For two Kronecker matrices with the same axis groups,A @ Breturns a newComposedKroneckerStencilMatrix. This mirrorsLinearOperator @ LinearOperator → ComposedLinearOperator, and the new class subclassesComposedLinearOperator:multiplicandsare the operands. Chains are flattened on both sides, so(A @ B) @ CandA @ (B @ C)both have(A, B, C).dotapplies them from right to left, as inComposedLinearOperator.matsholds the exact Kronecker factorsC_k = A_k @ B_kof the product. They are computed once at construction in scipy sparse format and stored on process-local spaces with wider pads (p_A + p_B). If the operands are distributed, the rows ofB_kowned by other processes are gathered, so construction is collective.tosparseandtoarrayuse the exactC_k, which is also correct in parallel. AKroneckerLinearSolverfor the product can be built frommats.copy,transposeand further products with Kronecker matrices keep the type. Any other operand (another operator type, other axis groups) gives a plainComposedLinearOperator.matsalways means the Kronecker factors (A₁ ⊗ A₂ ⊗ …), andmultiplicandsthe operands of the matrix product (A · B · …).KroneckerStencilMatrixitself takes no extra argument for products. Its factor pads must always fit into the domain pads, sodotandtostencilalways work.multiplicants→multiplicands.ComposedLinearOperator.multiplicantsis renamed tomultiplicands, the correct spelling.multiplicantsremains as a deprecated alias that emits aDeprecationWarning. All uses in feectools (api/fem_common.py,api/fem_bilinear_form.py) are updated.KroneckerLinearSolver/kronecker_solve. New optional argumentfactor_ndims(default: all 1, so existing behaviour is unchanged). A solver of a factor with several axes receives the vectors flattened over these axes in C order, as inStencilMatrix.tosparse. The axes of such a factor must not be distributed across processes; otherwiseNotImplementedErroris raised. 1d factors are solved in parallel as before.KroneckerSumSolver(fast diagonalization). A new solver for sums of Kronecker products,with symmetric 1d matrices
S_dand symmetric positive definite 1d matricesM_d, e.g. a Laplacian on a tensor-product grid. Such a sum is not a single Kronecker product, soKroneckerLinearSolvercan't invert it.S_d U_d = M_d U_d Λ_dis solved (withU_dᵀ M_d U_d = I). ThenA⁻¹ = U Λ⁻¹ Uᵀ, withU = U_1 ⊗ … ⊗ U_nand the diagonalΛ = Λ_1 ⊕ … ⊕ Λ_n + σ.UᵀandUare applied withKroneckerLinearSolver, through a smallLinearSolverthat multiplies by a dense matrix, as infft.py. So it also works along distributed axes.Λ⁻¹is applied locally.σ = 0.None.It is used by the new
StiffnessPreconditionerin struphy (see the paired PR).DirectionalDerivativeOperatorgets the public propertiesdiffdir,negativeandtransposed.Other
KroneckerStencilMatrix.__imul__, which assigned into a tuple and raisedTypeError.kronecker_solve.devel-tiny, which is merged into this branch, is already at 0.6.0).compile_psydac.mk:pyccel compilewithout-v, for a shorter compile output.Tests
New tests in
linalg/tests/test_kron_stencil_matrix.pycompare against dense matrices for the groupings (1,1,1), (2,1), (1,2) and (3,). They cover:dot,tosparse,__getitem__, scaling andtostencilfor grouped factors;@: type, values,multiplicands, chains on both sides, scaling, copy, transpose, the fallback toComposedLinearOperator, and theValueErrorcase;factor_ndims, including a solver for a product;multiplicantsalias;KroneckerSumSolveragainst dense solves, regular and singular (pseudo-inverse), with and without a direction that has no stiffness term.Run locally:
mpirun -n 2and-n 4: all Kronecker andtest_fftMPI tests pass, including distributed axes forKroneckerSumSolver;DeprecationWarnings turned into errors: 9101 passed.The four collection errors in
ddm/tests/test_cart_2d.py,test_cart_3d.py,linalg/tests/test_block.pyandtest_toarray.py(unregisteredparallelmarker) also occur without this change.Transposes are tested in serial only: in parallel, the process-local factors don't hold the rows of other processes. This was already the case before this PR.
Paired struphy PR (runs the struphy tests against this branch): struphy-hub/struphy#731
🤖 Generated with Claude Code