Skip to content

Set up sorting boxes on the CuPy backend - #736

Open
max-models wants to merge 1 commit into
develfrom
726-sorting-boxes-cupy
Open

max-models wants to merge 1 commit into
develfrom
726-sorting-boxes-cupy

Conversation

@max-models

Copy link
Copy Markdown
Member

Closes #726

What changed

  • SortingBoxes._set_boxes (pic/sorting.py): n_cols is computed with math.sqrt (host bookkeeping). _neighbours is built on the host with the pyccel initialize_neighbours and moved to the active backend once with xp.to_cunumpy (setup data).
  • Particles setup with boxes (pic/base.py):
    • n_boxes uses math.prod instead of xp.prod(xp.array(...)).
    • is_domain_boundary holds Python bools read from a host copy of domain_array.
    • _get_neighbouring_proc (SPH only) reads a host copy of domain_array. Before this, the pyccel distance got 0-d device arrays and ParticlesSPH setup failed on CuPy.
  • put_particles_in_boxes / do_sort: assign_box_to_each_particle, assign_particles_to_boxes and sort_boxed_particles are wrapped in cunumpy.kernels.PyccelKernel with declared outputs. On NumPy they run the kernel directly, as before. On CuPy this is an explicit host fallback: every call copies device → host → device.

Time loop of non-SPH models

For a non-SPH model (e.g. VlasovAmpereOneSpecies), setting boxes_per_dim alone does not sort anything during the time loop. Pushers only call put_particles_in_boxes when sorting_boxes.communicate is true, which is SPH only. Box sorting runs only:

  • during setup, via draw_markers(sort=True), when SortingParameters(do_sort=True);
  • every sort_step steps in sim.py, via do_sort(), when EnvironmentOptions.sort_step > 0.

Both paths now work on CuPy through the host fallback above. The fallback still has to be ported (#694).

Remaining for #694

  • Device-side box sorting: CUDA versions of the three sorting kernels, to replace the PyccelKernel host round trip.
  • _check_and_assign_particles_to_boxes still compares a device max_in_box in a Python if, which forces an implicit sync.
  • SPH evaluation and ghost-box communication on the device. I checked that ParticlesSPH setup works on CuPy. Drawing markers with a domain needs CUDA launches, which the fake CuPy in cunumpy 0.6.1 cannot run, so nothing after setup was exercised for SPH.

Tests

New check_sorting_boxes_on_cupy in pic/tests/test_sorting.py:

  • creates Particles6D(boxes_per_dim=(4, 3, 2)) on NumPy and on CuPy, and checks that neighbours, boxes, next_index and cumul_next_index match;
  • runs do_sort() on the same markers on both backends and checks that the markers and box arrays match;
  • creates ParticlesSPH with boxes and checks that is_domain_boundary (Python bools), the boundary boxes, the neighbouring procs and neighbours match NumPy.

It runs in two tests:

  • test_sorting_boxes_on_cupy uses the real CuPy and is skipped without a GPU.
  • test_sorting_boxes_on_cupy_fake_cupy uses cunumpy's fake CuPy in a serial subprocess with CUNUMPY_FAKE_CUPY=1, on rank 0 only. It follows the pattern of test_cuda_args_domain_fake_cupy.

Results:

  • Before the fix, the fake-CuPy test fails with TypeError: type ndarray doesn't define __round__ method. With only the setup fix, it fails with argument must be numpy.ndarray in the sorting kernels.
  • pytest src/struphy/pic/tests/test_sorting.py: 53 passed, 3 skipped (MPI-only tests and the GPU test).
  • mpirun -n 2 pytest --with-mpi src/struphy/pic/tests/test_sorting.py: 55 passed, 1 skipped (the GPU test).

Merge order

Based on devel. It does not depend on the unmerged CUDA PRs: on devel, Particles can already be created on the fake CuPy backend.

🤖 Generated with Claude Code

SortingBoxes could not be created on CuPy: n_cols used xp.sqrt (a 0-d
device array that round() rejects) and the pyccel initialize_neighbours
was called with a device array.

- SortingBoxes._set_boxes: n_cols with math.sqrt; neighbours built on
  the host with initialize_neighbours and moved to the backend once.
- Particles: n_boxes with math.prod; is_domain_boundary and
  _get_neighbouring_proc (SPH) read a host copy of domain_array.
- put_particles_in_boxes / do_sort: the pyccel box-sorting kernels are
  wrapped in PyccelKernel, an explicit host fallback on CuPy (device ->
  host -> device per call) until #694 ports them.
- Test: Particles6D and ParticlesSPH with boxes_per_dim on CuPy (real or
  fake CuPy in a serial subprocess on rank 0) match NumPy, including one
  do_sort.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SortingBoxes cannot be created on the CuPy backend

1 participant