Repository navigation
Set up sorting boxes on the CuPy backend - #736
Open
max-models wants to merge 1 commit into
Open
max-models wants to merge 1 commit into
max-models wants to merge 1 commit into
Conversation
SortingBoxes could not be created on CuPy: n_cols used xp.sqrt (a 0-d device array that round() rejects) and the pyccel initialize_neighbours was called with a device array. - SortingBoxes._set_boxes: n_cols with math.sqrt; neighbours built on the host with initialize_neighbours and moved to the backend once. - Particles: n_boxes with math.prod; is_domain_boundary and _get_neighbouring_proc (SPH) read a host copy of domain_array. - put_particles_in_boxes / do_sort: the pyccel box-sorting kernels are wrapped in PyccelKernel, an explicit host fallback on CuPy (device -> host -> device per call) until #694 ports them. - Test: Particles6D and ParticlesSPH with boxes_per_dim on CuPy (real or fake CuPy in a serial subprocess on rank 0) match NumPy, including one do_sort. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #726
What changed
SortingBoxes._set_boxes(pic/sorting.py):n_colsis computed withmath.sqrt(host bookkeeping)._neighboursis built on the host with the pyccelinitialize_neighboursand moved to the active backend once withxp.to_cunumpy(setup data).Particlessetup with boxes (pic/base.py):n_boxesusesmath.prodinstead ofxp.prod(xp.array(...)).is_domain_boundaryholds Python bools read from a host copy ofdomain_array._get_neighbouring_proc(SPH only) reads a host copy ofdomain_array. Before this, the pycceldistancegot 0-d device arrays andParticlesSPHsetup failed on CuPy.put_particles_in_boxes/do_sort:assign_box_to_each_particle,assign_particles_to_boxesandsort_boxed_particlesare wrapped incunumpy.kernels.PyccelKernelwith declaredoutputs. On NumPy they run the kernel directly, as before. On CuPy this is an explicit host fallback: every call copies device → host → device.Time loop of non-SPH models
For a non-SPH model (e.g.
VlasovAmpereOneSpecies), settingboxes_per_dimalone does not sort anything during the time loop. Pushers only callput_particles_in_boxeswhensorting_boxes.communicateis true, which is SPH only. Box sorting runs only:draw_markers(sort=True), whenSortingParameters(do_sort=True);sort_stepsteps insim.py, viado_sort(), whenEnvironmentOptions.sort_step > 0.Both paths now work on CuPy through the host fallback above. The fallback still has to be ported (#694).
Remaining for #694
PyccelKernelhost round trip._check_and_assign_particles_to_boxesstill compares a devicemax_in_boxin a Pythonif, which forces an implicit sync.ParticlesSPHsetup works on CuPy. Drawing markers with a domain needs CUDA launches, which the fake CuPy in cunumpy 0.6.1 cannot run, so nothing after setup was exercised for SPH.Tests
New
check_sorting_boxes_on_cupyinpic/tests/test_sorting.py:Particles6D(boxes_per_dim=(4, 3, 2))on NumPy and on CuPy, and checks thatneighbours,boxes,next_indexandcumul_next_indexmatch;do_sort()on the same markers on both backends and checks that the markers and box arrays match;ParticlesSPHwith boxes and checks thatis_domain_boundary(Python bools), the boundary boxes, the neighbouring procs andneighboursmatch NumPy.It runs in two tests:
test_sorting_boxes_on_cupyuses the real CuPy and is skipped without a GPU.test_sorting_boxes_on_cupy_fake_cupyuses cunumpy's fake CuPy in a serial subprocess withCUNUMPY_FAKE_CUPY=1, on rank 0 only. It follows the pattern oftest_cuda_args_domain_fake_cupy.Results:
TypeError: type ndarray doesn't define __round__ method. With only the setup fix, it fails withargument must be numpy.ndarrayin the sorting kernels.pytest src/struphy/pic/tests/test_sorting.py: 53 passed, 3 skipped (MPI-only tests and the GPU test).mpirun -n 2 pytest --with-mpi src/struphy/pic/tests/test_sorting.py: 55 passed, 1 skipped (the GPU test).Merge order
Based on
devel. It does not depend on the unmerged CUDA PRs: on devel,Particlescan already be created on the fake CuPy backend.🤖 Generated with Claude Code