Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Added
- `kernel_testing.emulated_launches()` runs every `CudaKernel` launch in a block
on the CPU, on the host buffers of the fake CuPy arrays, so code that launches
kernels can be tested without a GPU. `emulate_cuda_kernel` now accepts struct
parameters (a mapping of field values, a `CudaStructValue`, or an object with
an attribute per field). `kernel_testing.host_buffer(array)` returns the NumPy
array behind a fake CuPy array.
- `CUNUMPY_REQUIRE_CUDA=1` makes the GPU markers of `kernel_testing` fail instead
of skipping: `requires_cupy`, the `cupy` run of the `backend` fixture (which
activates CuPy strictly) and `assert_kernels_agree`. New `kernel_testing.cuda_required()`.
- `count_transfers()` records `sync` events (the host waiting for the device):
`xp.synchronize()`, the waits of the MPI helpers and of the CUDA debug mode, and,
on the fake CuPy, scalar reads of device arrays. They are in `counter.syncs`
and the report but not in `total`; `assert_no_transfers(syncs=True)` rejects them.
The real CuPy's own `float(a)` cannot be observed from Python and is not counted.
- `xp.algorithms.compact_by_mask(mask, *arrays)` moves the masked rows of arrays to
the front, in place and in order, and returns their number.
- Kernel outputs: a name in `PyccelKernel(outputs=...)` also finds a positional
argument and an index a keyword argument, using the parameter names of the
function (or the new `parameters=`); a `Kernel` supplies those of its host
function. `xp.kernels.outputs_from_annotations()` reads the outputs from the
annotations (not `Final`, `const` or a scalar), and `Kernel.from_folder()` /
`KernelCatalog.from_package()` take `outputs=` (names, indices or
`"annotations"`). `assert_kernels_agree(outputs=...)` takes parameter names and
`"name.field"` to compare only some fields of a struct argument.
- `xp.kernels.MetalKernel` runs a Metal Shading Language kernel on the GPU of an
Apple silicon Mac through MLX (`pip install 'cunumpy[metal]'`). It takes and
fills NumPy arrays, is float32 only (`float64="cast"` computes float64 data in
Expand Down
51 changes: 41 additions & 10 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ The submodules are named so that they do not hide a NumPy name (`rng`, not
| `cunumpy.arguments` | CUDA only | `CudaArguments`, `CudaStruct`, `CudaStructArguments`, `CudaStructValue`, `write_cuda_header` |
| `cunumpy.cuda` | CUDA only | device selection and memory, `stream`, streams/events, `pin_memory`, debug mode, CUDA headers and source tools (`cuda_include_dir`, `parse_cuda_signature`) |
| `cunumpy.rng` | both | `random_streams`, `get_rng`, `philox_*` |
| `cunumpy.algorithms` | both | `morton_*`, `sort_by_key`, `segment_sum` |
| `cunumpy.algorithms` | both | `morton_*`, `sort_by_key`, `segment_sum`, `compact_by_mask` |
| `cunumpy.mpi` | both | `mpi_buffer`, CUDA-aware MPI detection, `local_rank`, `synchronize_for_mpi` |
| `cunumpy.profiling` | both | `timed_region`, `nvtx_range`, `count_transfers`, `assert_no_transfers` |
| `cunumpy.memory` | both | `HostStaging`, `DeviceMirror` |
Expand Down Expand Up @@ -267,6 +267,21 @@ keys, order, positions, charges = xp.algorithms.sort_by_key(keys, positions, cha
Returns `(keys[order], order, *(a[order] for a in arrays))`, `order` as
`int64`. Equal keys keep their order, so the result is reproducible.

### `algorithms.compact_by_mask(mask, *arrays)`

Moves the rows where the boolean `mask` is True to the front of every array, in
place and in their original order, and returns how many there are. Typical use:
keep the live particles at the front of the marker arrays.

```python
n = xp.algorithms.compact_by_mask(alive, markers, weights)
markers, weights = markers[:n], weights[:n]
```

The rows after the first `n` are unspecified. The count is needed on the host,
so on CuPy each call synchronizes once. The mask and the arrays must be on the
same backend.

## Count transfers

A transfer inside a time loop is the classic performance bug of a GPU port:
Expand All @@ -286,7 +301,7 @@ with xp.profiling.count_transfers() as counter:
assert counter.total == 0, counter.report()
```

Five kinds of events are recorded:
Six kinds of events are recorded:

* `to_host`: an actual device-to-host copy through conversion, mirror, staging,
serial MPI, or host kernel helpers;
Expand All @@ -298,7 +313,12 @@ Five kinds of events are recorded:
* `fallback`: a `Kernel` without CUDA kernel calling its host kernel on the
CuPy backend (`missing_cuda="fallback"`), one event per call, naming the
kernel. Physical copies and a `kernel_conversion` marker are recorded separately;
* `device_copy`: a device-only dtype/layout conversion through CuNumpy helpers.
* `device_copy`: a device-only dtype/layout conversion through CuNumpy helpers;
* `sync`: the host waited for the device: `xp.synchronize()`, the waits of the MPI
helpers (`synchronize_for_mpi()`, `mpi_buffer()` staging) and of the CUDA debug
mode, and, on the fake CuPy, a scalar read of a device array (`float(a)`,
`int(a)`, `bool(a)`, `a.item()`, `a.tolist()`). They are in `counter.syncs` and in
the report, but not in `total`.

Only real transfers count: `to_numpy()` of a NumPy array or `to_cupy()` of a
CuPy array records nothing. The counter has the attributes `to_host`,
Expand Down Expand Up @@ -327,16 +347,17 @@ not thread-safe.

**Limitation:** only transfers made through CuNumpy are seen. Raw
`cupy.ndarray.get()`, `cupy.asarray(numpy_array)`, `numpy.asarray(cupy_array)`,
`float(device_array)`, forwarded backend calls such as `xp.asarray()`, and
implicit conversions inside other libraries are
not counted. Use `nsys` (or CuPy's profiling hooks) to find those.
forwarded backend calls such as `xp.asarray()`, and implicit conversions inside
other libraries are not counted. Neither is `float(device_array)` (an implicit
sync) with the real CuPy, which cannot be observed from Python; the fake CuPy
counts it. Use `nsys` (or CuPy's profiling hooks) to find those.

### `profiling.assert_no_transfers()`
### `profiling.assert_no_transfers(*, syncs=False)`

Context manager that raises `AssertionError` with the counter's `report()` if
the block makes a host/device transfer or host fallback through CuNumpy.
Device-only dtype/layout conversions are allowed. It yields the `TransferCounter`
too. An exception raised inside the block propagates as it is:
Device-only dtype/layout conversions are allowed, and so are syncs unless
`syncs=True`. It yields the `TransferCounter` too. An exception raised inside the block propagates as it is:

```python
def test_time_step_stays_on_the_device():
Expand Down Expand Up @@ -808,8 +829,18 @@ CuPy arrays.
* `is_array`: predicate for host array values to convert back to CuPy. The
default is `isinstance(value, numpy.ndarray)`.
* `outputs`: sequence of arguments the kernel may write to. Entries are
positional indices or keyword names. If omitted, every converted array is
indices or parameter names; a name also finds a positional argument, and an
index a keyword argument, when the parameter names are known (from the
Python signature, or `parameters`). If omitted, every converted array is
copied back.
* `parameters`: the names of the positional parameters, or a function that
returns them, for compiled kernels without a Python signature. A `Kernel`
supplies the names of its host function.

`kernels.outputs_from_annotations(function)` returns the parameters a kernel may
write to, read from its annotations: everything that is not `Final`, `const` or a
scalar. Use it as `outputs=outputs_from_annotations(push)`, or pass
`outputs="annotations"` to `Kernel.from_folder()` / `KernelCatalog.from_package()`.

### Example and output declarations

Expand Down
24 changes: 24 additions & 0 deletions docs/source/guides/data-movement.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,3 +165,27 @@ with xp.cuda.stream():
Pinned memory is a limited system resource; use it for large, repeatedly
transferred buffers after a profile shows transfers matter.
`pin_memory()` requires CuPy.

## Host-only code: `host_call`

Some code can only run on the host: a SciPy spline, a file reader, an external
equilibrium code. `xp.host_call(fun, *args, **kwargs)` calls it with arguments of
either backend. Device arrays are copied to the host, `fun` runs on the NumPy
backend, and array results are copied back, once per call. `@xp.evaluate_on_host`
does the same for a method, and `@xp.setup_on_host` runs an `__init__` on the NumPy
backend, so the object holds only host data. The copies are counted by
`count_transfers()`. Use it for setup and diagnostics, not in a time loop.

```python
values = xp.host_call(spline, x) # x on the device -> values on the device


class Equilibrium:
@xp.setup_on_host
def __init__(self, path):
self.spline = read_spline(path) # NumPy and SciPy

@xp.evaluate_on_host
def pressure(self, x):
return self.spline(x)
```
15 changes: 14 additions & 1 deletion docs/source/guides/execution-helpers.md
Original file line number Diff line number Diff line change
Expand Up @@ -142,10 +142,23 @@ per device; catalog setup catches CUDA compiler failures before timesteps.
checks actual block/grid, kernel thread, and static-plus-dynamic shared-memory
limits. These execution checks are independent of scope-profiler.

GPU CI requires a real CUDA device (`CUNUMPY_REQUIRE_CUDA=1`) and runs focused
GPU CI requires a real CUDA device (`CUNUMPY_REQUIRE_CUDA=1`; the GPU markers of
`cunumpy.kernel_testing` then fail instead of skipping) and runs focused
`memcheck`, `racecheck`, and `synccheck` jobs with nonzero sanitizer error exits.
Numerical parity tests are separate from performance comparisons.

## Keep the live particles: `compact_by_mask`

```python
n = xp.algorithms.compact_by_mask(alive, markers, weights)
markers, weights = markers[:n], weights[:n]
```

Moves the rows where the boolean mask is True to the front of every array, in
place and in order, and returns their number. The rows after the first `n` are
unspecified. The count is needed on the host, so on CuPy each call
synchronizes once.

## Prepare cell ranges and segment reductions

```python
Expand Down
14 changes: 10 additions & 4 deletions docs/source/kernels/pyccel-kernel.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,10 +50,16 @@ xp.kernels.PyccelKernel(solve, outputs=("out",)) # solve(a, b, out=out)
xp.kernels.PyccelKernel(norm, outputs=()) # writes nothing
```

* Positional arguments are declared by index (negative indices count from the
end), keyword arguments by name. The two forms are not interchangeable,
because compiled functions usually do not expose a Python signature that
would map names to positions.
* Arguments are declared by index (negative indices count from the end) or by
name. A name finds the argument also when it is passed positionally, and an
index also when it is passed as a keyword, as long as the parameter names are
known: from the Python signature, from `parameters=[...]`, or, for a kernel in a
`Kernel`, from its host function. For a compiled function without any of those
the two forms are not interchangeable.
* `xp.kernels.outputs_from_annotations(function)` reads the outputs from the
annotations: every parameter that is not `Final`, `const` or a scalar. A
`Kernel.from_folder(..., outputs="annotations")` (and `KernelCatalog.from_package`)
applies it to the host function of each kernel folder.
* A container or object declared as output has all its arrays copied back.
* **A missing declaration is a silent bug**: if the function writes an
argument that is not declared, the device array keeps its old values. When in
Expand Down
54 changes: 52 additions & 2 deletions docs/source/kernels/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,8 +95,11 @@ Things to know:
* **Build random data on the host.** NumPy and CuPy generators produce
different sequences from the same seed, so use `numpy.random.default_rng` and
convert with `to_cunumpy()`, as above.
* **Which arguments are compared**: `outputs=(2,)` selects them by index;
otherwise the host kernel's declared `outputs` are used, and if there are
* **Which arguments are compared**: `outputs=(2,)` selects them by index and
`outputs=("markers",)` by parameter name (the names of the host function).
`"markers.positions"` compares only that field of a struct or argument
object, leaving out fields the two kernels fill differently (scratch buffers,
for instance); otherwise the host kernel's declared `outputs` are used, and if there are
none, every array argument. Arrays held by argument objects (one level deep,
e.g. a `CudaArguments` object or a list) are compared too. A
`CudaStructArguments` object is read through its struct fields, so its
Expand Down Expand Up @@ -216,6 +219,40 @@ intrinsics; kernels using the latter are refused with `NotImplementedError`, so
those still need a GPU run. The compiler may fuse multiply-adds as NVRTC does,
so compare with a tolerance of a few ulp.

### Struct parameters and `emulated_launches()`

A kernel with struct parameters takes, for each struct, a dictionary of field
names to values, a `CudaStructValue`, or any object with an attribute per field
(a host argument class, a `CudaStructArguments`); the arrays in it are updated in
place:

```python
emulate_cuda_kernel(
push,
{"markers": markers, "alive": alive, "n": 4}, # the struct Particles
0.5,
total,
n_threads=4,
)
```

To test code that *launches* kernels (a `Kernel` on the CuPy backend, a
propagator), wrap it in `emulated_launches()`. With the fake CuPy (below) every
`CudaKernel` launch in the block then runs through the emulation on the host
buffers of the fake arrays, so the device arrays hold the results:

```python
from cunumpy.kernel_testing import emulated_launches, host_buffer

with emulated_launches():
propagator(dt) # CuPy backend = the fake CuPy
np.testing.assert_allclose(host_buffer(markers), expected)
```

`host_buffer(array)` is the NumPy array behind a fake CuPy array (not a copy).
Launches are serial and compile the kernel as C++ every time, so use small
problems.

## Test `__device__` helpers: `device_function_kernel`

Helpers such as B-spline evaluation or coordinate maps are `__device__`
Expand Down Expand Up @@ -314,6 +351,13 @@ This catches `to_numpy()` calls, `PyccelKernel` conversions and `Kernel`
fallbacks that crept into the step. It does not see copies made outside
CuNumpy (see [Data movement](../guides/data-movement.md)).

Syncs (the host waiting for the device) are recorded as well, in
`counter.syncs`, and are accepted unless you ask for none:
`assert_no_transfers(syncs=True)`. They include `xp.synchronize()` and the waits
of the MPI helpers; on the fake CuPy also `float(a)`, `int(a)`, `bool(a)`,
`a.item()` and `a.tolist()`, which stall the real CuPy too but cannot be
observed there from Python.

## Test generated headers

When struct headers are generated with `write_cuda_header()` and committed,
Expand All @@ -324,6 +368,12 @@ offsets.

## CI setup

* With `CUNUMPY_REQUIRE_CUDA=1` the GPU markers fail instead of skipping:
`requires_cupy` (as an error when the test is set up), the `cupy` run of the
`backend` fixture (which also activates CuPy strictly, never falling back to
NumPy) and `assert_kernels_agree`. Set it on the GPU CI job, so a broken CuPy
or driver cannot pass as a set of skipped tests.

* Run the suite on a normal CPU runner: everything on NumPy runs, GPU cases
are reported as skipped, and `emulate_cuda_kernel` tests check the CUDA
kernels' arithmetic (the runner needs a C++ compiler, which Linux images
Expand Down
4 changes: 4 additions & 0 deletions src/cunumpy/LLM_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,10 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository.
| call an existing NumPy-only kernel with GPU arrays (slow, correct) | `xp.kernels.PyccelKernel(fn, outputs=(...))` |
| launch a hand-written CUDA C kernel | `xp.kernels.CudaKernel(source, "name")` / `CudaKernel.from_file(path)` |
| run a Metal (MSL) kernel on an Apple silicon GPU, NumPy float32 in and out (float64 raises; `float64="cast"` computes in float32); not part of `Kernel` dispatch | `xp.kernels.MetalKernel(body, inputs=[...], outputs=[...])(*args, out=arrays, n_threads=n)`; check `xp.kernels.metal_available()` |
| call host-only code (SciPy, file readers) with arguments of either backend | `xp.host_call(fun, *args)`; `@xp.evaluate_on_host` on a method; `@xp.setup_on_host` on `__init__` |
| keep the live rows of particle arrays at the front | `n = xp.algorithms.compact_by_mask(alive, markers, weights)` |
| run CUDA kernel launches on the CPU in a test (fake CuPy) | `with kernel_testing.emulated_launches(): ...`; `kernel_testing.host_buffer(a)` reads a fake array; struct arguments are read through their fields |
| find the arguments a host kernel writes (copy only those back) | `PyccelKernel(fn, outputs=("out",))` (names work positionally); `xp.kernels.outputs_from_annotations(fn)`; `Kernel.from_folder(..., outputs="annotations")` |
| host kernel + CUDA port, chosen by backend | `xp.kernels.Kernel(host_fn, cuda_kernel_or_None)` |
| many kernels in a package, ported incrementally | `xp.kernels.KernelCatalog.from_package(__name__, missing_cuda="fallback")` |
| host kernels compiled at first call (your compile function), NumPy fallback | `from_package(..., host_suffix="_pyccel", compile_host=my_compile, host_fallback={...})` -> `xp.kernels.CompiledHostKernel` |
Expand Down
46 changes: 46 additions & 0 deletions src/cunumpy/_algorithms.py
Original file line number Diff line number Diff line change
Expand Up @@ -251,3 +251,49 @@ def sort_by_key(keys: Any, *arrays: Any) -> tuple[Any, ...]:
)
order = xpm.argsort(keys, kind="stable").astype(xpm.int64, copy=False)
return (keys[order], order, *(array[order] for array in arrays))


def compact_by_mask(mask: Any, *arrays: Any) -> int:
"""Move the rows where `mask` is True to the front of every array, in place.

The usual step after particles left the domain or were absorbed: keep the
live ones at the front of the marker array (and of the arrays that go with
it) and continue with ``markers[:n]``. The order of the kept rows is
preserved, so the result is reproducible::

n = xp.algorithms.compact_by_mask(alive, markers, weights)
markers, weights = markers[:n], weights[:n]

Parameters
----------
mask : array of bool, shape (n,)
True for the rows to keep.
*arrays : arrays
Arrays with ``n`` rows (any further axes), on the backend of `mask`.
Rows ``[:count]`` hold the kept rows afterwards; the rows after them
are unspecified, so ignore them (or overwrite them).

Returns
-------
int
The number of kept rows. Its value is needed on the host, so on CuPy
the call synchronizes once per call (counted by
:func:`~cunumpy.profiling.count_transfers` where it can be seen).
"""
assert_same_backend(mask, *arrays)
xpm = get_array_module(mask)
mask = xpm.asarray(mask)
if mask.ndim != 1 or mask.dtype != np.bool_:
raise TypeError(
f"mask must be a 1D boolean array, got dtype {mask.dtype}, {mask.ndim}D",
)
for i, array in enumerate(arrays):
if array.ndim < 1 or array.shape[0] != mask.shape[0]:
raise ValueError(
f"array {i} has shape {array.shape}, expected {mask.shape[0]} rows",
)
rows = xpm.nonzero(mask)[0]
n_kept = int(rows.size)
for array in arrays:
array[:n_kept] = array[rows] # the right side is a copy: no overlap problem
return n_kept
5 changes: 5 additions & 0 deletions src/cunumpy/_cuda_kernel.py
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,9 @@ class is the one definition of the arguments.

import numpy as np

from cunumpy._transfers import _ACTIVE as _COUNTERS
from cunumpy._transfers import _record_sync

__all__ = [
"DEBUG_OPTIONS",
"CudaArguments",
Expand Down Expand Up @@ -2532,6 +2535,8 @@ def _synchronize_after_launch(
stream = cp.cuda.get_current_stream()
if _is_capturing(stream):
return
if _COUNTERS:
_record_sync(f"debug synchronization after kernel {self.expression!r}")
try:
stream.synchronize()
except Exception as error:
Expand Down
Loading
Loading