Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions .github/workflows/macos.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
name: Tests (macOS)

on:
push:
branches:
- main
- devel
pull_request:
branches:
- main
- devel
workflow_dispatch:

jobs:
build:
# macos-latest is Apple silicon (arm64): checks the NumPy backend, the
# compiled host kernels and the MLX import on the platform of MetalKernel.
runs-on: macos-latest

strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.13"]

steps:
- name: Checkout code
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
cache: pip
cache-dependency-path: pyproject.toml

# Pyccel is built from source here (no wheel for macOS arm64) and needs a
# Fortran compiler, which the macOS runners do not have.
- name: Install gfortran
run: |
brew install gcc
gfortran --version

- name: Install project
run: |
pip install --upgrade pip
pip install ".[test-compiled,metal]"

# Hosted macOS runners are virtual machines and usually have no Metal GPU,
# in which case the MetalKernel launch tests are skipped. This step shows
# which case this run is in.
- name: Report Metal availability
run: |
python -c "import platform, cunumpy as xp; print(platform.machine(), 'metal_available =', xp.kernels.metal_available())"

- name: Run tests
run: pytest . -rs
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Added
- `xp.kernels.MetalKernel` runs a Metal Shading Language kernel on the GPU of an
Apple silicon Mac through MLX (`pip install 'cunumpy[metal]'`). It takes and
fills NumPy arrays, is float32 only (`float64="cast"` computes float64 data in
float32), and its copies are counted by `xp.profiling.count_transfers()`.
`xp.kernels.metal_available()` tells whether it can run.
- `xp.host_call`, `xp.evaluate_on_host` and `xp.setup_on_host` run host-only code
(SciPy splines, file readers, external libraries) with arguments of either
backend: device arrays are copied to the host, the call runs on the NumPy
Expand Down
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ never hide a NumPy name:

| Submodule | Contents |
|---|---|
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, host implementations, `fuse` |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `MetalKernel`, host implementations, `fuse` |
| `xp.arguments` | CUDA only: `CudaStruct`, `CudaStructArguments`, `CudaArguments` |
| `xp.cuda` | CUDA only: devices, streams, debug mode, CUDA headers |
| `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
Expand All @@ -36,6 +36,8 @@ never hide a NumPy name:
| `cunumpy.kernel_testing` | pytest helpers for host/CUDA kernel pairs |

Everything except `xp.cuda`, `xp.arguments` and `CudaKernel` works on both backends.
`MetalKernel` runs Metal kernels on the GPU of an Apple silicon Mac (MLX, float32, NumPy arrays;
`pip install 'cunumpy[metal]'`, see [the guide](docs/source/kernels/metal-kernel.md)).

## Install

Expand Down
47 changes: 46 additions & 1 deletion docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ The submodules are named so that they do not hide a NumPy name (`rng`, not
| Submodule | Backends | Contents |
|---|---|---|
| `cunumpy` | both | NumPy/CuPy namespace, backend selection, array inspection and conversion, `synchronize`, `host_call`, `evaluate_on_host`, `setup_on_host`, `scipy`, `require_version` |
| `cunumpy.kernels` | both | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `CudaKernelVariants`, host implementations, `as_kernel_array`, `kernel_output`, `fuse` |
| `cunumpy.kernels` | both | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `CudaKernelVariants`, `MetalKernel`, `metal_available`, host implementations, `as_kernel_array`, `kernel_output`, `fuse` |
| `cunumpy.arguments` | CUDA only | `CudaArguments`, `CudaStruct`, `CudaStructArguments`, `CudaStructValue`, `write_cuda_header` |
| `cunumpy.cuda` | CUDA only | device selection and memory, `stream`, streams/events, `pin_memory`, debug mode, CUDA headers and source tools (`cuda_include_dir`, `parse_cuda_signature`) |
| `cunumpy.rng` | both | `random_streams`, `get_rng`, `philox_*` |
Expand Down Expand Up @@ -857,6 +857,51 @@ lists) are converted back using `is_array`; dictionaries in return values are
not recursively converted. On the NumPy path, the original return value and
normal Python mutation and exception behavior are preserved.

## `kernels.MetalKernel`

A Metal Shading Language kernel for the GPU of an Apple silicon Mac, run with
[MLX](https://github.com/ml-explore/mlx) (`pip install 'cunumpy[metal]'`). It
takes NumPy arrays and fills the output arrays you pass, so no backend switch
is needed.

```python
import numpy as np
import cunumpy as xp

scale = xp.kernels.MetalKernel(
"uint i = thread_position_in_grid.x; y[i] = a[0] * x[i];",
inputs=["x", "a"],
outputs=["y"],
)
x = np.arange(8, dtype=np.float32)
y = np.empty_like(x)
scale(x, 2.0, out=y)
```

`MetalKernel(source, inputs, outputs, *, name="cunumpy_kernel", header="",
threadgroup=256, float64="error", atomic_outputs=False, init_value=None)`

- `source` is the body of the kernel function. MLX generates the signature: each
name in `inputs` and `outputs` is a pointer to the flat, row-major data of that
array, so `x[i]` is the flat index. `thread_position_in_grid` and the other
Metal attributes used in the body are added automatically. `header` goes before
the function (includes, defines, helper functions).
- Calling the kernel: `kernel(*inputs, out=array_or_arrays, n_threads=None,
template=None)`. `n_threads` is the total thread count (default: first axis of
the first output). `template` gives compile-time constants, e.g.
`template={"NSTEPS": 200}`. It returns the output array, or a tuple of them.
- Outputs are uninitialized: write every element, pass the old array as an input
too if the kernel reads it, or set `init_value`.
- **float32 only.** The Apple GPU has no float64: a float64 input or output
raises `TypeError`. With `float64="cast"` float64 data is computed in float32
(a push over 200 steps agreed with float64 to about 3e-5).
- Every call copies the inputs to MLX arrays and the results back, counted as
`to_device` and `to_host` transfers by `count_transfers()`. On a 4M-particle
push these copies were about 7 ms next to an 11.5 ms kernel.
- `xp.kernels.metal_available()` is True if MLX is installed and a Metal GPU is
present. Without them, calling a `MetalKernel` raises `ImportError` or
`RuntimeError`.

## `kernels.CudaKernel`

### Constructor
Expand Down
1 change: 1 addition & 0 deletions docs/source/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,7 @@ array-api-compat
kernels/overview
kernels/pyccel-kernel
kernels/cuda-kernel
kernels/metal-kernel
kernels/dispatch
kernels/arguments
kernels/accumulation
Expand Down
16 changes: 16 additions & 0 deletions docs/source/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,10 +39,26 @@ print("active backend:", xp.get_backend()) # 'cupy' if the GPU works
If `cupy_available()` is `False`, CuNumpy quietly falls back to NumPy when CuPy
is requested. See [Troubleshooting](troubleshooting.md) for the usual causes.

## Apple silicon GPUs

CuPy does not run on Macs. The GPU of an Apple silicon Mac can run
[Metal kernels](kernels/metal-kernel.md) on NumPy float32 arrays through MLX:

```bash
python -m pip install 'cunumpy[metal]'
```

```python
import cunumpy as xp

print("Metal usable:", xp.kernels.metal_available())
```

## Optional extras

| Extra | Installs | Use it for |
| --- | --- | --- |
| `cunumpy[metal]` | `mlx` (Apple silicon Macs only) | [`MetalKernel`](kernels/metal-kernel.md) on the Mac GPU |
| `cunumpy[test]` | `pytest`, `coverage` | running the test suite, using `cunumpy.kernel_testing` |
| `cunumpy[test-compiled]` | the above plus `pyccel` | tests that compile host kernels with Pyccel |
| `cunumpy[docs]` | Sphinx, MyST, the book theme | building this documentation |
Expand Down
105 changes: 105 additions & 0 deletions docs/source/kernels/metal-kernel.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
# Metal kernels on Apple silicon

`MetalKernel` runs a kernel written in the Metal Shading Language (MSL) on the
GPU of an Apple silicon Mac. It uses [MLX](https://github.com/ml-explore/mlx),
Apple's array framework, to compile and launch the kernel, so no Objective-C or
Xcode project is needed. It is the Mac counterpart of
[`CudaKernel`](cuda-kernel.md), with two differences:

* it works on **NumPy arrays**: there is no Metal backend, and the kernel copies
its arguments to MLX arrays and the results back into your output arrays;
* it is **float32 only**, because the Apple GPU has no float64.

Install it with `pip install 'cunumpy[metal]'` (MLX is installed on arm64 Macs
only). `xp.kernels.metal_available()` is `True` when MLX is installed and a Metal
GPU is present.

## A first kernel

```python
import numpy as np
import cunumpy as xp

scale = xp.kernels.MetalKernel(
"uint i = thread_position_in_grid.x; y[i] = a[0] * x[i];",
inputs=["x", "a"],
outputs=["y"],
)

x = np.arange(8, dtype=np.float32)
y = np.empty_like(x)
scale(x, 2.0, out=y)
```

The essentials:

* `source` is only the **body** of the kernel function. MLX writes the
signature from `inputs` and `outputs`: each name is a pointer to the flat,
row-major data of that array, so a 2D array `a` of shape `(n, 3)` is read as
`a[3 * i + j]`. Attributes such as `thread_position_in_grid` are added to the
signature when the body uses them.
* Inputs are passed in the order of `inputs`. A Python `float` or `int` becomes
a one-element `float32` or `int32` array; read it as `a[0]`.
* `out` is one array, or one array per name in `outputs`. Their shapes and dtypes
define the outputs, and they are filled in place. The call returns them.
* `n_threads` is the **total** number of threads, not the number of threadgroups
(the opposite of the CUDA `grid`). It defaults to the first axis of the first
output. Threads beyond the data must return early, as for CUDA, if the thread
count is not a multiple of the threadgroup size.
* Outputs start uninitialized. Write every element, or pass the old array as an
input too when the kernel updates it, or give `init_value=0.0`.
* Compilation happens at the first call and MLX caches the result.

## Compile-time constants

`template` gives constants that are known when the kernel is compiled, which
lets the compiler unroll loops. Each name is available in the source:

```python
push = xp.kernels.MetalKernel(
"""
uint i = thread_position_in_grid.x;
float x = pos[i], v = vel[i];
for (int k = 0; k < NSTEPS; ++k) { x += v * dt[0]; }
pos_out[i] = x;
""",
inputs=["pos", "vel", "dt"],
outputs=["pos_out"],
)
push(pos, vel, 0.01, out=pos_out, template={"NSTEPS": 200})
```

A different value of `NSTEPS` compiles another variant. `header` takes
`#include`s, `#define`s and helper functions that go before the kernel.

## float32 and float64

An Apple GPU computes in float32. A float64 array passed to a `MetalKernel`
raises `TypeError` that names the argument, so no precision is lost silently. If
float32 is accurate enough for the kernel, create it with `float64="cast"`:
float64 inputs are computed in float32 and float64 outputs are filled from the
float32 results. A particle push over 200 steps agreed with a float64 reference
to about 3e-5, but whether that is acceptable depends on the physics, so check it
against the host kernel with
[`assert_kernels_agree`](testing.md) using a float32 tolerance.

## Cost of the copies

Every call copies the inputs to MLX arrays and the results back. The copies are
counted as `to_device` and `to_host` by
[`count_transfers()`](../guides/profiling.md), so they can be found in a profile.
On an M1, 4 million particles (3 float32 coordinates each) took about 5 ms to
copy to the GPU and 2 ms to copy back, next to a push kernel of 11.5 ms. A kernel
that runs once per time step on fresh NumPy arrays therefore pays a noticeable
share of its time in copies, and a kernel with little work per element (an
`a * x + y` update) is no faster than NumPy. The GPU pays off for kernels with
much arithmetic per element, such as pushers and interpolation.

## What it does not do

* It is not part of the [`Kernel`](dispatch.md) dispatch: a `Kernel` pairs a host
kernel with a CUDA kernel only, so call a `MetalKernel` directly.
* It does not keep data on the GPU between calls.
* It cannot be tested on GitHub's hosted macOS runners, which are virtual
machines without a Metal GPU. Its launch tests are skipped there and run on a
Mac with `pytest tests/unit/test_metal_kernel.py -rs`.
1 change: 1 addition & 0 deletions docs/source/kernels/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ unchanged.
| --- | --- | --- |
| `PyccelKernel` | calls a host kernel with CuPy arrays by copying them to the host and back | [Host kernels with GPU data](pyccel-kernel.md) |
| `CudaKernel` | wraps a CUDA C kernel, checks every call against its signature | [Writing CUDA kernels](cuda-kernel.md) |
| `MetalKernel` | runs a Metal kernel on the GPU of an Apple silicon Mac, on NumPy float32 arrays | [Metal kernels on Apple silicon](metal-kernel.md) |
| `Kernel` | a host kernel plus its CUDA kernel; calls the one matching the backend | [Pairing host and CUDA kernels](dispatch.md) |
| `KernelCatalog` | all `Kernel`s of a package, found by folder convention | [Pairing host and CUDA kernels](dispatch.md) |
| `CudaArguments`, `CudaStruct`, `CudaStructArguments` | pass a group of arrays and scalars as one argument | [Kernel arguments and structs](arguments.md) |
Expand Down
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,7 @@ optional-dependencies.docs = [
"sphinx",
"sphinx-book-theme",
]
optional-dependencies.metal = [ "mlx; sys_platform == 'darwin' and platform_machine == 'arm64'" ]
optional-dependencies.test = [ "coverage", "pytest", "scipy" ]
optional-dependencies.test-compiled = [ "cunumpy[test]", "pyccel" ]
urls."Source" = "https://github.com/max-models/cunumpy"
Expand Down
1 change: 1 addition & 0 deletions src/cunumpy/LLM_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,7 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository.
| hand data to SciPy/matplotlib/h5py | `xp.to_numpy(a)` |
| call an existing NumPy-only kernel with GPU arrays (slow, correct) | `xp.kernels.PyccelKernel(fn, outputs=(...))` |
| launch a hand-written CUDA C kernel | `xp.kernels.CudaKernel(source, "name")` / `CudaKernel.from_file(path)` |
| run a Metal (MSL) kernel on an Apple silicon GPU, NumPy float32 in and out (float64 raises; `float64="cast"` computes in float32); not part of `Kernel` dispatch | `xp.kernels.MetalKernel(body, inputs=[...], outputs=[...])(*args, out=arrays, n_threads=n)`; check `xp.kernels.metal_available()` |
| host kernel + CUDA port, chosen by backend | `xp.kernels.Kernel(host_fn, cuda_kernel_or_None)` |
| many kernels in a package, ported incrementally | `xp.kernels.KernelCatalog.from_package(__name__, missing_cuda="fallback")` |
| host kernels compiled at first call (your compile function), NumPy fallback | `from_package(..., host_suffix="_pyccel", compile_host=my_compile, host_fallback={...})` -> `xp.kernels.CompiledHostKernel` |
Expand Down
Loading
Loading