By Qingfeng Xia
API Docs | Developer Docs | Benchmark | Architecture | Issues
numpy-ascend brings the CuPy ndarray API to Huawei Ascend NPUs. It is a fork
of CuPy (v14 lineage, NumPy 2.x) with the CUDA backend replaced by an Ascend
backend built on CANN's aclnn operator library, so that existing NumPy/CuPy
code and the wider CuPy/SciPy ecosystem run on the NPU with import cupy.
- Default dtype
float32(same as CuPy on GPU) — 10x–200x NPU acceleration for math workloads. float64andcomplex128are supported, with three processing modes (see below) because the NPU has no native double-precision throughput and someaclnnops rejectDOUBLE/COMPLEX128outright.int64andint32have hardware acceleration for add/subtract/multiply/divide.- API-compatible with CuPy (GPU) and mostly compatible with NumPy (CPU); works with the CuPy and SciPy ecosystem.
Selected via the environment variable CUPY_ASCEND_FLOAT64_MODE (read once at
first use; legacy switch CUPY_ASCEND_ENABLE_FLOAT64_TO_FLOAT32=1 is equivalent
to float32):
| Mode | Behavior | Precision | Speed |
|---|---|---|---|
float32 (default) |
Demote float64/complex128 operands to float32/complex64 at the dispatch layer, compute on NPU, cast results back. |
loses ~9 decimal digits per operation | fast (NPU) |
cpu |
Whole-op interception: D2H → NumPy in true double precision → H2D. Works for ufuncs, reductions and most general ops (unsupported ones fail loudly instead of silently). | exact (matches NumPy) | slow (PCIe transfers per op) |
off |
No demotion, no fallback: ops that cannot run in double precision raise an error. | exact or error | — |
export CUPY_ASCEND_FLOAT64_MODE=cpu # exact double precision, host-computed
export CUPY_ASCEND_FLOAT64_MODE=float32 # default: NPU speed, demoted precision
export CUPY_ASCEND_FLOAT64_MODE=off # never silently lose precisionWhy the default float32 mode can be wrong for you — error accumulation:
float32carries a 24-bit mantissa: ~7 significant decimal digits (float64: ~16). Each arithmetic op rounds to the nearest representable value, and the rounding errors accumulate across the computation.- Reductions are the worst case: summing
Nvalues has worst-case relative error ~N·eps— summing 10^8 randomfloat64samples infloat32mode can drift by 1e-2 relative, versus ~1e-16 incpumode. Long chains of elementwise ops (iterative solvers,cumsum, normalization) compound the same way; values above 2^24 (≈1.7e7) lose integer precision entirely. complex128demotes tocomplex64, so both real and imaginary parts share the same 24-bit mantissa — the relative error applies to magnitude and phase alike.- Rule of thumb: for ML training/inference and image data,
float32mode is the intended trade-off. For numerical verification, financial/scientific accumulation (large sums, variance over big arrays, eigen/solvers where the residual matters), usecpumode — it is exact and easy to switch per process; or keepfloat64data in NumPy and move only thefloat32-hot path to the NPU.
OS: Linux only. Any Linux distribution should work provided the C++
runtime (libstdc++, glibc >= 2.17) is new enough; CI is done on x86_64 Linux.
| CANN | x86_64 | aarch64 | Notes |
|---|---|---|---|
| 9.5 | source | source | not tested; may work, no CI coverage |
| 9.0 | binary wheel + source | source | primary test target (Ascend 910B) |
| 8.5 | binary wheel + source | source | primary test target (Ascend 910B) |
| 8.2 | source build | source build | not tested; may work, no CI coverage |
| Python | Binary wheel | Source build |
|---|---|---|
| 3.10 | yes | yes |
| 3.11 | yes | yes |
| 3.12 | no package yet | yes |
- Binary wheels: see the Releases page.
Wheel naming:
numpy_ascend_cann<XY>-<ver>-cp3XX-cp3XX-manylinux_<glibc>_<arch>.whl— pick the wheel matching your interpreter (cp311= Python 3.11), CPU architecture and CANN release train. - The wheel does not bundle the CANN SDK: the target machine needs a CANN
toolkit of the same release train plus
set_env.sh; see docs/ascend/Package.md for the hard runtime prerequisites (libstdc++ version, optional ops-fft, etc.). - Only one package providing
cupymay be installed per environment: never co-install this with upstreamcupy/cupy-cudaXXor a different CANN train — they would overwrite each other's files.
pip install numpy_ascend_cann90-...whl # pick from the Releases page
source /usr/local/Ascend/ascend-toolkit/set_env.sh
export ASCEND_HOME_PATH=/usr/local/Ascend/ascend-toolkit/latest
python -c "import cupy; print(cupy.__version__, cupy.backends.ascend.check_cann_version())"git clone git@github.com:qingfengxia/numpy-ascend.git
cd numpy-ascend && git checkout ascend # the working branch is ascend, not main
export CUPY_INSTALL_USE_ASCEND=1
python setup.py build_ext --inplace # incremental in-place build (dev)
python -c "import cupy._core" # L2/L3 check: link + op registration OK
python -m build --wheel # optional wheel (cannX.Y tag auto-detected)
python -m build --sdist # optional sdist (one file, CPython/CANN agnostic)Environment setup (CANN 8.2/8.5/9.0 installation and switching, development on machines without an NPU, porting to new platforms, troubleshooting) is documented in DeveloperNotes.md §1–§3; wheel and sdist packaging strategy is in docs/ascend/Package.md.
pytest tests/ascend -q # no NPU needed: registry / composed ops / kernel parsing
pytest tests/cupy_tests -q # needs NPU; unsupported dtypes skipped by default
pytest tests/cupy_tests -q --ascend-dtype-filter=off # full dtype matrix (shows real failures)
python tools/benchmark.py --list # no NPU needed: print the op x dtype matrix
python tools/benchmark.py --csv result.csv # needs NPU
# minimal numerical check (needs NPU)
python -c "
import numpy as np, cupy as cp
x = np.random.rand(1000).astype(np.float32)
assert np.allclose(cp.asnumpy(cp.asarray(x).sum()), x.sum(), rtol=1e-5); print('ok')
"Verification levels: L1 Cython generates .cpp → L2 links → L3
import cupy + op-registry audit → L4 numerical results on real hardware.
A machine without an NPU can only reach L3; L4 must run on a 910B.
See Progress.md for the full snapshot:
- Python Array API standard: > 98% covered (2 gaps:
i0— NumPy itself recommendsscipy.special.i0— andnextafter). - CuPy top-level API: > 98% covered.
- Operator-level facts are auto-generated by tools/scan_ops.py into tools/cst_db.md.
| Suite | Pass rate |
|---|---|
tests/cupy_tests/math_tests (math_test) |
100% |
full tests/cupy_tests + tests/ascend |
90% |
Failures concentrate in the documented limitation areas below, not in the math kernels.
- Custom kernels (AscendC, and Triton-Ascend in Python).
- All major CuPy features: custom kernels, profiler, stream/device management — except operator fusion (use a Triton kernel instead).
- scipy ecosystem support via array_api, see scipy benchmark with numpy-ascend
uint64is not supported (int64is used, with hardware acceleration for add/multiply).- Operator fusion.
- Sparse array/matrix and some
cupy.randomdistributions are supported only partially (low priority). - Multi-NPU is not ported/tested yet.
| Tool | Purpose |
|---|---|
tools/benchmark.py |
CLI benchmark: NumPy (CPU) vs numpy-ascend (NPU) speedup over an op × dtype matrix (unary/binary/reduction/matmul/manipulation/sort). --list prints the matrix without an NPU; --category, --dtype, --repeat, --csv, --strict filter and report. Exit code 2/3 signal import/device errors. |
tools/numpy_ascend_migration_helper/ |
Static migration analyzer for NumPy/CuPy source: reports unsupported APIs/dtypes, semantic differences and CPU-fallback candidates (AST-based, not string matching); --fix applies AUTO_SAFE rewrites, --fail-on-error gates CI. See its README.md. |
tools/scan_ops.py |
Scans the built backend and generates the operator/coverage database (tools/cst_db.md). |
https://github.com/data-apis/array-api
import cupy as cp
import cupy.array_api as cpx # CuPy's Array API namespace
x_gpu = cpx.asarray([1, 2, 3, 4], device='cuda') # device slot, NPU here
y_gpu = cpx.reshape(x_gpu, (2, 2))
z_gpu = cpx.matmul(y_gpu, y_gpu)If you are writing new data-processing code and do not need NumPy/CuPy
compatibility, torch_npu's array API (torch._numpy) is an alternative.
| Document | Content |
|---|---|
| Progress.md | Status snapshot: coverage, operator counts, milestones, remaining gaps |
| Memory.md | Developer handbook: command cheat-sheet, architecture key points, verification levels, pitfalls |
| DeveloperNotes.md | Environment setup (incl. no-NPU development), CANN 8.2/8.5/9.0 install, conditional compilation, FFT |
| docs/ascend/Package.md | Wheel packaging (cann tag), RPATH strategy, runtime hard prerequisites |
| install/README.md | Build system architecture (features / backends) and op registration flow |
| docs/ascend/ | Custom AscendC kernels, FFT, matmul, code reviews, NumPy/CuPy/PyTorch API diffs |
| tools/cst_db.md | Auto-generated operator/coverage database (python tools/scan_ops.py) |
| Roadmap.md · TODO.md | Roadmap and task list |
Issues and PRs are welcome at
https://github.com/qingfengxia/numpy-ascend. For a code change: rebuild with
CUPY_INSTALL_USE_ASCEND=1, verify at least to L3 (pytest tests/ascend -q),
and state honestly which verification level was reached (see
Memory.md). Commit messages use the ASCEND [AI]: prefix on the
ascend branch.
MIT — see LICENSE. This project forks CuPy; upstream CuPy is MIT-licensed and its copyright notice is preserved in the source tree.
简体中文说明见 Readme_ZH.md。