Skip to content

fix(cuda): align forward bound management in single and batch kernels - #803

Closed
Zhaoxian-Wu wants to merge 1 commit into
IBM:masterfrom
Zhaoxian-Wu:fix/cuda-forward-bound-management
Closed

Zhaoxian-Wu wants to merge 1 commit into
IBM:masterfrom
Zhaoxian-Wu:fix/cuda-forward-bound-management

Conversation

@Zhaoxian-Wu

@Zhaoxian-Wu Zhaoxian-Wu commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Related issues

N/A

Description

CUDA forward bound management returned different results for identical inputs when a single-vector kernel or a batch kernel was selected.

  • Without noise management, the batch path advanced the bound-management scale as 1, 2, 8, ... instead of 1, 2, 4, .... The regression input should return 3.0 for every vector.
  • The single-vector output path inverted bm_test_negative_bound. For a raw output of -2.25 and an output bound of 1.0, enabling negative-bound testing should retry and return -2.25; disabling it should accept the clipped -1.0.

Details

With noise management, the scale-update kernel reconstructs an absolute scale from the noise-management value, so it needs the cumulative bound-management reduction. Without noise management, it multiplies the previous scale, so applying that cumulative reduction repeats earlier factors. The fix passes bound_management_factor_ for the latter path and retains the cumulative reduction for the former.

The single-vector output kernel now counts negative saturation when test_neg is enabled, matching the batch kernel and the meaning of bm_test_negative_bound.

The CUDA regression cases live alongside other simulator tile forward tests in tests/test_simulator_tiles.py. They compare single-vector and batch outputs for deterministic inputs and assert the expected values.

Minimal Working Example

This is the regression code from tests/test_simulator_tiles.py (imports shown for clarity):

from pytest import mark
from torch import Tensor, full, ones
from torch.testing import assert_close

from aihwkit.simulator.configs.configs import SingleRPUConfig
from aihwkit.simulator.configs.devices import ConstantStepDevice
from aihwkit.simulator.parameters.enums import BoundManagementType, NoiseManagementType
from aihwkit.simulator.parameters.io import IOParameters
from aihwkit.simulator.tiles.analog import AnalogTile

from .helpers.testcases import SKIP_CUDA_TESTS


def _forward_bm_cuda_tile(
    weights: Tensor, inp_res: float = 0.0, test_negative_bound: bool = False
) -> AnalogTile:
    """Build a deterministic native CUDA tile with iterative forward BM."""
    forward = IOParameters(
        bound_management=BoundManagementType.ITERATIVE,
        noise_management=NoiseManagementType.NONE,
        inp_res=inp_res,
        out_res=0.0,
        out_noise=0.0,
        inp_bound=1.0,
        out_bound=1.0,
        bm_test_negative_bound=test_negative_bound,
    )
    config = SingleRPUConfig(
        device=ConstantStepDevice(w_min=-1.0, w_max=1.0, w_min_dtod=0.0, w_max_dtod=0.0),
        forward=forward,
    )
    tile = AnalogTile(weights.shape[0], weights.shape[1], config)
    tile.set_weights(weights)
    return tile.cuda()


def _forward_bm_output(tile: AnalogTile, inputs: Tensor) -> Tensor:
    """Run native forward on CUDA and return its result on CPU."""
    return tile.tile.forward(inputs.to(tile.device)).cpu()


@mark.skipif(SKIP_CUDA_TESTS, reason="CUDA unavailable")
def test_cuda_batch_uses_incremental_bm_factor_without_nm() -> None:
    """Identical vectors must not depend on selecting the single or batch CUDA kernel."""
    weights = full((1, 8), 0.75)
    tile = _forward_bm_cuda_tile(weights, inp_res=1 / 16)
    vector = full((1, 8), 0.625)

    single = _forward_bm_output(tile, vector)
    batch = _forward_bm_output(tile, vector.repeat(4, 1))

    # DAC step delta = 2 * inp_bound * inp_res = 1 / 8. For BM scale s,
    # y_raw(s) = 8 * 0.75 * Q_delta(0.625 / s). The first two attempts saturate:
    #   s=1: y_raw = 8 * 0.75 * 0.625 = 3.75 > 1,
    #   s=2: y_raw = 8 * 0.75 * 0.375 = 2.25 > 1.
    # The third attempt succeeds and restores its scale in the final output:
    #   s=4: y_raw = 8 * 0.75 * 0.125 = 0.75; y = s * y_raw = 3.0.
    assert_close(single, full((1, 1), 3.0), atol=0, rtol=0)
    assert_close(batch, single.repeat(4, 1), atol=0, rtol=0)


@mark.skipif(SKIP_CUDA_TESTS, reason="CUDA unavailable")
@mark.parametrize("test_negative_bound,expected", [(True, -2.25), (False, -1.0)])
def test_cuda_single_vector_tests_negative_bound(
    test_negative_bound: bool, expected: float
) -> None:
    """Negative-bound handling must agree between single-vector and batch CUDA kernels."""
    weights = full((8, 3), -0.75)
    tile = _forward_bm_cuda_tile(weights, test_negative_bound=test_negative_bound)
    vector = ones(1, 3)

    single = _forward_bm_output(tile, vector)
    batch = _forward_bm_output(tile, vector.repeat(4, 1))

    # Before the ADC bound, each result is y = 3 * 1 * (-0.75) = -2.25.
    # When enabled, negative-bound testing retries and restores -2.25;
    # otherwise the first pass is accepted with its ADC-clipped value of -1.0.
    assert_close(batch, full((4, 8), expected), atol=1e-6, rtol=0)
    assert_close(single, batch[:1], atol=1e-6, rtol=0)

Run the regression cases with:

python3 -m pytest -p no:cacheprovider \
  'tests/test_simulator_tiles.py::test_cuda_batch_uses_incremental_bm_factor_without_nm' \
  'tests/test_simulator_tiles.py::test_cuda_single_vector_tests_negative_bound'

The fixed revision passes all three parameterized cases. On the upstream baseline, all three fail with assertion failures: the no-noise-management case has a greatest absolute difference of 3.0, and each negative-bound case has a greatest absolute difference of 1.25.

The full Python suite on the fixed code completed with 4,363 passed and 727 skipped; pycodestyle, mypy, and pylint also passed (Pylint 10.00/10).

…nels

Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
@Zhaoxian-Wu
Zhaoxian-Wu force-pushed the fix/cuda-forward-bound-management branch from 70dcd67 to ce06cf8 Compare October 3, 2026 01:00
@Zhaoxian-Wu Zhaoxian-Wu changed the title fix(cuda): align forward bound management across single and batch ker… fix(cuda): align forward bound management in single and batch kernels Oct 3, 2026
@Zhaoxian-Wu Zhaoxian-Wu closed this Oct 3, 2026
@Zhaoxian-Wu
Zhaoxian-Wu deleted the fix/cuda-forward-bound-management branch October 3, 2026 03:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant