Skip to content

fix(rpucuda): preserve forward bound-management state across kernels and retries - #806

Open
Zhaoxian-Wu wants to merge 3 commits into
IBM:masterfrom
Zhaoxian-Wu:fix/cuda-forward-BM
Open

Zhaoxian-Wu wants to merge 3 commits into
IBM:masterfrom
Zhaoxian-Wu:fix/cuda-forward-BM

Conversation

@Zhaoxian-Wu

Copy link
Copy Markdown
Contributor

Related issues

N/A

Description

Forward bound management had four inconsistent failure paths across the native CPU and CUDA implementations:

  1. CUDA batch processing without noise management multiplied the cumulative BM factor into an already scaled value, producing factors such as 1, 2, 8 instead of 1, 2, 4.
  2. The CUDA single-vector output kernel interpreted bm_test_negative_bound in the opposite way from the batch kernel.
  3. CPU forward processing allowed an ignored negative saturation to clear a previously detected positive saturation, so BM could stop even though a positive output had clipped.
  4. CUDA split MVMs replaced the I/O manager input-buffer pointer. A subsequent BM retry could reuse that scratch buffer as both source and destination and lose one sign of the input.

The CPU mixed-bound case returned [1, -1] instead of [6, -6]. Both CUDA split-MVM modes returned -0.75 instead of 4.5.

Details

When CUDA batch BM runs without noise management, the scale-update kernel now applies only the incremental factor to its existing scale. With noise management enabled, it continues rebuilding the absolute scale from the NM value.

The CUDA single-vector output kernel now uses the same bm_test_negative_bound meaning as the batch kernel.

The CPU bound-test result is now monotonic within one output pass: ignoring a negative saturation cannot change an earlier positive-bound failure back to success.

The CUDA forward retry loop saves its original input buffer and restores it before every retry. Split positive/negative MVMs can therefore use their scratch buffers without changing where the next input-management pass writes its data.

Regression coverage checks:

  • single-vector versus batch CUDA scaling without noise management;
  • negative-bound flag consistency between CUDA kernels;
  • mixed positive and negative CPU saturation;
  • retries for POS_NEG_SEPARATE;
  • retries for POS_NEG_SEPARATE_DIGITAL_SUM.

The complete simulator tile test file passes with 499 passed, 54 skipped.

Minimal Working Example

The relevant regression cases use deterministic I/O parameters:

from pytest import mark
from torch import Tensor, full, ones
from torch.testing import assert_close

from aihwkit.simulator.configs import SingleRPUConfig
from aihwkit.simulator.configs.devices import ConstantStepDevice
from aihwkit.simulator.parameters.enums import (
    AnalogMVType,
    BoundManagementType,
    NoiseManagementType,
)
from aihwkit.simulator.parameters.io import IOParameters
from aihwkit.simulator.tiles.analog import AnalogTile


def _forward_bm_tile(
    weights: Tensor,
    inp_res: float = 0.0,
    test_negative_bound: bool = False,
    mv_type: AnalogMVType = AnalogMVType.ONE_PASS,
    use_cuda: bool = True,
) -> AnalogTile:
    forward = IOParameters(
        bound_management=BoundManagementType.ITERATIVE,
        noise_management=NoiseManagementType.NONE,
        inp_res=inp_res,
        out_res=0.0,
        out_noise=0.0,
        inp_bound=1.0,
        out_bound=1.0,
        bm_test_negative_bound=test_negative_bound,
        mv_type=mv_type,
    )
    config = SingleRPUConfig(
        device=ConstantStepDevice(
            w_min=-1.0,
            w_max=1.0,
            w_min_dtod=0.0,
            w_max_dtod=0.0,
        ),
        forward=forward,
    )
    tile = AnalogTile(weights.shape[0], weights.shape[1], config)
    tile.set_weights(weights)
    return tile.cuda() if use_cuda else tile


def test_cpu_ignored_negative_bound_does_not_hide_positive_clipping() -> None:
    weights = full((2, 8), 0.75)
    weights[1] *= -1
    tile = _forward_bm_tile(weights, use_cuda=False)
    inputs = ones(1, 8)

    actual = tile.tile.forward(inputs)

    # The first pass clips [6, -6] to [1, -1]. The negative saturation is
    # ignored, but the preceding positive saturation must still trigger BM.
    assert_close(actual, Tensor([[6.0, -6.0]]), atol=1e-6, rtol=0)


@mark.parametrize(
    "mv_type",
    [AnalogMVType.POS_NEG_SEPARATE, AnalogMVType.POS_NEG_SEPARATE_DIGITAL_SUM],
)
def test_cuda_split_mvm_retries_restore_input_buffer(mv_type: AnalogMVType) -> None:
    weights = full((3, 8), 0.75)
    tile = _forward_bm_tile(weights, mv_type=mv_type)
    inputs = ones(2, 8)
    inputs[:, -1] = -1.0

    actual = tile.tile.forward(inputs.to(tile.device)).cpu()

    # Each output is (7 * 1 + 1 * -1) * 0.75 = 4.5.
    assert_close(actual, full((2, 3), 4.5), atol=1e-6, rtol=0)

Run the regression groups with:

python3 -m pytest -p no:cacheprovider \
  'tests/test_simulator_tiles.py::test_cuda_batch_uses_incremental_bm_factor_without_nm' \
  'tests/test_simulator_tiles.py::test_cuda_single_vector_tests_negative_bound' \
  'tests/test_simulator_tiles.py::test_cpu_ignored_negative_bound_does_not_hide_positive_clipping' \
  'tests/test_simulator_tiles.py::test_cuda_split_mvm_retries_restore_input_buffer'

The old-versus-new comparisons used independently built native extensions for each worktree. For the follow-up mixed-bound and split-MVM cases:

  • fixed revision fdaf78a: 3 passed in 6.11s;
  • upstream revision 10fdaff: 3 failed with AssertionError in 79.24s;
  • CPU greatest absolute difference: 5.0;
  • CUDA split-MVM greatest absolute difference: 5.25 for both modes.

The earlier single/batch factor and negative-bound comparisons likewise pass on the fix branch and reproduce assertion failures on upstream. Neither comparison failed because of collection, import, CUDA availability, or CUDA OOM errors.

…nels

Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
@Zhaoxian-Wu

Copy link
Copy Markdown
Contributor Author

Another implementation issue is observed: restore NPSum scaling in the forward bound management

Description

Forward bound management can return a clipped output when the DAC resolution is unlimited and NPSum should provide a larger input scale. With eight inputs of 1, weights of 0.5, and an output bound of 1, the expected MVM output is 4. On the upstream baseline, CUDA AbsMaxNPSum returns 1; IterativeWorstCase with initially disabled noise management returns 2 on CPU and CUDA.

Details

The CUDA NPSum kernel capped its scale at AbsMax even when inp_res <= 0. The single-vector CUDA kernel also selected noise management from the initial configuration after a BM retry had activated NPSum. On CPU, the forward retry computed NPSum but continued treating noise management as disabled.

The fix removes the DAC-resolution cap when resolution is unlimited, selects the CUDA kernel from the current NoiseManager scale state, and applies the newly computed NPSum scale on the CPU retry.

The regression test covers CPU and CUDA, batches of 1 and 4, direct NPSum, and a worst-case retry that starts with noise management disabled. Each revision's native CUDA extension was built separately. The fix worktree (4056f4b) passed all 8 cases, including all 4 CUDA cases; upstream (10fdaff) passed 2 and failed 6 with AssertionError.

Minimal Working Example

The same regression case from tests/test_simulator_tiles.py was run against both revisions:

from pytest import mark, param
from torch import full, ones
from torch.testing import assert_close

from aihwkit.simulator.configs.configs import SingleRPUConfig
from aihwkit.simulator.configs.devices import ConstantStepDevice
from aihwkit.simulator.parameters.enums import BoundManagementType, NoiseManagementType
from aihwkit.simulator.parameters.io import IOParameters
from aihwkit.simulator.tiles.analog import AnalogTile
from tests.helpers.testcases import SKIP_CUDA_TESTS


@mark.parametrize(
    "use_cuda",
    [False, param(True, marks=mark.skipif(SKIP_CUDA_TESTS, reason="CUDA unavailable"))],
)
@mark.parametrize("batch", [1, 4])
@mark.parametrize(
    "bm_type,nm_type",
    [
        (BoundManagementType.NONE, NoiseManagementType.ABS_MAX_NP_SUM),
        (BoundManagementType.ITERATIVE_WORST_CASE, NoiseManagementType.NONE),
    ],
)
def test_forward_npsum_without_dac_resolution(
    use_cuda: bool, batch: int, bm_type: BoundManagementType, nm_type: NoiseManagementType
) -> None:
    """NPSum scaling works directly and on a worst-case retry, including without initial NM."""
    forward = IOParameters(
        bound_management=bm_type,
        noise_management=nm_type,
        inp_res=0.0,
        out_res=0.0,
        inp_bound=1.0,
        out_bound=1.0,
        out_noise=0.0,
        max_bm_factor=1,
    )
    config = SingleRPUConfig(
        device=ConstantStepDevice(w_min=-1.0, w_max=1.0, w_min_dtod=0.0, w_max_dtod=0.0),
        forward=forward,
    )
    tile = AnalogTile(3, 8, config)
    tile.set_weights(full((3, 8), 0.5))
    if use_cuda:
        tile = tile.cuda()

    actual = tile.joint_forward(ones(batch, 8, device=tile.device)).cpu()
    assert_close(actual, full((batch, 3), 4.0), atol=1e-6, rtol=0)

Build the native CUDA extension separately on each revision and use an available GPU. Run:

python3 -m pytest -p no:cacheprovider \
  tests/test_simulator_tiles.py::test_forward_npsum_without_dac_resolution

The dedicated worktree comparison yielded 8 passed, 0 skipped on the fix and 2 passed, 6 failed on upstream. On upstream, direct AbsMaxNPSum returned 1 instead of 4 on CUDA for both batch sizes; IterativeWorstCase with initially disabled noise management returned 2 instead of 4 on CPU and CUDA for both batch sizes. The fixed revision returned 4 in every case.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant