fix(rpucuda): preserve forward bound-management state across kernels and retries - #806
Zhaoxian-Wu wants to merge 3 commits into
Conversation
…nels Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
Signed-off-by: Zhaoxian Wu <wuzhaoxian97@gmail.com>
|
Another implementation issue is observed: restore NPSum scaling in the forward bound management DescriptionForward bound management can return a clipped output when the DAC resolution is unlimited and NPSum should provide a larger input scale. With eight inputs of 1, weights of 0.5, and an output bound of 1, the expected MVM output is 4. On the upstream baseline, CUDA AbsMaxNPSum returns 1; IterativeWorstCase with initially disabled noise management returns 2 on CPU and CUDA. DetailsThe CUDA NPSum kernel capped its scale at AbsMax even when The fix removes the DAC-resolution cap when resolution is unlimited, selects the CUDA kernel from the current NoiseManager scale state, and applies the newly computed NPSum scale on the CPU retry. The regression test covers CPU and CUDA, batches of 1 and 4, direct NPSum, and a worst-case retry that starts with noise management disabled. Each revision's native CUDA extension was built separately. The fix worktree ( Minimal Working ExampleThe same regression case from from pytest import mark, param
from torch import full, ones
from torch.testing import assert_close
from aihwkit.simulator.configs.configs import SingleRPUConfig
from aihwkit.simulator.configs.devices import ConstantStepDevice
from aihwkit.simulator.parameters.enums import BoundManagementType, NoiseManagementType
from aihwkit.simulator.parameters.io import IOParameters
from aihwkit.simulator.tiles.analog import AnalogTile
from tests.helpers.testcases import SKIP_CUDA_TESTS
@mark.parametrize(
"use_cuda",
[False, param(True, marks=mark.skipif(SKIP_CUDA_TESTS, reason="CUDA unavailable"))],
)
@mark.parametrize("batch", [1, 4])
@mark.parametrize(
"bm_type,nm_type",
[
(BoundManagementType.NONE, NoiseManagementType.ABS_MAX_NP_SUM),
(BoundManagementType.ITERATIVE_WORST_CASE, NoiseManagementType.NONE),
],
)
def test_forward_npsum_without_dac_resolution(
use_cuda: bool, batch: int, bm_type: BoundManagementType, nm_type: NoiseManagementType
) -> None:
"""NPSum scaling works directly and on a worst-case retry, including without initial NM."""
forward = IOParameters(
bound_management=bm_type,
noise_management=nm_type,
inp_res=0.0,
out_res=0.0,
inp_bound=1.0,
out_bound=1.0,
out_noise=0.0,
max_bm_factor=1,
)
config = SingleRPUConfig(
device=ConstantStepDevice(w_min=-1.0, w_max=1.0, w_min_dtod=0.0, w_max_dtod=0.0),
forward=forward,
)
tile = AnalogTile(3, 8, config)
tile.set_weights(full((3, 8), 0.5))
if use_cuda:
tile = tile.cuda()
actual = tile.joint_forward(ones(batch, 8, device=tile.device)).cpu()
assert_close(actual, full((batch, 3), 4.0), atol=1e-6, rtol=0)Build the native CUDA extension separately on each revision and use an available GPU. Run: python3 -m pytest -p no:cacheprovider \
tests/test_simulator_tiles.py::test_forward_npsum_without_dac_resolutionThe dedicated worktree comparison yielded 8 passed, 0 skipped on the fix and 2 passed, 6 failed on upstream. On upstream, direct AbsMaxNPSum returned 1 instead of 4 on CUDA for both batch sizes; IterativeWorstCase with initially disabled noise management returned 2 instead of 4 on CPU and CUDA for both batch sizes. The fixed revision returned 4 in every case. |
Related issues
N/A
Description
Forward bound management had four inconsistent failure paths across the native CPU and CUDA implementations:
1, 2, 8instead of1, 2, 4.bm_test_negative_boundin the opposite way from the batch kernel.The CPU mixed-bound case returned
[1, -1]instead of[6, -6]. Both CUDA split-MVM modes returned-0.75instead of4.5.Details
When CUDA batch BM runs without noise management, the scale-update kernel now applies only the incremental factor to its existing scale. With noise management enabled, it continues rebuilding the absolute scale from the NM value.
The CUDA single-vector output kernel now uses the same
bm_test_negative_boundmeaning as the batch kernel.The CPU bound-test result is now monotonic within one output pass: ignoring a negative saturation cannot change an earlier positive-bound failure back to success.
The CUDA forward retry loop saves its original input buffer and restores it before every retry. Split positive/negative MVMs can therefore use their scratch buffers without changing where the next input-management pass writes its data.
Regression coverage checks:
POS_NEG_SEPARATE;POS_NEG_SEPARATE_DIGITAL_SUM.The complete simulator tile test file passes with
499 passed, 54 skipped.Minimal Working Example
The relevant regression cases use deterministic I/O parameters:
Run the regression groups with:
The old-versus-new comparisons used independently built native extensions for each worktree. For the follow-up mixed-bound and split-MVM cases:
fdaf78a:3 passed in 6.11s;10fdaff:3 failedwithAssertionErrorin79.24s;5.0;5.25for both modes.The earlier single/batch factor and negative-bound comparisons likewise pass on the fix branch and reproduce assertion failures on upstream. Neither comparison failed because of collection, import, CUDA availability, or CUDA OOM errors.