Skip to content

perf(builder): replace hash-map dedup with sort-based CSR assembly - #8

Merged
deltaeecs merged 1 commit into
masterfrom
fix/issue-7-builder-sort-dedup
Sep 7, 2026
Merged

deltaeecs merged 1 commit into
masterfrom
fix/issue-7-builder-sort-dedup

Conversation

@deltaeecs

Copy link
Copy Markdown
Owner

摘要

cCSRMatrixBuilder::Build() 原先通过全局 std::unordered_map 去重 triplet:每个唯一条目一个堆节点(~48-64B)、每条 triplet 一次 hash 查找、按 triplet 总数 reserve() 桶数组、最后再全量 O(n log n) 排序。对数千万级 triplet 的近场装配(issue #7 中 3λ 场景 CSR build 210s),这既是时间瓶颈也是内存峰值主因(~4 倍最终 CSR 体积)。

改为流式四遍管线,公开 API 不变:

  1. 按行计数在范围内的条目 → rowPtr 前缀和(O(N))
  2. 将 (col, value) scatter 到行分段槽位(O(N),行内保持插入序)
  3. 每行 std::stable_sort 按列排序 + 相邻等列归并(累加 / 取最后)
  4. 归并时回写最终行偏移,输出 CSR 数组

std::stable_sort 保证等列重复项保持插入顺序,因此累加顺序与原 hash-map 实现完全一致——结果逐位相同(浮点累加顺序敏感)。

实测数据(MinGW g++ -O2,近场模式 + 10% 重复)

规模 旧实现 新实现 加速
7.7M triplets 1.06s 0.106s ~10×
31M triplets ~4.3s(外推) 0.439s ~71M triplets/s

对照 issue #7 目标(3λ CSR build < 30s):新实现吞吐 ~71M triplets/s,30M triplet 场景 < 1s 级别,远超目标。

测试

  • 现有 17 个 CSRBuilderTest 全部原样通过(接口未变)
  • 新增 5 个 bit-exact 对拍测试:在测试内保留旧 hash-map 管线作为参考实现,随机 triplets(累加/取最后两种模式)、近场模式(含越界行过滤)、边界情形、复数标量——结构与数值位模式完全一致
  • 本机 ctest:86/86 通过(2 个预先存在的 MPI skip 项与本次无关)
  • 新增 bench/bench_builder.cpp 复现近场装配模式(csr4mpi_bench_builder 目标)

Fixes #7

cCSRMatrixBuilder::Build() previously deduplicated triplets through a
global std::unordered_map: one heap node per unique entry (~48-64B),
a hashed lookup per triplet, a reserve() sized on the full triplet
count, and a final global O(n log n) sort. For near-field assemblies
with tens of millions of triplets this dominated ILU-preconditioner
build time (issue #7: 210s at the 3-lambda case) and peak memory
(~4x the final CSR size).

Replace it with a streaming four-pass pipeline:
  1. count in-range entries per row -> rowPtr prefix sum      O(N)
  2. scatter (col, value) into row-segmented slots            O(N)
  3. per-row std::stable_sort by column + merge of equal-
     column runs (accumulate or keep-last)
  4. final row offsets while emitting the CSR arrays

std::stable_sort preserves insertion order within equal columns, so
duplicates accumulate in the same order - and to the same bits - as
the hash-map implementation. The public API is unchanged.

Measured (MinGW g++ -O2, near-field pattern, 10% duplicates):
  7.7M triplets: 1.06s -> 0.106s (~10x)
  31M triplets:  0.439s (~71M triplets/s)

Tests: all 17 existing CSRBuilderTest cases pass unchanged; add five
bit-exact equivalence tests against a verbatim copy of the old
hash-map pipeline (accumulate/keep-last, near-field pattern with
out-of-range rows, edge cases, complex scalars). Add
bench/bench_builder.cpp reproducing the near-field assembly.

Fixes #7
@deltaeecs
deltaeecs merged commit 904948b into master Sep 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CSR4MPI 性能瓶颈分析

1 participant