perf(builder): replace hash-map dedup with sort-based CSR assembly - #8
Merged
Merged
Conversation
cCSRMatrixBuilder::Build() previously deduplicated triplets through a global std::unordered_map: one heap node per unique entry (~48-64B), a hashed lookup per triplet, a reserve() sized on the full triplet count, and a final global O(n log n) sort. For near-field assemblies with tens of millions of triplets this dominated ILU-preconditioner build time (issue #7: 210s at the 3-lambda case) and peak memory (~4x the final CSR size). Replace it with a streaming four-pass pipeline: 1. count in-range entries per row -> rowPtr prefix sum O(N) 2. scatter (col, value) into row-segmented slots O(N) 3. per-row std::stable_sort by column + merge of equal- column runs (accumulate or keep-last) 4. final row offsets while emitting the CSR arrays std::stable_sort preserves insertion order within equal columns, so duplicates accumulate in the same order - and to the same bits - as the hash-map implementation. The public API is unchanged. Measured (MinGW g++ -O2, near-field pattern, 10% duplicates): 7.7M triplets: 1.06s -> 0.106s (~10x) 31M triplets: 0.439s (~71M triplets/s) Tests: all 17 existing CSRBuilderTest cases pass unchanged; add five bit-exact equivalence tests against a verbatim copy of the old hash-map pipeline (accumulate/keep-last, near-field pattern with out-of-range rows, edge cases, complex scalars). Add bench/bench_builder.cpp reproducing the near-field assembly. Fixes #7
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
摘要
cCSRMatrixBuilder::Build()原先通过全局std::unordered_map去重 triplet:每个唯一条目一个堆节点(~48-64B)、每条 triplet 一次 hash 查找、按 triplet 总数reserve()桶数组、最后再全量 O(n log n) 排序。对数千万级 triplet 的近场装配(issue #7 中 3λ 场景 CSR build 210s),这既是时间瓶颈也是内存峰值主因(~4 倍最终 CSR 体积)。改为流式四遍管线,公开 API 不变:
std::stable_sort按列排序 + 相邻等列归并(累加 / 取最后)std::stable_sort保证等列重复项保持插入顺序,因此累加顺序与原 hash-map 实现完全一致——结果逐位相同(浮点累加顺序敏感)。实测数据(MinGW g++ -O2,近场模式 + 10% 重复)
对照 issue #7 目标(3λ CSR build < 30s):新实现吞吐 ~71M triplets/s,30M triplet 场景 < 1s 级别,远超目标。
测试
CSRBuilderTest全部原样通过(接口未变)bench/bench_builder.cpp复现近场装配模式(csr4mpi_bench_builder目标)Fixes #7