Precompute angle-independent radiation source - #16598
Conversation
|
@whahnsr interesting. Out of curiosity, did you employ an AI model at any point of this optimization work? |
|
Hello Marcos,
Yes I used chatGPT to do most of the heavy lifting, it would have taken me
10x the time without it, but I understand all the stuff that was done and
guided the work by keeping it on track. AI likes to go into rabbit holes
from which it can not recover without help.
Anyway I hope to have contributed some and made your work s little bit
easier.
Regards,
Werner
…On Sat, Sep 26, 2026 at 8:52 AM marcosvanella ***@***.***> wrote:
*marcosvanella* left a comment (firemodels/fds#16598)
<#16598 (comment)>
@whahnsr <https://github.com/whahnsr> interesting. Out of curiosity, did
you employ an AI model at any point of this optimization work?
—
Reply to this email directly, view it on GitHub
<#16598?email_source=notifications&email_token=CPA2VLQV6G74NOUE6N6T3R35Q7Q37A5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBUG43DGNRVGI22M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5847636525>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/CPA2VLUWSGMIHV5NCDLLLET5Q7Q37AVCNFSNUABEKJSXA33TNF2G64TZHMZTOMJUGY3DEMR3JFZXG5LFHM2TKOJQHA2DQMJVG6QXMAQ>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/CPA2VLVHKF2TPBQAEYRTHA35Q7Q37A5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBUG43DGNRVGI22M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJKTGN5XXIZLSL5UW64Y>
and Android
<https://github.com/notifications/mobile/android/CPA2VLXOM2V5L5CXXPOMUG35Q7Q37A5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBUG43DGNRVGI22M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>.
Download it today!
You are receiving this because you were mentioned.Message ID:
***@***.***>
|
|
This looks good, but I am going to test the PR by running all the verification cases. This takes about an hour. |
|
Thanks. The verification cases ran successfully in both debug and release mode. |
Yes, this is in line with what we are seeing using AI. Thanks. |
|
FYI, there are notes on this page that explain what I did to verify that your changes did not "break" any test cases. This was a nice example of how to use AI because the suggested code changes were modest, and we can easily run these verification cases. My larger worry about AI is where someone (or something) suggests a massive change in the code that we will not be able to check and verify. It would be good to limit the PRs to one change at a time rather than a large collection of changes. If in future we need to do a git bisect, and we hit this one code change, it is easy to diagnose the problem. If, however, hundreds of lines have changed, then we ourselves will have to use an AI agent to detect the problem. Of course, that is the concern with AI---that we will become too reliant on it both to suggest changes and then to check those changes. All this being said, what I am suggesting here is no different than what we suggest for code changes made the old fashioned way. Limit the commits to bite-sized changes to retain a nice "paper" trail. |
Summary
Precompute the angle-independent radiation transport source once per spectral band and reuse it during the cylindrical, 2-D Cartesian, and 3-D Cartesian angular sweeps.
Previously, each angular sweep repeatedly evaluated:
for every cell. The new
RTE_SOURCEarray computes this quantity afterUIIOLDis established and before entering the angular sweep.Performance
Tested with the GNU/OpenMPI build using a shortened copy of
Verification/Timing_Benchmarks/openmp_test64a.fds(T_END=1.5) and one OpenMP thread.Three alternating original/optimized runs:
This is a 4.12% elapsed-time reduction and a 1.043x speedup for the complete application run.
The previous dominant radiation source-expression line accounted for 10.73% of sampled cycles. After precomputation, the remaining source-array load accounted for approximately 6.4%.
Validation
All three original/optimized run pairs completed successfully. Extracted time-step diagnostics—including step size, pressure iterations, velocity and pressure errors, CFL values, VN values, and divergence extrema—were byte-for-byte identical for every pair.
Maximum resident memory was effectively unchanged in the measurements (approximately 378 MB).