Conversation
|
Here is a build log with access violation — expand Capture Windows OpenBLAS diagnostic crashes. Codex working its magic with WindDbg from the artifact, not sure what to think of it:
|
|
In the past I discovered CPU architecture specific openblas bugs which could explain why you cannot reproduce on a local VM. However, it's a bit weird because this code is mostly about thread management operations rather than linear algebra compute kernels. |
Summary for the 50 runs with both with and without ProcDump commit messagesIn every new failing build, both attempts crashed: without ProcDump (exit 139), then with ProcDump (dump captured). These add 18 crashing diagnostic attempts and 9 dumps. Results across all 50 comparison runs
Observations
Failing buildsAll commit messages are “both with and without ProcDump N”.
|
|
With Codex I looked at whether this was a known issue in OpenBLAS and it seems like it is. I need to look at it closer but it seems reasonable as a first glance and if right it would mean that the OpenBLAS issues have been fixed in the For example: OpenMathLib/OpenBLAS#6013 has been fixed by OpenMathLib/OpenBLAS#6014 but the fix hasn't been released yet. OpenBLAS #6013 reports OpenBLAS 0.3.34 Windows crashes with this split:
They traced it to a stack-buffer overrun. New Zen 4/5 tuning increased the DGEMM block size to 512 when particular cache characteristics and AVX-512 were detected. However, the Windows Zen/Haswell kernel’s buffer was sized for 128. Their single-threaded crash even occurred at This is a much stronger explanation than my earlier scheduling hypotheses. We still need to verify AVX-512 exposure on our runners to confirm the connection. The architecture-detection issue is separate:
The buffer-overrun bug is not exclusively Windows-specific either. The follow-up reproduced it on Linux when forcing I would now prioritise:
They traced it to a stack-buffer overrun. New Zen 4/5 tuning increased the DGEMM block size to 512 when particular cache characteristics and AVX-512 were detected. However, the Windows Zen/Haswell kernel’s buffer was sized for 128. Their single-threaded crash even occurred at This is a much stronger explanation than my earlier scheduling hypotheses. We still need to verify AVX-512 exposure on our runners to confirm the connection. The architecture-detection issue is separate:
The buffer-overrun bug is not exclusively Windows-specific either. The follow-up reproduced it on Linux when forcing I would now prioritise:
|
Debug #261 since I am unable to reproduce locally the problem in my Windows VM.