Skip to content

Debug Windows access violation - #262

Draft
lesteve wants to merge 84 commits into
joblib:masterfrom
lesteve:debug-win-access-violation
Draft

lesteve wants to merge 84 commits into
joblib:masterfrom
lesteve:debug-win-access-violation

Conversation

@lesteve

@lesteve lesteve commented Oct 1, 2026

Copy link
Copy Markdown
Member

Debug #261 since I am unable to reproduce locally the problem in my Windows VM.

@lesteve

lesteve commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

Here is a build log with access violation — expand Capture Windows OpenBLAS diagnostic crashes.

Codex working its magic with WindDbg from the artifact, not sure what to think of it:

  • CI captured an access violation in OpenBLAS’s dgemm_kernel_ZEN + 0x16e, during A.dot(A).
  • The faulting instruction writes to invalid memory. Disassembly suggests a stack scratch-buffer overrun, but the root cause remains unconfirmed.
  • WinDbg recovered only the faulting frame; it could not unwind further.
  • CI and the VM use identical OpenBLAS DLLs. All 30 local default-kernel trials and three forced-Zen trials passed.

@ogrisel

ogrisel commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

In the past I discovered CPU architecture specific openblas bugs which could explain why you cannot reproduce on a local VM.

However, it's a bit weird because this code is mostly about thread management operations rather than linear algebra compute kernels.

@lesteve

lesteve commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

Summary for the 50 runs with both with and without ProcDump commit messages

In every new failing build, both attempts crashed: without ProcDump (exit 139), then with ProcDump (dump captured). These add 18 crashing diagnostic attempts and 9 dumps.

Results across all 50 comparison runs

CPU model OpenBLAS target Both passed Both crashed
AMD EPYC 7763 Zen 31 0
AMD EPYC 9V74 Zen 5 4
AMD EPYC 9V45 Haswell 0 5
Intel Xeon Platinum 8573C Haswell 1 0
Intel Xeon 6973P-C Haswell 2 0
Intel Xeon Platinum 8370C Haswell 2 0
Total 41 9

Observations

  • All nine failures occurred on EPYC 9V74 or 9V45, while all 31 EPYC 7763 runs passed.
  • EPYC 9V74 also passed five times, so the reported CPU model alone does not determine the outcome.
  • The attempts with and without ProcDump agreed in every run. This points toward something shared by the runner/environment and argues against ProcDump being the main factor suppressing the crash.
  • The crashes are not exclusive to the Zen target: all five EPYC 9V45 failures selected Haswell.

Failing builds

All commit messages are “both with and without ProcDump N”.

N CPU model OpenBLAS target Build log
8 AMD EPYC 9V45 Haswell Log
11 AMD EPYC 9V74 Zen Log
18 AMD EPYC 9V74 Zen Log
23 AMD EPYC 9V45 Haswell Log
29 AMD EPYC 9V74 Zen Log
41 AMD EPYC 9V45 Haswell Log
44 AMD EPYC 9V45 Haswell Log
48 AMD EPYC 9V74 Zen Log
49 AMD EPYC 9V45 Haswell Log

@lesteve

lesteve commented Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

With Codex I looked at whether this was a known issue in OpenBLAS and it seems like it is. I need to look at it closer but it seems reasonable as a first glance and if right it would mean that the OpenBLAS issues have been fixed in the develop branch and that the next OpenBLAS release will fix it. Given that it's a 20% chance of "random" failure, maybe we can live with it for now?

For example: OpenMathLib/OpenBLAS#6013 has been fixed by OpenMathLib/OpenBLAS#6014 but the fix hasn't been released yet.


OpenBLAS #6013 reports OpenBLAS 0.3.34 Windows crashes with this split:

CPU AVX-512 exposed to VM Selected kernel Result
EPYC 9V74 Yes Zen 2/2 crashed
EPYC 9V74 No Zen 8/8 passed
EPYC 7763 No Zen 5/5 passed
Xeon 8573C Yes Haswell Passed

They traced it to a stack-buffer overrun. New Zen 4/5 tuning increased the DGEMM block size to 512 when particular cache characteristics and AVX-512 were detected. However, the Windows Zen/Haswell kernel’s buffer was sized for 128. Their single-threaded crash even occurred at dgemm_kernel_ZEN + 0x16e, matching our reported location. Root-cause analysis.

This is a much stronger explanation than my earlier scheduling hypotheses. We still need to verify AVX-512 exposure on our runners to confirm the connection.

The architecture-detection issue is separate:

  • 9V74 → Zen: consistent with the affected Windows build having AVX-512 kernels disabled at build time.
  • 9V45 → Haswell: the Zen 5 family was missing from explicit runtime detection, so it used feature-based fallback.
  • Issue #6029 reports Zen 5 misidentification on FreeBSD, so detection problems are not Windows-specific. PR #6033, merged September 15, adds the missing Zen 5/5c/6 detection.

The buffer-overrun bug is not exclusively Windows-specific either. The follow-up reproduced it on Linux when forcing Zen or Haswell; normal AVX-512 builds generally selected an unaffected kernel. Follow-up reproduction. PR #6027, merged September 13, restricts the larger blocks to appropriate AVX-512 kernels.

I would now prioritise:

  1. Log NumPy’s CPU features before each attempt:

    from numpy._core._multiarray_umath import __cpu_features__
    print(__cpu_features__, flush=True)
    
  2. Compare failing and passing runners’ AVX512F.

  3. Test OpenBLAS 0.3.33, which passed in the upstream report, or a build verified to contain the fix.

OPENBLAS_L2_SIZE=2048 also bypassed the faulty tuning in the follow-up, making it a useful diagnostic experiment. Forcing OPENBLAS_CORETYPE=ZEN is not a fix for this bug.

[OpenBLAS \#6013]() reports OpenBLAS 0.3.34 Windows crashes with this split:
CPU AVX-512 exposed to VM Selected kernel Result
EPYC 9V74 Yes Zen 2/2 crashed
EPYC 9V74 No Zen 8/8 passed
EPYC 7763 No Zen 5/5 passed
Xeon 8573C Yes Haswell Passed

They traced it to a stack-buffer overrun. New Zen 4/5 tuning increased the DGEMM block size to 512 when particular cache characteristics and AVX-512 were detected. However, the Windows Zen/Haswell kernel’s buffer was sized for 128. Their single-threaded crash even occurred at dgemm_kernel_ZEN + 0x16e, matching our reported location. [Root-cause analysis](OpenMathLib/OpenBLAS#6013 (comment)).

This is a much stronger explanation than my earlier scheduling hypotheses. We still need to verify AVX-512 exposure on our runners to confirm the connection.

The architecture-detection issue is separate:

  • 9V74 → Zen: consistent with the affected Windows build having AVX-512 kernels disabled at build time.
  • 9V45 → Haswell: the Zen 5 family was missing from explicit runtime detection, so it used feature-based fallback.
  • Issue #6029 reports Zen 5 misidentification on FreeBSD, so detection problems are not Windows-specific. PR #6033, merged September 15, adds the missing Zen 5/5c/6 detection.

The buffer-overrun bug is not exclusively Windows-specific either. The follow-up reproduced it on Linux when forcing Zen or Haswell; normal AVX-512 builds generally selected an unaffected kernel. [Follow-up reproduction](OpenMathLib/OpenBLAS#6021 (comment)). PR #6027, merged September 13, restricts the larger blocks to appropriate AVX-512 kernels.

I would now prioritise:

  1. Log NumPy’s CPU features before each attempt:

    from numpy._core._multiarray_umath import __cpu_features__
    print(__cpu_features__, flush=True)
    
  2. Compare failing and passing runners’ AVX512F.

  3. Test OpenBLAS 0.3.33, which passed in the upstream report, or a build verified to contain the fix.

OPENBLAS_L2_SIZE=2048 also bypassed the faulty tuning in the follow-up, making it a useful diagnostic experiment. Forcing OPENBLAS_CORETYPE=ZEN is not a fix for this bug.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants