Qwen3-MoE: fused-expert quantisation, BF16 router and expert-wise GPTQ - #324
Open
Shreyas8612 wants to merge 4 commits into
Open
Shreyas8612 wants to merge 4 commits into
Shreyas8612 wants to merge 4 commits into
Conversation
added 4 commits
September 16, 2026 19:13
Qwen3-30B-A3B stores its routed experts in two fused 3-D Parameters, which the per-Linear quantisation path cannot reach. Add expert modules with independent prefill/decode weight banks, a BF16 router kept as a safety island with FP32 softmax and top-k, an MX router-precision ablation lane, phase-aware RMSNorm and decoder-layer residual rounding, and an import-time guard pinning the fused transformers 5.5.0 expert ABI. Attention and MLP are rebased onto the shared Llama phase mixins, adding the MXFP rotate variant and requiring identical K/V cache precision.
Register Qwen3MoeDecoderLayer, Experts, RMSNorm, SparseMoeBlock and
TopKRouter in the prefix and from_self maps so the modify helper can
resolve them, and carry the fused expert decode banks across wholesale
replacement via adopt_decode_{gptq,fp}_expert_weights. The decoder layer
joins the runtime phase pre-hook list. Rotation search gains the MoE
rotate attention classes and excludes up, gate and down projections on
fused experts, which have no rotation lowerer; the scope is recorded in
the results and cached decisions are rejected on mismatch.
Add per-expert GPTQ over the fused gate_up and down banks: a forward hook on the experts module collects each expert's routed token slice, quantises slice by slice and falls back to RTN below min_expert_calibration_hits, recording hits and coverage on the model. The FP snapshot and decode/prefill hand-off now cover the fused Parameters as well as Linears. A device_map_aware mode keeps layers where accelerate placed them with CPU-side activation buffers, the calibration catcher uses a private exception instead of ValueError, and clip-search activations are only retained when clip_search_y is set.
Seventeen tests over a tiny Qwen3MoeConfig: the transformers ABI pin, fused expert prefill/decode banks and the GPTQ decode hand-off end to end, rotation scope exclusion of fused experts, phase-aware attention with tied K/V precision, module replacement plus decoder phase hooks and the BF16 router island. The router-precision file checks MXINT8, E4M3 and E5M2 against explicit matrix products, FP32 route agreement and rejection of unplanned formats.
There was a problem hiding this comment.
🟡 Changes recommended
Critical import-time compatibility and expert-output consistency issues, plus moderate calibration and cache-compatibility issues, remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Adds phase-aware Qwen3-MoE quantization, fused-expert GPTQ, router precision variants, module registration, and compatibility handling.
Changes:
- Adds fused expert quantization and per-expert GPTQ.
- Adds BF16/MX routing and phase-aware modules.
- Updates rotation search, module replacement, and tests.
File summaries
| File | Summary |
|---|---|
test/passes/module/transforms/quantize/test_qwen3_moe_router_precision.py |
Router precision tests |
test/passes/module/transforms/quantize/test_qwen3_moe_phase_quantize.py |
Phase quantization tests |
src/chop/passes/module/transforms/quantize/rotation_search.py |
Rotation exclusions and cache scope |
src/chop/passes/module/transforms/quantize/quantize.py |
Qwen3-MoE quantization hooks |
src/chop/passes/module/transforms/gptq/run.py |
Expert GPTQ and device-aware calibration |
src/chop/passes/module/module_modify_helper.py |
Module replacement and phase-bank transfer |
src/chop/nn/quantized/modules/qwen3_moe/router.py |
BF16 router modules |
src/chop/nn/quantized/modules/qwen3_moe/router_precision.py |
Router precision variants |
src/chop/nn/quantized/modules/qwen3_moe/rms_norm.py |
Phase-aware RMSNorm |
src/chop/nn/quantized/modules/qwen3_moe/mlp.py |
Phase-aware MLP wrappers |
src/chop/nn/quantized/modules/qwen3_moe/experts.py |
Fused expert dispatch and quantization |
src/chop/nn/quantized/modules/qwen3_moe/decoder_layer.py |
Decoder residual handling |
src/chop/nn/quantized/modules/qwen3_moe/compat.py |
Transformers ABI guard |
src/chop/nn/quantized/modules/qwen3_moe/attention.py |
Phase-aware attention |
src/chop/nn/quantized/modules/qwen3_moe/__init__.py |
Qwen3-MoE exports |
src/chop/nn/quantized/modules/phase_config.py |
Phase-bank configuration |
src/chop/nn/quantized/modules/__init__.py |
Module registration |
Review details
Suppressed comments (2)
src/chop/passes/module/transforms/gptq/run.py:93
- In
device_map_awaremode this moves the rotary-embedding module to the layer device for every calibration sample. With layers sharded across devices, that repeatedly copies its buffers/parameters and can make calibration prohibitively slow (and retain allocations on each device); move/cache the rotary embedding once per layer instead of calling.to()inside the per-sample helper.
local_rope = rope.to(layer_device)
src/chop/passes/module/transforms/quantize/rotation_search.py:520
- This makes every pre-existing rotation-search cache unusable: cache files produced before
rotation_scopewas added have no such key, so even a generic Llama/Qwen3 cache now raises instead of being reused. Treat a missing scope as the legacy generic scope (while still rejecting missing scope for fused Qwen3-MoE), or invalidate/upgrade the cache explicitly before comparing it.
if cached.get("rotation_scope") != rotation_scope:
raise ValueError(
- Files reviewed: 17/17 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| raise ImportError("Qwen3MoeTopKRouter does not expose the expected router ABI") | ||
|
|
||
|
|
||
| require_qwen3_moe_fused_abi() |
Comment on lines
+346
to
+350
| # the combine: the loop accumulates expert partials into the BF16 output in | ||
| # expert-index order, the gathered path sums the weighted partials of each | ||
| # token in FP32 and rounds once. ``MASE_MOE_EXPERT_DISPATCH=loop`` restores | ||
| # the original loop; ``gather`` forces the gathered path; the default uses | ||
| # the gathered path below ``MASE_MOE_GATHER_MAX_ASSIGNMENTS`` assignments |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on the phase-split quantisation PR. Qwen3-30B-A3B stores its routed
experts as two fused 3-D parameters that the per-Linear quantisation path
cannot reach. This adds qwen3_moe modules: fused experts with independent
prefill/decode weight banks, a BF16 router with FP32 softmax/top-k plus an MX
router-precision ablation lane, phase-aware attention (MXFP rotate variant,
tied K/V precision), RMSNorm and decoder-layer residual rounding, and an
import-time guard for the fused-expert ABI of transformers 5.5. The modules
are registered for module replacement and phase hooks; rotation search
excludes fused expert projections, which have no rotation lowerer. GPTQ gains
a per-expert path over the fused gate_up/down banks (routed-token slices, RTN
fallback below a calibration-hit threshold, coverage recorded on the model)
and a device_map-aware mode for sharded models.
Testing: 21 new CPU tests over a tiny Qwen3MoeConfig;
pytest test/passes/module/transforms/quantize/ passes (66 tests). Requires a
transformers release with the fused Qwen3-MoE expert layout (verified with
5.5.0; the compat guard refuses older layouts).
Provenance: the MoE expert GPTQ path (commit ec32347) reworks Yuxuan Han's
yx/qwen3-moe-gptq (shared helper names, roughly 40% shared lines).