Skip to content

Qwen3-MoE: fused-expert quantisation, BF16 router and expert-wise GPTQ - #324

Open
Shreyas8612 wants to merge 4 commits into
feat/phase-split-quantisationfrom
feat/qwen3-moe-quantisation
Open

Shreyas8612 wants to merge 4 commits into
feat/phase-split-quantisationfrom
feat/qwen3-moe-quantisation

Conversation

@Shreyas8612

Copy link
Copy Markdown
Collaborator

Stacked on the phase-split quantisation PR. Qwen3-30B-A3B stores its routed
experts as two fused 3-D parameters that the per-Linear quantisation path
cannot reach. This adds qwen3_moe modules: fused experts with independent
prefill/decode weight banks, a BF16 router with FP32 softmax/top-k plus an MX
router-precision ablation lane, phase-aware attention (MXFP rotate variant,
tied K/V precision), RMSNorm and decoder-layer residual rounding, and an
import-time guard for the fused-expert ABI of transformers 5.5. The modules
are registered for module replacement and phase hooks; rotation search
excludes fused expert projections, which have no rotation lowerer. GPTQ gains
a per-expert path over the fused gate_up/down banks (routed-token slices, RTN
fallback below a calibration-hit threshold, coverage recorded on the model)
and a device_map-aware mode for sharded models.

Testing: 21 new CPU tests over a tiny Qwen3MoeConfig;
pytest test/passes/module/transforms/quantize/ passes (66 tests). Requires a
transformers release with the fused Qwen3-MoE expert layout (verified with
5.5.0; the compat guard refuses older layouts).

Provenance: the MoE expert GPTQ path (commit ec32347) reworks Yuxuan Han's
yx/qwen3-moe-gptq (shared helper names, roughly 40% shared lines).

Shreyas8612 added 4 commits September 16, 2026 19:13
Qwen3-30B-A3B stores its routed experts in two fused 3-D Parameters,
which the per-Linear quantisation path cannot reach. Add expert modules
with independent prefill/decode weight banks, a BF16 router kept as a
safety island with FP32 softmax and top-k, an MX router-precision
ablation lane, phase-aware RMSNorm and decoder-layer residual rounding,
and an import-time guard pinning the fused transformers 5.5.0 expert ABI.
Attention and MLP are rebased onto the shared Llama phase mixins, adding
the MXFP rotate variant and requiring identical K/V cache precision.
Register Qwen3MoeDecoderLayer, Experts, RMSNorm, SparseMoeBlock and
TopKRouter in the prefix and from_self maps so the modify helper can
resolve them, and carry the fused expert decode banks across wholesale
replacement via adopt_decode_{gptq,fp}_expert_weights. The decoder layer
joins the runtime phase pre-hook list. Rotation search gains the MoE
rotate attention classes and excludes up, gate and down projections on
fused experts, which have no rotation lowerer; the scope is recorded in
the results and cached decisions are rejected on mismatch.
Add per-expert GPTQ over the fused gate_up and down banks: a forward hook
on the experts module collects each expert's routed token slice,
quantises slice by slice and falls back to RTN below
min_expert_calibration_hits, recording hits and coverage on the model.
The FP snapshot and decode/prefill hand-off now cover the fused
Parameters as well as Linears. A device_map_aware mode keeps layers where
accelerate placed them with CPU-side activation buffers, the calibration
catcher uses a private exception instead of ValueError, and clip-search
activations are only retained when clip_search_y is set.
Seventeen tests over a tiny Qwen3MoeConfig: the transformers ABI pin,
fused expert prefill/decode banks and the GPTQ decode hand-off end to
end, rotation scope exclusion of fused experts, phase-aware attention
with tied K/V precision, module replacement plus decoder phase hooks and
the BF16 router island. The router-precision file checks MXINT8, E4M3
and E5M2 against explicit matrix products, FP32 route agreement and
rejection of unplanned formats.
Copilot AI lite review requested due to automatic review settings September 17, 2026 01:57

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Critical import-time compatibility and expert-output consistency issues, plus moderate calibration and cache-compatibility issues, remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Adds phase-aware Qwen3-MoE quantization, fused-expert GPTQ, router precision variants, module registration, and compatibility handling.

Changes:

  • Adds fused expert quantization and per-expert GPTQ.
  • Adds BF16/MX routing and phase-aware modules.
  • Updates rotation search, module replacement, and tests.
File summaries
File Summary
test/passes/module/transforms/quantize/test_qwen3_moe_router_precision.py Router precision tests
test/passes/module/transforms/quantize/test_qwen3_moe_phase_quantize.py Phase quantization tests
src/chop/passes/module/transforms/quantize/rotation_search.py Rotation exclusions and cache scope
src/chop/passes/module/transforms/quantize/quantize.py Qwen3-MoE quantization hooks
src/chop/passes/module/transforms/gptq/run.py Expert GPTQ and device-aware calibration
src/chop/passes/module/module_modify_helper.py Module replacement and phase-bank transfer
src/chop/nn/quantized/modules/qwen3_moe/router.py BF16 router modules
src/chop/nn/quantized/modules/qwen3_moe/router_precision.py Router precision variants
src/chop/nn/quantized/modules/qwen3_moe/rms_norm.py Phase-aware RMSNorm
src/chop/nn/quantized/modules/qwen3_moe/mlp.py Phase-aware MLP wrappers
src/chop/nn/quantized/modules/qwen3_moe/experts.py Fused expert dispatch and quantization
src/chop/nn/quantized/modules/qwen3_moe/decoder_layer.py Decoder residual handling
src/chop/nn/quantized/modules/qwen3_moe/compat.py Transformers ABI guard
src/chop/nn/quantized/modules/qwen3_moe/attention.py Phase-aware attention
src/chop/nn/quantized/modules/qwen3_moe/__init__.py Qwen3-MoE exports
src/chop/nn/quantized/modules/phase_config.py Phase-bank configuration
src/chop/nn/quantized/modules/__init__.py Module registration
Review details

Suppressed comments (2)

src/chop/passes/module/transforms/gptq/run.py:93

  • In device_map_aware mode this moves the rotary-embedding module to the layer device for every calibration sample. With layers sharded across devices, that repeatedly copies its buffers/parameters and can make calibration prohibitively slow (and retain allocations on each device); move/cache the rotary embedding once per layer instead of calling .to() inside the per-sample helper.
    local_rope = rope.to(layer_device)

src/chop/passes/module/transforms/quantize/rotation_search.py:520

  • This makes every pre-existing rotation-search cache unusable: cache files produced before rotation_scope was added have no such key, so even a generic Llama/Qwen3 cache now raises instead of being reused. Treat a missing scope as the legacy generic scope (while still rejecting missing scope for fused Qwen3-MoE), or invalidate/upgrade the cache explicitly before comparing it.
        if cached.get("rotation_scope") != rotation_scope:
            raise ValueError(
  • Files reviewed: 17/17 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

raise ImportError("Qwen3MoeTopKRouter does not expose the expected router ABI")


require_qwen3_moe_fused_abi()
Comment on lines +346 to +350
# the combine: the loop accumulates expert partials into the BF16 output in
# expert-index order, the gathered path sums the weighted partials of each
# token in FP32 and rounds once. ``MASE_MOE_EXPERT_DISPATCH=loop`` restores
# the original loop; ``gather`` forces the gathered path; the default uses
# the gathered path below ``MASE_MOE_GATHER_MAX_ASSIGNMENTS`` assignments

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants