Skip to content

Work around grouped-expert LoRA export and checkpoint load in pinned Miles/Bridge - #16

Draft
kevintli wants to merge 4 commits into
devin/1790812059-execute-sample-nonpreemptiblefrom
devin/1790903668-expert-lora-compat
Draft

kevintli wants to merge 4 commits into
devin/1790812059-execute-sample-nonpreemptiblefrom
devin/1790903668-expert-lora-compat

Conversation

@kevintli

@kevintli kevintli commented Oct 2, 2026 •

Copy link
Copy Markdown

Update: bump Megatron-Bridge, keep only the load workaround

Bug 1 turned out to be fixed upstream already: radixark/Megatron-Bridge's bridge branch was rebuilt on NVIDIA main, which doesn't have the gpt-oss export override, so the generic export path now publishes the full (32, ...) expert LoRA tensors. Our Miles commit (5510af6) already expects that Bridge (2e09c234), but Spindle pinned the older 582783a. This PR now:

  • Bumps the Miles image's BRIDGE_REVISION from 582783a to 2e09c234, and MEGATRON_REVISION from 8c1e057 to 8a5dbe5, since the new Bridge imports megatron.core.transformer.mla_qk_norm_config (radixark/Megatron-LM@8a5dbe5)
  • Drops the checkpoint-write half of the workaround (step 1 below), along with publish_noop and its tests
  • Keeps the load-time restore (step 2 below), since Miles main still throws away the result of dist_checkpointing.load

The new Bridge also includes NVIDIA-NeMo/Megatron-Bridge#5376. On the old pin, the grouped-expert linear_fc1 merge on load reordered the gate/up rows, so every run that resumed from a checkpoint (including the "after" runs below) trained with permuted expert gate/up LoRA B after the resume. That also covers the FP32 masters and Adam moments, so it persisted past the first optimizer step.

Validation: resumed the F4 run from its update-16 checkpoint on the new pins for 3 updates (batches 16-18). The restore runs (tensors=24), the expert linear_fc1 LoRA B moves by the same ~6.7% as linear_fc2 from update 16 to 17 with no gate/up permutation, the published gate_up_proj LoRA B is (32, 5760, 32) and bit-identical to the trainer checkpoint, and KL stays at ~0.002 like the earlier "after" run.

The rest of this description is from the original version of this PR.

Summary

While trying to get gpt-oss-20b working on Spindle, I noticed a few bugs in Miles and Megatron-Bridge which caused our sampler policy to diverge from the trainer policy over the course of training. This showed up as unusually large KL and worse reward curves compared to the original blog post results. This PR implements a workaround on the Spindle side before we upstream the actual bugfixes to those repos.

Observed bugs

  1. Export (Megatron-Bridge gpt-oss integration): Due to a bug with expert parallelism shape handling, Megatron-Bridge was exporting bad expert LoRA weights and SGLang would throw them away silently at load time. This means rollout workers were essentially running with untrained expert weights.
  2. Load (Miles lora/checkpoint.load_slot): Miles unintentionally throws away the result of dist_checkpointing.load, which means the first training step uses uninitialized weights (until optim_step copies over the correct ones from the optimizer state), causing a temporary spike in KL / instability.

Workaround

This PR implements miles_runtime/expert_lora_compat.py, which runs only when EXPERT_LORA_COMPAT=1. That compatibility wrapper does the following:

  1. At checkpoint write time: all-gather the expert LoRA tensors over EP groups and rewrite the checkpoint in the correct format
  2. At load time: Rebuild the state dict and correctly copy the params into the live weights

Safety checks:

  • The rewrite requires the published tensor to be bit-identical to a prefix of the gathered experts, with an optional leading singleton axis. It raises on any other layout.
  • It leaves tensors alone that are already (E, ...), and logs publish_noop and restore_noop when upstream is already correct.
  • It asserts PP=1 and expert-TP=1.

Removal: once upstream fixes both bugs, delete the module and the two enabled() branches.

Validation

Setup: gpt-oss-20b, TP=8 EP=8, sec-search-rl, full 24-update runs

The plots below show kl_sample_train_v2 before and after this bug fix. Notice that without the fix:

  • KL grows steadily over time, since the trainer diverges from the untrained-expert-LoRA rollout weights
  • When we restart the run from a checkpoint, KL temporarily drops before climbing steadily -- this is because with bug 2 above, the trainer temporarily has untrained LoRAs (matching the bug on the LoRA side) before it gets fixed on the first optimizer step

F4 KL before vs after
traj-recall KL before vs after
format-penalty KL before vs after

Tests:

  • tests/backends/test_expert_lora_compat.py covers the 4→32 rewrite, the no-op on complete exports, rejection of mismatched or missing tensors, the param copy, and the master refresh
  • Full pytest (with torch installed) gives 714 passed, 1 skipped
  • Ruff check and format are clean on the new files

Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli

@devin-ai-integration

Copy link
Copy Markdown

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

@kevintli
kevintli added this pull request to stack #13 October 2, 2026 22:05
@devin-ai-integration
devin-ai-integration Bot removed this pull request from stack #13 October 2, 2026 23:45
kevintli and others added 2 commits October 2, 2026 23:45
…Miles/Bridge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1790903668-expert-lora-compat branch from 1b9f474 to 84cc6e6 Compare October 2, 2026 23:45
@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #20 October 2, 2026 23:45
kevintli and others added 2 commits October 3, 2026 02:04
…karound

The new Bridge has no gpt-oss export override (per-expert LoRA exports as
(E, ...) through the generic path) and includes NVIDIA-NeMo/Megatron-Bridge#5376,
which keeps grouped-expert SwiGLU gate/up order on checkpoint load. The
checkpoint-load restore is still needed because Miles load_slot drops the
factory-merged weights.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Bridge 2e09c234 imports megatron.core.transformer.mla_qk_norm_config,
which radixark/Megatron-LM added in 8a5dbe5.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant