Skip to content

Work around grouped-expert LoRA export and checkpoint load in pinned Miles/Bridge - #16

Draft
kevintli wants to merge 6 commits into
devin/1791181625-miles-bridge-pin-bumpfrom
devin/1790903668-expert-lora-compat
Draft

kevintli wants to merge 6 commits into
devin/1791181625-miles-bridge-pin-bumpfrom
devin/1790903668-expert-lora-compat

Conversation

@kevintli

@kevintli kevintli commented Oct 2, 2026 •

Copy link
Copy Markdown

Summary

While trying to get gpt-oss-20b working on Spindle, I noticed a few bugs in Miles and Megatron-Bridge which caused our sampler policy to diverge from the trainer policy over the course of training. This showed up as unusually large KL and worse reward curves compared to the original blog post results. This PR implements a workaround on the Spindle side before we upstream the actual bugfixes to those repos.

Observed bugs

  1. Export (Megatron-Bridge gpt-oss integration): Due to a bug with expert parallelism shape handling, Megatron-Bridge was exporting bad expert LoRA weights and SGLang would throw them away silently at load time. This means rollout workers were essentially running with untrained expert weights.
  2. Load (Miles lora/checkpoint.load_slot): Miles unintentionally throws away the result of dist_checkpointing.load, which means the first training step uses uninitialized weights (until optim_step copies over the correct ones from the optimizer state), causing a temporary spike in KL / instability.

Workaround

This PR implements miles_runtime/expert_lora_compat.py, which runs only when EXPERT_LORA_COMPAT=1. That compatibility wrapper does the following:

  1. At checkpoint write time: all-gather the expert LoRA tensors over EP groups and rewrite the checkpoint in the correct format
  2. At load time: Rebuild the state dict and correctly copy the params into the live weights

Safety checks:

  • The rewrite requires the published tensor to be bit-identical to a prefix of the gathered experts, with an optional leading singleton axis. It raises on any other layout.
  • It leaves tensors alone that are already (E, ...), and logs publish_noop and restore_noop when upstream is already correct.
  • It asserts PP=1 and expert-TP=1.

Removal: once upstream fixes both bugs, delete the module and the two enabled() branches.

Validation

Setup: gpt-oss-20b, TP=8 EP=8, sec-search-rl, full 24-update runs

The plots below show kl_sample_train_v2 before and after this bug fix. Notice that without the fix:

  • KL grows steadily over time, since the trainer diverges from the untrained-expert-LoRA rollout weights
  • When we restart the run from a checkpoint, KL temporarily drops before climbing steadily -- this is because with bug 2 above, the trainer temporarily has untrained LoRAs (matching the bug on the LoRA side) before it gets fixed on the first optimizer step

F4 KL before vs after
traj-recall KL before vs after
format-penalty KL before vs after

Tests:

  • tests/backends/test_expert_lora_compat.py covers the 4→32 rewrite, the no-op on complete exports, rejection of mismatched or missing tensors, the param copy, and the master refresh
  • Full pytest (with torch installed) gives 714 passed, 1 skipped
  • Ruff check and format are clean on the new files

Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli

@devin-ai-integration

Copy link
Copy Markdown

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

@kevintli
kevintli added this pull request to stack #13 October 2, 2026 22:05
@devin-ai-integration
devin-ai-integration Bot removed this pull request from stack #13 October 2, 2026 23:45
kevintli and others added 2 commits October 2, 2026 23:45
…Miles/Bridge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1790903668-expert-lora-compat branch from 1b9f474 to 84cc6e6 Compare October 2, 2026 23:45
@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #20 October 2, 2026 23:45
kevintli and others added 4 commits October 3, 2026 02:04
…karound

The new Bridge has no gpt-oss export override (per-expert LoRA exports as
(E, ...) through the generic path) and includes NVIDIA-NeMo/Megatron-Bridge#5376,
which keeps grouped-expert SwiGLU gate/up order on checkpoint load. The
checkpoint-load restore is still needed because Miles load_slot drops the
factory-merged weights.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Bridge 2e09c234 imports megatron.core.transformer.mla_qk_norm_config,
which radixark/Megatron-LM added in 8a5dbe5.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…668-expert-lora-compat

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/backends/test_miles_actor.py
@devin-ai-integration
devin-ai-integration Bot removed this pull request from stack #20 October 5, 2026 08:08
@devin-ai-integration
devin-ai-integration Bot changed the base branch from devin/1790812059-execute-sample-nonpreemptible to devin/1791181625-miles-bridge-pin-bump October 5, 2026 08:08
@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #25 October 5, 2026 08:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant