Skip to content

Build the Miles image on the v0.1.1 release and its locked Megatron stack - #24

Open
kevintli wants to merge 3 commits into
devin/1790812059-execute-sample-nonpreemptiblefrom
devin/1791181625-miles-bridge-pin-bump
Open

kevintli wants to merge 3 commits into
devin/1790812059-execute-sample-nonpreemptiblefrom
devin/1791181625-miles-bridge-pin-bump

Conversation

@kevintli

@kevintli kevintli commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Bumps our Miles trainer image to the fixed Miles release radixark/miles:v0.1.1 and updates the corresponding Megatron-LM / Megatron-Bridge versions.

Moving forward, we will use Megatron-LM and Megatron-Bridge as-is (with the exact versions included in the Miles release), instead of having separate pinned versions for each of the three repos. This ensures that we use combos that upstream (radixark) has already tested with, instead of potentially incompatible versions with bugs.

Motivation

The reason we needed this particular version bump is that on our old pins (Bridge 582783a, Megatron-LM 8c1e057), we hit two bugs that caused our sampler/trainer policies to diverge for MoE LoRA on gpt-oss-20b:

  1. Export (gpt-oss): a stale gpt-oss override exported expert LoRA as (1, E_local, ...), which SGLang silently threw away, so rollout workers ran with untrained expert LoRA weights.
  2. Load (grouped experts): the grouped-expert linear_fc1 merge on load reordered the gate/up rows (fixed by NVIDIA-NeMo/Megatron-Bridge#5376), so any resumed run kept training with permuted expert gate/up LoRA B, including the optimizer state.

v0.1.1 contains both fixes. For full completeness on the sec-search-rl task, we also had to implement #16 as a Spindle-side workaround on top.

Changes

  • Bumped Miles/Megatron-LM/Megatron-Bridge versions

  • RELEASE_CHECK image step: fails the build unless all three of the following hold:

    • /root/miles HEAD equals MILES_COMMIT;
    • /root/Megatron-LM HEAD equals release-lock.json["megatron_commit"];
    • the installed megatron-bridge direct_url.json commit equals the Megatron-Bridge.git@<sha> in Miles' docker/Dockerfile.

    For v0.1.1 this resolves to Megatron-LM f148a32b and Bridge 8cd3466d.

  • Base image stays on CUDA 13 (v0.1.0 was already CUDA 13.0.1; v0.1.1 is 13.0.3, same driver floor). The release does move torch 2.11 → 2.13 (cu130) and the SGLang base 0.5.16 → 0.5.20.

  • miles_runtime/qwen3_vl_cp.py: the new Bridge raises on CP-pre-sharded packed inputs unless the caller passes explicit MRoPE position_ids. That broke the Qwen3.8 CP presets on their first forward. This change computes the ids with Miles' existing helpers and passes them in.

The native megatron backend has its own image and Bridge pin. It only runs full fine-tunes with raw torch.save/torch.load checkpoints, so neither LoRA bug applies to it.

Validation

  • The v0.1.1 image builds, and RELEASE_CHECK passes on H100.
  • tests/providers: 130 passed.
  • Revalidation on v0.1.1 is still pending. The table below was measured on the intermediate pins (Bridge 2e09c234, Megatron-LM 8a5dbe5). v0.1.1 adds 16 Miles commits on top of those, including a multi-LoRA fix to lora/bridge.py. The same checks, plus the Add multi-LoRA GPU CI using a shared RL example #28 multi-client correctness mode, need rerunning on this image.

Setup: 4 train steps, save_state + restore, then 1 more step. Numbers are kl_v2. The last column restores weights only (fresh Adam), as a control for what a broken optimizer restore looks like.

Config Pins sampler vs. trainer (step 4) restored vs. live (step 5) restored without optimizer state (step 5)
Qwen3.5-9B LoRA 16k intermediate 0.0004 0.0005 0.26
Qwen3.6-35B-A3B LoRA 32k (EP=8) old 0.0015 0.143 1.27
Qwen3.6-35B-A3B LoRA 32k (EP=8) intermediate 0.0038 0.0091 0.42
Qwen3.8-27B LoRA 64k (CP=2) old 0.0006 0.0014 0.31
Qwen3.8-27B LoRA 64k (CP=2) intermediate 0.0003 0.0008 0.15

^ Interpretation of the table above:

  • This PR removes the inconsistency between sampler ("restored") and trainer ("live") weights, as measured by KL divergence, in the LoRA MoE case.
  • All other cases did not have this bug to begin with, and KL remains low as expected

End-to-end training test with gpt-oss-20b on sec-search-rl: see #16 for further details

Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #25 October 5, 2026 08:08
@devin-ai-integration devin-ai-integration Bot changed the title Bump Miles Megatron-LM and Megatron-Bridge pins Build the Miles image on the v0.1.1 release and its locked Megatron stack Oct 5, 2026
@devin-ai-integration
devin-ai-integration Bot marked this pull request as ready for review October 7, 2026 00:04
BRIDGE_REPOSITORY = "https://github.com/radixark/Megatron-Bridge.git"
BRIDGE_REVISION = "582783a05442245647239e4c5e7d733d7f0e00ea"
BRIDGE_PATH = "/root/Megatron-Bridge"
RELEASE_CHECK = "; ".join(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move this into a normal Python runction then run it with Modal's Image.run_function?

Something like

def _check_release(miles_commit: str) -> None:
    lock = json.loads(Path(MILES_PATH, "release-lock.json").read_text())
    assert _head(MILES_PATH) == miles_commit, "Miles is not at MILES_COMMIT"
    assert _head(MEGATRON_PATH) == lock["megatron_commit"], "Megatron-LM differs from release-lock.json"
    ...

image = (
    modal.Image.from_registry(BASE_IMAGE)
    ...
    .run_function(_check_release, kwargs={"miles_commit": MILES_COMMIT})

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is fine for now but can we make an upstream PR to Miles?

Comment on lines +13 to -21
# The release tag pins Miles together with the Megatron-LM and Megatron-Bridge
# revisions its image was built and tested with. Update all three by moving to a
# newer release tag, digest, and commit together, then refresh the trainer app.
MILES_RELEASE = "v0.1.1"
BASE_IMAGE = (
f"radixark/miles:{MILES_RELEASE}"
"@sha256:6355834f16bacd35d5d40c43f142e3758376f7b2e8d678bccfe870c092bd96bf"
)
MILES_COMMIT = "2806267d060d51b1d3b62f85a1f9b145047aeef9"
MILES_PATH = "/root/miles"
MEGATRON_REPOSITORY = "https://github.com/radixark/Megatron-LM.git"
MEGATRON_REVISION = "8c1e05747eb612b382df2632783df5c83a853646"
MEGATRON_PATH = "/root/Megatron-LM"
BRIDGE_REPOSITORY = "https://github.com/radixark/Megatron-Bridge.git"
BRIDGE_REVISION = "582783a05442245647239e4c5e7d733d7f0e00ea"
BRIDGE_PATH = "/root/Megatron-Bridge"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

solid change btw, good find

@micahtyong
micahtyong force-pushed the devin/1791181625-miles-bridge-pin-bump branch from 351ee32 to 5748b2f Compare October 7, 2026 01:38
kevintli and others added 3 commits October 7, 2026 17:35
Megatron-Bridge 582783a -> 2e09c234 (what Miles 5510af6 builds against) and Megatron-LM 8c1e057 -> 8a5dbe5, which adds the Megatron-Core APIs the new Bridge imports.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tack

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@micahtyong
micahtyong force-pushed the devin/1791181625-miles-bridge-pin-bump branch from 5748b2f to 7e6504a Compare October 7, 2026 17:35

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants