Repository navigation
Conversation
Contributor
|
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
|
micahtyong
approved these changes
Oct 7, 2026
| BRIDGE_REPOSITORY = "https://github.com/radixark/Megatron-Bridge.git" | ||
| BRIDGE_REVISION = "582783a05442245647239e4c5e7d733d7f0e00ea" | ||
| BRIDGE_PATH = "/root/Megatron-Bridge" | ||
| RELEASE_CHECK = "; ".join( |
Collaborator
There was a problem hiding this comment.
Can we move this into a normal Python runction then run it with Modal's Image.run_function?
Something like
def _check_release(miles_commit: str) -> None:
lock = json.loads(Path(MILES_PATH, "release-lock.json").read_text())
assert _head(MILES_PATH) == miles_commit, "Miles is not at MILES_COMMIT"
assert _head(MEGATRON_PATH) == lock["megatron_commit"], "Megatron-LM differs from release-lock.json"
...
image = (
modal.Image.from_registry(BASE_IMAGE)
...
.run_function(_check_release, kwargs={"miles_commit": MILES_COMMIT})
Collaborator
There was a problem hiding this comment.
This is fine for now but can we make an upstream PR to Miles?
micahtyong
reviewed
Oct 7, 2026
Comment on lines
+13
to
-21
| # The release tag pins Miles together with the Megatron-LM and Megatron-Bridge | ||
| # revisions its image was built and tested with. Update all three by moving to a | ||
| # newer release tag, digest, and commit together, then refresh the trainer app. | ||
| MILES_RELEASE = "v0.1.1" | ||
| BASE_IMAGE = ( | ||
| f"radixark/miles:{MILES_RELEASE}" | ||
| "@sha256:6355834f16bacd35d5d40c43f142e3758376f7b2e8d678bccfe870c092bd96bf" | ||
| ) | ||
| MILES_COMMIT = "2806267d060d51b1d3b62f85a1f9b145047aeef9" | ||
| MILES_PATH = "/root/miles" | ||
| MEGATRON_REPOSITORY = "https://github.com/radixark/Megatron-LM.git" | ||
| MEGATRON_REVISION = "8c1e05747eb612b382df2632783df5c83a853646" | ||
| MEGATRON_PATH = "/root/Megatron-LM" | ||
| BRIDGE_REPOSITORY = "https://github.com/radixark/Megatron-Bridge.git" | ||
| BRIDGE_REVISION = "582783a05442245647239e4c5e7d733d7f0e00ea" | ||
| BRIDGE_PATH = "/root/Megatron-Bridge" |
Collaborator
There was a problem hiding this comment.
solid change btw, good find
micahtyong
force-pushed
the
devin/1791181625-miles-bridge-pin-bump
branch
from
October 7, 2026 01:38
351ee32 to
5748b2f
Compare
Megatron-Bridge 582783a -> 2e09c234 (what Miles 5510af6 builds against) and Megatron-LM 8c1e057 -> 8a5dbe5, which adds the Megatron-Core APIs the new Bridge imports. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tack Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
micahtyong
force-pushed
the
devin/1791181625-miles-bridge-pin-bump
branch
from
October 7, 2026 17:35
5748b2f to
7e6504a
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bumps our Miles trainer image to the fixed Miles release
radixark/miles:v0.1.1and updates the corresponding Megatron-LM / Megatron-Bridge versions.Moving forward, we will use Megatron-LM and Megatron-Bridge as-is (with the exact versions included in the Miles release), instead of having separate pinned versions for each of the three repos. This ensures that we use combos that upstream (radixark) has already tested with, instead of potentially incompatible versions with bugs.
Motivation
The reason we needed this particular version bump is that on our old pins (Bridge
582783a, Megatron-LM8c1e057), we hit two bugs that caused our sampler/trainer policies to diverge for MoE LoRA ongpt-oss-20b:(1, E_local, ...), which SGLang silently threw away, so rollout workers ran with untrained expert LoRA weights.linear_fc1merge on load reordered the gate/up rows (fixed by NVIDIA-NeMo/Megatron-Bridge#5376), so any resumed run kept training with permuted expert gate/up LoRA B, including the optimizer state.v0.1.1 contains both fixes. For full completeness on the sec-search-rl task, we also had to implement #16 as a Spindle-side workaround on top.
Changes
Bumped Miles/Megatron-LM/Megatron-Bridge versions
RELEASE_CHECKimage step: fails the build unless all three of the following hold:/root/milesHEAD equalsMILES_COMMIT;/root/Megatron-LMHEAD equalsrelease-lock.json["megatron_commit"];megatron-bridgedirect_url.jsoncommit equals theMegatron-Bridge.git@<sha>in Miles'docker/Dockerfile.For v0.1.1 this resolves to Megatron-LM
f148a32band Bridge8cd3466d.Base image stays on CUDA 13 (v0.1.0 was already CUDA 13.0.1; v0.1.1 is 13.0.3, same driver floor). The release does move torch 2.11 → 2.13 (cu130) and the SGLang base 0.5.16 → 0.5.20.
miles_runtime/qwen3_vl_cp.py: the new Bridge raises on CP-pre-sharded packed inputs unless the caller passes explicit MRoPEposition_ids. That broke the Qwen3.8 CP presets on their first forward. This change computes the ids with Miles' existing helpers and passes them in.The native
megatronbackend has its own image and Bridge pin. It only runs full fine-tunes with rawtorch.save/torch.loadcheckpoints, so neither LoRA bug applies to it.Validation
RELEASE_CHECKpasses on H100.tests/providers: 130 passed.2e09c234, Megatron-LM8a5dbe5). v0.1.1 adds 16 Miles commits on top of those, including a multi-LoRA fix tolora/bridge.py. The same checks, plus the Add multi-LoRA GPU CI using a shared RL example #28 multi-client correctness mode, need rerunning on this image.Setup: 4 train steps,
save_state+ restore, then 1 more step. Numbers arekl_v2. The last column restores weights only (fresh Adam), as a control for what a broken optimizer restore looks like.^ Interpretation of the table above:
End-to-end training test with gpt-oss-20b on sec-search-rl: see #16 for further details
Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli