feat: long convergence resiliency for release tests - #4335
Merged
ko3n1g merged 5 commits intoApr 16, 2026
Conversation
Add LAUNCHER: ft_launcher to all 16 release-type model_config.yaml files so that release runs use ft_launcher instead of torchrun. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: oliver könig <okoenig@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
/ok to test |
Add golden_values_dev_dgx_gb200.json for the deepseekv3_proxy_flex_tp2pp2emp16etp1cp1_gb_200_release test case. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: oliver könig <okoenig@nvidia.com>
Contributor
Author
|
/ok to test |
Downloaded from GitLab pipeline 47719739. Skipped files with only 100 steps (not representative of full release runs). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: oliver könig <okoenig@nvidia.com>
Contributor
Author
|
/ok to test |
Add send_slack_alert() to launch_jet_workload.py that calls notify.py at three checkpoints exclusive to release-type tests: training finished, pipeline failed (retrying), and max attempts exhausted. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: oliver könig <okoenig@nvidia.com>
Contributor
Author
|
/ok to test |
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: oliver könig <okoenig@nvidia.com>
Contributor
Author
|
/ok to test |
chtruong814
approved these changes
Apr 16, 2026
ko3n1g
enabled auto-merge
April 16, 2026 12:36
Contributor
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24514520678 |
yangbofun
pushed a commit
to xlm-research/Megatron-LM
that referenced
this pull request
May 22, 2026
Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
yhgalaxy
pushed a commit
to yhgalaxy/Megatron-LM
that referenced
this pull request
Jun 17, 2026
Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
jon-barker
pushed a commit
to jon-barker/Megatron-LM
that referenced
this pull request
Jul 10, 2026
Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
terminator123
pushed a commit
to 021ai/Megatron-LM
that referenced
this pull request
Aug 3, 2026
Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR improves the resiliency and observability of long-running release test pipelines.
Changes
ft_launcherfor all release tests — addsLAUNCHER: ft_launcherto all 16_release`model_config.yaml` files so release runs use `ft_launcher` instead of `torchrun`Affected test cases (ft_launcher)
Example diff (ft_launcher)
```yaml
TEST_TYPE: "release"
+LAUNCHER: ft_launcher
MODEL_ARGS:
```
Slack alert flow
```
release run iteration
├── training finished → notify.py (pipeline-context: "<test_case> | iteration=N | attempt=N | training finished")
├── pipeline failed, retry → notify.py (pipeline-context: "<test_case> | iteration=N | attempt=N | pipeline failed, retrying")
└── max attempts exhausted → notify.py (pipeline-context: "<test_case> | iteration=N | attempt=N | max attempts exhausted")
```
Test plan