Skip to content

ci: add configurable launcher support for functional tests (ft_launcher / torchrun) - #4298

Merged
ko3n1g merged 2 commits into
NVIDIA:mainfrom
ko3n1g:ko3n1g/feat/ft-launcher-ci
Apr 14, 2026
Merged

ci: add configurable launcher support for functional tests (ft_launcher / torchrun)#4298
ko3n1g merged 2 commits into
NVIDIA:mainfrom
ko3n1g:ko3n1g/feat/ft-launcher-ci

Conversation

@ko3n1g

@ko3n1g ko3n1g commented Apr 14, 2026

Copy link
Copy Markdown
Contributor

Summary

Introduces a LAUNCHER top-level key in model_config.yaml that lets individual CI test cases opt into ft_launcher (from nvidia-resiliency-ext) instead of the default torchrun.

  • Default is torchrun — all existing test configs continue to work without any changes.
  • Set LAUNCHER: ft_launcher in a test's model_config.yaml to switch to the fault-tolerance launcher.
  • All distributed arguments (--nproc_per_node, --nnodes, --master_addr, etc.) are forwarded unchanged to both launchers since ft_launcher is a drop-in torchrun replacement.

Example model_config.yaml snippet to opt into ft_launcher:

LAUNCHER: ft_launcher
MODEL_ARGS:
  --enable-ft-package: true
  # ... rest of args

Resolves AUT-360

Test plan

  • Existing tests with no LAUNCHER key continue to use torchrun (backward-compatible default via // "torchrun" yq fallback)
  • A test with LAUNCHER: ft_launcher invokes ft_launcher instead of python -m torch.distributed.run
  • Both NeMo and non-NeMo code paths are covered

🤖 Generated with Claude Code

@copy-pr-bot

copy-pr-bot Bot commented Apr 14, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

Introduce a top-level `LAUNCHER` key in `model_config.yaml` (default:
`torchrun`) that allows individual test cases to opt into `ft_launcher`
(nvidia-resiliency-ext fault-tolerance launcher) as a drop-in torchrun
replacement.  All distributed arguments are forwarded unchanged; only the
launcher binary is switched.

Resolves AUT-360

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
@ko3n1g
ko3n1g force-pushed the ko3n1g/feat/ft-launcher-ci branch from 6c8d3cd to e516859 Compare April 14, 2026 14:14
@ko3n1g

ko3n1g commented Apr 14, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

@svcnvidia-nemo-ci svcnvidia-nemo-ci added this to the Core 0.16 milestone Apr 14, 2026
@ko3n1g
ko3n1g marked this pull request as ready for review April 14, 2026 15:52
@ko3n1g
ko3n1g requested a review from a team as a code owner April 14, 2026 15:52
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team April 14, 2026 15:52
Enable ft_launcher (nvidia-resiliency-ext fault-tolerance launcher) for
the following distributed optimizer / checkpoint-resume test cases:

GPT:
- gpt3_mcore_te_tp1_pp4_vp1_resume_torch_dist_dist_optimizer_overlap_grad_reduce_param_gather
- gpt3_mcore_te_tp2_pp1_resume_torch_dist_multi_dist_optimizer_instances
- gpt3_mcore_te_tp4_pp1_resume_torch_dist_dist_optimizer_overlap_grad_reduce_param_gather

MoE:
- gpt3_mcore_te_tp1_pp2_resume_torch_dist_reshard_2x1x4_te_8experts2parallel_dist_optimizer
- gpt3_mcore_te_tp2_pp1_resume_torch_dist_te_8experts2parallel_multi_dist_optimizer_instances
- gpt3_moe_mcore_te_tp4_ep2_etp2_pp2_resume_torch_dist_dist_optimizer

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
@ko3n1g

ko3n1g commented Apr 14, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

@svcnvidia-nemo-ci svcnvidia-nemo-ci added the Approved All necessary approvals have been made label Apr 14, 2026
@ko3n1g
ko3n1g enabled auto-merge April 14, 2026 21:52
@ko3n1g
ko3n1g added this pull request to the merge queue Apr 14, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24424777504

Merged via the queue into NVIDIA:main with commit 6636eb0 Apr 14, 2026
172 of 174 checks passed
@ko3n1g
ko3n1g deleted the ko3n1g/feat/ft-launcher-ci branch April 14, 2026 22:32
yangbofun pushed a commit to xlm-research/Megatron-LM that referenced this pull request May 22, 2026
…er / torchrun) (NVIDIA#4298)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
yhgalaxy pushed a commit to yhgalaxy/Megatron-LM that referenced this pull request Jun 17, 2026
…er / torchrun) (NVIDIA#4298)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
jon-barker pushed a commit to jon-barker/Megatron-LM that referenced this pull request Jul 10, 2026
…er / torchrun) (NVIDIA#4298)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
…er / torchrun) (NVIDIA#4298)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: low Run functional tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants