Skip to content

Perf Config for 1 node GB200 DSV3 - #3545

Closed
gautham-kollu wants to merge 61 commits into
mainfrom
gk/1node_gb200_dsv3
Closed

Perf Config for 1 node GB200 DSV3#3545
gautham-kollu wants to merge 61 commits into
mainfrom
gk/1node_gb200_dsv3

Conversation

@gautham-kollu

Copy link
Copy Markdown
Contributor

What does this PR do ?

Perf Config for 1 node GB200 DSV3 with all the perk knobs

Changelog

  • Add specific line by line info of high level changes in this PR.

GitHub Actions CI

See the CI sectionin the Contributing doc for how to trigger the CI. A Nvidia developer will need to approve and trigger the CI for external contributors.

Before your PR is "Ready for review"

Pre checks:

  • Make sure you read and followed Contributor guidelines
  • Did you write any new necessary tests?
  • Did you add or update any necessary documentation?
  • Does the PR affect components that are optional to install? (Ex: Numba, Pynini, Apex etc)
    • Reviewer: Does the PR have correct import guards for all optional libraries?

If you haven't finished some of the above items you can still open "Draft" PR.

Additional Information

  • Related to # (issue)

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Apr 27, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@gautham-kollu
gautham-kollu marked this pull request as ready for review April 27, 2026 22:20
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test b55883a

@cspades cspades left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already LGTM but I think you turned on an argument that isn't relevant to MFSDP.

Comment thread tests/functional_tests/test_groups/recipes/test_deepseek_recipes_pretrain_perf.py Outdated
@cspades

cspades commented Apr 27, 2026

Copy link
Copy Markdown
Contributor

Also, let's wait until this PR merges and you can use the precision-aware optimization path: NVIDIA/Megatron-LM#4427

Loss should go down, if you can triple-check for me!

@rapatel

rapatel commented Apr 27, 2026

Copy link
Copy Markdown
Contributor

LGTM. Thanks @gautham-kollu

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test 8ea7c14

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu
gautham-kollu requested a review from a team as a code owner April 27, 2026 23:41
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test 77cfa8c

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test 0146b09

cspades
cspades previously approved these changes Apr 28, 2026

@cspades cspades left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Comment thread src/megatron/bridge/training/config.py Outdated

# Validate reuse_grad_buf_for_mxfp8_param_ag when FSDP is not enabled
is_fsdp = self.dist.use_megatron_fsdp or self.ddp.use_megatron_fsdp
if not is_fsdp and self.mixed_precision.fp8_param_gather and self.mixed_precision.fp8_recipe == "mxfp8":

@cspades cspades Apr 28, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

self.mixed_precision.fp8_param_gather -> self.ddp.fp8_param_gather IIRC?

@cspades
cspades dismissed their stale review April 28, 2026 01:27

Config valid failure

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test eb0f20b

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
Signed-off-by: gautham-kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

/ok to test cf09f4b

ko3n1g and others added 25 commits May 12, 2026 15:50
)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
…or FLOPs calculator (#3695)

Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
…256x gpu scale expendable segments addition due to CUDA OOM issue (#3759)

Signed-off-by: Rahul Salagame <rsalagame@nvidia.com>
Signed-off-by: Oliver Koenig <okoenig@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
1
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
Signed-off-by: Malay Nagda <malayn@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
…d generation (#3758)

Signed-off-by: Chen Cui <chcui@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Lianglipeng <lianglipeng@didiglobal.com>
Signed-off-by: hbhflw2000 <417911774@qq.com>
Co-authored-by: Huy Vu <86480512+huvunvidia@users.noreply.github.com>
Co-authored-by: Yuekai Zhang <zhangyuekai@foxmail.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
…workflow (#3779)

Signed-off-by: Yu Yao <yaoyu.094@gmail.com>
Signed-off-by: Malay Nagda <malayn@nvidia.com>
Signed-off-by: gautham-kollu <gkollu@nvidia.com>
1
Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu
gautham-kollu force-pushed the gk/1node_gb200_dsv3 branch from eda5670 to d9f7c49 Compare May 12, 2026 22:50
gautham-kollu added a commit that referenced this pull request May 13, 2026
Squashed version of PR #3545 — perf config for 1-node GB200 DeepSeek V3
with all perk knobs.

- Add GB200 DSV3 perf functional test (L0_Launch_recipes_deepseek_perf.sh,
  test_deepseek_recipes_pretrain_perf.py, gb200 golden values)
- Wire reuse_grad_buf_for_mxfp8_param_ag for FSDP and handle None
  mixed_precision case in training/config.py
- Drop now-unused branch in training/mixed_precision.py
- Update unit tests for the above
- Extend scripts/performance argument_parser and evaluate utilities

Signed-off-by: Gautham Kollu <gkollu@nvidia.com>
@gautham-kollu

Copy link
Copy Markdown
Contributor Author

Closing this in favor of #3796

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:perf Performance optimizations and benchmarking area:recipe Training recipes and launch configs feature New capabilities, enhancements, or enablement work full-test-suite needs-review PR is ready for code review and waiting on a reviewer

Projects

None yet

Development

Successfully merging this pull request may close these issues.