Skip to content

Add ability to perform local gradient accumulation in FP32 for a subset of parameters in the model - #4028

Merged
deepakn94 merged 3 commits into
NVIDIA:mainfrom
deepakn94:dnarayanan/local_grad_accumulation_in_fp32
Mar 25, 2026
Merged

Add ability to perform local gradient accumulation in FP32 for a subset of parameters in the model#4028
deepakn94 merged 3 commits into
NVIDIA:mainfrom
deepakn94:dnarayanan/local_grad_accumulation_in_fp32

Conversation

@deepakn94

Copy link
Copy Markdown
Contributor

No description provided.

deepakn94 and others added 2 commits March 25, 2026 12:15
…accumulation

- Add separate copy to communication buffer before collective is launched (taking care to do this with and without overlapping enabled)
- Zero out extra main_grad tensors when reset() is called
- Log main_grad.dtype for every parameter for debugging
- Nit: Log params in right order
- [From Cursor / Claude Opus 4.5] Copy from communication buffer back to .main_grad to make sure optimizer uses reduced weights in optimizer step
- Scale extra main_grad if it exists if using --calculate-per-token-loss
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@deepakn94
deepakn94 requested review from a team as code owners March 25, 2026 19:20
@svcnvidia-nemo-ci
svcnvidia-nemo-ci marked this pull request as draft March 25, 2026 19:21
@github-actions

Copy link
Copy Markdown
Contributor

This PR has been automatically converted to draft because all PRs must start as drafts.

When you are ready for review, click Ready for Review to begin the review process. This will:

  1. Add the oncall reviewer (optional reviewer)
  2. Add required review teams based on your changes

See the contribution guide for more details.

@copy-pr-bot

copy-pr-bot Bot commented Mar 25, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@svcnvidia-nemo-ci svcnvidia-nemo-ci added this to the Core 0.16 milestone Mar 25, 2026
@deepakn94 deepakn94 changed the title Add ability to perform local gradient accumulation in FP32 for a subset of layers Add ability to perform local gradient accumulation in FP32 for a subset of parameters in the model Mar 25, 2026
@deepakn94
deepakn94 marked this pull request as ready for review March 25, 2026 19:21
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team March 25, 2026 19:21
@deepakn94

Copy link
Copy Markdown
Contributor Author

/claude review

Comment thread megatron/core/distributed/distributed_data_parallel_config.py Outdated
Comment thread megatron/core/distributed/distributed_data_parallel_config.py Outdated
@svcnvidia-nemo-ci svcnvidia-nemo-ci added the Approved All necessary approvals have been made label Mar 25, 2026
@deepakn94

Copy link
Copy Markdown
Contributor Author

/ok to test ee2463f

@deepakn94
deepakn94 force-pushed the dnarayanan/local_grad_accumulation_in_fp32 branch from ee2463f to 1df3f83 Compare March 25, 2026 21:07
@deepakn94
deepakn94 enabled auto-merge March 25, 2026 21:10
@deepakn94

Copy link
Copy Markdown
Contributor Author

/ok to test 1df3f83

@deepakn94
deepakn94 added this pull request to the merge queue Mar 25, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/23566193634

Merged via the queue into NVIDIA:main with commit c586f6d Mar 25, 2026
62 of 63 checks passed
@deepakn94
deepakn94 deleted the dnarayanan/local_grad_accumulation_in_fp32 branch March 25, 2026 22:26
yangbofun pushed a commit to xlm-research/Megatron-LM that referenced this pull request May 22, 2026
…et of parameters in the model (NVIDIA#4028)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
yhgalaxy pushed a commit to yhgalaxy/Megatron-LM that referenced this pull request Jun 17, 2026
…et of parameters in the model (NVIDIA#4028)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
jon-barker pushed a commit to jon-barker/Megatron-LM that referenced this pull request Jul 10, 2026
…et of parameters in the model (NVIDIA#4028)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
…et of parameters in the model (NVIDIA#4028)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: medium

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants