Skip to content

[feat] Generalized Tensor Parallelism (GTP) - #4967

Merged
fanshiqing merged 100 commits into
NVIDIA:mainfrom
fanshiqing:gtp_release
Jul 25, 2026
Merged

[feat] Generalized Tensor Parallelism (GTP)#4967
fanshiqing merged 100 commits into
NVIDIA:mainfrom
fanshiqing:gtp_release

Conversation

@fanshiqing

@fanshiqing fanshiqing commented May 25, 2026

Copy link
Copy Markdown
Member
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do ?

Generalized Tensor Parallelism (GTP) is a light-weight, high-performance and memory-efficient distributed training strategy implemented in Megatron-LM and TransformerEngine. It shards weight tensors across an GTP process group and reconstructs them on-demand via async all-gather, enabling training of larger models without sacrificing throughput by overlapping communication with computation.

GTP's Architecture (Mcore + TE)

• TE ships self-contained: Protocol + dispatcher no-op.
• GTP lives 100% in Mcore as one protocol implementer.

image

Changes Summary

image

Convergence Test (07/25)

image

Config: 1/4 of NT3 Ultra,27Layers512Experts, 32xGB200, GTP16_EP8EGTP2_m1g32_mxfp8_CG, Adam Optimizer, forceBalance, vs. 3D parallel baseline with parallelism=TP8_EP16DP2_m1g32, ~10k steps, dataset=blend_files/1T-phase1var-moresft.json

GTP matches well with the 3d baseline in lm loss curve:

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Linked issue:

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

Co-authored-by: Jieming Zhang <jiemingz@nvidia.com>
Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing
fanshiqing requested review from a team as code owners May 25, 2026 07:25
@copy-pr-bot

copy-pr-bot Bot commented May 25, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@svcnvidia-nemo-ci
svcnvidia-nemo-ci marked this pull request as draft May 25, 2026 07:25
@github-actions

Copy link
Copy Markdown
Contributor

This PR has been automatically converted to draft because all PRs must start as drafts.

When you are ready for review, click Ready for Review to begin the review process. This will:

  1. Add the oncall reviewer (optional reviewer)
  2. Add required review teams based on your changes

See the contribution guide for more details.

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing
fanshiqing marked this pull request as ready for review May 25, 2026 08:11
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team May 25, 2026 08:12
Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing fanshiqing changed the title Generalized Tensor Parallelism (GTP) [feat] Generalized Tensor Parallelism (GTP) May 25, 2026
@fanshiqing

Copy link
Copy Markdown
Member Author

/claude review

Comment thread megatron/core/extensions/transformer_engine.py Outdated
@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test 34d420c

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test 9870010

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test 10fa249

Comment thread megatron/core/optimizer/emerging_optimizers.py
Comment thread megatron/core/optimizer/emerging_optimizers.py

@mkhona-nvidia mkhona-nvidia left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please leave TODOs in the 2 places where a comment has been added

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test dea1206

@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test 269949e

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/30157876620

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/30159727680

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
Buckets were split by the layout's num_optimizer_shards (full DP world) while
the collective runs over the intra-instance data_parallel_group, leaving half
the params with no owning rank. Take the count from the group again.

Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
@fanshiqing

Copy link
Copy Markdown
Member Author

/ok to test 19f8482

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/30171324705

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.