Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
80 commits
Select commit Hold shift + click to select a range
0136876
Supporting inference when called within an asyncio loop (#2816)
shanmugamr1992 Jan 23, 2026
03e0915
Update type hints and doc strings for moe_utils.py (#2821)
JavaZeroo Jan 23, 2026
10c6f01
Remove calculation of padding token in moe routing loss (#2142)
HaochenYuan Jan 23, 2026
029f48f
Bug fix with --no-use-tokenizer-from-checkpoint-args (#3049)
jon-barker Jan 23, 2026
0683679
Revert "Bug fix with --no-use-tokenizer-from-checkpoint-args (#3049)"…
thomasdhc Jan 23, 2026
93567e8
Add health endpoint to dynamic text gen server (#3009)
santhnm2 Jan 23, 2026
3593301
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jan 24, 2026
30dea5d
ci: Skip test_precision_aware_optimizer (#3062)
thomasdhc Jan 23, 2026
485ed18
Support multimodule communication (#2031)
yaoyu-33 Jan 24, 2026
0972f02
Revert "Support multimodule communication (#2031)" (#3068)
ko3n1g Jan 24, 2026
4cfaa7d
Revert "Remove calculation of padding token in moe routing loss (#214…
ko3n1g Jan 24, 2026
dbde759
Add ability to save wgrads and dgrads (#3032)
deepakn94 Jan 24, 2026
55fe705
ci: Mark test_mode_partial_cudagraph unit tests as flaky (#3064)
chtruong814 Jan 24, 2026
3a7d74d
Keep FSDP's and DDP's finish_grad_sync API identical (#3070)
deepakn94 Jan 24, 2026
389436b
(REPLAY) Bug fix with --no-use-tokenizer-from-checkpoint-args (#3059)
jon-barker Jan 24, 2026
53a2b19
Optimizing post-processing of requests (#2920)
sidsingh-nvidia Jan 24, 2026
369e0eb
Fix broken functional tests in #2920 (#3071)
sidsingh-nvidia Jan 25, 2026
b2e9390
fix ep weight gradnorm/num_zero calculation error for muon (#3024)
FDecaYed Jan 26, 2026
497d42d
[training migration] Add LoggerConfig dataclass (#2414)
maanug-nv Jan 26, 2026
06d0f46
Added --ft-num-warmup-iters option. (#3052)
hexinw-nvidia Jan 26, 2026
642fdd9
Reapply "Various CUDA graph improvements on capture time, replay time…
jiemingz Jan 26, 2026
bb42a00
fix(fsdp): add CLI argument for outer_dp_sharding_strategy (#3053)
liuyun7345 Jan 26, 2026
94d8186
ci: Log node name (#3081)
ko3n1g Jan 26, 2026
528cb2e
add all_gather process-group for overlapping in fsdp disributed train…
jeffnvidia Jan 26, 2026
23a76d1
docs: Release docs (#3055)
ko3n1g Jan 26, 2026
b47c376
Support NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 FP8/NVFP4 PTQ in example …
ChenhanYu Jan 27, 2026
35e85a6
Reapply "Support multimodule communication (#2031)" (#3068)
ko3n1g Jan 27, 2026
db6b895
Add router replay for MoE models (#2101)
litianjian Jan 26, 2026
7031953
ci: Disable gpt_dynamic_inference_tp1_pp1_dp8_583m_throughputtest_zmq…
ko3n1g Jan 27, 2026
dea21a0
ci: Repeat func tests, save logs of unit tests and lessen debug outpu…
ko3n1g Jan 27, 2026
dd83fc6
ci: Update improvement of step-time (#3104)
ko3n1g Jan 27, 2026
2bdf7e1
ci: Add GPU health checks (#3100)
ko3n1g Jan 27, 2026
d68721b
Harden GRPO functional tests (#3065)
jon-barker Jan 27, 2026
4015ff1
Inference functional tests: Write outputs to INFERENCE_OUTPUT_PATH in…
mathemakitten Jan 27, 2026
0888a06
Update moe readme. (#2830)
Jan 27, 2026
2b02a28
build: Bump to TE2.12 (#3086)
ko3n1g Jan 27, 2026
6cf285b
Logging cleanup (only log on rank 0 if possible) (#3036)
deepakn94 Jan 27, 2026
65217aa
Move all bert and t5 tests to nightly (#3106)
Phlip79 Jan 27, 2026
4fb549f
Create greptile.json (#3087)
Phlip79 Jan 28, 2026
6273d74
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jan 28, 2026
33224cc
Fix bug of reuse_grad_buf_for_mxfp8_param_ag (#2802)
kunlunl Jan 28, 2026
fb6a592
Fix for Hybrid CP (#3091)
parthmannan Jan 27, 2026
f6c8a61
Fix GRPO re-fit functional test (#3113)
jon-barker Jan 28, 2026
991138e
Minimize README contents (#3020)
megnvidia Jan 28, 2026
964c902
Add end-to-end tests for M-FSDP and ND-Parallel (#3031)
shjwudp Jan 28, 2026
38cd9fc
Fix Multimodal Dockerfile (#3006)
faradawn Jan 28, 2026
b6b49e7
[M-FSDP] Fix double buffering not working with activation recompute (…
shjwudp Jan 28, 2026
d4f9347
[training migration] Add CheckpointConfig dataclass (#2431)
maanug-nv Jan 28, 2026
d5cac80
chore: rotate oncall schedule
github-actions[bot] Jan 28, 2026
008926a
[training migration] Add StragglerDetectionConfig dataclass (#2435)
maanug-nv Jan 28, 2026
1453f94
Standardize RL unit tests (#3088)
tdene Jan 28, 2026
71c49b5
Fix for PR-2142 (#3096)
HaochenYuan Jan 28, 2026
93ddc24
Use the latest hybrid-ep (#3093)
Autumn1998 Jan 28, 2026
fc6969f
remove retro (#3001)
dimapihtar Jan 28, 2026
c22615e
ci: Mark test_compatible_with_nd_parallel as flaky (#3122)
ko3n1g Jan 28, 2026
a883e96
build: Use merge-commit-sha for container (#3123)
ko3n1g Jan 28, 2026
42986ac
Refactor `rl_offload_kv_cache_during_training` to offload KV cache to…
mathemakitten Jan 28, 2026
d41bf66
Disable Greptile status comments (#3127)
Phlip79 Jan 28, 2026
e2ff203
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jan 29, 2026
50132f2
ci: Add unit tests to merge queue (#3125)
ko3n1g Jan 28, 2026
9f05aac
Create CodeRabbit config (#3131)
Phlip79 Jan 29, 2026
f0b1cb2
build: Explicitly set minimum torch version to >= 2.6.0 (#3085)
chtruong814 Jan 29, 2026
190f5b6
Move kitchen extension file to private kitchen repository (#2779)
kwyss-nvidia Jan 29, 2026
287d2f4
Fix RL optimizer offload (#3112)
jon-barker Jan 29, 2026
3955c49
Revert "Fix RL optimizer offload (#3112)" (#3141)
ko3n1g Jan 29, 2026
4913c46
Revise and move KD docs (#3108)
AAnoosheh Jan 29, 2026
558fdaf
build: Bump FLA (#3139)
ko3n1g Jan 29, 2026
409af92
ci: Add job timeouts (#3142)
ko3n1g Jan 29, 2026
f4af1bf
ci: Set NODE_RANK (#3143)
ko3n1g Jan 29, 2026
0b619c2
Multiturn rollout support prep (#2966)
yobibyte Jan 29, 2026
36411dd
Reapply 3955c49ed9af5e5b38dccdd30c1323c00b9bcd29 (#3146)
jon-barker Jan 29, 2026
dbd8dda
Revert "Multiturn rollout support prep (#2966)" (#3153)
ko3n1g Jan 29, 2026
f58b6d6
Fix coderabbit instructions error (#3150)
Phlip79 Jan 29, 2026
063624b
Force input ids generated by mock dataset are < vocab_size (#2945)
asolergi-nv Jan 29, 2026
4652e7b
Add a check to make sure we are distributing all the layers when usin…
asolergi-nv Jan 29, 2026
18deeff
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jan 30, 2026
67f3515
Automatically choose available ports in ZMQ (#2278)
tdene Jan 29, 2026
639c08a
Generate arguments from TransformerConfig (#2896)
maanug-nv Jan 30, 2026
a9fb6c8
Merge branch 'main' into deyuf/dev_pull_main_260130
FDecaYed Jan 30, 2026
20e8ac8
fix merge main issues
FDecaYed Jan 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions .coderabbit.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json
language: "en-US"

# Only comment on Critical/Major bugs. No Minor, Trivial, or style comments.
tone_instructions: "Only comment on Critical or Major bugs. Never comment on Minor issues, style, refactoring, or suggestions. When in doubt, stay silent."

reviews:
# Use chill profile - filters out nitpicks automatically
profile: "chill"

# Disable all summary features
high_level_summary: false
high_level_summary_in_walkthrough: false

# Disable walkthrough comment entirely
collapse_walkthrough: true
changed_files_summary: false
sequence_diagrams: false

# Disable status/effort estimates
review_status: false
commit_status: false
estimate_code_review_effort: false

# Disable auto-suggestions for labels/reviewers
suggested_labels: false
suggested_reviewers: false

# Disable related issues/PRs lookup
assess_linked_issues: false
related_issues: false
related_prs: false

# Auto-review disabled - only review when explicitly requested via @coderabbitai review
auto_review:
enabled: false

chat:
auto_reply: true
32 changes: 32 additions & 0 deletions .github/actions/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,38 @@ runs:
shell: bash -x -e -u -o pipefail {0}
run: echo "node_name=$NODE_NAME" | tee -a "$GITHUB_OUTPUT"

- name: GPU Sanity Check
shell: bash -x -e -u -o pipefail {0}
run: |
echo "Starting GPU Sanity Check..."

# 1. Check for active Compute Processes
# query-compute-apps returns a list of PIDs using the GPU. If empty, we are good.
OPEN_PROCESSES=$(docker run --rm --gpus all ubuntu nvidia-smi --query-compute-apps=pid,process_name --format=csv,noheader)

if [ -n "$OPEN_PROCESSES" ]; then
echo "::error::❌ GPU is not clean! Found active processes:"
echo "$OPEN_PROCESSES"
else
echo "✅ No active compute processes found."
fi

# 2. Check VRAM Usage (Optional but recommended)
# We allow a small buffer (e.g., < 300MiB) for driver overhead/Xorg,
# though on headless K8s nodes this should be very close to 0.

MEMORY_USAGES=$(docker run --rm --gpus all ubuntu nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits)

# Check each GPU visible to the container
for MEMORY in $MEMORY_USAGES; do
if [ "$MEMORY" -gt 300 ]; then
echo "::error::❌ GPU VRAM usage is suspiciously high: ${MEMORY} MiB"
fi
done

echo "✅ GPU Memory is clean (all < 300 MiB)."
echo "Ready to start workflow."

- name: Checkout repository
uses: actions/checkout@v2

Expand Down
2 changes: 1 addition & 1 deletion .github/copy-pr-bot.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
enabled: true
auto_sync_draft: false
auto_sync_ready: true
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "ChenhanYu", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Phlip79", "QiZhangNV", "ShriyaRishab", "Victarry", "Wohox", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "asolergi-nv", "buptzyb", "chtruong814", "cspades", "cuichenx", "deepakn94", "dimapihtar", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "frsun-nvda", "gautham-kollu", "gdengk", "guyueh1", "hxbai", "jalbericiola", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kanz-nv", "kevalmorabia97", "ko3n1g", "kunlunl", "kvareddy", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mehraakash", "mkhona-nvidia", "pablo-garay", "parthmannan", "pthombre", "rogerwaleffe", "sanandaraj5597", "santhnm2", "sbak5", "shanmugamr1992", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "trintamaki", "tylerpoon", "wdykas", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yuzhongw-nvidia", "zhongbozhu"]
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "ChenhanYu", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Phlip79", "QiZhangNV", "ShriyaRishab", "Victarry", "Wohox", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "asolergi-nv", "buptzyb", "chtruong814", "cspades", "cuichenx", "deepakn94", "dimapihtar", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "frsun-nvda", "gautham-kollu", "gdengk", "guyueh1", "hxbai", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kanz-nv", "kevalmorabia97", "ko3n1g", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mehraakash", "mkhona-nvidia", "parthmannan", "prajwal1210", "pthombre", "rogerwaleffe", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "trintamaki", "tylerpoon", "wdykas", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yuzhongw-nvidia", "zhongbozhu"]
8 changes: 4 additions & 4 deletions .github/oncall_schedule.json
Original file line number Diff line number Diff line change
@@ -1,8 +1,4 @@
[
{
"user": "maanug-nv",
"date": "2026-01-21"
},
{
"user": "dimapihtar",
"date": "2026-01-28"
Expand Down Expand Up @@ -46,5 +42,9 @@
{
"user": "maanug-nv",
"date": "2026-04-08"
},
{
"user": "BoxiangW",
"date": "2026-04-15"
}
]
29 changes: 23 additions & 6 deletions .github/workflows/cicd-main.yml
Original file line number Diff line number Diff line change
Expand Up @@ -204,8 +204,28 @@ jobs:
&& needs.pre-flight.outputs.is_merge_group == 'false'
&& !cancelled()
steps:
- name: Get PR info
id: get-pr-info
if: startsWith(github.ref, 'refs/heads/pull-request/')
uses: nv-gha-runners/get-pr-info@main

- name: Get merge commit sha
shell: bash -x -e -u -o pipefail {0}
id: sha
env:
IS_PR: ${{ startsWith(github.ref, 'refs/heads/pull-request/') }}
run: |
if [[ "$IS_PR" == "true" ]]; then
SHA=${{ fromJSON(steps.get-pr-info.outputs.pr-info || '{}').merge_commit_sha }}
else
SHA=${GITHUB_SHA}
fi
echo "main=${SHA}" | tee -a "$GITHUB_OUTPUT"

- name: Checkout
uses: actions/checkout@v4
with:
ref: ${{ steps.sha.outputs.main }}

- name: Setup python
uses: actions/setup-python@v5
Expand All @@ -218,11 +238,6 @@ jobs:
apt-get update
apt-get install -y gh

- name: Get PR info
id: get-pr-info
if: startsWith(github.ref, 'refs/heads/pull-request/')
uses: nv-gha-runners/get-pr-info@main

- name: Has lts label
id: has-lts-label
env:
Expand Down Expand Up @@ -347,6 +362,7 @@ jobs:
- cicd-container-build
- cicd-parse-unit-tests
runs-on: ${{ needs.is-not-external-contributor.outputs.selected_runner }}
timeout-minutes: 60
name: "${{ matrix.bucket }} - latest"
if: |
(
Expand Down Expand Up @@ -376,6 +392,7 @@ jobs:

cicd-parse-integration-tests:
runs-on: ubuntu-latest
timeout-minutes: 60
needs:
- pre-flight
- cicd-wait-in-queue
Expand All @@ -387,8 +404,8 @@ jobs:
success()
|| needs.pre-flight.outputs.is_ci_workload == 'true'
|| needs.pre-flight.outputs.force_run_all == 'true'
|| needs.pre-flight.outputs.is_merge_group == 'true'
)
&& needs.pre-flight.outputs.is_merge_group == 'false'
&& !cancelled()
outputs:
integration-tests: ${{ steps.main.outputs.integration-tests }}
Expand Down
74 changes: 74 additions & 0 deletions .github/workflows/release-docs.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Copyright (c) 2025, NVIDIA CORPORATION.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
name: Release docs
on:
workflow_dispatch:
inputs:
dry-run:
description: Whether to run the workflow in dry-run mode
required: true
type: boolean
default: true
version-number:
description: Version number to release this as (use `latest` for main branch)
required: true
type: string
notify-emails:
description: Email addresses to send the notification to. Format as "me@me.com,you@you.com".
required: true
type: string
aws-region:
description: AWS region
required: false
type: string
default: us-east-1

jobs:
build-docs:
uses: NVIDIA-NeMo/FW-CI-templates/.github/workflows/_build_docs.yml@v0.67.0

publish-docs:
runs-on: ubuntu-latest
needs: [build-docs]
steps:
- uses: actions/checkout@v6
with:
repository: NVIDIA-NeMo/FW-CI-templates
ref: v0.67.2
path: FW-CI-templates

- uses: ./FW-CI-templates/.github/actions/publish-docs
# This workflow runs either on main, or on a version tag. Any other git ref will lead
# to an error.
# If its on main, it will publish to "latest" directory in Akamai.
# If its on a versioned tag, it will extract the version number from the tag (strip `v` prefix)
# and publish to the versioned directory in Akamai.
with:
dry-run: ${{ inputs.dry-run }}
artifacts-name: docs-html
artifacts-path: _build/html
emails-csv: ${{ inputs.notify-emails && format('{0},{1}', vars.docs_release_emails, inputs.notify-emails) || vars.docs_release_emails }}
overwrite-latest-on-tag: false
run-on-version-tag-only: ${{ github.ref_name != 'main' }}
request-name: megatron-core-publish-docs-${{ github.run_id }}
aws-region: ${{ inputs.aws-region }}
aws-role-to-assume: ${{ secrets.AWS_ASSUME_ROLE_ARN }}
aws-access-key-id: ${{ secrets.AWS_ACCESS_KEY_ID }}
aws-secret-access-key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
akamai-host: ${{ secrets.AKAMAI_HOST }}
akamai-client-token: ${{ secrets.AKAMAI_CLIENT_TOKEN }}
akamai-client-secret: ${{ secrets.AKAMAI_CLIENT_SECRET }}
akamai-access-token: ${{ secrets.AKAMAI_ACCESS_TOKEN }}
s3-target-root: ${{ secrets.S3_BUCKET_NAME }}
s3-target-path: megatron-core/developer-guide
3 changes: 0 additions & 3 deletions .gitlab/labeler-config.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,6 @@ BERT:
GPT:
- megatron/core/models/gpt/**

RETRO:
- megatron/core/models/retro/**

Dist-Ckpt:
- megatron/core/dist_checkpointing

Expand Down
2 changes: 1 addition & 1 deletion docker/Dockerfile.ci.dev
Original file line number Diff line number Diff line change
Expand Up @@ -70,7 +70,7 @@ RUN bash -ex <<"EOF"

git clone --branch hybrid-ep https://github.com/deepseek-ai/DeepEP.git
pushd DeepEP
git checkout 83e0d156807f31abed4ea55c2fa6eb4b62a11b82
git checkout eb9cee7de5a24193bf09500668d3a619d3d3f3fb
patch -p1 < /workspace/deepep.patch
popd
TORCH_CUDA_ARCH_LIST="9.0 10.0 12.0" uv pip install --no-build-isolation -v DeepEP/.
Expand Down
2 changes: 1 addition & 1 deletion docs/api-guide/models/models.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# models package

This package contains most of the popular LLMs . Currently we have support for GPT, Bert, T5 and Retro . This is an ever growing list so keep an eye out.
This package contains most of the popular LLMs . Currently we have support for GPT, Bert, and T5 . This is an ever growing list so keep an eye out.

## Subpackages

Expand Down
Loading