Skip to content

ci: prune old artifacts on cluster lustre during weekly/release runs - #5084

Merged
ko3n1g merged 15 commits into
NVIDIA:mainfrom
ko3n1g:ko3n1g/ci/cleanup-old-artifacts
Jun 1, 2026
Merged

ci: prune old artifacts on cluster lustre during weekly/release runs#5084
ko3n1g merged 15 commits into
NVIDIA:mainfrom
ko3n1g:ko3n1g/ci/cleanup-old-artifacts

Conversation

@ko3n1g

@ko3n1g ko3n1g commented Jun 1, 2026

Copy link
Copy Markdown
Contributor
Claude summary

Summary

  • Adds functional:cleanup_dgx_{a100,h100,gb200} jobs (one per cluster) to .gitlab/stages/04.functional-tests.yml, gated on FUNCTIONAL_TEST_SCOPE in [release, weekly]. Each submits a JET workload that deletes weekly-* and release-* directories older than 28 days at depth 1 of each cluster-side base dir.
  • New JET cluster configs in dl/jet/ci on branches mcore/{coreweave,oci-hsg,draco-oci-iad,eos}-cpu — partition cpu (Eos: interactive, no dedicated CPU partition), no gpus_per_node, no exclusive. The cleanup never reserves a GPU.

How it works

Each functional:run_* job has an optional: true needs: on the matching functional:cleanup_dgx_<platform>:

  • mr / nightly pipelines — cleanup excluded by rule; functional run starts immediately.
  • weekly / release pipelines — cleanup runs first, then functional jobs.

.functional_cleanup derives CLEANUP_CLUSTER="cpu_${GPU_CLUSTER#dgx*_}" and submits via launch_jet_workload.py --scope cleanup. The launch script forwards CLEANUP_PATHS into the JET workload environment.

The recipe walks each CLEANUP_PATHS base dir at depth 1 only and matches directories whose names start with weekly- or release-. Matched directories older than 28 days are removed with rm -rf. The name pattern is the safety boundary — nothing outside those two prefixes is touched.

Required GitLab CI variables (already configured)

Set in project Settings → CI/CD → Variables (newline-separated paths, raw: true):

Variable Purpose
CLEANUP_PATHS_DGX_{A100,H100,GB200} Base dirs whose direct weekly-* / release-* children are eligible

Recipe exits cleanly with "nothing to do" if CLEANUP_PATHS is empty.

Recipe example

# tests/test_utils/recipes/_cleanup.yaml
spec:
  model: cleanup
  build: mcore-pyt-{environment}
  gpus: 0                       # CPU partition handles the rest
  script: |-
    THRESHOLD_DAYS=28
    while IFS= read -r path; do
        find "$path" -mindepth 1 -maxdepth 1 -type d \
            \( -name 'weekly-*' -o -name 'release-*' \) \
            -mtime +$THRESHOLD_DAYS -print0 \
          | xargs -0 -r rm -rf
    done <<<"$CLEANUP_PATHS"
products:
  - test_case: [cleanup_old_artifacts]
    products:
      - scope: [cleanup]
        platforms: [dgx_a100, dgx_h100, dgx_gb200]

Test plan

  • Trigger weekly-scope pipeline; verify functional:cleanup_dgx_* runs before functional:run_* on the matching cluster.
  • Each cleanup job log lists deletions of weekly-* / release-* directories older than 28 days under each base dir, and nothing else.
  • mr-scope pipeline does not run cleanup jobs (gated by scope rule).
  • Post-run, sanity-check ls on each base dir and confirm only fresh weekly-* / release-* entries remain; all other siblings (e.g. mcore_ci) intact.

ko3n1g added 11 commits June 1, 2026 07:25
…runs

Adds a JET workload (`_cleanup.yaml`, scope `cleanup`, `gpus: 0`) that runs
`find -mtime +14 | xargs rm -rf` against the parents of `{assets_dir}` and
`{artifacts_dir}` on each cluster.

Wires per-cluster `functional:cleanup_dgx_{a100,h100,gb200}` jobs in
`.gitlab/stages/04.functional-tests.yml`, gated on
`FUNCTIONAL_TEST_SCOPE in [release, weekly]`. Each `functional:run_*` job
adds the matching cleanup as `optional: true` so MR/nightly pipelines are
unaffected.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Per-cluster CPU launchers now live in dl/jet/ci:
`mcore/{coreweave,oci-hsg,draco-oci-iad,eos}-cpu`. They mirror the GPU
sibling but use partition `cpu` (Eos: `interactive`, since Eos has no
dedicated CPU partition), drop `gpus_per_node`, `exclusive`, and the
`srun ntasks/cpus_per_task` block.

- `resolve_cluster_config` maps `cpu_dgx*` cluster ids to the matching
  `*-cpu` branches.
- `.functional_cleanup` derives `CLEANUP_CLUSTER=cpu_${GPU_CLUSTER}` from
  the active per-platform GPU cluster, so cleanup follows whichever
  cluster the H100/A100/GB200 functional run resolved to.

The recipe still declares `gpus: 0` defensively; the partition switch is
what actually prevents GPU reservation.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Prints the entries that would be removed under the JET assets/artifacts
roots without actually deleting them. To re-enable deletion, swap the
final `find ... -print` back to a `find ... -print0 | xargs -0 -r rm -rf`.

Signed-off-by: oliver könig <okoenig@nvidia.com>
JET caps launcher keys at 20 chars (pydantic ValidationError on the long
form `cpu_dgxh100_coreweave`). Renamed all four cleanup clusters to drop
the redundant `dgx<arch>_` prefix, matching the existing
`dlalgo-ci/draco-oci-ord-cpu` precedent. `.functional_cleanup` now derives
the cleanup cluster by stripping the `dgx*_` prefix from the active GPU
cluster name.

Signed-off-by: oliver könig <okoenig@nvidia.com>
The chmod-g+w paths in each GPU mcore sibling's before_script (e.g.
mcore/coreweave: `/lustre/fsw/.../release-testing` and `.../mcore_ci`) are
the canonical artifact roots. Every CPU mcore branch mounts the cluster's
lustre at `/lustre/fsw/coreai_dlalgo_mcore` inside the container, so the
container-side paths are uniform across clusters. The `if [[ ! -d ]]`
guard makes the recipe safely degrade on clusters where a root happens
not to exist.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Recipe no longer hardcodes lustre roots. Per-cluster CI/CD variables
(`CLEANUP_PATHS_DGX_A100` / `_H100` / `_GB200`, set in GitLab project
Settings) hold the newline-separated path list; `.functional_cleanup`
forwards the active platform's variable as `CLEANUP_PATHS` and
`launch_jet_workload.py` propagates it into the JET workload env.

The recipe exits cleanly with "nothing to do" when the variable is
unset, so a pipeline can run before the variables are configured
without failing.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Previous behavior: scan the listed paths and delete their contents.
That's wrong for the configured roots — those directories contain
long-lived assets (datasets, caches) that must not be pruned, only the
*siblings* around them in the parent should age out.

New behavior: CLEANUP_PATHS is the protect list. The recipe walks each
unique parent directory, lists entries older than 14 days, and
explicitly excludes the protected basenames via `find ! -name`. Parents
matching `/` are refused as a safety guard.

Signed-off-by: oliver könig <okoenig@nvidia.com>
CLEANUP_PATHS is again the GC root list — the recipe lists direct
children older than THRESHOLD_DAYS inside each listed root. Protect-by-
exclude logic is gone; protection is implicit (anything not listed is
untouched). The GitLab CI variables now hold only the GC roots
(release-testing); mcore_ci is no longer in the list.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Without the depth cap, `find -mindepth 1 -mtime +14` traverses the full
tree under each CLEANUP_PATHS root, listing every entry whose mtime is
older than the threshold. When this flips from dry-run to actual delete,
leaves are pruned without leaving recent-parent directories tagged.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Adds `-type f` so the recursive traversal lists files only — directories
(whose mtime updates on any child change) won't show up as candidates,
and an actual delete pass would touch leaves without removing the
directory skeleton.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Second pass per cleanup root: `find -mindepth 1 -depth -type d -empty
-print`. Depth-first traversal so deeper empty dirs surface before their
parents — when this pairs with real deletion, parents that become empty
through the recursion are also pruned in the same sweep.

In dry-run, this lists only the directories that are already empty
(since the file pass hasn't actually removed anything yet).

Signed-off-by: oliver könig <okoenig@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Jun 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@ko3n1g
ko3n1g marked this pull request as ready for review June 1, 2026 10:33
@ko3n1g
ko3n1g requested a review from a team as a code owner June 1, 2026 10:33
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team June 1, 2026 10:33
CLEANUP_PATHS becomes the scan-root list, descended recursively, and a
new CLEANUP_PROTECT env var holds newline-separated subtrees that
`find -path X -prune` skips during traversal. Both file-find and
empty-dir sweep apply the prune.

Per-cluster CI variables: `CLEANUP_PATHS_DGX_<X>` (scan roots) and
`CLEANUP_PROTECT_DGX_<X>` (skip list); both forwarded to JET via
`launch_jet_workload.py` and the cleanup gitlab job.

Signed-off-by: oliver könig <okoenig@nvidia.com>
Switch from `-print` to `-print -delete` for both file and empty-dir
passes. Because `-delete` implies `-depth` (which makes `-prune` a
silent no-op), the protected subtrees are now filtered via
`! -path X ! -path X/*` predicates instead of `-prune`. find still
descends into the protected subtree but never matches anything inside
it — verified by local bash test against a synthetic tree.

THRESHOLD_DAYS bumped from 14 to 28.

Signed-off-by: oliver könig <okoenig@nvidia.com>
@svcnvidia-nemo-ci svcnvidia-nemo-ci added the Approved All necessary approvals have been made label Jun 1, 2026
- Recipe now targets only `weekly-*` and `release-*` directories at
  depth 1 of each `CLEANUP_PATHS` base dir; everything else is left alone.
- Pattern match replaces the `CLEANUP_PROTECT` traversal filter as the
  safety boundary. Drop `CLEANUP_PROTECT` plumbing from
  `launch_jet_workload.py` and `04.functional-tests.yml`.
- Switch from per-file `find -delete` + empty-dir sweep to a single
  `find ... -print0 | xargs -0 -r rm -rf` over matched directory trees.

Signed-off-by: oliver könig <okoenig@nvidia.com>
@ko3n1g

ko3n1g commented Jun 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test ffedb81

The cleanup recipe runs with gpus: 0 and no torchrun, so
extract_torchrunlogs_to_string returns empty even when the JET pipeline
succeeds. The previous "No logs found, retry" guard tripped purely on
the absent per-rank logs and looped until n_attempts hit 9. Treat the
run as having logs when either allranks OR mainrank output is non-empty;
the cleanup workload still produces output_script-0.log, so the launcher
now exits on JET pipeline status.

Signed-off-by: oliver könig <okoenig@nvidia.com>
@ko3n1g

ko3n1g commented Jun 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 80c4ed3

@ko3n1g
ko3n1g merged commit 595f697 into NVIDIA:main Jun 1, 2026
36 checks passed
copy-pr-bot Bot pushed a commit that referenced this pull request Jun 12, 2026
yhgalaxy pushed a commit to yhgalaxy/Megatron-LM that referenced this pull request Jun 17, 2026
…VIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
shjwudp added a commit to shjwudp/Megatron-LM that referenced this pull request Jul 6, 2026
* test: re-enable test_pp2_create_cudagraphs_first_stage on TE 2.15+ (NVIDIA#4985)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tests): initialize num_microbatches calculator in vision cudagraph tests (NVIDIA#4986)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci: Add allow_failure flag to gpt and moe recipes that are failing in nightlies (NVIDIA#4905)

* Drain predecessor reduce-scatter at dispatch time (NVIDIA#4940)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* nightly(ci): Update golden values for functional t5 tests (NVIDIA#4995)

* chore: rotate oncall schedule

* [main] Refactor and Improve MoE Logginginit commit (NVIDIA#3431)

Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>

* ci: validate release branch-rules (NVIDIA#4929)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Megatron-FSDP] Add conditional param.grad dereferencing logic to support full-iteration (FWD-BWD) CUDA graphability. (NVIDIA#4663)

Signed-off-by: Cory Ye <cye@nvidia.com>

* test: restrict iter-time comparison to steady-state window (NVIDIA#5010)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Add mHC support for HybridModel on dev (NVIDIA#4949)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* fix(test): pin eval-global-batch-size on 15b gb200 release configs (NVIDIA#5022)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* [fix] Release MTP assertion when EP overlap with PP=1 (NVIDIA#4796)

* fix(test): widen iter-time steady-state window for short tests (NVIDIA#5023)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Perf fix (NVIDIA#4996)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add dev-feature preservation gate and change schedule (NVIDIA#4773)

* chore(test): remove orphan nemotron3_super_release_g200 dir (NVIDIA#5024)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Ignore Claude worktree directory (NVIDIA#5020)

* Update copy-pr-bot.yaml [skip ci]

* ci: update CI workflow conditions for integration tests (NVIDIA#4658)

* Add NVSkills CI request workflow (NVIDIA#5033)

* DDP wrap pg size fixes (NVIDIA#5006)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* fix(layer_wise): tag MTP-stage word_embeddings as is_embedding_or_output_parameter (NVIDIA#5034)

* [Dev] Add Qwen3 30B MoE recipes (NVIDIA#5012)

* [DEV] fix(megatron-fsdp): reduce padding for grouped expert weights (NVIDIA#5013)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4908)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [dev] [fix] [DeepSeek-v4] fix dense loss and rope type in DSv4 Hybrid Attention (NVIDIA#5018)

Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Move LTS dependencies from pyproject.toml to Dockerfile.ci.lts (NVIDIA#4877)

* Use shared ModelOpt calibration loop on 0.45+ with 0.44 fallback fix (NVIDIA#4881)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* test(release): skip golden comparison on intermediate resume windows (NVIDIA#5040)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [mimo] Thread position_ids through MimoModel for multimodal RoPE (NVIDIA#4938)

Signed-off-by: Li Ding <liding@nvidia.com>

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5039)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Fix: Import unwrap_model from megatron.core.utils in modelopt examples (NVIDIA#5045)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Simple and stable Inference APIs (NVIDIA#4697)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: Add notification step for MBridge downstream test results (NVIDIA#5028)

* [dev] [5/5] Qwen3.5 support: Qwen3.5-VL training example (NVIDIA#4751)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: Update Docker image version to 26.04-py3 on dev (NVIDIA#5051)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Delete output tensor early (NVIDIA#4742)

* Support ScaledSReLU in TE grouped MLP fuser (NVIDIA#4859)

Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: a <a>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>

* Skip gradient updates when grad norm exceeds threshold (NVIDIA#3460)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Add 9 user skills (NVIDIA#5066)

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>

* test(nemotron): align nemotron3 super GB200 goldens with exit-interval 4768 (NVIDIA#5069)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* chore: Update transformer-engine dependency to version 2.16.0 (NVIDIA#4992)

* Update energon version requirement (NVIDIA#4572)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Fix test failures for new inference APIs (NVIDIA#5068)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): set PYTHONUNBUFFERED=1 in JET workload env (NVIDIA#5072)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Preserve non-FSDP-unit buckets across AllGatherPipeline reset (NVIDIA#4717)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add opt-in MXFP8 LM-head output projection (NVIDIA#4825)

Co-authored-by: a <a>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-01)

* fix(ci): bound JET pipeline polling with a watchdog to prevent indefinite hangs (NVIDIA#5076)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: prune old artifacts on cluster lustre during weekly/release runs (NVIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci(test): isolate ckpt-resume tensorboard per phase (NVIDIA#5074)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: unmark EP A2A activation offload test flaky (NVIDIA#5009)

* Change ownership groups (NVIDIA#5021)

* test: skip mfsdp_fully_shard cases when world_size < mesh size (NVIDIA#4487)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix mimo optimizer checkpoint metadata restore (NVIDIA#4791)

Signed-off-by: Li Ding <liding@nvidia.com>

* [mimo] Support bridge fan-out for variable modality tokens (NVIDIA#5062)

Signed-off-by: Li Ding <liding@nvidia.com>

* Add separate mtp_grad_scale_func for MTP loss scaling (NVIDIA#3459)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [training migration] Migrate GPT builder (NVIDIA#4741)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Make Mamba conv params direct mixer params (NVIDIA#4899)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Update copy-pr-bot.yaml [skip ci]

* Update oncall reviewer assignment (NVIDIA#5093)

* Pass explicit process groups to hybrid logging (NVIDIA#4781)

* Clean up top-level repository files (NVIDIA#5097)

* [main] fix(moe): Fix several bugs for DSA rope and spec. (NVIDIA#3026)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* [Dev] Skip identity alltoall chunk sort (NVIDIA#5102)

* Move MIMO unit tests into models/mimo (NVIDIA#5063)

* test: update DeepSeek FSDP2 GB200 memory golden (NVIDIA#5094)

* Remove DeepEP hardware limit check (NVIDIA#4846)

* Update transformer-engine dependency to revision 4220403 (NVIDIA#5112)

* ci: make CI resilient to pip/uv network timeouts (NVIDIA#5118)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: treat docker container-removal conflict as flaky (NVIDIA#5120)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix GDN DTensor splitting for FSDP checkpointing (NVIDIA#4843)

Signed-off-by: conver334 <conver334@gmail.com>

* Fix MoE aux_loss / z_loss gradient scaling with TP > 1 (NVIDIA#5047)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update copy-pr-bot.yaml [skip ci]

* Update Claude copy workflow to enforce user restrictions and improve error messages (NVIDIA#5117)

* Add advisory process group guidance to Claude reviews (NVIDIA#5111)

* build: cap pydantic<2.14 in transformer-engine dependency metadata (NVIDIA#5125)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* fix(test): skip scalar-less tensorboard event files in resume checks (NVIDIA#5121)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: rotate oncall schedule

* docs: fix contributor guide typo (NVIDIA#4858)

Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* ci(unit-tests): split slow unit-test buckets over 15min SLA (NVIDIA#5133)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix DSA indexer loss not averaged across micro-batches (NVIDIA#4070)

Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Update MINOR version to 19 (NVIDIA#5096)

* Fix Muon QKV split for gated attention (NVIDIA#4728)

Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Roll input IDs for MTP labels (NVIDIA#3457)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Refactor: Move paged stashing Triton kernels (NVIDIA#5003)

* Adding blackwell tests (NVIDIA#5113)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Relax atol for test_router_gating_linear router_dtype=torch.float32 (NVIDIA#4915)

Signed-off-by: Aditya Singh <adisin650@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Fix incorrect inference metadata tensor dtypes (NVIDIA#4855)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: sraman-rgb <sraman@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>

* Disable TE cross entropy loss fusion (NVIDIA#5115)

Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>

* fix(optimizer): gate ChainedOptimizer MXFP8 defer-sync on DDP-level overlap_param_gather (NVIDIA#4982)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Pass TP group to unfused cross entropy (NVIDIA#5128)

* fix: correct dsv4_hybrid Q-up FLOPs by using args.v_head_dim (NVIDIA#5142)

* [dev] [follow-up] Qwen3.5 support: MoE aux loss padding_mask (NVIDIA#4776)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>

* [Dev] Support isolated MTP loss (NVIDIA#5080)

Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>

* test(elastification): quarantine flaky test_gumbel_determinism as flaky_in_dev (NVIDIA#5156)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(notify): mention mcore-oncall and philipp on critical CI events (NVIDIA#5152)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Change the cudagraph distribution from linearly to exponentially-decreasing + grid for mixed prefill (NVIDIA#3509)

* ci: Disable a few gb200 test cases to support 2 branches. (NVIDIA#5151)

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5164)

* Add MTP acceptance rate metrics (NVIDIA#3458)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Nemotron Ultra config for ModelOpt examples (NVIDIA#5159)

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>

* Make MTP / prefix cache stats persist for engine lifetime (NVIDIA#4101)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* Restore Greptile configuration (NVIDIA#5166)

* chore: bump `_code_freeze` workflow to `v1.4.2` (NVIDIA#5132)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [dev] moe(fix): Avoid TE cuda graph dummy attention masks (NVIDIA#5131)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* ci: Remove docs build test in favor of release test (NVIDIA#5182)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Move TE cross entropy guard to training args (NVIDIA#5162)

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>

* Fix error in deepseek parser (NVIDIA#5136)

* Fix logprob slicing for 0 generated token case (NVIDIA#5167)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* [Perf] Fold frozen linear dgrad matmul (NVIDIA#5092)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* Clamp `max_new_tokens` in MInf to mirror vllm (NVIDIA#5181)

* build: add managed = true to [tool.uv] (NVIDIA#5190)

Signed-off-by: Kajal Jain <kajalj@nvidia.com>

* Stabilize GB200 inference perf tests against cold-start noise (NVIDIA#5171)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections (docs dup fields, nvrx version gate, utils re-export)

- Remove merge-introduced duplicate definitions of use_transformer_engine_op_fuser
  and moe_expert_rank_capacity_factor (transformer_config.py) and the duplicate
  _post_param_sync method (param_and_grad_buffer.py) — tripped the docs build.
- Relax is_nvrx_min_version() to compare release segments so dev's pinned
  nvidia-resiliency-ext pre-release (0.6.0.dev33) satisfies main's '>= 0.6.0'
  assertion (the required-symbol check remains the authoritative guard).
- Re-export unwrap_model from megatron/training/utils/__init__.py. main refactored
  the monolithic training/utils.py into a package whose __init__ dropped this
  re-export; dev's training.py (and others) import it via 'from .utils import',
  which broke tests/unit_tests/conftest.py's import chain (cascaded to all suites).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] [DeepSeek-v4] Add ClampedSwiGLU to MoE mlp_op_fuser and add force balance to hash routing (NVIDIA#5130)

* nvidia style guide audit for getting started folder (NVIDIA#5168)

Signed-off-by: meg miranda <mmiranda@nvidia.com>

* AI aided audit for Nvidia Style guidance (NVIDIA#5141)

Signed-off-by: meg miranda <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>

* Enable selective recompute for `norm_out` in GDN layers  (NVIDIA#4715)

* fix(elastification): align with get_batch + utils refactors (NVIDIA#5194)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4909)

* Minor improvements for Dynamic-cp (NVIDIA#4226)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>

* fix: post-CI corrections

* chore(beep boop 🤖): Bump  (main) (2026-06-08)

* Avoid stat syscall in rerun result validation (NVIDIA#5107)

Signed-off-by: dimapihtar <dpykhtar@nvidia.com>

* docs: Update Latest News in README.md (NVIDIA#3790)

* Fix bug with Megatron-FSDP zero counter not working with decoupled gradients. (NVIDIA#4802)

Signed-off-by: Cory Ye <cye@nvidia.com>

* varlendataset for thd e2e and benchmark (NVIDIA#4832)

Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* ci: add smoke tests (NVIDIA#5143)

* Add mtp_detach_heads config to detach MTP head inputs (NVIDIA#3456)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: wrap mtp logging comment

* docs: fix install guide NGC container anchor (NVIDIA#5224)

* [Dev] Add separate toggle for varlen input padding for HybridEP in THD training  (NVIDIA#5048)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* Fuse per-sequence AlltoAll into a unified one in GDN forward (NVIDIA#4913)

* Apply MIMO SP/CP sharding with explicit groups and enable THD in non-colocated path (NVIDIA#5150)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix CUDA IMA in fsdp_double_buffer when an FSDP unit's bucket doesn't fit the pool (NVIDIA#4810)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add named layouts to HyperCommGrid for heterogeneous parallelism (NVIDIA#5148)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix wgrad race condition when using double buffers. (NVIDIA#5222)

Signed-off-by: Cory Ye <cye@nvidia.com>

* Move uneven DTensor distributed fixture to conftest (NVIDIA#5237)

* Route bridge communicator cross-grid P2P through a dedicated process group (NVIDIA#5234)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix test_split_tensor_along_last_dim to actually assert correctness (NVIDIA#4710)

Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* chore: rotate oncall schedule

* Add optional group= to common_utils model/data-parallel reduction helpers (NVIDIA#5251)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement to Contribution guide (NVIDIA#5278)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Add MIMO hetero topology + distributed bootstrap (examples/mimo training-loop folder) (NVIDIA#5260)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)

Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>

* [Dev] DeepSeek-V4-Flash recipe 20260610 (NVIDIA#5266)

* [Dev] Generalized fix for mxfp8 param gather  (NVIDIA#4994)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Cherry-pick MTP detach heads (NVIDIA#5223)

* [TE] Restore original CP group after dynamic CP forward in TEDotProductAttention (NVIDIA#5215)

Co-authored-by: rionawang <rionawang@tencent.com>

* [examples] Add dynamic context parallel benchmark example (NVIDIA#5123)

* Remove duplicate nccl_allocator import (NVIDIA#5057)

* [dev] Add experimental Megatron Lite as agentic exploration (NVIDIA#4885)

Co-authored-by: Deyu Fu <deyuf@nvidia.com>

* fix(ci): resolve t5 dataloader stall + GRPO cudagraph-memory regression (CI-validated) (NVIDIA#5280)

Signed-off-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix Dockerfile warnings (NVIDIA#4856)

* Fix fused MLA delayed weight grad hooks (NVIDIA#5273)

Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>

* ci: limit retries on unsuccessful test launches (NVIDIA#5275)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Thread pg_collection into get_model DDP bucket sizing (NVIDIA#5250)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Enable non-deterministic results in model configuration for nemotron tests (NVIDIA#5239)

* Stabilize hybrid nanov3 gb200 perf (NVIDIA#5295)

* Clip mtp grads separately when mtp_detach_heads=True (NVIDIA#4116)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Thread pg_collection into train_step reductions (NVIDIA#5259)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement (NVIDIA#5305)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>

* fix: restore dev-only training.py features dropped by the main override

The merge takes main's training.py wholesale (per the sync skill's override
list), which silently reverted four dev-only features whose supporting core
code the merge keeps at dev's version. Each surfaced as a CI failure:

1. Dynamic context-parallel API rename. Dev renamed
   get_hybrid_data_context_parallel_groups -> get_dynamic_data_context_parallel_groups
   (identical signature) and args.hybrid_context_parallel -> dynamic_context_parallel.
   Point training.py at dev's names. Fixes the conftest.py ImportError that
   cascaded to every unit-test bucket.

2. Dynamic-CP / sequence-packing data loading. Replace main's setup-time
   HybridCPDataLoaderWrapper wrap with dev's per-step wrap_data_iterator
   (gated on config.sequence_packing_scheduler), which returns a
   RerunDataIterator-compatible iterator. Fixes the *_cp4_dcp RerunDataIterator
   assertion.

3. MTP loss logging scale. Dev's MTPLossLoggingHelper stores raw loss sums and
   token counts and computes the per-token loss after reduction, so the log
   scale must be 1.0; main's 1/get_num_microbatches() divided the reported
   mtp_N loss by num_microbatches. Fixes moe gpt3_..._scoped_cudagraph mtp_1
   loss (was 16x too small: 0.682 vs golden 10.915).

4. DSA indexer loss cross-PP reduction. dsa.py requires num_layers (and
   csa_compress_ratios) so first-pipeline-stage ranks lazily initialize the
   tracker and join the cross-PP all_reduce in reduce_loss_in_tracker; main's
   call omitted them, so stage 0 returned early and the last stage hung on an
   unmatched all_reduce. Fixes the gpt3_..._dsv4_hybrid_mhc_mtp NCCL timeout.

Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>

* [dev]: faster implementation of mHC fused kernels (NVIDIA#4624)

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>

* Allow for pre-bound socket to be passed in server (NVIDIA#5301)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Offline Logits-Based Knowledge Distillation (NVIDIA#5019)

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

* Handle None values in sampling parameters (NVIDIA#5300)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Add moe loss normalization for RL SFT (NVIDIA#3956)

Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Add code owners for optimizer-related files (NVIDIA#5297)

Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* Fix EP=1 inference by allocating buffers anyway (NVIDIA#5233)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Fix crash due to tool call at sequence length (NVIDIA#5302)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Inference: Cudagraph-aware admission gating in prefill scheduler (NVIDIA#4870)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Account for reasoning token stripping (NVIDIA#5313)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Thread pg_collection through wrap_model_chunks_with_ddp (NVIDIA#5328)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Dev] Add MoE recipe performance summary (NVIDIA#5289)

Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>

* [Dev] Add DeepEP v2 flex dispatcher backend (NVIDIA#4793)

Signed-off-by: tongliu <tongliu@nvidia.com>
Co-authored-by: Dennis(Zhenhuan) Liu <denliu@nvidia.com>

* fix tflops calculation when sequence_packing_scheduler is not none (NVIDIA#5342)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-15)

* [dev] bump emerging optimizers to v0.3.0 (NVIDIA#5320)

Signed-off-by: Deyu Fu <deyuf@nvidia.com>

* Fix LatentMoE theoretical memory estimate (NVIDIA#5145)

Signed-off-by: Shijie Wang <jaywan@nvidia.com>

* Add zstandard package to Docker LTS requirements. Fix nightly failures (NVIDIA#5347)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* fix: preserve seqlen stats in train_step

Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>

* Thread MIMO support through the stock training loop (schedule + optimizer) (NVIDIA#5333)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (1/N) (NVIDIA#5042)

Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(ssm): whole-module 'gdn' selective recompute for GatedDeltaNet (NVIDIA#5296)

Signed-off-by: jinliangl <975761915@qq.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Fix] Fix optimizer parameter override bugs. (NVIDIA#5213)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Add Megatron-FSDP weight prefetch for full recompute (NVIDIA#5175)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* ci: default functional test time limit to 4h for release/weekly scopes (NVIDIA#5360)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Dev] add cuda graph support for thd format training. (NVIDIA#4359)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>

* Fix memory leak with log_max_attention_logit (NVIDIA#4699) (NVIDIA#5067)

Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>

* Clean up pretrain_gpt.py and pretrain_hybrid.py formatting and remove module globals (NVIDIA#5351)

Signed-off-by: ilml <tolong@nvidia.com>

* Add full model cuda graph support for MTP inference (NVIDIA#4950)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Expand the Mamba prefix caching memory safety check to include scratch space buffers (NVIDIA#5348)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Make Megatron RL only materialize last token logit (NVIDIA#4551)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Profiling  (NVIDIA#3110)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Support fused MLA QKV checkpoint reload (NVIDIA#5310)

Signed-off-by: sraman <sraman@nvidia.com>

* Add minimal DBuffer implementation (NVIDIA#4835)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [Dev] restore DSv4 tflops calc in training and fix the packed seq case (NVIDIA#5358)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* [split 1/5] Fix packed THD RoPE under CP (NVIDIA#5243)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* Update copy-pr-bot.yaml [skip ci]

* Document agent PR commit sign-off and signing (NVIDIA#5381)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Remove unused distributed pytest markers (NVIDIA#5380)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [feat] Support fine-grained activation offloading in fused group mlp (NVIDIA#5082)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* Thread tensor-parallel group into the RADIO patch embedder (NVIDIA#5371)

Signed-off-by: ykarnati <ykarnati@nvidia.com>

* Add MimoModel.zero_grad_buffer delegating to active DDP submodules (NVIDIA#5372)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [split 3/5] Refactor absorbed MLA projection handling (NVIDIA#5245)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* [dev] moe(perf): Restore fused GDN THD all-to-all on dev (NVIDIA#5389)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* chore: rotate oncall schedule

* ci: Remove sync skills workflow (NVIDIA#5091)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Add flaky marker to fine-grained activation offloading test (NVIDIA#5350) (NVIDIA#5368)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Revert "Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)" (NVIDIA#5366)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Update goldens for weekly tests after pytorch and TE bumps. (NVIDIA#5399)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Add MIMO runtime setup: per-role RNG seeding and DDP wrapping (NVIDIA#5285)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Add --mamba-training-ssm-states-dtype argument (NVIDIA#5309)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-22)

* Fix Mamba prefix match for chunked prefill (NVIDIA#4758)

Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>

* Disag MR2: Refit into multiple destination pools and tied-embedding + UVM fixes (NVIDIA#5187)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Disag MR1: Add inference shard specs and pg-collection building (NVIDIA#5186)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [Fix] Fix MoE router z-loss compatibility with TE CUDA Graph capture. (NVIDIA#5401)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>

* test: mark ep_a2a_overlap activation-offloading test flaky_in_dev (NVIDIA#5450)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark TestParallelTransformerBlockCudagraphs::test_gpu_cudagraph flaky_in_dev (NVIDIA#5475)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark gated_delta_net selective-recompute test flaky_in_dev (NVIDIA#5476)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections for sync test/impl reconciliation

1) test_optimizer.py: reverted to dev. The sync kept dev's
   multi_latent_attention.py (split q/kv down-proj, no
   _synthesize_fused_qkv_down_weight), but auto-merged main's
   test asserting the fused linear_qkv_down_proj.weight key.

2) training.py: guard the dev-only sequence_packing_scheduler config
   access with getattr (lines in train_step and train()). main's new
   MIMO schedule-plumbing test (NVIDIA#5333) passes an empty SimpleNamespace
   config; the reconciled training.py keeps dev's packing path, so the
   access must tolerate a config lacking the attribute. Real configs are
   unaffected (getattr returns the same value).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>

* fix sequence packing wrapper for eval (NVIDIA#5483)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* [dev] Megatron Lite (4/4) shared attention (NVIDIA#5427)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* fix: restore fused group MLP offload in main2dev sync (NVIDIA#5493)

Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>

* [dev] [DeepSeek-v4] Packed Sequence (THD) support for DSv4 Hybrid Attention (NVIDIA#5011)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* Improve default dynamic CP packing scheduler (NVIDIA#5154)

Signed-off-by: tailaim <tailaim@nvidia.com>

* [dev] moe(perf): Pre-GDR kernel fusion (NVIDIA#5361)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Preserve DSA output across fused inverse RoPE (NVIDIA#5526)

Signed-off-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (2/N) (NVIDIA#5485)

Signed-off-by: guihong-nv <guihongl@nvidia.com>

* [dev] Add experimental decoupled compact LayerWise DDP layout for Muon (NVIDIA#5388)

Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] Sync Megatron Lite with the latest implementation (NVIDIA#5577)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* [Dev] fix padding mask docstring (NVIDIA#5598)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>

* handle split_dtensor

---------

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Cory Ye <cye@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Li Ding <liding@nvidia.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Jingyue Wu <wujingyue@gmail.com>
Signed-off-by: conver334 <conver334@gmail.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Signed-off-by: Aditya Singh <adisin650@gmail.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Signed-off-by: meg miranda <mmiranda@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com>
Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Signed-off-by: Shivanjan Chakravorty <shivanjanc@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Xiaowei Ren <xren@nvidia.com>
Signed-off-by: yuzhongw <yuzhongw@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: ksivamani <ksivamani@nvidia.com>
Signed-off-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: Pavel Gein <pavel.gein@gmail.com>
Signed-off-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Signed-off-by: jinliangl <jinliangl@nvidia.com>
Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: ykarnati <ykarnati@nvidia.com>
Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com>
Signed-off-by: Shijie Wang <jaywan@nvidia.com>
Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Signed-off-by: jinliangl <975761915@qq.com>
Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Signed-off-by: ilml <tolong@nvidia.com>
Signed-off-by: sraman <sraman@nvidia.com>
Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>
Signed-off-by: Hollow Man <hollowman@opensuse.org>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com>
Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zijie Yan <zijiey@nvidia.com>
Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>
Co-authored-by: Cory Ye <44509866+cspades@users.noreply.github.com>
Co-authored-by: Yan Xu <45385219+Connor-XY@users.noreply.github.com>
Co-authored-by: Pingtian Li <158665726+Wohox@users.noreply.github.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: xuwchen <xuwenc@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Li Ding <liding@nvidia.com>
Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com>
Co-authored-by: Li Jinliang <jinliangl@nvidia.com>
Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com>
Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: Jingyue Wu <wujingyue@gmail.com>
Co-authored-by: Gao Deng <160076886+gdengk@users.noreply.github.com>
Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com>
Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com>
Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>
Co-authored-by: janEbert <janpabloe@nvidia.com>
Co-authored-by: Simiao Zhang <56124251+conver334@users.noreply.github.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Co-authored-by: Shaurya Singh <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
Co-authored-by: Kevin <23258141+kaimo455@users.noreply.github.com>
Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Moozy <108285604+Moozy23232@users.noreply.github.com>
Co-authored-by: Aditya Singh <60082699+adityasingh2400@users.noreply.github.com>
Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@nvidia.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>
Co-authored-by: Dingqing Yang <dingqingy@nvidia.com>
Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>
Co-authored-by: mathemakitten <helenn@nvidia.com>
Co-authored-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Yu Yao <54727607+yaoyu-33@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com>
Co-authored-by: kajalj22 <kajalj@nvidia.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: megnvidia <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>
Co-authored-by: Xuanteng Huang <44627253+xuantengh@users.noreply.github.com>
Co-authored-by: Tailai Ma <58548582+xiaoyao0115@users.noreply.github.com>
Co-authored-by: Santosh Bhavani <santosh.bhavani@live.com>
Co-authored-by: Zhongbo Zhu <42691305+zhongbozhu@users.noreply.github.com>
Co-authored-by: Nick Schank <nick@reflection.ai>
Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Shivanjan Chakravorty <schakravorty846@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Tuomas Rintamaki <trintamaki@nvidia.com>
Co-authored-by: Tyler Poon <tylerpoon@gmail.com>
Co-authored-by: Collin McCarthy <cmccarthy@nvidia.com>
Co-authored-by: Matthieu Le <matthieul@nvidia.com>
Co-authored-by: Piotr Zelasko <pzelasko@nvidia.com>
Co-authored-by: Ehsan Hosseini Asl <ehosseiniasl@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>
Co-authored-by: Xiaowei Ren <103958965+xrennvidia@users.noreply.github.com>
Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com>
Co-authored-by: William Dykas <wdykas@oci-hsg-cs-001-vscode-03.cm.cluster>
Co-authored-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Xuesong Ye <xuesongyey@gmail.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Daisy Gao <daisyg@nvidia.com>
Co-authored-by: lichenlu <lichenlu_8618@163.com>
Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Gao Deng <gdeng@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Pavel Gein <pavelgejn@yandex.ru>
Co-authored-by: Gao Deng <gdeng@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Eric Harper <eharper@nvidia.com>
Co-authored-by: Abhishree Thittenamane <47577437+athitten@users.noreply.github.com>
Co-authored-by: Abhishree Thittenamane <athittenaman@cw-dfw-cs-001-login-01.cm.cluster>
Co-authored-by: root <root@pool0-01849.cm.cluster>
Co-authored-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Co-authored-by: Nan Zheng <80790206+nanz-nv@users.noreply.github.com>
Co-authored-by: Li Tao <lit@nvidia.com>
Co-authored-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Robert Kirby <ArEsKay3@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Cole Hawkins <chawkins@nvidia.com>
Co-authored-by: Qi Zhang <qizhang@nvidia.com>
Co-authored-by: Vasudevan Rengasamy <vrengasamy@nvidia.com>
Co-authored-by: tongliu <tongliu@nvidia.com>
Co-authored-by: shurkat-nvidia <shurkat@nvidia.com>
Co-authored-by: Rui WANG <51357798+rui23@users.noreply.github.com>
Co-authored-by: rionawang <rionawang@tencent.com>
Co-authored-by: Tom Long <tolong@nvidia.com>
Co-authored-by: Lintch <44701395+returnL@users.noreply.github.com>
Co-authored-by: Yan Bai <bayan@nvidia.com>
Co-authored-by: Deyu Fu <deyuf@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Co-authored-by: jingqiny-99 <jingqiny@nvidia.com>
Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Pranav Thombre <pthombre@nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Shijie <505749828@qq.com>
Co-authored-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Li Jinliang <975761915@qq.com>
Co-authored-by: Baibaifan <39549453+Baibaifan@users.noreply.github.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: HaochenYuan <106647990+HaochenYuan@users.noreply.github.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>
Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org>
Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com>
Co-authored-by: Kunlun Li <94586211+kunlunl@users.noreply.github.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>
jon-barker pushed a commit to jon-barker/Megatron-LM that referenced this pull request Jul 10, 2026
…VIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
shjwudp added a commit to shjwudp/Megatron-LM that referenced this pull request Jul 16, 2026
* test: re-enable test_pp2_create_cudagraphs_first_stage on TE 2.15+ (NVIDIA#4985)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tests): initialize num_microbatches calculator in vision cudagraph tests (NVIDIA#4986)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci: Add allow_failure flag to gpt and moe recipes that are failing in nightlies (NVIDIA#4905)

* Drain predecessor reduce-scatter at dispatch time (NVIDIA#4940)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* nightly(ci): Update golden values for functional t5 tests (NVIDIA#4995)

* chore: rotate oncall schedule

* [main] Refactor and Improve MoE Logginginit commit (NVIDIA#3431)

Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>

* ci: validate release branch-rules (NVIDIA#4929)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Megatron-FSDP] Add conditional param.grad dereferencing logic to support full-iteration (FWD-BWD) CUDA graphability. (NVIDIA#4663)

Signed-off-by: Cory Ye <cye@nvidia.com>

* test: restrict iter-time comparison to steady-state window (NVIDIA#5010)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Add mHC support for HybridModel on dev (NVIDIA#4949)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* fix(test): pin eval-global-batch-size on 15b gb200 release configs (NVIDIA#5022)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* [fix] Release MTP assertion when EP overlap with PP=1 (NVIDIA#4796)

* fix(test): widen iter-time steady-state window for short tests (NVIDIA#5023)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Perf fix (NVIDIA#4996)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add dev-feature preservation gate and change schedule (NVIDIA#4773)

* chore(test): remove orphan nemotron3_super_release_g200 dir (NVIDIA#5024)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Ignore Claude worktree directory (NVIDIA#5020)

* Update copy-pr-bot.yaml [skip ci]

* ci: update CI workflow conditions for integration tests (NVIDIA#4658)

* Add NVSkills CI request workflow (NVIDIA#5033)

* DDP wrap pg size fixes (NVIDIA#5006)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* fix(layer_wise): tag MTP-stage word_embeddings as is_embedding_or_output_parameter (NVIDIA#5034)

* [Dev] Add Qwen3 30B MoE recipes (NVIDIA#5012)

* [DEV] fix(megatron-fsdp): reduce padding for grouped expert weights (NVIDIA#5013)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4908)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [dev] [fix] [DeepSeek-v4] fix dense loss and rope type in DSv4 Hybrid Attention (NVIDIA#5018)

Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Move LTS dependencies from pyproject.toml to Dockerfile.ci.lts (NVIDIA#4877)

* Use shared ModelOpt calibration loop on 0.45+ with 0.44 fallback fix (NVIDIA#4881)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* test(release): skip golden comparison on intermediate resume windows (NVIDIA#5040)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [mimo] Thread position_ids through MimoModel for multimodal RoPE (NVIDIA#4938)

Signed-off-by: Li Ding <liding@nvidia.com>

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5039)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Fix: Import unwrap_model from megatron.core.utils in modelopt examples (NVIDIA#5045)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Simple and stable Inference APIs (NVIDIA#4697)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: Add notification step for MBridge downstream test results (NVIDIA#5028)

* [dev] [5/5] Qwen3.5 support: Qwen3.5-VL training example (NVIDIA#4751)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: Update Docker image version to 26.04-py3 on dev (NVIDIA#5051)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Delete output tensor early (NVIDIA#4742)

* Support ScaledSReLU in TE grouped MLP fuser (NVIDIA#4859)

Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: a <a>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>

* Skip gradient updates when grad norm exceeds threshold (NVIDIA#3460)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Add 9 user skills (NVIDIA#5066)

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>

* test(nemotron): align nemotron3 super GB200 goldens with exit-interval 4768 (NVIDIA#5069)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* chore: Update transformer-engine dependency to version 2.16.0 (NVIDIA#4992)

* Update energon version requirement (NVIDIA#4572)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Fix test failures for new inference APIs (NVIDIA#5068)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): set PYTHONUNBUFFERED=1 in JET workload env (NVIDIA#5072)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Preserve non-FSDP-unit buckets across AllGatherPipeline reset (NVIDIA#4717)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add opt-in MXFP8 LM-head output projection (NVIDIA#4825)

Co-authored-by: a <a>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-01)

* fix(ci): bound JET pipeline polling with a watchdog to prevent indefinite hangs (NVIDIA#5076)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: prune old artifacts on cluster lustre during weekly/release runs (NVIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci(test): isolate ckpt-resume tensorboard per phase (NVIDIA#5074)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: unmark EP A2A activation offload test flaky (NVIDIA#5009)

* Change ownership groups (NVIDIA#5021)

* test: skip mfsdp_fully_shard cases when world_size < mesh size (NVIDIA#4487)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix mimo optimizer checkpoint metadata restore (NVIDIA#4791)

Signed-off-by: Li Ding <liding@nvidia.com>

* [mimo] Support bridge fan-out for variable modality tokens (NVIDIA#5062)

Signed-off-by: Li Ding <liding@nvidia.com>

* Add separate mtp_grad_scale_func for MTP loss scaling (NVIDIA#3459)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [training migration] Migrate GPT builder (NVIDIA#4741)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Make Mamba conv params direct mixer params (NVIDIA#4899)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Update copy-pr-bot.yaml [skip ci]

* Update oncall reviewer assignment (NVIDIA#5093)

* Pass explicit process groups to hybrid logging (NVIDIA#4781)

* Clean up top-level repository files (NVIDIA#5097)

* [main] fix(moe): Fix several bugs for DSA rope and spec. (NVIDIA#3026)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* [Dev] Skip identity alltoall chunk sort (NVIDIA#5102)

* Move MIMO unit tests into models/mimo (NVIDIA#5063)

* test: update DeepSeek FSDP2 GB200 memory golden (NVIDIA#5094)

* Remove DeepEP hardware limit check (NVIDIA#4846)

* Update transformer-engine dependency to revision 4220403 (NVIDIA#5112)

* ci: make CI resilient to pip/uv network timeouts (NVIDIA#5118)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: treat docker container-removal conflict as flaky (NVIDIA#5120)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix GDN DTensor splitting for FSDP checkpointing (NVIDIA#4843)

Signed-off-by: conver334 <conver334@gmail.com>

* Fix MoE aux_loss / z_loss gradient scaling with TP > 1 (NVIDIA#5047)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update copy-pr-bot.yaml [skip ci]

* Update Claude copy workflow to enforce user restrictions and improve error messages (NVIDIA#5117)

* Add advisory process group guidance to Claude reviews (NVIDIA#5111)

* build: cap pydantic<2.14 in transformer-engine dependency metadata (NVIDIA#5125)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* fix(test): skip scalar-less tensorboard event files in resume checks (NVIDIA#5121)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: rotate oncall schedule

* docs: fix contributor guide typo (NVIDIA#4858)

Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* ci(unit-tests): split slow unit-test buckets over 15min SLA (NVIDIA#5133)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix DSA indexer loss not averaged across micro-batches (NVIDIA#4070)

Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Update MINOR version to 19 (NVIDIA#5096)

* Fix Muon QKV split for gated attention (NVIDIA#4728)

Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Roll input IDs for MTP labels (NVIDIA#3457)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Refactor: Move paged stashing Triton kernels (NVIDIA#5003)

* Adding blackwell tests (NVIDIA#5113)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Relax atol for test_router_gating_linear router_dtype=torch.float32 (NVIDIA#4915)

Signed-off-by: Aditya Singh <adisin650@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Fix incorrect inference metadata tensor dtypes (NVIDIA#4855)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: sraman-rgb <sraman@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>

* Disable TE cross entropy loss fusion (NVIDIA#5115)

Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>

* fix(optimizer): gate ChainedOptimizer MXFP8 defer-sync on DDP-level overlap_param_gather (NVIDIA#4982)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Pass TP group to unfused cross entropy (NVIDIA#5128)

* fix: correct dsv4_hybrid Q-up FLOPs by using args.v_head_dim (NVIDIA#5142)

* [dev] [follow-up] Qwen3.5 support: MoE aux loss padding_mask (NVIDIA#4776)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>

* [Dev] Support isolated MTP loss (NVIDIA#5080)

Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>

* test(elastification): quarantine flaky test_gumbel_determinism as flaky_in_dev (NVIDIA#5156)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(notify): mention mcore-oncall and philipp on critical CI events (NVIDIA#5152)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Change the cudagraph distribution from linearly to exponentially-decreasing + grid for mixed prefill (NVIDIA#3509)

* ci: Disable a few gb200 test cases to support 2 branches. (NVIDIA#5151)

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5164)

* Add MTP acceptance rate metrics (NVIDIA#3458)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Nemotron Ultra config for ModelOpt examples (NVIDIA#5159)

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>

* Make MTP / prefix cache stats persist for engine lifetime (NVIDIA#4101)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* Restore Greptile configuration (NVIDIA#5166)

* chore: bump `_code_freeze` workflow to `v1.4.2` (NVIDIA#5132)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [dev] moe(fix): Avoid TE cuda graph dummy attention masks (NVIDIA#5131)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* ci: Remove docs build test in favor of release test (NVIDIA#5182)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Move TE cross entropy guard to training args (NVIDIA#5162)

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>

* Fix error in deepseek parser (NVIDIA#5136)

* Fix logprob slicing for 0 generated token case (NVIDIA#5167)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* [Perf] Fold frozen linear dgrad matmul (NVIDIA#5092)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* Clamp `max_new_tokens` in MInf to mirror vllm (NVIDIA#5181)

* build: add managed = true to [tool.uv] (NVIDIA#5190)

Signed-off-by: Kajal Jain <kajalj@nvidia.com>

* Stabilize GB200 inference perf tests against cold-start noise (NVIDIA#5171)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections (docs dup fields, nvrx version gate, utils re-export)

- Remove merge-introduced duplicate definitions of use_transformer_engine_op_fuser
  and moe_expert_rank_capacity_factor (transformer_config.py) and the duplicate
  _post_param_sync method (param_and_grad_buffer.py) — tripped the docs build.
- Relax is_nvrx_min_version() to compare release segments so dev's pinned
  nvidia-resiliency-ext pre-release (0.6.0.dev33) satisfies main's '>= 0.6.0'
  assertion (the required-symbol check remains the authoritative guard).
- Re-export unwrap_model from megatron/training/utils/__init__.py. main refactored
  the monolithic training/utils.py into a package whose __init__ dropped this
  re-export; dev's training.py (and others) import it via 'from .utils import',
  which broke tests/unit_tests/conftest.py's import chain (cascaded to all suites).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] [DeepSeek-v4] Add ClampedSwiGLU to MoE mlp_op_fuser and add force balance to hash routing (NVIDIA#5130)

* nvidia style guide audit for getting started folder (NVIDIA#5168)

Signed-off-by: meg miranda <mmiranda@nvidia.com>

* AI aided audit for Nvidia Style guidance (NVIDIA#5141)

Signed-off-by: meg miranda <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>

* Enable selective recompute for `norm_out` in GDN layers  (NVIDIA#4715)

* fix(elastification): align with get_batch + utils refactors (NVIDIA#5194)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4909)

* Minor improvements for Dynamic-cp (NVIDIA#4226)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>

* fix: post-CI corrections

* chore(beep boop 🤖): Bump  (main) (2026-06-08)

* Avoid stat syscall in rerun result validation (NVIDIA#5107)

Signed-off-by: dimapihtar <dpykhtar@nvidia.com>

* docs: Update Latest News in README.md (NVIDIA#3790)

* Fix bug with Megatron-FSDP zero counter not working with decoupled gradients. (NVIDIA#4802)

Signed-off-by: Cory Ye <cye@nvidia.com>

* varlendataset for thd e2e and benchmark (NVIDIA#4832)

Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* ci: add smoke tests (NVIDIA#5143)

* Add mtp_detach_heads config to detach MTP head inputs (NVIDIA#3456)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: wrap mtp logging comment

* docs: fix install guide NGC container anchor (NVIDIA#5224)

* [Dev] Add separate toggle for varlen input padding for HybridEP in THD training  (NVIDIA#5048)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* Fuse per-sequence AlltoAll into a unified one in GDN forward (NVIDIA#4913)

* Apply MIMO SP/CP sharding with explicit groups and enable THD in non-colocated path (NVIDIA#5150)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix CUDA IMA in fsdp_double_buffer when an FSDP unit's bucket doesn't fit the pool (NVIDIA#4810)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add named layouts to HyperCommGrid for heterogeneous parallelism (NVIDIA#5148)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix wgrad race condition when using double buffers. (NVIDIA#5222)

Signed-off-by: Cory Ye <cye@nvidia.com>

* Move uneven DTensor distributed fixture to conftest (NVIDIA#5237)

* Route bridge communicator cross-grid P2P through a dedicated process group (NVIDIA#5234)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix test_split_tensor_along_last_dim to actually assert correctness (NVIDIA#4710)

Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* chore: rotate oncall schedule

* Add optional group= to common_utils model/data-parallel reduction helpers (NVIDIA#5251)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement to Contribution guide (NVIDIA#5278)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Add MIMO hetero topology + distributed bootstrap (examples/mimo training-loop folder) (NVIDIA#5260)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)

Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>

* [Dev] DeepSeek-V4-Flash recipe 20260610 (NVIDIA#5266)

* [Dev] Generalized fix for mxfp8 param gather  (NVIDIA#4994)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Cherry-pick MTP detach heads (NVIDIA#5223)

* [TE] Restore original CP group after dynamic CP forward in TEDotProductAttention (NVIDIA#5215)

Co-authored-by: rionawang <rionawang@tencent.com>

* [examples] Add dynamic context parallel benchmark example (NVIDIA#5123)

* Remove duplicate nccl_allocator import (NVIDIA#5057)

* [dev] Add experimental Megatron Lite as agentic exploration (NVIDIA#4885)

Co-authored-by: Deyu Fu <deyuf@nvidia.com>

* fix(ci): resolve t5 dataloader stall + GRPO cudagraph-memory regression (CI-validated) (NVIDIA#5280)

Signed-off-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix Dockerfile warnings (NVIDIA#4856)

* Fix fused MLA delayed weight grad hooks (NVIDIA#5273)

Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>

* ci: limit retries on unsuccessful test launches (NVIDIA#5275)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Thread pg_collection into get_model DDP bucket sizing (NVIDIA#5250)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Enable non-deterministic results in model configuration for nemotron tests (NVIDIA#5239)

* Stabilize hybrid nanov3 gb200 perf (NVIDIA#5295)

* Clip mtp grads separately when mtp_detach_heads=True (NVIDIA#4116)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Thread pg_collection into train_step reductions (NVIDIA#5259)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement (NVIDIA#5305)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>

* fix: restore dev-only training.py features dropped by the main override

The merge takes main's training.py wholesale (per the sync skill's override
list), which silently reverted four dev-only features whose supporting core
code the merge keeps at dev's version. Each surfaced as a CI failure:

1. Dynamic context-parallel API rename. Dev renamed
   get_hybrid_data_context_parallel_groups -> get_dynamic_data_context_parallel_groups
   (identical signature) and args.hybrid_context_parallel -> dynamic_context_parallel.
   Point training.py at dev's names. Fixes the conftest.py ImportError that
   cascaded to every unit-test bucket.

2. Dynamic-CP / sequence-packing data loading. Replace main's setup-time
   HybridCPDataLoaderWrapper wrap with dev's per-step wrap_data_iterator
   (gated on config.sequence_packing_scheduler), which returns a
   RerunDataIterator-compatible iterator. Fixes the *_cp4_dcp RerunDataIterator
   assertion.

3. MTP loss logging scale. Dev's MTPLossLoggingHelper stores raw loss sums and
   token counts and computes the per-token loss after reduction, so the log
   scale must be 1.0; main's 1/get_num_microbatches() divided the reported
   mtp_N loss by num_microbatches. Fixes moe gpt3_..._scoped_cudagraph mtp_1
   loss (was 16x too small: 0.682 vs golden 10.915).

4. DSA indexer loss cross-PP reduction. dsa.py requires num_layers (and
   csa_compress_ratios) so first-pipeline-stage ranks lazily initialize the
   tracker and join the cross-PP all_reduce in reduce_loss_in_tracker; main's
   call omitted them, so stage 0 returned early and the last stage hung on an
   unmatched all_reduce. Fixes the gpt3_..._dsv4_hybrid_mhc_mtp NCCL timeout.

Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>

* [dev]: faster implementation of mHC fused kernels (NVIDIA#4624)

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>

* Allow for pre-bound socket to be passed in server (NVIDIA#5301)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Offline Logits-Based Knowledge Distillation (NVIDIA#5019)

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

* Handle None values in sampling parameters (NVIDIA#5300)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Add moe loss normalization for RL SFT (NVIDIA#3956)

Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Add code owners for optimizer-related files (NVIDIA#5297)

Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* Fix EP=1 inference by allocating buffers anyway (NVIDIA#5233)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Fix crash due to tool call at sequence length (NVIDIA#5302)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Inference: Cudagraph-aware admission gating in prefill scheduler (NVIDIA#4870)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Account for reasoning token stripping (NVIDIA#5313)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Thread pg_collection through wrap_model_chunks_with_ddp (NVIDIA#5328)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Dev] Add MoE recipe performance summary (NVIDIA#5289)

Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>

* [Dev] Add DeepEP v2 flex dispatcher backend (NVIDIA#4793)

Signed-off-by: tongliu <tongliu@nvidia.com>
Co-authored-by: Dennis(Zhenhuan) Liu <denliu@nvidia.com>

* fix tflops calculation when sequence_packing_scheduler is not none (NVIDIA#5342)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-15)

* [dev] bump emerging optimizers to v0.3.0 (NVIDIA#5320)

Signed-off-by: Deyu Fu <deyuf@nvidia.com>

* Fix LatentMoE theoretical memory estimate (NVIDIA#5145)

Signed-off-by: Shijie Wang <jaywan@nvidia.com>

* Add zstandard package to Docker LTS requirements. Fix nightly failures (NVIDIA#5347)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* fix: preserve seqlen stats in train_step

Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>

* Thread MIMO support through the stock training loop (schedule + optimizer) (NVIDIA#5333)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (1/N) (NVIDIA#5042)

Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(ssm): whole-module 'gdn' selective recompute for GatedDeltaNet (NVIDIA#5296)

Signed-off-by: jinliangl <975761915@qq.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Fix] Fix optimizer parameter override bugs. (NVIDIA#5213)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Add Megatron-FSDP weight prefetch for full recompute (NVIDIA#5175)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* ci: default functional test time limit to 4h for release/weekly scopes (NVIDIA#5360)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Dev] add cuda graph support for thd format training. (NVIDIA#4359)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>

* Fix memory leak with log_max_attention_logit (NVIDIA#4699) (NVIDIA#5067)

Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>

* Clean up pretrain_gpt.py and pretrain_hybrid.py formatting and remove module globals (NVIDIA#5351)

Signed-off-by: ilml <tolong@nvidia.com>

* Add full model cuda graph support for MTP inference (NVIDIA#4950)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Expand the Mamba prefix caching memory safety check to include scratch space buffers (NVIDIA#5348)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Make Megatron RL only materialize last token logit (NVIDIA#4551)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Profiling  (NVIDIA#3110)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Support fused MLA QKV checkpoint reload (NVIDIA#5310)

Signed-off-by: sraman <sraman@nvidia.com>

* Add minimal DBuffer implementation (NVIDIA#4835)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [Dev] restore DSv4 tflops calc in training and fix the packed seq case (NVIDIA#5358)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* [split 1/5] Fix packed THD RoPE under CP (NVIDIA#5243)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* Update copy-pr-bot.yaml [skip ci]

* Document agent PR commit sign-off and signing (NVIDIA#5381)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Remove unused distributed pytest markers (NVIDIA#5380)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [feat] Support fine-grained activation offloading in fused group mlp (NVIDIA#5082)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* Thread tensor-parallel group into the RADIO patch embedder (NVIDIA#5371)

Signed-off-by: ykarnati <ykarnati@nvidia.com>

* Add MimoModel.zero_grad_buffer delegating to active DDP submodules (NVIDIA#5372)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [split 3/5] Refactor absorbed MLA projection handling (NVIDIA#5245)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* [dev] moe(perf): Restore fused GDN THD all-to-all on dev (NVIDIA#5389)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* chore: rotate oncall schedule

* ci: Remove sync skills workflow (NVIDIA#5091)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Add flaky marker to fine-grained activation offloading test (NVIDIA#5350) (NVIDIA#5368)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Revert "Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)" (NVIDIA#5366)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Update goldens for weekly tests after pytorch and TE bumps. (NVIDIA#5399)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Add MIMO runtime setup: per-role RNG seeding and DDP wrapping (NVIDIA#5285)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Add --mamba-training-ssm-states-dtype argument (NVIDIA#5309)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-22)

* Fix Mamba prefix match for chunked prefill (NVIDIA#4758)

Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>

* Disag MR2: Refit into multiple destination pools and tied-embedding + UVM fixes (NVIDIA#5187)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Disag MR1: Add inference shard specs and pg-collection building (NVIDIA#5186)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [Fix] Fix MoE router z-loss compatibility with TE CUDA Graph capture. (NVIDIA#5401)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>

* test: mark ep_a2a_overlap activation-offloading test flaky_in_dev (NVIDIA#5450)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark TestParallelTransformerBlockCudagraphs::test_gpu_cudagraph flaky_in_dev (NVIDIA#5475)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark gated_delta_net selective-recompute test flaky_in_dev (NVIDIA#5476)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections for sync test/impl reconciliation

1) test_optimizer.py: reverted to dev. The sync kept dev's
   multi_latent_attention.py (split q/kv down-proj, no
   _synthesize_fused_qkv_down_weight), but auto-merged main's
   test asserting the fused linear_qkv_down_proj.weight key.

2) training.py: guard the dev-only sequence_packing_scheduler config
   access with getattr (lines in train_step and train()). main's new
   MIMO schedule-plumbing test (NVIDIA#5333) passes an empty SimpleNamespace
   config; the reconciled training.py keeps dev's packing path, so the
   access must tolerate a config lacking the attribute. Real configs are
   unaffected (getattr returns the same value).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>

* fix sequence packing wrapper for eval (NVIDIA#5483)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* [dev] Megatron Lite (4/4) shared attention (NVIDIA#5427)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* fix: restore fused group MLP offload in main2dev sync (NVIDIA#5493)

Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>

* [dev] [DeepSeek-v4] Packed Sequence (THD) support for DSv4 Hybrid Attention (NVIDIA#5011)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* Improve default dynamic CP packing scheduler (NVIDIA#5154)

Signed-off-by: tailaim <tailaim@nvidia.com>

* [dev] moe(perf): Pre-GDR kernel fusion (NVIDIA#5361)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Preserve DSA output across fused inverse RoPE (NVIDIA#5526)

Signed-off-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (2/N) (NVIDIA#5485)

Signed-off-by: guihong-nv <guihongl@nvidia.com>

* [dev] Add experimental decoupled compact LayerWise DDP layout for Muon (NVIDIA#5388)

Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] Sync Megatron Lite with the latest implementation (NVIDIA#5577)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* [Dev] fix padding mask docstring (NVIDIA#5598)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>

* handle split_dtensor

---------

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Cory Ye <cye@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Li Ding <liding@nvidia.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Jingyue Wu <wujingyue@gmail.com>
Signed-off-by: conver334 <conver334@gmail.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Signed-off-by: Aditya Singh <adisin650@gmail.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Signed-off-by: meg miranda <mmiranda@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com>
Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Signed-off-by: Shivanjan Chakravorty <shivanjanc@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Xiaowei Ren <xren@nvidia.com>
Signed-off-by: yuzhongw <yuzhongw@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: ksivamani <ksivamani@nvidia.com>
Signed-off-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: Pavel Gein <pavel.gein@gmail.com>
Signed-off-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Signed-off-by: jinliangl <jinliangl@nvidia.com>
Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: ykarnati <ykarnati@nvidia.com>
Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com>
Signed-off-by: Shijie Wang <jaywan@nvidia.com>
Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Signed-off-by: jinliangl <975761915@qq.com>
Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Signed-off-by: ilml <tolong@nvidia.com>
Signed-off-by: sraman <sraman@nvidia.com>
Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>
Signed-off-by: Hollow Man <hollowman@opensuse.org>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com>
Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zijie Yan <zijiey@nvidia.com>
Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>
Co-authored-by: Cory Ye <44509866+cspades@users.noreply.github.com>
Co-authored-by: Yan Xu <45385219+Connor-XY@users.noreply.github.com>
Co-authored-by: Pingtian Li <158665726+Wohox@users.noreply.github.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: xuwchen <xuwenc@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Li Ding <liding@nvidia.com>
Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com>
Co-authored-by: Li Jinliang <jinliangl@nvidia.com>
Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com>
Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: Jingyue Wu <wujingyue@gmail.com>
Co-authored-by: Gao Deng <160076886+gdengk@users.noreply.github.com>
Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com>
Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com>
Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>
Co-authored-by: janEbert <janpabloe@nvidia.com>
Co-authored-by: Simiao Zhang <56124251+conver334@users.noreply.github.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Co-authored-by: Shaurya Singh <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
Co-authored-by: Kevin <23258141+kaimo455@users.noreply.github.com>
Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Moozy <108285604+Moozy23232@users.noreply.github.com>
Co-authored-by: Aditya Singh <60082699+adityasingh2400@users.noreply.github.com>
Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@nvidia.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>
Co-authored-by: Dingqing Yang <dingqingy@nvidia.com>
Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>
Co-authored-by: mathemakitten <helenn@nvidia.com>
Co-authored-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Yu Yao <54727607+yaoyu-33@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com>
Co-authored-by: kajalj22 <kajalj@nvidia.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: megnvidia <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>
Co-authored-by: Xuanteng Huang <44627253+xuantengh@users.noreply.github.com>
Co-authored-by: Tailai Ma <58548582+xiaoyao0115@users.noreply.github.com>
Co-authored-by: Santosh Bhavani <santosh.bhavani@live.com>
Co-authored-by: Zhongbo Zhu <42691305+zhongbozhu@users.noreply.github.com>
Co-authored-by: Nick Schank <nick@reflection.ai>
Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Shivanjan Chakravorty <schakravorty846@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Tuomas Rintamaki <trintamaki@nvidia.com>
Co-authored-by: Tyler Poon <tylerpoon@gmail.com>
Co-authored-by: Collin McCarthy <cmccarthy@nvidia.com>
Co-authored-by: Matthieu Le <matthieul@nvidia.com>
Co-authored-by: Piotr Zelasko <pzelasko@nvidia.com>
Co-authored-by: Ehsan Hosseini Asl <ehosseiniasl@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>
Co-authored-by: Xiaowei Ren <103958965+xrennvidia@users.noreply.github.com>
Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com>
Co-authored-by: William Dykas <wdykas@oci-hsg-cs-001-vscode-03.cm.cluster>
Co-authored-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Xuesong Ye <xuesongyey@gmail.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Daisy Gao <daisyg@nvidia.com>
Co-authored-by: lichenlu <lichenlu_8618@163.com>
Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Gao Deng <gdeng@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Pavel Gein <pavelgejn@yandex.ru>
Co-authored-by: Gao Deng <gdeng@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Eric Harper <eharper@nvidia.com>
Co-authored-by: Abhishree Thittenamane <47577437+athitten@users.noreply.github.com>
Co-authored-by: Abhishree Thittenamane <athittenaman@cw-dfw-cs-001-login-01.cm.cluster>
Co-authored-by: root <root@pool0-01849.cm.cluster>
Co-authored-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Co-authored-by: Nan Zheng <80790206+nanz-nv@users.noreply.github.com>
Co-authored-by: Li Tao <lit@nvidia.com>
Co-authored-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Robert Kirby <ArEsKay3@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Cole Hawkins <chawkins@nvidia.com>
Co-authored-by: Qi Zhang <qizhang@nvidia.com>
Co-authored-by: Vasudevan Rengasamy <vrengasamy@nvidia.com>
Co-authored-by: tongliu <tongliu@nvidia.com>
Co-authored-by: shurkat-nvidia <shurkat@nvidia.com>
Co-authored-by: Rui WANG <51357798+rui23@users.noreply.github.com>
Co-authored-by: rionawang <rionawang@tencent.com>
Co-authored-by: Tom Long <tolong@nvidia.com>
Co-authored-by: Lintch <44701395+returnL@users.noreply.github.com>
Co-authored-by: Yan Bai <bayan@nvidia.com>
Co-authored-by: Deyu Fu <deyuf@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Co-authored-by: jingqiny-99 <jingqiny@nvidia.com>
Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Pranav Thombre <pthombre@nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Shijie <505749828@qq.com>
Co-authored-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Li Jinliang <975761915@qq.com>
Co-authored-by: Baibaifan <39549453+Baibaifan@users.noreply.github.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: HaochenYuan <106647990+HaochenYuan@users.noreply.github.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>
Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org>
Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com>
Co-authored-by: Kunlun Li <94586211+kunlunl@users.noreply.github.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

Signed-off-by: Jianbin Chang <jianbinc@nvidia.com>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
svcnvidia-nemo-ci pushed a commit to dimapihtar/Megatron-LM that referenced this pull request Aug 4, 2026
…VIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: low

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants