Skip to content

Perf fix - #4996

Merged
shanmugamr1992 merged 18 commits into
NVIDIA:mainfrom
shanmugamr1992:perf-fix
May 27, 2026
Merged

Perf fix#4996
shanmugamr1992 merged 18 commits into
NVIDIA:mainfrom
shanmugamr1992:perf-fix

Conversation

@shanmugamr1992

Copy link
Copy Markdown
Contributor
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do ?

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact @NVIDIA/mcore-oncall.

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Linked issue:

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

shanmugamr1992 and others added 16 commits May 14, 2026 12:21
… GPT 16B

Adds tests/performance_tests/, a server+client benchmark harness that spins
up `tools.run_dynamic_text_generation_server`, sweeps batch sizes against the
OpenAI-compatible `/v1/completions` endpoint, and compares throughput / avg-
p50-p99 latency / TPOT against checked-in baseline_values.json with a 10%
tolerance. Vendored from /Users/shanmugamr/inference-bench and adapted to
work inside the existing cog + nemo-run CI flow.

Three test cases shipped with baselines captured on cw-dfw H100:
- gpt/gpt_583m_perf: 4 batches (1, 8, 32, 128), ~5 min
- hybrid/hybrid_2b_perf: 4 batches (1, 8, 32, 128), ~8 min
- gpt/gpt_16b_perf: 3 batches (1, 8, 32), ~13 min

CI recipes in tests/test_utils/recipes/h100/{gpt,hybrid}-perf.yaml trigger on
every MR via the existing mr-github scope (alongside functional tests).

Skill + slash command for human-driven runs live under
skills/run-performance-tests/ and .claude/commands/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
These were committed by mistake in 1404dbe; they're local-only Claude
Code workflow files and should not live in the repo.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…dd +20% upper bound

Before this commit, throughput coverage for the dynamic-inference 583M dp=8
config lived in two places:

  - tests/functional_tests/test_cases/gpt/gpt_dynamic_inference_tp1_pp1_dp8_583m_throughputtest_zmq/
    (with a pytest at tests/functional_tests/python_test_utils/test_inference_regular_pipeline.py
    enforcing throughput within 10% and +20% upper-bound)
  - tests/performance_tests/test_cases/gpt/gpt_583m_perf/ (the new perf harness,
    DP=1, throughput-only regression check)

The functional-test entry was already commented out in
tests/test_utils/recipes/h100/gpt-dynamic-inference-with-coordinator.yaml, so
it had been dormant in CI. To keep all perf tests in one place, this commit:

1. Deletes the dormant functional test directory + its commented recipe stub.
2. Switches the perf-harness gpt_583m_perf test from DP=1 to DP=8 so it
   exercises the same multi-rank dynamic-batching + ZMQ coordinator path the
   deleted test used to cover. Re-recorded its baseline_values.json on cw-dfw
   H100 (batch_128 → 5643 tok/s, p99 latency 2901 ms, tpot 22.7 ms/tok).
3. Adds a +20% upper-bound throughput check to
   tests/performance_tests/shell_test_utils/compare_to_baseline.py, matching
   the rule the deleted pytest applied (improvements beyond UPPER_TOLERANCE_PCT
   fail with "refresh baseline" so the regression floor doesn't go stale).
   New `UPPER_TOLERANCE_PCT` config field, default 20.

Test-case changes:

  - DELETED: tests/functional_tests/test_cases/gpt/gpt_dynamic_inference_tp1_pp1_dp8_583m_throughputtest_zmq/
  - DELETED: commented entry from
    tests/test_utils/recipes/h100/gpt-dynamic-inference-with-coordinator.yaml
  - MODIFIED: tests/performance_tests/test_cases/gpt/gpt_583m_perf/model_config.yaml
    (DP 1 → 8; same batch sweep 1,8,32,128; same ISL/OSL 512/128)
  - MODIFIED: tests/performance_tests/test_cases/gpt/gpt_583m_perf/baseline_values.json
    (re-recorded for DP=8 on cw-dfw H100, 2026-05-19)

Recipe changes:

  - SPLIT: tests/test_utils/recipes/h100/gpt-perf.yaml now holds 1-GPU tests
    only (gpt_16b_perf). 8-GPU tests move to the new
    tests/test_utils/recipes/h100/gpt-perf-dp8.yaml (gpt_583m_perf at
    gpus: 8, ntasks-per-node: 1).

Runner changes (tests/performance_tests/shell_test_utils/run_perf_test.sh):

  - Added optional EP schema field in model_config.yaml.
    world_size = TP * PP * max(DP, EP). EP defaults to 1 so existing dense
    configs (583m, 2b, 16b) are unaffected.
  - Reordered the torchrun arg vector so SERVER_COMMON_ARGS come before
    MODEL_ARGS. With argparse's last-wins semantics, this lets a model's
    .args file override default server flags
    (e.g. --inference-dynamic-batching-buffer-size-gb,
    --inference-dynamic-batching-max-requests) without having to edit the
    runner.

Verified (cw-dfw H100, 2026-05-19): gpt_583m_perf (DP=8), hybrid_2b_perf,
gpt_16b_perf — all three pass compare_to_baseline.py against their checked-in
baselines within tolerance.

Not addressed in this PR (filed as follow-ups in skills/run-performance-tests/SKILL.md):
  - mem-max-allocated-bytes regression check (deleted pytest had ±5% on this;
    bringing it back via this harness needs a server-side /v1/stats endpoint
    or server.log parser).
  - Real-prompt-from-disk mode (deleted pytest loaded sharegpt-vicuna; our
    client uses synthetic "hello " * N prompts).
  - nanov3 perf coverage: 7 bootstrap attempts on cw-dfw via cog all hit a
    Triton incompatibility in megatron.core.inference.moe.permute kernels
    (`'constexpr_type' object has no attribute 'is_ptr'`) against the
    Triton 3.6.0 shipped in the cog dev image. The existing nanov3
    functional test uses a different CI-built image. Will revisit once a
    Triton-compatible image is available.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Resolves 5 add/add conflicts on perf-tests scaffolding that landed
separately on main (with the original DP=1 baseline) and on this branch
(with DP=8 + +20% upper-bound check). Took OUR side for all five since
this branch's content is the consolidation work.

  - tests/performance_tests/shell_test_utils/compare_to_baseline.py
  - tests/performance_tests/shell_test_utils/run_perf_test.sh
  - tests/performance_tests/test_cases/gpt/gpt_583m_perf/baseline_values.json
  - tests/performance_tests/test_cases/gpt/gpt_583m_perf/model_config.yaml
  - tests/test_utils/recipes/h100/gpt-perf.yaml

Also retains the deletion of tests/functional_tests/test_cases/gpt/
gpt_dynamic_inference_tp1_pp1_dp8_583m_throughputtest_zmq/ and the
corresponding commented stub removal in
tests/test_utils/recipes/h100/gpt-dynamic-inference-with-coordinator.yaml.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…back)

Reviewer feedback on PR NVIDIA#4917: MoE models should not benchmark with
synthetic `"hello " * N` prompts. Identical token IDs route to the same
expert subset every step, so router/dispatcher load is uniform → misleading
throughput. Same concern applies to mamba/hybrid models where SSM state
evolution is content-sensitive.

Changes:

  - Vendor 256 gsm8k prompts at tests/performance_tests/client/data/gsm8k_prompts.jsonl
    (from openai/gsm8k test split, ~64 KB). Keeps the test data-independent
    of cluster-side mounts.

  - Add `--dataset {synthetic,gsm8k}` flag to static_benchmark.py. gsm8k
    mode cycles through the vendored prompts deterministically across
    requests/iterations and ignores --num-input-tokens (prompts have natural
    variable length). results.json now records `num_input_tokens_avg`
    (measured) and `dataset` instead of the assumed fixed-length field.

  - Add `DATASET` field to model_config.yaml schema (default: synthetic for
    back-compat). Wired through run_perf_test.sh.

  - Switch gpt_16b_perf and hybrid_2b_perf to `DATASET: gsm8k`.
    gpt_583m_perf (dense) stays on synthetic — the MoE/hybrid concern
    doesn't apply.

  - Re-record baselines for gpt_16b_perf and hybrid_2b_perf on cw-dfw H100.
    The new MoE numbers tell the real story: gpt_16b at batch=32 measures
    317 tok/s with 101 ms/tok (vs the prior synthetic 1356 tok/s @ 24 ms/tok
    — ~4× artifact from uniform expert routing).

Also restores the .pth-based mamba-ssm shim in run_perf_test.sh that was
inadvertently reverted by the merge from main: the older PYTHONPATH-prepend
form shadowed the cog venv's nvrx 0.6.0 with /opt/venv's 0.6.0.dev69 and
broke hybrid model load. The .pth approach appends /opt/venv *after* the
venv on sys.path, so newer cog-venv packages win.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Reviewer feedback on PR NVIDIA#4917: OSL=128 lets prefill dominate the per-token
average and inflates TPOT; the production MoE workload (Nemo-RL rollouts) is
decode-heavy. Bumping to OSL=2048 to reflect that regime.

Recorded fresh baseline on cw-dfw H100 (1 GPU, gsm8k prompts, ~1h 12m total
wall time across batches 1/8/32). The actual finding is mild:

  OSL=128  batch=32  ->  317 tok/s, tpot=101 ms/tok
  OSL=2048 batch=32  ->  326 tok/s, tpot= 98 ms/tok

TPOT only dropped 3 ms/tok. Prefill *was* a small share of total work (ISL
~60 tokens vs OSL 128), and ~95-100 ms/tok is the actual steady-state MoE
decode rate, not a prefill artifact. The longer OSL still gives a cleaner
decode-only signal, which is what we want for catching regressions in the
expert-dispatch path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
New 8-GPU inference perf test mirroring the nanov3 3B hybrid MoE
functional config (chunked prefill + local CUDA graphs + nvls dispatcher
+ inference_optimized transformer impl + mamba state dtypes fp32).
Uses gsm8k prompts since both the mamba state path and the MoE
router/dispatcher are content-sensitive.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…flag list)

Reviewer asked to add to gpt_16b:
  --transformer-impl inference_optimized
  --inference-grouped-gemm-backend vllm
  --inference-moe-token-dispatcher-type nvls
  --moe-shared-expert-overlap
  --cuda-graph-impl local
  --cuda-graph-scope full_iteration_inference
  --inference-dynamic-batching-num-cuda-graphs -1
  --enable-chunked-prefill

Only the last one is compatible with this model:

  - `inference_optimized` rejects --swiglu, but the 16B deepseek checkpoint
    requires SwiGLU (encoded in its FFN weight layout). Dropping --swiglu
    would silently corrupt model outputs.
  - vllm-backend MoE / nvls dispatcher / shared-expert overlap all *require*
    inference_optimized, so they're out by transitive closure.
  - CUDA graph capture at full_iteration_inference scope crashes because the
    alltoall MoE dispatcher (transformer_engine path) does
    `d2h_event.synchronize()` inside its forward, which is illegal during
    stream capture.

Comment block in gpt_16b.args records the incompatibility so reviewers don't
re-request these flags. The inference_optimized stack will become available
for this model once it supports SwiGLU (groundwork exists — see commit
6815c0f "Combine GEMM + SwiGLU fused MLP PRs" — but the route isn't wired
through inference_optimized yet).

Re-recorded baseline on cw-dfw H100. Chunked prefill is a ~no-op in our
decode-heavy regime (OSL=2048, avg ISL=60 means decode is 97% of work):
batch=1 10.5 -> 10.2 tok/s (-3%), batch=8 83.4 -> 80.9 (-3%),
batch=32 325.7 -> 320.7 (-1.5%) — all within run-to-run noise. Keeping the
flag for the prefill-warm-up portion and for prompt-mix tests we might add
later.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Drops the per-CI-run wall time for gpt_16b_perf from ~1h 12m down to ~10m
without losing the steady-state decode signal. The numbers barely move:

  OSL=2048 batch=32:  326 tok/s, tpot=98 ms/tok, avg_lat=201 s
  OSL=256  batch=32:  325 tok/s, tpot=99 ms/tok, avg_lat= 25 s

Prefill (avg ISL=60 in gsm8k mode) is amortized over 256 output tokens,
which is enough that the TPOT delta from OSL=2048 is well inside run-to-run
noise (~1 ms). The decode-dominated regime the reviewer cares about is
already visible at OSL=256.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
With NUM_TIMED_ITERS=5, p99 is literally max-of-5 samples — not a real
percentile — so a single transient iter spike trips the regression check
even when throughput / avg / p50 are stable. p99 is still recorded in
results.json for visibility, just not gated on.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@shanmugamr1992
shanmugamr1992 requested a review from a team as a code owner May 26, 2026 19:07
@copy-pr-bot

copy-pr-bot Bot commented May 26, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@svcnvidia-nemo-ci
svcnvidia-nemo-ci marked this pull request as draft May 26, 2026 19:07
@github-actions

Copy link
Copy Markdown
Contributor

This PR has been automatically converted to draft because all PRs must start as drafts.

When you are ready for review, click Ready for Review to begin the review process. This will:

  1. Add the oncall reviewer (optional reviewer)
  2. Add required review teams based on your changes

See the contribution guide for more details.

shanmugamr1992 and others added 2 commits May 26, 2026 12:08
# Conflicts:
#	tests/performance_tests/shell_test_utils/run_perf_test.sh
These three torchrun args were added by NVIDIA#4951 on main but lost when
perf-fix branched off perf-tests (which predates NVIDIA#4951). The merge of
main into perf-fix did not pick them up cleanly. Restoring so the file
matches main exactly — the PR no longer touches run_perf_test.sh.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@shanmugamr1992
shanmugamr1992 marked this pull request as ready for review May 26, 2026 23:04
@shanmugamr1992
shanmugamr1992 enabled auto-merge May 26, 2026 23:04
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team May 26, 2026 23:04
@shanmugamr1992

Copy link
Copy Markdown
Contributor Author

/ok to test 59ba2fd

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/26535310603

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks May 27, 2026
@shanmugamr1992
shanmugamr1992 added this pull request to the merge queue May 27, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/26538368647

Merged via the queue into NVIDIA:main with commit 805e24e May 27, 2026
82 checks passed
@shanmugamr1992
shanmugamr1992 deleted the perf-fix branch May 27, 2026 21:45
mathemakitten pushed a commit to mathemakitten/Megatron-LM that referenced this pull request Jun 12, 2026
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
yhgalaxy pushed a commit to yhgalaxy/Megatron-LM that referenced this pull request Jun 17, 2026
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
shjwudp added a commit to shjwudp/Megatron-LM that referenced this pull request Jul 6, 2026
* test: re-enable test_pp2_create_cudagraphs_first_stage on TE 2.15+ (NVIDIA#4985)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tests): initialize num_microbatches calculator in vision cudagraph tests (NVIDIA#4986)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci: Add allow_failure flag to gpt and moe recipes that are failing in nightlies (NVIDIA#4905)

* Drain predecessor reduce-scatter at dispatch time (NVIDIA#4940)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* nightly(ci): Update golden values for functional t5 tests (NVIDIA#4995)

* chore: rotate oncall schedule

* [main] Refactor and Improve MoE Logginginit commit (NVIDIA#3431)

Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>

* ci: validate release branch-rules (NVIDIA#4929)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Megatron-FSDP] Add conditional param.grad dereferencing logic to support full-iteration (FWD-BWD) CUDA graphability. (NVIDIA#4663)

Signed-off-by: Cory Ye <cye@nvidia.com>

* test: restrict iter-time comparison to steady-state window (NVIDIA#5010)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Add mHC support for HybridModel on dev (NVIDIA#4949)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* fix(test): pin eval-global-batch-size on 15b gb200 release configs (NVIDIA#5022)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* [fix] Release MTP assertion when EP overlap with PP=1 (NVIDIA#4796)

* fix(test): widen iter-time steady-state window for short tests (NVIDIA#5023)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Perf fix (NVIDIA#4996)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add dev-feature preservation gate and change schedule (NVIDIA#4773)

* chore(test): remove orphan nemotron3_super_release_g200 dir (NVIDIA#5024)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Ignore Claude worktree directory (NVIDIA#5020)

* Update copy-pr-bot.yaml [skip ci]

* ci: update CI workflow conditions for integration tests (NVIDIA#4658)

* Add NVSkills CI request workflow (NVIDIA#5033)

* DDP wrap pg size fixes (NVIDIA#5006)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* fix(layer_wise): tag MTP-stage word_embeddings as is_embedding_or_output_parameter (NVIDIA#5034)

* [Dev] Add Qwen3 30B MoE recipes (NVIDIA#5012)

* [DEV] fix(megatron-fsdp): reduce padding for grouped expert weights (NVIDIA#5013)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4908)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [dev] [fix] [DeepSeek-v4] fix dense loss and rope type in DSv4 Hybrid Attention (NVIDIA#5018)

Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Move LTS dependencies from pyproject.toml to Dockerfile.ci.lts (NVIDIA#4877)

* Use shared ModelOpt calibration loop on 0.45+ with 0.44 fallback fix (NVIDIA#4881)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* test(release): skip golden comparison on intermediate resume windows (NVIDIA#5040)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [mimo] Thread position_ids through MimoModel for multimodal RoPE (NVIDIA#4938)

Signed-off-by: Li Ding <liding@nvidia.com>

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5039)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Fix: Import unwrap_model from megatron.core.utils in modelopt examples (NVIDIA#5045)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Simple and stable Inference APIs (NVIDIA#4697)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: Add notification step for MBridge downstream test results (NVIDIA#5028)

* [dev] [5/5] Qwen3.5 support: Qwen3.5-VL training example (NVIDIA#4751)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: Update Docker image version to 26.04-py3 on dev (NVIDIA#5051)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Delete output tensor early (NVIDIA#4742)

* Support ScaledSReLU in TE grouped MLP fuser (NVIDIA#4859)

Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: a <a>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>

* Skip gradient updates when grad norm exceeds threshold (NVIDIA#3460)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Add 9 user skills (NVIDIA#5066)

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>

* test(nemotron): align nemotron3 super GB200 goldens with exit-interval 4768 (NVIDIA#5069)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* chore: Update transformer-engine dependency to version 2.16.0 (NVIDIA#4992)

* Update energon version requirement (NVIDIA#4572)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Fix test failures for new inference APIs (NVIDIA#5068)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): set PYTHONUNBUFFERED=1 in JET workload env (NVIDIA#5072)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Preserve non-FSDP-unit buckets across AllGatherPipeline reset (NVIDIA#4717)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add opt-in MXFP8 LM-head output projection (NVIDIA#4825)

Co-authored-by: a <a>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-01)

* fix(ci): bound JET pipeline polling with a watchdog to prevent indefinite hangs (NVIDIA#5076)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: prune old artifacts on cluster lustre during weekly/release runs (NVIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci(test): isolate ckpt-resume tensorboard per phase (NVIDIA#5074)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: unmark EP A2A activation offload test flaky (NVIDIA#5009)

* Change ownership groups (NVIDIA#5021)

* test: skip mfsdp_fully_shard cases when world_size < mesh size (NVIDIA#4487)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix mimo optimizer checkpoint metadata restore (NVIDIA#4791)

Signed-off-by: Li Ding <liding@nvidia.com>

* [mimo] Support bridge fan-out for variable modality tokens (NVIDIA#5062)

Signed-off-by: Li Ding <liding@nvidia.com>

* Add separate mtp_grad_scale_func for MTP loss scaling (NVIDIA#3459)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [training migration] Migrate GPT builder (NVIDIA#4741)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Make Mamba conv params direct mixer params (NVIDIA#4899)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Update copy-pr-bot.yaml [skip ci]

* Update oncall reviewer assignment (NVIDIA#5093)

* Pass explicit process groups to hybrid logging (NVIDIA#4781)

* Clean up top-level repository files (NVIDIA#5097)

* [main] fix(moe): Fix several bugs for DSA rope and spec. (NVIDIA#3026)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* [Dev] Skip identity alltoall chunk sort (NVIDIA#5102)

* Move MIMO unit tests into models/mimo (NVIDIA#5063)

* test: update DeepSeek FSDP2 GB200 memory golden (NVIDIA#5094)

* Remove DeepEP hardware limit check (NVIDIA#4846)

* Update transformer-engine dependency to revision 4220403 (NVIDIA#5112)

* ci: make CI resilient to pip/uv network timeouts (NVIDIA#5118)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: treat docker container-removal conflict as flaky (NVIDIA#5120)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix GDN DTensor splitting for FSDP checkpointing (NVIDIA#4843)

Signed-off-by: conver334 <conver334@gmail.com>

* Fix MoE aux_loss / z_loss gradient scaling with TP > 1 (NVIDIA#5047)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update copy-pr-bot.yaml [skip ci]

* Update Claude copy workflow to enforce user restrictions and improve error messages (NVIDIA#5117)

* Add advisory process group guidance to Claude reviews (NVIDIA#5111)

* build: cap pydantic<2.14 in transformer-engine dependency metadata (NVIDIA#5125)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* fix(test): skip scalar-less tensorboard event files in resume checks (NVIDIA#5121)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: rotate oncall schedule

* docs: fix contributor guide typo (NVIDIA#4858)

Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* ci(unit-tests): split slow unit-test buckets over 15min SLA (NVIDIA#5133)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix DSA indexer loss not averaged across micro-batches (NVIDIA#4070)

Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Update MINOR version to 19 (NVIDIA#5096)

* Fix Muon QKV split for gated attention (NVIDIA#4728)

Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Roll input IDs for MTP labels (NVIDIA#3457)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Refactor: Move paged stashing Triton kernels (NVIDIA#5003)

* Adding blackwell tests (NVIDIA#5113)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Relax atol for test_router_gating_linear router_dtype=torch.float32 (NVIDIA#4915)

Signed-off-by: Aditya Singh <adisin650@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Fix incorrect inference metadata tensor dtypes (NVIDIA#4855)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: sraman-rgb <sraman@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>

* Disable TE cross entropy loss fusion (NVIDIA#5115)

Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>

* fix(optimizer): gate ChainedOptimizer MXFP8 defer-sync on DDP-level overlap_param_gather (NVIDIA#4982)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Pass TP group to unfused cross entropy (NVIDIA#5128)

* fix: correct dsv4_hybrid Q-up FLOPs by using args.v_head_dim (NVIDIA#5142)

* [dev] [follow-up] Qwen3.5 support: MoE aux loss padding_mask (NVIDIA#4776)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>

* [Dev] Support isolated MTP loss (NVIDIA#5080)

Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>

* test(elastification): quarantine flaky test_gumbel_determinism as flaky_in_dev (NVIDIA#5156)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(notify): mention mcore-oncall and philipp on critical CI events (NVIDIA#5152)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Change the cudagraph distribution from linearly to exponentially-decreasing + grid for mixed prefill (NVIDIA#3509)

* ci: Disable a few gb200 test cases to support 2 branches. (NVIDIA#5151)

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5164)

* Add MTP acceptance rate metrics (NVIDIA#3458)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Nemotron Ultra config for ModelOpt examples (NVIDIA#5159)

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>

* Make MTP / prefix cache stats persist for engine lifetime (NVIDIA#4101)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* Restore Greptile configuration (NVIDIA#5166)

* chore: bump `_code_freeze` workflow to `v1.4.2` (NVIDIA#5132)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [dev] moe(fix): Avoid TE cuda graph dummy attention masks (NVIDIA#5131)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* ci: Remove docs build test in favor of release test (NVIDIA#5182)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Move TE cross entropy guard to training args (NVIDIA#5162)

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>

* Fix error in deepseek parser (NVIDIA#5136)

* Fix logprob slicing for 0 generated token case (NVIDIA#5167)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* [Perf] Fold frozen linear dgrad matmul (NVIDIA#5092)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* Clamp `max_new_tokens` in MInf to mirror vllm (NVIDIA#5181)

* build: add managed = true to [tool.uv] (NVIDIA#5190)

Signed-off-by: Kajal Jain <kajalj@nvidia.com>

* Stabilize GB200 inference perf tests against cold-start noise (NVIDIA#5171)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections (docs dup fields, nvrx version gate, utils re-export)

- Remove merge-introduced duplicate definitions of use_transformer_engine_op_fuser
  and moe_expert_rank_capacity_factor (transformer_config.py) and the duplicate
  _post_param_sync method (param_and_grad_buffer.py) — tripped the docs build.
- Relax is_nvrx_min_version() to compare release segments so dev's pinned
  nvidia-resiliency-ext pre-release (0.6.0.dev33) satisfies main's '>= 0.6.0'
  assertion (the required-symbol check remains the authoritative guard).
- Re-export unwrap_model from megatron/training/utils/__init__.py. main refactored
  the monolithic training/utils.py into a package whose __init__ dropped this
  re-export; dev's training.py (and others) import it via 'from .utils import',
  which broke tests/unit_tests/conftest.py's import chain (cascaded to all suites).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] [DeepSeek-v4] Add ClampedSwiGLU to MoE mlp_op_fuser and add force balance to hash routing (NVIDIA#5130)

* nvidia style guide audit for getting started folder (NVIDIA#5168)

Signed-off-by: meg miranda <mmiranda@nvidia.com>

* AI aided audit for Nvidia Style guidance (NVIDIA#5141)

Signed-off-by: meg miranda <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>

* Enable selective recompute for `norm_out` in GDN layers  (NVIDIA#4715)

* fix(elastification): align with get_batch + utils refactors (NVIDIA#5194)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4909)

* Minor improvements for Dynamic-cp (NVIDIA#4226)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>

* fix: post-CI corrections

* chore(beep boop 🤖): Bump  (main) (2026-06-08)

* Avoid stat syscall in rerun result validation (NVIDIA#5107)

Signed-off-by: dimapihtar <dpykhtar@nvidia.com>

* docs: Update Latest News in README.md (NVIDIA#3790)

* Fix bug with Megatron-FSDP zero counter not working with decoupled gradients. (NVIDIA#4802)

Signed-off-by: Cory Ye <cye@nvidia.com>

* varlendataset for thd e2e and benchmark (NVIDIA#4832)

Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* ci: add smoke tests (NVIDIA#5143)

* Add mtp_detach_heads config to detach MTP head inputs (NVIDIA#3456)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: wrap mtp logging comment

* docs: fix install guide NGC container anchor (NVIDIA#5224)

* [Dev] Add separate toggle for varlen input padding for HybridEP in THD training  (NVIDIA#5048)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* Fuse per-sequence AlltoAll into a unified one in GDN forward (NVIDIA#4913)

* Apply MIMO SP/CP sharding with explicit groups and enable THD in non-colocated path (NVIDIA#5150)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix CUDA IMA in fsdp_double_buffer when an FSDP unit's bucket doesn't fit the pool (NVIDIA#4810)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add named layouts to HyperCommGrid for heterogeneous parallelism (NVIDIA#5148)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix wgrad race condition when using double buffers. (NVIDIA#5222)

Signed-off-by: Cory Ye <cye@nvidia.com>

* Move uneven DTensor distributed fixture to conftest (NVIDIA#5237)

* Route bridge communicator cross-grid P2P through a dedicated process group (NVIDIA#5234)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix test_split_tensor_along_last_dim to actually assert correctness (NVIDIA#4710)

Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* chore: rotate oncall schedule

* Add optional group= to common_utils model/data-parallel reduction helpers (NVIDIA#5251)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement to Contribution guide (NVIDIA#5278)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Add MIMO hetero topology + distributed bootstrap (examples/mimo training-loop folder) (NVIDIA#5260)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)

Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>

* [Dev] DeepSeek-V4-Flash recipe 20260610 (NVIDIA#5266)

* [Dev] Generalized fix for mxfp8 param gather  (NVIDIA#4994)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Cherry-pick MTP detach heads (NVIDIA#5223)

* [TE] Restore original CP group after dynamic CP forward in TEDotProductAttention (NVIDIA#5215)

Co-authored-by: rionawang <rionawang@tencent.com>

* [examples] Add dynamic context parallel benchmark example (NVIDIA#5123)

* Remove duplicate nccl_allocator import (NVIDIA#5057)

* [dev] Add experimental Megatron Lite as agentic exploration (NVIDIA#4885)

Co-authored-by: Deyu Fu <deyuf@nvidia.com>

* fix(ci): resolve t5 dataloader stall + GRPO cudagraph-memory regression (CI-validated) (NVIDIA#5280)

Signed-off-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix Dockerfile warnings (NVIDIA#4856)

* Fix fused MLA delayed weight grad hooks (NVIDIA#5273)

Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>

* ci: limit retries on unsuccessful test launches (NVIDIA#5275)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Thread pg_collection into get_model DDP bucket sizing (NVIDIA#5250)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Enable non-deterministic results in model configuration for nemotron tests (NVIDIA#5239)

* Stabilize hybrid nanov3 gb200 perf (NVIDIA#5295)

* Clip mtp grads separately when mtp_detach_heads=True (NVIDIA#4116)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Thread pg_collection into train_step reductions (NVIDIA#5259)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement (NVIDIA#5305)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>

* fix: restore dev-only training.py features dropped by the main override

The merge takes main's training.py wholesale (per the sync skill's override
list), which silently reverted four dev-only features whose supporting core
code the merge keeps at dev's version. Each surfaced as a CI failure:

1. Dynamic context-parallel API rename. Dev renamed
   get_hybrid_data_context_parallel_groups -> get_dynamic_data_context_parallel_groups
   (identical signature) and args.hybrid_context_parallel -> dynamic_context_parallel.
   Point training.py at dev's names. Fixes the conftest.py ImportError that
   cascaded to every unit-test bucket.

2. Dynamic-CP / sequence-packing data loading. Replace main's setup-time
   HybridCPDataLoaderWrapper wrap with dev's per-step wrap_data_iterator
   (gated on config.sequence_packing_scheduler), which returns a
   RerunDataIterator-compatible iterator. Fixes the *_cp4_dcp RerunDataIterator
   assertion.

3. MTP loss logging scale. Dev's MTPLossLoggingHelper stores raw loss sums and
   token counts and computes the per-token loss after reduction, so the log
   scale must be 1.0; main's 1/get_num_microbatches() divided the reported
   mtp_N loss by num_microbatches. Fixes moe gpt3_..._scoped_cudagraph mtp_1
   loss (was 16x too small: 0.682 vs golden 10.915).

4. DSA indexer loss cross-PP reduction. dsa.py requires num_layers (and
   csa_compress_ratios) so first-pipeline-stage ranks lazily initialize the
   tracker and join the cross-PP all_reduce in reduce_loss_in_tracker; main's
   call omitted them, so stage 0 returned early and the last stage hung on an
   unmatched all_reduce. Fixes the gpt3_..._dsv4_hybrid_mhc_mtp NCCL timeout.

Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>

* [dev]: faster implementation of mHC fused kernels (NVIDIA#4624)

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>

* Allow for pre-bound socket to be passed in server (NVIDIA#5301)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Offline Logits-Based Knowledge Distillation (NVIDIA#5019)

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

* Handle None values in sampling parameters (NVIDIA#5300)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Add moe loss normalization for RL SFT (NVIDIA#3956)

Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Add code owners for optimizer-related files (NVIDIA#5297)

Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* Fix EP=1 inference by allocating buffers anyway (NVIDIA#5233)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Fix crash due to tool call at sequence length (NVIDIA#5302)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Inference: Cudagraph-aware admission gating in prefill scheduler (NVIDIA#4870)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Account for reasoning token stripping (NVIDIA#5313)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Thread pg_collection through wrap_model_chunks_with_ddp (NVIDIA#5328)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Dev] Add MoE recipe performance summary (NVIDIA#5289)

Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>

* [Dev] Add DeepEP v2 flex dispatcher backend (NVIDIA#4793)

Signed-off-by: tongliu <tongliu@nvidia.com>
Co-authored-by: Dennis(Zhenhuan) Liu <denliu@nvidia.com>

* fix tflops calculation when sequence_packing_scheduler is not none (NVIDIA#5342)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-15)

* [dev] bump emerging optimizers to v0.3.0 (NVIDIA#5320)

Signed-off-by: Deyu Fu <deyuf@nvidia.com>

* Fix LatentMoE theoretical memory estimate (NVIDIA#5145)

Signed-off-by: Shijie Wang <jaywan@nvidia.com>

* Add zstandard package to Docker LTS requirements. Fix nightly failures (NVIDIA#5347)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* fix: preserve seqlen stats in train_step

Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>

* Thread MIMO support through the stock training loop (schedule + optimizer) (NVIDIA#5333)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (1/N) (NVIDIA#5042)

Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(ssm): whole-module 'gdn' selective recompute for GatedDeltaNet (NVIDIA#5296)

Signed-off-by: jinliangl <975761915@qq.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Fix] Fix optimizer parameter override bugs. (NVIDIA#5213)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Add Megatron-FSDP weight prefetch for full recompute (NVIDIA#5175)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* ci: default functional test time limit to 4h for release/weekly scopes (NVIDIA#5360)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Dev] add cuda graph support for thd format training. (NVIDIA#4359)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>

* Fix memory leak with log_max_attention_logit (NVIDIA#4699) (NVIDIA#5067)

Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>

* Clean up pretrain_gpt.py and pretrain_hybrid.py formatting and remove module globals (NVIDIA#5351)

Signed-off-by: ilml <tolong@nvidia.com>

* Add full model cuda graph support for MTP inference (NVIDIA#4950)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Expand the Mamba prefix caching memory safety check to include scratch space buffers (NVIDIA#5348)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Make Megatron RL only materialize last token logit (NVIDIA#4551)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Profiling  (NVIDIA#3110)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Support fused MLA QKV checkpoint reload (NVIDIA#5310)

Signed-off-by: sraman <sraman@nvidia.com>

* Add minimal DBuffer implementation (NVIDIA#4835)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [Dev] restore DSv4 tflops calc in training and fix the packed seq case (NVIDIA#5358)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* [split 1/5] Fix packed THD RoPE under CP (NVIDIA#5243)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* Update copy-pr-bot.yaml [skip ci]

* Document agent PR commit sign-off and signing (NVIDIA#5381)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Remove unused distributed pytest markers (NVIDIA#5380)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [feat] Support fine-grained activation offloading in fused group mlp (NVIDIA#5082)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* Thread tensor-parallel group into the RADIO patch embedder (NVIDIA#5371)

Signed-off-by: ykarnati <ykarnati@nvidia.com>

* Add MimoModel.zero_grad_buffer delegating to active DDP submodules (NVIDIA#5372)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [split 3/5] Refactor absorbed MLA projection handling (NVIDIA#5245)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* [dev] moe(perf): Restore fused GDN THD all-to-all on dev (NVIDIA#5389)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* chore: rotate oncall schedule

* ci: Remove sync skills workflow (NVIDIA#5091)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Add flaky marker to fine-grained activation offloading test (NVIDIA#5350) (NVIDIA#5368)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Revert "Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)" (NVIDIA#5366)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Update goldens for weekly tests after pytorch and TE bumps. (NVIDIA#5399)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Add MIMO runtime setup: per-role RNG seeding and DDP wrapping (NVIDIA#5285)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Add --mamba-training-ssm-states-dtype argument (NVIDIA#5309)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-22)

* Fix Mamba prefix match for chunked prefill (NVIDIA#4758)

Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>

* Disag MR2: Refit into multiple destination pools and tied-embedding + UVM fixes (NVIDIA#5187)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Disag MR1: Add inference shard specs and pg-collection building (NVIDIA#5186)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [Fix] Fix MoE router z-loss compatibility with TE CUDA Graph capture. (NVIDIA#5401)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>

* test: mark ep_a2a_overlap activation-offloading test flaky_in_dev (NVIDIA#5450)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark TestParallelTransformerBlockCudagraphs::test_gpu_cudagraph flaky_in_dev (NVIDIA#5475)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark gated_delta_net selective-recompute test flaky_in_dev (NVIDIA#5476)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections for sync test/impl reconciliation

1) test_optimizer.py: reverted to dev. The sync kept dev's
   multi_latent_attention.py (split q/kv down-proj, no
   _synthesize_fused_qkv_down_weight), but auto-merged main's
   test asserting the fused linear_qkv_down_proj.weight key.

2) training.py: guard the dev-only sequence_packing_scheduler config
   access with getattr (lines in train_step and train()). main's new
   MIMO schedule-plumbing test (NVIDIA#5333) passes an empty SimpleNamespace
   config; the reconciled training.py keeps dev's packing path, so the
   access must tolerate a config lacking the attribute. Real configs are
   unaffected (getattr returns the same value).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>

* fix sequence packing wrapper for eval (NVIDIA#5483)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* [dev] Megatron Lite (4/4) shared attention (NVIDIA#5427)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* fix: restore fused group MLP offload in main2dev sync (NVIDIA#5493)

Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>

* [dev] [DeepSeek-v4] Packed Sequence (THD) support for DSv4 Hybrid Attention (NVIDIA#5011)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* Improve default dynamic CP packing scheduler (NVIDIA#5154)

Signed-off-by: tailaim <tailaim@nvidia.com>

* [dev] moe(perf): Pre-GDR kernel fusion (NVIDIA#5361)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Preserve DSA output across fused inverse RoPE (NVIDIA#5526)

Signed-off-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (2/N) (NVIDIA#5485)

Signed-off-by: guihong-nv <guihongl@nvidia.com>

* [dev] Add experimental decoupled compact LayerWise DDP layout for Muon (NVIDIA#5388)

Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] Sync Megatron Lite with the latest implementation (NVIDIA#5577)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* [Dev] fix padding mask docstring (NVIDIA#5598)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>

* handle split_dtensor

---------

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Cory Ye <cye@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Li Ding <liding@nvidia.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Jingyue Wu <wujingyue@gmail.com>
Signed-off-by: conver334 <conver334@gmail.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Signed-off-by: Aditya Singh <adisin650@gmail.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Signed-off-by: meg miranda <mmiranda@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com>
Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Signed-off-by: Shivanjan Chakravorty <shivanjanc@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Xiaowei Ren <xren@nvidia.com>
Signed-off-by: yuzhongw <yuzhongw@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: ksivamani <ksivamani@nvidia.com>
Signed-off-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: Pavel Gein <pavel.gein@gmail.com>
Signed-off-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Signed-off-by: jinliangl <jinliangl@nvidia.com>
Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: ykarnati <ykarnati@nvidia.com>
Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com>
Signed-off-by: Shijie Wang <jaywan@nvidia.com>
Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Signed-off-by: jinliangl <975761915@qq.com>
Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Signed-off-by: ilml <tolong@nvidia.com>
Signed-off-by: sraman <sraman@nvidia.com>
Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>
Signed-off-by: Hollow Man <hollowman@opensuse.org>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com>
Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zijie Yan <zijiey@nvidia.com>
Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>
Co-authored-by: Cory Ye <44509866+cspades@users.noreply.github.com>
Co-authored-by: Yan Xu <45385219+Connor-XY@users.noreply.github.com>
Co-authored-by: Pingtian Li <158665726+Wohox@users.noreply.github.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: xuwchen <xuwenc@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Li Ding <liding@nvidia.com>
Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com>
Co-authored-by: Li Jinliang <jinliangl@nvidia.com>
Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com>
Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: Jingyue Wu <wujingyue@gmail.com>
Co-authored-by: Gao Deng <160076886+gdengk@users.noreply.github.com>
Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com>
Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com>
Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>
Co-authored-by: janEbert <janpabloe@nvidia.com>
Co-authored-by: Simiao Zhang <56124251+conver334@users.noreply.github.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Co-authored-by: Shaurya Singh <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
Co-authored-by: Kevin <23258141+kaimo455@users.noreply.github.com>
Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Moozy <108285604+Moozy23232@users.noreply.github.com>
Co-authored-by: Aditya Singh <60082699+adityasingh2400@users.noreply.github.com>
Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@nvidia.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>
Co-authored-by: Dingqing Yang <dingqingy@nvidia.com>
Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>
Co-authored-by: mathemakitten <helenn@nvidia.com>
Co-authored-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Yu Yao <54727607+yaoyu-33@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com>
Co-authored-by: kajalj22 <kajalj@nvidia.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: megnvidia <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>
Co-authored-by: Xuanteng Huang <44627253+xuantengh@users.noreply.github.com>
Co-authored-by: Tailai Ma <58548582+xiaoyao0115@users.noreply.github.com>
Co-authored-by: Santosh Bhavani <santosh.bhavani@live.com>
Co-authored-by: Zhongbo Zhu <42691305+zhongbozhu@users.noreply.github.com>
Co-authored-by: Nick Schank <nick@reflection.ai>
Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Shivanjan Chakravorty <schakravorty846@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Tuomas Rintamaki <trintamaki@nvidia.com>
Co-authored-by: Tyler Poon <tylerpoon@gmail.com>
Co-authored-by: Collin McCarthy <cmccarthy@nvidia.com>
Co-authored-by: Matthieu Le <matthieul@nvidia.com>
Co-authored-by: Piotr Zelasko <pzelasko@nvidia.com>
Co-authored-by: Ehsan Hosseini Asl <ehosseiniasl@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>
Co-authored-by: Xiaowei Ren <103958965+xrennvidia@users.noreply.github.com>
Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com>
Co-authored-by: William Dykas <wdykas@oci-hsg-cs-001-vscode-03.cm.cluster>
Co-authored-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Xuesong Ye <xuesongyey@gmail.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Daisy Gao <daisyg@nvidia.com>
Co-authored-by: lichenlu <lichenlu_8618@163.com>
Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Gao Deng <gdeng@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Pavel Gein <pavelgejn@yandex.ru>
Co-authored-by: Gao Deng <gdeng@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Eric Harper <eharper@nvidia.com>
Co-authored-by: Abhishree Thittenamane <47577437+athitten@users.noreply.github.com>
Co-authored-by: Abhishree Thittenamane <athittenaman@cw-dfw-cs-001-login-01.cm.cluster>
Co-authored-by: root <root@pool0-01849.cm.cluster>
Co-authored-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Co-authored-by: Nan Zheng <80790206+nanz-nv@users.noreply.github.com>
Co-authored-by: Li Tao <lit@nvidia.com>
Co-authored-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Robert Kirby <ArEsKay3@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Cole Hawkins <chawkins@nvidia.com>
Co-authored-by: Qi Zhang <qizhang@nvidia.com>
Co-authored-by: Vasudevan Rengasamy <vrengasamy@nvidia.com>
Co-authored-by: tongliu <tongliu@nvidia.com>
Co-authored-by: shurkat-nvidia <shurkat@nvidia.com>
Co-authored-by: Rui WANG <51357798+rui23@users.noreply.github.com>
Co-authored-by: rionawang <rionawang@tencent.com>
Co-authored-by: Tom Long <tolong@nvidia.com>
Co-authored-by: Lintch <44701395+returnL@users.noreply.github.com>
Co-authored-by: Yan Bai <bayan@nvidia.com>
Co-authored-by: Deyu Fu <deyuf@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Co-authored-by: jingqiny-99 <jingqiny@nvidia.com>
Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Pranav Thombre <pthombre@nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Shijie <505749828@qq.com>
Co-authored-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Li Jinliang <975761915@qq.com>
Co-authored-by: Baibaifan <39549453+Baibaifan@users.noreply.github.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: HaochenYuan <106647990+HaochenYuan@users.noreply.github.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>
Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org>
Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com>
Co-authored-by: Kunlun Li <94586211+kunlunl@users.noreply.github.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>
jon-barker pushed a commit to jon-barker/Megatron-LM that referenced this pull request Jul 10, 2026
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
shjwudp added a commit to shjwudp/Megatron-LM that referenced this pull request Jul 16, 2026
* test: re-enable test_pp2_create_cudagraphs_first_stage on TE 2.15+ (NVIDIA#4985)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tests): initialize num_microbatches calculator in vision cudagraph tests (NVIDIA#4986)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci: Add allow_failure flag to gpt and moe recipes that are failing in nightlies (NVIDIA#4905)

* Drain predecessor reduce-scatter at dispatch time (NVIDIA#4940)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* nightly(ci): Update golden values for functional t5 tests (NVIDIA#4995)

* chore: rotate oncall schedule

* [main] Refactor and Improve MoE Logginginit commit (NVIDIA#3431)

Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>

* ci: validate release branch-rules (NVIDIA#4929)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Megatron-FSDP] Add conditional param.grad dereferencing logic to support full-iteration (FWD-BWD) CUDA graphability. (NVIDIA#4663)

Signed-off-by: Cory Ye <cye@nvidia.com>

* test: restrict iter-time comparison to steady-state window (NVIDIA#5010)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Add mHC support for HybridModel on dev (NVIDIA#4949)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* fix(test): pin eval-global-batch-size on 15b gb200 release configs (NVIDIA#5022)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* [fix] Release MTP assertion when EP overlap with PP=1 (NVIDIA#4796)

* fix(test): widen iter-time steady-state window for short tests (NVIDIA#5023)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Perf fix (NVIDIA#4996)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add dev-feature preservation gate and change schedule (NVIDIA#4773)

* chore(test): remove orphan nemotron3_super_release_g200 dir (NVIDIA#5024)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Ignore Claude worktree directory (NVIDIA#5020)

* Update copy-pr-bot.yaml [skip ci]

* ci: update CI workflow conditions for integration tests (NVIDIA#4658)

* Add NVSkills CI request workflow (NVIDIA#5033)

* DDP wrap pg size fixes (NVIDIA#5006)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* fix(layer_wise): tag MTP-stage word_embeddings as is_embedding_or_output_parameter (NVIDIA#5034)

* [Dev] Add Qwen3 30B MoE recipes (NVIDIA#5012)

* [DEV] fix(megatron-fsdp): reduce padding for grouped expert weights (NVIDIA#5013)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4908)

Co-authored-by: Xin Yao <xiny@nvidia.com>

* [dev] [fix] [DeepSeek-v4] fix dense loss and rope type in DSv4 Hybrid Attention (NVIDIA#5018)

Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Move LTS dependencies from pyproject.toml to Dockerfile.ci.lts (NVIDIA#4877)

* Use shared ModelOpt calibration loop on 0.45+ with 0.44 fallback fix (NVIDIA#4881)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* test(release): skip golden comparison on intermediate resume windows (NVIDIA#5040)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [mimo] Thread position_ids through MimoModel for multimodal RoPE (NVIDIA#4938)

Signed-off-by: Li Ding <liding@nvidia.com>

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5039)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Fix: Import unwrap_model from megatron.core.utils in modelopt examples (NVIDIA#5045)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Simple and stable Inference APIs (NVIDIA#4697)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: Add notification step for MBridge downstream test results (NVIDIA#5028)

* [dev] [5/5] Qwen3.5 support: Qwen3.5-VL training example (NVIDIA#4751)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: Update Docker image version to 26.04-py3 on dev (NVIDIA#5051)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Delete output tensor early (NVIDIA#4742)

* Support ScaledSReLU in TE grouped MLP fuser (NVIDIA#4859)

Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: a <a>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>

* Skip gradient updates when grad norm exceeds threshold (NVIDIA#3460)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Add 9 user skills (NVIDIA#5066)

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>

* test(nemotron): align nemotron3 super GB200 goldens with exit-interval 4768 (NVIDIA#5069)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* chore: Update transformer-engine dependency to version 2.16.0 (NVIDIA#4992)

* Update energon version requirement (NVIDIA#4572)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Fix test failures for new inference APIs (NVIDIA#5068)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): set PYTHONUNBUFFERED=1 in JET workload env (NVIDIA#5072)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Preserve non-FSDP-unit buckets across AllGatherPipeline reset (NVIDIA#4717)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add opt-in MXFP8 LM-head output projection (NVIDIA#4825)

Co-authored-by: a <a>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-01)

* fix(ci): bound JET pipeline polling with a watchdog to prevent indefinite hangs (NVIDIA#5076)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: prune old artifacts on cluster lustre during weekly/release runs (NVIDIA#5084)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* ci(test): isolate ckpt-resume tensorboard per phase (NVIDIA#5074)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: unmark EP A2A activation offload test flaky (NVIDIA#5009)

* Change ownership groups (NVIDIA#5021)

* test: skip mfsdp_fully_shard cases when world_size < mesh size (NVIDIA#4487)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix mimo optimizer checkpoint metadata restore (NVIDIA#4791)

Signed-off-by: Li Ding <liding@nvidia.com>

* [mimo] Support bridge fan-out for variable modality tokens (NVIDIA#5062)

Signed-off-by: Li Ding <liding@nvidia.com>

* Add separate mtp_grad_scale_func for MTP loss scaling (NVIDIA#3459)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* [training migration] Migrate GPT builder (NVIDIA#4741)

Signed-off-by: Maanu Grover <maanug@nvidia.com>

* Make Mamba conv params direct mixer params (NVIDIA#4899)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Update copy-pr-bot.yaml [skip ci]

* Update oncall reviewer assignment (NVIDIA#5093)

* Pass explicit process groups to hybrid logging (NVIDIA#4781)

* Clean up top-level repository files (NVIDIA#5097)

* [main] fix(moe): Fix several bugs for DSA rope and spec. (NVIDIA#3026)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* [Dev] Skip identity alltoall chunk sort (NVIDIA#5102)

* Move MIMO unit tests into models/mimo (NVIDIA#5063)

* test: update DeepSeek FSDP2 GB200 memory golden (NVIDIA#5094)

* Remove DeepEP hardware limit check (NVIDIA#4846)

* Update transformer-engine dependency to revision 4220403 (NVIDIA#5112)

* ci: make CI resilient to pip/uv network timeouts (NVIDIA#5118)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: treat docker container-removal conflict as flaky (NVIDIA#5120)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix GDN DTensor splitting for FSDP checkpointing (NVIDIA#4843)

Signed-off-by: conver334 <conver334@gmail.com>

* Fix MoE aux_loss / z_loss gradient scaling with TP > 1 (NVIDIA#5047)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update copy-pr-bot.yaml [skip ci]

* Update Claude copy workflow to enforce user restrictions and improve error messages (NVIDIA#5117)

* Add advisory process group guidance to Claude reviews (NVIDIA#5111)

* build: cap pydantic<2.14 in transformer-engine dependency metadata (NVIDIA#5125)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* fix(test): skip scalar-less tensorboard event files in resume checks (NVIDIA#5121)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: rotate oncall schedule

* docs: fix contributor guide typo (NVIDIA#4858)

Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* ci(unit-tests): split slow unit-test buckets over 15min SLA (NVIDIA#5133)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix DSA indexer loss not averaged across micro-batches (NVIDIA#4070)

Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Update MINOR version to 19 (NVIDIA#5096)

* Fix Muon QKV split for gated attention (NVIDIA#4728)

Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Roll input IDs for MTP labels (NVIDIA#3457)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Refactor: Move paged stashing Triton kernels (NVIDIA#5003)

* Adding blackwell tests (NVIDIA#5113)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Relax atol for test_router_gating_linear router_dtype=torch.float32 (NVIDIA#4915)

Signed-off-by: Aditya Singh <adisin650@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* Fix incorrect inference metadata tensor dtypes (NVIDIA#4855)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: sraman-rgb <sraman@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>

* Disable TE cross entropy loss fusion (NVIDIA#5115)

Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>

* fix(optimizer): gate ChainedOptimizer MXFP8 defer-sync on DDP-level overlap_param_gather (NVIDIA#4982)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Pass TP group to unfused cross entropy (NVIDIA#5128)

* fix: correct dsv4_hybrid Q-up FLOPs by using args.v_head_dim (NVIDIA#5142)

* [dev] [follow-up] Qwen3.5 support: MoE aux loss padding_mask (NVIDIA#4776)

Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>

* [Dev] Support isolated MTP loss (NVIDIA#5080)

Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>

* test(elastification): quarantine flaky test_gumbel_determinism as flaky_in_dev (NVIDIA#5156)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(notify): mention mcore-oncall and philipp on critical CI events (NVIDIA#5152)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Change the cudagraph distribution from linearly to exponentially-decreasing + grid for mixed prefill (NVIDIA#3509)

* ci: Disable a few gb200 test cases to support 2 branches. (NVIDIA#5151)

* build: Switch DSv3 on H100 to HybridEP (NVIDIA#5164)

* Add MTP acceptance rate metrics (NVIDIA#3458)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Nemotron Ultra config for ModelOpt examples (NVIDIA#5159)

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>

* Make MTP / prefix cache stats persist for engine lifetime (NVIDIA#4101)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* Restore Greptile configuration (NVIDIA#5166)

* chore: bump `_code_freeze` workflow to `v1.4.2` (NVIDIA#5132)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [dev] moe(fix): Avoid TE cuda graph dummy attention masks (NVIDIA#5131)

Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>

* ci: Remove docs build test in favor of release test (NVIDIA#5182)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Move TE cross entropy guard to training args (NVIDIA#5162)

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>

* Fix error in deepseek parser (NVIDIA#5136)

* Fix logprob slicing for 0 generated token case (NVIDIA#5167)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* [Perf] Fold frozen linear dgrad matmul (NVIDIA#5092)

Signed-off-by: Chen Cui <chcui@nvidia.com>

* Clamp `max_new_tokens` in MInf to mirror vllm (NVIDIA#5181)

* build: add managed = true to [tool.uv] (NVIDIA#5190)

Signed-off-by: Kajal Jain <kajalj@nvidia.com>

* Stabilize GB200 inference perf tests against cold-start noise (NVIDIA#5171)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections (docs dup fields, nvrx version gate, utils re-export)

- Remove merge-introduced duplicate definitions of use_transformer_engine_op_fuser
  and moe_expert_rank_capacity_factor (transformer_config.py) and the duplicate
  _post_param_sync method (param_and_grad_buffer.py) — tripped the docs build.
- Relax is_nvrx_min_version() to compare release segments so dev's pinned
  nvidia-resiliency-ext pre-release (0.6.0.dev33) satisfies main's '>= 0.6.0'
  assertion (the required-symbol check remains the authoritative guard).
- Re-export unwrap_model from megatron/training/utils/__init__.py. main refactored
  the monolithic training/utils.py into a package whose __init__ dropped this
  re-export; dev's training.py (and others) import it via 'from .utils import',
  which broke tests/unit_tests/conftest.py's import chain (cascaded to all suites).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] [DeepSeek-v4] Add ClampedSwiGLU to MoE mlp_op_fuser and add force balance to hash routing (NVIDIA#5130)

* nvidia style guide audit for getting started folder (NVIDIA#5168)

Signed-off-by: meg miranda <mmiranda@nvidia.com>

* AI aided audit for Nvidia Style guidance (NVIDIA#5141)

Signed-off-by: meg miranda <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>

* Enable selective recompute for `norm_out` in GDN layers  (NVIDIA#4715)

* fix(elastification): align with get_batch + utils refactors (NVIDIA#5194)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(combined-1f1b): release loss-node input storage after combined backward (NVIDIA#4909)

* Minor improvements for Dynamic-cp (NVIDIA#4226)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>

* fix: post-CI corrections

* chore(beep boop 🤖): Bump  (main) (2026-06-08)

* Avoid stat syscall in rerun result validation (NVIDIA#5107)

Signed-off-by: dimapihtar <dpykhtar@nvidia.com>

* docs: Update Latest News in README.md (NVIDIA#3790)

* Fix bug with Megatron-FSDP zero counter not working with decoupled gradients. (NVIDIA#4802)

Signed-off-by: Cory Ye <cye@nvidia.com>

* varlendataset for thd e2e and benchmark (NVIDIA#4832)

Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* ci: add smoke tests (NVIDIA#5143)

* Add mtp_detach_heads config to detach MTP head inputs (NVIDIA#3456)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: wrap mtp logging comment

* docs: fix install guide NGC container anchor (NVIDIA#5224)

* [Dev] Add separate toggle for varlen input padding for HybridEP in THD training  (NVIDIA#5048)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* Fuse per-sequence AlltoAll into a unified one in GDN forward (NVIDIA#4913)

* Apply MIMO SP/CP sharding with explicit groups and enable THD in non-colocated path (NVIDIA#5150)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix CUDA IMA in fsdp_double_buffer when an FSDP unit's bucket doesn't fit the pool (NVIDIA#4810)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add named layouts to HyperCommGrid for heterogeneous parallelism (NVIDIA#5148)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix wgrad race condition when using double buffers. (NVIDIA#5222)

Signed-off-by: Cory Ye <cye@nvidia.com>

* Move uneven DTensor distributed fixture to conftest (NVIDIA#5237)

* Route bridge communicator cross-grid P2P through a dedicated process group (NVIDIA#5234)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix test_split_tensor_along_last_dim to actually assert correctness (NVIDIA#4710)

Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>

* chore: rotate oncall schedule

* Add optional group= to common_utils model/data-parallel reduction helpers (NVIDIA#5251)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement to Contribution guide (NVIDIA#5278)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Update copy-pr-bot.yaml [skip ci]

* Add MIMO hetero topology + distributed bootstrap (examples/mimo training-loop folder) (NVIDIA#5260)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)

Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>

* [Dev] DeepSeek-V4-Flash recipe 20260610 (NVIDIA#5266)

* [Dev] Generalized fix for mxfp8 param gather  (NVIDIA#4994)

Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Cherry-pick MTP detach heads (NVIDIA#5223)

* [TE] Restore original CP group after dynamic CP forward in TEDotProductAttention (NVIDIA#5215)

Co-authored-by: rionawang <rionawang@tencent.com>

* [examples] Add dynamic context parallel benchmark example (NVIDIA#5123)

* Remove duplicate nccl_allocator import (NVIDIA#5057)

* [dev] Add experimental Megatron Lite as agentic exploration (NVIDIA#4885)

Co-authored-by: Deyu Fu <deyuf@nvidia.com>

* fix(ci): resolve t5 dataloader stall + GRPO cudagraph-memory regression (CI-validated) (NVIDIA#5280)

Signed-off-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix Dockerfile warnings (NVIDIA#4856)

* Fix fused MLA delayed weight grad hooks (NVIDIA#5273)

Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>

* ci: limit retries on unsuccessful test launches (NVIDIA#5275)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Thread pg_collection into get_model DDP bucket sizing (NVIDIA#5250)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Enable non-deterministic results in model configuration for nemotron tests (NVIDIA#5239)

* Stabilize hybrid nanov3 gb200 perf (NVIDIA#5295)

* Clip mtp grads separately when mtp_detach_heads=True (NVIDIA#4116)

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Thread pg_collection into train_step reductions (NVIDIA#5259)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: Allow DCO check in merge queue and add DCO requirement (NVIDIA#5305)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>

* fix: restore dev-only training.py features dropped by the main override

The merge takes main's training.py wholesale (per the sync skill's override
list), which silently reverted four dev-only features whose supporting core
code the merge keeps at dev's version. Each surfaced as a CI failure:

1. Dynamic context-parallel API rename. Dev renamed
   get_hybrid_data_context_parallel_groups -> get_dynamic_data_context_parallel_groups
   (identical signature) and args.hybrid_context_parallel -> dynamic_context_parallel.
   Point training.py at dev's names. Fixes the conftest.py ImportError that
   cascaded to every unit-test bucket.

2. Dynamic-CP / sequence-packing data loading. Replace main's setup-time
   HybridCPDataLoaderWrapper wrap with dev's per-step wrap_data_iterator
   (gated on config.sequence_packing_scheduler), which returns a
   RerunDataIterator-compatible iterator. Fixes the *_cp4_dcp RerunDataIterator
   assertion.

3. MTP loss logging scale. Dev's MTPLossLoggingHelper stores raw loss sums and
   token counts and computes the per-token loss after reduction, so the log
   scale must be 1.0; main's 1/get_num_microbatches() divided the reported
   mtp_N loss by num_microbatches. Fixes moe gpt3_..._scoped_cudagraph mtp_1
   loss (was 16x too small: 0.682 vs golden 10.915).

4. DSA indexer loss cross-PP reduction. dsa.py requires num_layers (and
   csa_compress_ratios) so first-pipeline-stage ranks lazily initialize the
   tracker and join the cross-PP all_reduce in reduce_loss_in_tracker; main's
   call omitted them, so stage 0 returned early and the last stage hung on an
   unmatched all_reduce. Fixes the gpt3_..._dsv4_hybrid_mhc_mtp NCCL timeout.

Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>

* [dev]: faster implementation of mHC fused kernels (NVIDIA#4624)

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>

* Allow for pre-bound socket to be passed in server (NVIDIA#5301)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Offline Logits-Based Knowledge Distillation (NVIDIA#5019)

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

* Handle None values in sampling parameters (NVIDIA#5300)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Add moe loss normalization for RL SFT (NVIDIA#3956)

Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Add code owners for optimizer-related files (NVIDIA#5297)

Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>

* Fix EP=1 inference by allocating buffers anyway (NVIDIA#5233)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Fix crash due to tool call at sequence length (NVIDIA#5302)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* Inference: Cudagraph-aware admission gating in prefill scheduler (NVIDIA#4870)

Signed-off-by: Helen Ngo <helenn@nvidia.com>

* Account for reasoning token stripping (NVIDIA#5313)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Thread pg_collection through wrap_model_chunks_with_ddp (NVIDIA#5328)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Dev] Add MoE recipe performance summary (NVIDIA#5289)

Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>

* [Dev] Add DeepEP v2 flex dispatcher backend (NVIDIA#4793)

Signed-off-by: tongliu <tongliu@nvidia.com>
Co-authored-by: Dennis(Zhenhuan) Liu <denliu@nvidia.com>

* fix tflops calculation when sequence_packing_scheduler is not none (NVIDIA#5342)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-15)

* [dev] bump emerging optimizers to v0.3.0 (NVIDIA#5320)

Signed-off-by: Deyu Fu <deyuf@nvidia.com>

* Fix LatentMoE theoretical memory estimate (NVIDIA#5145)

Signed-off-by: Shijie Wang <jaywan@nvidia.com>

* Add zstandard package to Docker LTS requirements. Fix nightly failures (NVIDIA#5347)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* fix: preserve seqlen stats in train_step

Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>

* Thread MIMO support through the stock training loop (schedule + optimizer) (NVIDIA#5333)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (1/N) (NVIDIA#5042)

Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(ssm): whole-module 'gdn' selective recompute for GatedDeltaNet (NVIDIA#5296)

Signed-off-by: jinliangl <975761915@qq.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Fix] Fix optimizer parameter override bugs. (NVIDIA#5213)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>

* [Dev] Add Megatron-FSDP weight prefetch for full recompute (NVIDIA#5175)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* ci: default functional test time limit to 4h for release/weekly scopes (NVIDIA#5360)

Signed-off-by: oliver könig <okoenig@nvidia.com>

* [Dev] add cuda graph support for thd format training. (NVIDIA#4359)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>

* Fix memory leak with log_max_attention_logit (NVIDIA#4699) (NVIDIA#5067)

Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>

* Clean up pretrain_gpt.py and pretrain_hybrid.py formatting and remove module globals (NVIDIA#5351)

Signed-off-by: ilml <tolong@nvidia.com>

* Add full model cuda graph support for MTP inference (NVIDIA#4950)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Expand the Mamba prefix caching memory safety check to include scratch space buffers (NVIDIA#5348)

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>

* Make Megatron RL only materialize last token logit (NVIDIA#4551)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Profiling  (NVIDIA#3110)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>

* Support fused MLA QKV checkpoint reload (NVIDIA#5310)

Signed-off-by: sraman <sraman@nvidia.com>

* Add minimal DBuffer implementation (NVIDIA#4835)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [Dev] restore DSv4 tflops calc in training and fix the packed seq case (NVIDIA#5358)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* [split 1/5] Fix packed THD RoPE under CP (NVIDIA#5243)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* Update copy-pr-bot.yaml [skip ci]

* Document agent PR commit sign-off and signing (NVIDIA#5381)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* Remove unused distributed pytest markers (NVIDIA#5380)

Signed-off-by: Jingyue Wu <wujingyue@gmail.com>

* [feat] Support fine-grained activation offloading in fused group mlp (NVIDIA#5082)

Signed-off-by: hongbinl <hongbinl@nvidia.com>

* Thread tensor-parallel group into the RADIO patch embedder (NVIDIA#5371)

Signed-off-by: ykarnati <ykarnati@nvidia.com>

* Add MimoModel.zero_grad_buffer delegating to active DDP submodules (NVIDIA#5372)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [split 3/5] Refactor absorbed MLA projection handling (NVIDIA#5245)

Signed-off-by: Hollow Man <hollowman@opensuse.org>

* [dev] moe(perf): Restore fused GDN THD all-to-all on dev (NVIDIA#5389)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* chore: rotate oncall schedule

* ci: Remove sync skills workflow (NVIDIA#5091)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>

* Add flaky marker to fine-grained activation offloading test (NVIDIA#5350) (NVIDIA#5368)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Revert "Remove checkpoint-time GPU cache reclaim workaround (NVIDIA#5170)" (NVIDIA#5366)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Update goldens for weekly tests after pytorch and TE bumps. (NVIDIA#5399)

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>

* Add MIMO runtime setup: per-role RNG seeding and DDP wrapping (NVIDIA#5285)

Signed-off-by: ykarnati <ykarnati@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Add --mamba-training-ssm-states-dtype argument (NVIDIA#5309)

Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>

* chore(beep boop 🤖): Bump  (main) (2026-06-22)

* Fix Mamba prefix match for chunked prefill (NVIDIA#4758)

Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>

* Disag MR2: Refit into multiple destination pools and tied-embedding + UVM fixes (NVIDIA#5187)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Disag MR1: Add inference shard specs and pg-collection building (NVIDIA#5186)

Signed-off-by: wdykas <wdykas@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* [Fix] Fix MoE router z-loss compatibility with TE CUDA Graph capture. (NVIDIA#5401)

Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>

* test: mark ep_a2a_overlap activation-offloading test flaky_in_dev (NVIDIA#5450)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark TestParallelTransformerBlockCudagraphs::test_gpu_cudagraph flaky_in_dev (NVIDIA#5475)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: mark gated_delta_net selective-recompute test flaky_in_dev (NVIDIA#5476)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: post-CI corrections for sync test/impl reconciliation

1) test_optimizer.py: reverted to dev. The sync kept dev's
   multi_latent_attention.py (split q/kv down-proj, no
   _synthesize_fused_qkv_down_weight), but auto-merged main's
   test asserting the fused linear_qkv_down_proj.weight key.

2) training.py: guard the dev-only sequence_packing_scheduler config
   access with getattr (lines in train_step and train()). main's new
   MIMO schedule-plumbing test (NVIDIA#5333) passes an empty SimpleNamespace
   config; the reconciled training.py keeps dev's packing path, so the
   access must tolerate a config lacking the attribute. Real configs are
   unaffected (getattr returns the same value).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>

* fix sequence packing wrapper for eval (NVIDIA#5483)

Signed-off-by: xiaoyao0115 <1804647152@qq.com>

* [dev] Megatron Lite (4/4) shared attention (NVIDIA#5427)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* fix: restore fused group MLP offload in main2dev sync (NVIDIA#5493)

Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>

* [dev] [DeepSeek-v4] Packed Sequence (THD) support for DSv4 Hybrid Attention (NVIDIA#5011)

Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>

* Improve default dynamic CP packing scheduler (NVIDIA#5154)

Signed-off-by: tailaim <tailaim@nvidia.com>

* [dev] moe(perf): Pre-GDR kernel fusion (NVIDIA#5361)

Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>

* Preserve DSA output across fused inverse RoPE (NVIDIA#5526)

Signed-off-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

* Enable Deepseek-v4 hybrid_model in dev branch Part (2/N) (NVIDIA#5485)

Signed-off-by: guihong-nv <guihongl@nvidia.com>

* [dev] Add experimental decoupled compact LayerWise DDP layout for Muon (NVIDIA#5388)

Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [dev] Sync Megatron Lite with the latest implementation (NVIDIA#5577)

Signed-off-by: Yan Bai <bayan@nvidia.com>

* [Dev] fix padding mask docstring (NVIDIA#5598)

Signed-off-by: HaochenYuan <haocheny@nvidia.com>

* handle split_dtensor

---------

Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Cory Ye <cye@nvidia.com>
Signed-off-by: Maanu Grover <maanug@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Li Ding <liding@nvidia.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Jingyue Wu <wujingyue@gmail.com>
Signed-off-by: conver334 <conver334@gmail.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: LeSingh1 <sshaurya914@gmail.com>
Signed-off-by: Aditya Singh <adisin650@gmail.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Signed-off-by: qiyuw <qiyuw@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Signed-off-by: meg miranda <mmiranda@nvidia.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com>
Signed-off-by: Zhongbo Zhu <zhongboz@nvidia.com>
Signed-off-by: Shivanjan Chakravorty <shivanjanc@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Xiaowei Ren <xren@nvidia.com>
Signed-off-by: yuzhongw <yuzhongw@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: ksivamani <ksivamani@nvidia.com>
Signed-off-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: Pavel Gein <pavel.gein@gmail.com>
Signed-off-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Signed-off-by: jinliangl <jinliangl@nvidia.com>
Signed-off-by: Skand Hurkat <shurkat@nvidia.com>
Signed-off-by: Yan Xu <yxu1@nvidia.com>
Signed-off-by: Siddhartha Raman Sundara Raman <270218152+sraman-rgb@users.noreply.github.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: ykarnati <ykarnati@nvidia.com>
Signed-off-by: Dennis Liu <denliu@denliu.nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com>
Signed-off-by: Shijie Wang <jaywan@nvidia.com>
Signed-off-by: guihong-nv <guihongl@nvidia.com>
Signed-off-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Signed-off-by: Guihong Li <guihongl@nvidia.com>
Signed-off-by: jinliangl <975761915@qq.com>
Signed-off-by: yangfan.bai <yangfan.bai@shopee.com>
Signed-off-by: hongbinl <hongbinl@nvidia.com>
Signed-off-by: HaochenYuan <haocheny@nvidia.com>
Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Signed-off-by: ilml <tolong@nvidia.com>
Signed-off-by: sraman <sraman@nvidia.com>
Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>
Signed-off-by: Hollow Man <hollowman@opensuse.org>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com>
Signed-off-by: pingtianl <pingtianl@nvidia.com>
Signed-off-by: Pingtian Li <pingtianl@nvidia.com>
Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay <abalasa@nvidia.com>
Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zijie Yan <zijiey@nvidia.com>
Co-authored-by: Robin Zhang <robinz@nvidia.com>
Co-authored-by: Dennis Liu <denliu@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Shifang Xu <shifangx@nvidia.com>
Co-authored-by: Cory Ye <44509866+cspades@users.noreply.github.com>
Co-authored-by: Yan Xu <45385219+Connor-XY@users.noreply.github.com>
Co-authored-by: Pingtian Li <158665726+Wohox@users.noreply.github.com>
Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com>
Co-authored-by: Maanu Grover <maanug@nvidia.com>
Co-authored-by: xuwchen <xuwenc@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: hx <hongxiaob@nvidia.com>
Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Li Ding <liding@nvidia.com>
Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com>
Co-authored-by: Li Jinliang <jinliangl@nvidia.com>
Co-authored-by: BestJuly <19769279+BestJuly@users.noreply.github.com>
Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com>
Co-authored-by: sraman <sraman@users.noreply.github.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: sraman-rgb <270218152+sraman-rgb@users.noreply.github.com>
Co-authored-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Gerald Shen <geshen@nvidia.com>
Co-authored-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Co-authored-by: Jingyue Wu <wujingyue@gmail.com>
Co-authored-by: Gao Deng <160076886+gdengk@users.noreply.github.com>
Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com>
Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com>
Co-authored-by: Yuzhong Wang <yuzhongw@computelab-frontend-3.nvidia.com>
Co-authored-by: janEbert <janpabloe@nvidia.com>
Co-authored-by: Simiao Zhang <56124251+conver334@users.noreply.github.com>
Co-authored-by: Chen Cui <chcui@nvidia.com>
Co-authored-by: Shaurya Singh <sshaurya914@gmail.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
Co-authored-by: Kevin <23258141+kaimo455@users.noreply.github.com>
Co-authored-by: mokai <mokai@baidu.com>
Co-authored-by: Moozy <108285604+Moozy23232@users.noreply.github.com>
Co-authored-by: Aditya Singh <60082699+adityasingh2400@users.noreply.github.com>
Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddhartha Raman S <sraman@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: gautham-kollu <gkollu@nvidia.com>
Co-authored-by: Qiyu Wan <39144338+WanZzzzzz@users.noreply.github.com>
Co-authored-by: Devil1716 <149754374+Devil1716@users.noreply.github.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@nvidia.com>
Co-authored-by: Mike Chrzanowski <mchrzanowski@gcp-nrt-cs-001-login-001.cm.cluster>
Co-authored-by: Dingqing Yang <dingqingy@nvidia.com>
Co-authored-by: liuzhenhai93 <liuzhenhai93@outlook.com>
Co-authored-by: mathemakitten <helenn@nvidia.com>
Co-authored-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Yu Yao <54727607+yaoyu-33@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com>
Co-authored-by: kajalj22 <kajalj@nvidia.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: megnvidia <mmiranda@nvidia.com>
Co-authored-by: Philip Petrakian <pgpetrak@gmail.com>
Co-authored-by: Xuanteng Huang <44627253+xuantengh@users.noreply.github.com>
Co-authored-by: Tailai Ma <58548582+xiaoyao0115@users.noreply.github.com>
Co-authored-by: Santosh Bhavani <santosh.bhavani@live.com>
Co-authored-by: Zhongbo Zhu <42691305+zhongbozhu@users.noreply.github.com>
Co-authored-by: Nick Schank <nick@reflection.ai>
Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com>
Co-authored-by: Shivanjan Chakravorty <schakravorty846@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Tuomas Rintamaki <trintamaki@nvidia.com>
Co-authored-by: Tyler Poon <tylerpoon@gmail.com>
Co-authored-by: Collin McCarthy <cmccarthy@nvidia.com>
Co-authored-by: Matthieu Le <matthieul@nvidia.com>
Co-authored-by: Piotr Zelasko <pzelasko@nvidia.com>
Co-authored-by: Ehsan Hosseini Asl <ehosseiniasl@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com>
Co-authored-by: Xiaowei Ren <103958965+xrennvidia@users.noreply.github.com>
Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com>
Co-authored-by: William Dykas <wdykas@oci-hsg-cs-001-vscode-03.cm.cluster>
Co-authored-by: kunlunl <kunlunl@nvidia.com>
Co-authored-by: Xuesong Ye <xuesongyey@gmail.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Daisy Gao <daisyg@nvidia.com>
Co-authored-by: lichenlu <lichenlu_8618@163.com>
Co-authored-by: peibli <lipeibao@126.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Gao Deng <gdeng@login-lyris02.lyris.clusters.nvidia.com>
Co-authored-by: Pavel Gein <pavelgejn@yandex.ru>
Co-authored-by: Gao Deng <gdeng@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Eric Harper <eharper@nvidia.com>
Co-authored-by: Abhishree Thittenamane <47577437+athitten@users.noreply.github.com>
Co-authored-by: Abhishree Thittenamane <athittenaman@cw-dfw-cs-001-login-01.cm.cluster>
Co-authored-by: root <root@pool0-01849.cm.cluster>
Co-authored-by: Kamran Jafari <kjafarisadeg@nvidia.com>
Co-authored-by: Nan Zheng <80790206+nanz-nv@users.noreply.github.com>
Co-authored-by: Li Tao <lit@nvidia.com>
Co-authored-by: Zhongbo Zhu <zhongboz@nvidia.com>
Co-authored-by: Robert Kirby <ArEsKay3@users.noreply.github.com>
Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Co-authored-by: Cole Hawkins <chawkins@nvidia.com>
Co-authored-by: Qi Zhang <qizhang@nvidia.com>
Co-authored-by: Vasudevan Rengasamy <vrengasamy@nvidia.com>
Co-authored-by: tongliu <tongliu@nvidia.com>
Co-authored-by: shurkat-nvidia <shurkat@nvidia.com>
Co-authored-by: Rui WANG <51357798+rui23@users.noreply.github.com>
Co-authored-by: rionawang <rionawang@tencent.com>
Co-authored-by: Tom Long <tolong@nvidia.com>
Co-authored-by: Lintch <44701395+returnL@users.noreply.github.com>
Co-authored-by: Yan Bai <bayan@nvidia.com>
Co-authored-by: Deyu Fu <deyuf@nvidia.com>
Co-authored-by: Anish Mahishi <amahishi@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@users.noreply.github.com>
Co-authored-by: jingqiny-99 <jingqiny@nvidia.com>
Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Pranav Thombre <pthombre@nvidia.com>
Co-authored-by: Dennis Liu <denliu@denliu.nvidia.com>
Co-authored-by: Shijie <505749828@qq.com>
Co-authored-by: Guihong Li <guihongl@nvidia.com>
Co-authored-by: Guihong Li <guihongl@oci-hsg-cs-001-vscode-02.cm.cluster>
Co-authored-by: Yan Xu <yxu1@nvidia.com>
Co-authored-by: Li Jinliang <975761915@qq.com>
Co-authored-by: Baibaifan <39549453+Baibaifan@users.noreply.github.com>
Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
Co-authored-by: HaochenYuan <106647990+HaochenYuan@users.noreply.github.com>
Co-authored-by: Haochen Yuan <haocheny@login-eos01.eos.clusters.nvidia.com>
Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org>
Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com>
Co-authored-by: Kunlun Li <94586211+kunlunl@users.noreply.github.com>
Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>

Signed-off-by: Jianbin Chang <jianbinc@nvidia.com>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
svcnvidia-nemo-ci pushed a commit to dimapihtar/Megatron-LM that referenced this pull request Aug 4, 2026
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants