Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
73 commits
Select commit Hold shift + click to select a range
e9e513a
feat(ci): add strict review mode to Claude review workflow (#4197)
Apr 14, 2026
cd03c3e
Fix stale approvals (#4280)
Phlip79 Apr 14, 2026
eb80b74
[MoE] Add a new score function to the router (#3673)
yaox12 Apr 14, 2026
ebfa138
[MoE] Improvement of shared expert overlap, support shared expert ove…
Apr 14, 2026
123645b
build: bump DeepEP to 34152ae (#4228)
ko3n1g Apr 14, 2026
3d7a701
ci: mark test_fused_indexer_loss_gradient_tp_consistency as flaky_in_…
ko3n1g Apr 14, 2026
4cef23c
Fix typo in PR4133. (#4277)
cspades Apr 14, 2026
bcec618
ci: add retry loop to apt-get update to handle transient mirror sync …
ko3n1g Apr 14, 2026
4e85d74
fix: enforce correct pass thresholds for deterministic and approximat…
ko3n1g Apr 14, 2026
d245c44
remove legacy biencoder and realm models (#4205)
dimapihtar Apr 14, 2026
6636eb0
ci: add configurable launcher support for functional tests (ft_launch…
ko3n1g Apr 14, 2026
d530a04
chore: document --target main for local Docker builds (#4307)
ko3n1g Apr 14, 2026
97aca2f
Extract args init to launch scripts (#4225)
maanug-nv Apr 14, 2026
c2d1a8f
[Main] Fix TE version check for retain_pinned_cpu_buffers in cpu offl…
BestJuly Apr 15, 2026
b342602
chore: rotate oncall schedule
github-actions[bot] Apr 15, 2026
1d344ae
Fix documented shape (#3486)
janEbert Apr 15, 2026
4a79536
ci: add sync-skills workflow, rename CLAUDE.md → AGENTS.md, move .cla…
ko3n1g Apr 15, 2026
57fc3ae
chore(beep boop 🤖): symlink skills/ → .claude/skills, .agents/skills …
github-actions[bot] Apr 16, 2026
23265d2
Get `device` correctly when module returns a dict instead of individu…
shifangx Apr 16, 2026
e69cf43
remove vision legacy code (#4202)
dimapihtar Apr 16, 2026
f098fe8
feat: long convergence resiliency for release tests (#4335)
ko3n1g Apr 16, 2026
8681ebb
ci(action): improve GitHub Actions output UX (#4337)
ko3n1g Apr 16, 2026
ceac269
build: bump TransformerEngine to release_v2.14 (#4331)
ko3n1g Apr 16, 2026
efbe7a1
feat: add create-issue skill (#4338)
ko3n1g Apr 16, 2026
97f9ab6
Set megatron-fsdp to 0.5.0
ko3n1g Apr 16, 2026
01eb7e8
M4 leftover for TE cuda graph (#3137)
shifangx Apr 16, 2026
260cba7
fix: wait for async P2P send before deallocating output tensor (#4047)
ZhiyuLi-Nvidia Apr 16, 2026
2aeaf56
ci(gb200): add 1-node mr-github functional test variants (#4334)
ko3n1g Apr 17, 2026
ded22f4
Fix potential coredump issue that occurs when saving a checkpoint (#1…
ezioliao Apr 17, 2026
30bc230
docs: bump versions1.json to 0.17.0 (latest) (#4360)
ko3n1g Apr 17, 2026
a00e944
Port DeepSeek Sparse Attention to `MambaModel` (#3553)
janEbert Apr 17, 2026
23663a8
Add tables and histogram for RL staleness (#4097)
tdene Apr 17, 2026
4ece77d
[docs] ci: use parent-relative json_url for version picker (#4367)
ko3n1g Apr 17, 2026
ed5de26
Fix bug with non-partial rollouts (#3964)
tdene Apr 17, 2026
e15ec3c
Add QK layernorm support for dot-product attention in MambaModel (#4067)
Phlip79 Apr 17, 2026
75a2878
Docs: improve docstrings and comments in example training loop (#4041)
DhineshPonnarasan Apr 17, 2026
86b7218
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Apr 18, 2026
9978968
feat(ckpt): add --async-ckpt-use-cpu-shm argument (#4355)
sbak5 Apr 18, 2026
664baa8
cp: Fix UT timeout (#4310) (#4373)
chtruong814 Apr 18, 2026
e4d3a4c
Fix RL reward due to stop token (#4096)
tdene Apr 18, 2026
76ac7c2
FA4 Inference (#4186)
wdykas Apr 18, 2026
3315c86
Make param_index_map always use unpacked (full numel) offsets (#4328)
deepakn94 Apr 18, 2026
8be1e79
Add activation logging and tokens per expert logging (#3842)
Mellonta Apr 18, 2026
98a51eb
Fix RL to once again work with --skip-train (#4249)
tdene Apr 18, 2026
afae25b
Fix Megatron initialization with extra_args_provider (#4327)
santhnm2 Apr 18, 2026
15e07a2
Rename MambaModel/MambaStack to HybridModel/HybridStack (#4099)
Phlip79 Apr 19, 2026
3046182
chore(beep boop 🤖): Bump (main) (2026-04-20)
github-actions[bot] Apr 20, 2026
9c210f7
fix(ci): wrap uv install in retry block (#4387)
ko3n1g Apr 20, 2026
7928a84
Call save_checkpoint_and_time() when saving checkpoint and compute el…
awsankur Apr 20, 2026
ef1888b
refactor(tests): move NCCL env vars from docker launcher to shell tra…
ko3n1g Apr 20, 2026
b562151
Remove packed_attention_mask unused parameter (#3859)
tdene Apr 20, 2026
859b66a
Second batch of audit edits (#4115)
megnvidia Apr 20, 2026
c9e03d0
Replace rampup batch size scheduler with custom step batch size sched…
mkhona-nvidia Apr 20, 2026
dc87858
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Apr 21, 2026
a52112d
revert: replace rampup batch size scheduler with custom step batch si…
ko3n1g Apr 21, 2026
0b9bc20
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Apr 21, 2026
532ad92
Replace rampup batch size scheduler with custom step batch size sched…
deepakn94 Apr 21, 2026
e5ec9ab
Megatron-FSDP: log mcore detection only after imports succeed (#4400)
wujingyue Apr 21, 2026
a550e0e
ci(gb200): re-enable tunable_overlap 1-node mr-github test (#4405)
ko3n1g Apr 21, 2026
77e2dd4
Fix local docs building (#4416)
Phlip79 Apr 21, 2026
e778967
RL: Onload optimizer after logprobs computation (#4235)
tdene Apr 22, 2026
bbc6b4d
chore: rotate oncall schedule
github-actions[bot] Apr 22, 2026
7597a0d
Add RL token throughput and packing metrics (#3877)
tdene Apr 22, 2026
9834c99
ci: remove publish:merge_into_dev job (#4421)
ko3n1g Apr 22, 2026
a6bfe1a
docs: add data loading best practices for large-scale training (#4236)
sbhavani Apr 22, 2026
384e618
Fix: Auto enable manual registration and enhance the docummentation (…
youngeunkwon0405 Apr 22, 2026
57005c8
Merge remote-tracking branch 'origin/main' into main2dev/22_04_2026
github-actions[bot] Apr 22, 2026
8add4e4
chore: post-merge fixes for nightly sync main into dev (22_04_2026)
github-actions[bot] Apr 22, 2026
baa3df4
fix: revert nvidia-resiliency-ext revision to match uv.lock
github-actions[bot] Apr 22, 2026
bbb06e2
fix: reformat 4 files with correct black==24.4.2 and isort==5.13.2
github-actions[bot] Apr 22, 2026
89798f3
fix: restore missing ArgumentGroupFactory import in arguments.py
svcnvidia-nemo-ci Apr 22, 2026
8cf5458
chore: keep CODEOWNERS unchanged in main→dev sync
Phlip79 Apr 23, 2026
fba3a80
chore: update gpt3_mcore_te_tp2_pp2_mhc golden values for main→dev sync
Phlip79 Apr 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 1 addition & 0 deletions .agents/skills
1 change: 1 addition & 0 deletions .claude/skills
8 changes: 6 additions & 2 deletions .github/actions/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,11 @@ runs:
echo "apt attempt $i failed, retrying..."
sleep 10
done
curl -LsSf https://astral.sh/uv/install.sh | UV_INSTALL_DIR=/usr/local/bin sh
for i in 1 2 3; do
curl -LsSf https://astral.sh/uv/install.sh | UV_INSTALL_DIR=/usr/local/bin sh && break
echo "uv install attempt $i failed, retrying..."
sleep 10
done

- name: Create run-script (unit test)
shell: bash -x -e -u -o pipefail {0}
Expand Down Expand Up @@ -210,7 +214,7 @@ runs:
LOG_BASE=$([[ "$IS_UNIT_TEST" == "true" ]] && echo "assets_dir/logs" || echo "assets_dir")
LATEST_LOG=""
if [[ -d "$LOG_BASE" ]]; then
LATEST_LOG=$(find "$LOG_BASE" -maxdepth 3 -name "*.log" ! -name "nccl_debug.log" -type f 2>/dev/null \
LATEST_LOG=$(find "$LOG_BASE" -name "*.log" ! -name "nccl_debug.log" -type f 2>/dev/null \
| xargs -r ls -t 2>/dev/null | head -1 || true)
fi
if [[ -n "$LATEST_LOG" ]]; then
Expand Down
2 changes: 1 addition & 1 deletion .github/copy-pr-bot.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
enabled: true
auto_sync_draft: false
auto_sync_ready: true
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "WanZzzzzz", "Wohox", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "asolergi-nv", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "frsun-nvda", "gautham-kollu", "gdengk", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kanz-nv", "keshavb96", "kevalmorabia97", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "parthmannan", "prajwal1210", "pthombre", "rhewett-nv", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "wdykas", "wplf", "wujingyue", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Mellonta", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "WanZzzzzz", "Wohox", "YangFei1990", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "aroshanghias-nvd", "asolergi-nv", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "frsun-nvda", "gautham-kollu", "gdengk", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kanz-nv", "keshavb96", "kevalmorabia97", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "parthmannan", "prajwal1210", "pthombre", "rhewett-nv", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "wdykas", "wplf", "wujingyue", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
12 changes: 4 additions & 8 deletions .github/oncall_schedule.json
Original file line number Diff line number Diff line change
@@ -1,12 +1,4 @@
[
{
"user": "ilml",
"date": "2026-04-08"
},
{
"user": "Phlip79",
"date": "2026-04-15"
},
{
"user": "asolergi-nv",
"date": "2026-04-22"
Expand Down Expand Up @@ -50,5 +42,9 @@
{
"user": "maanug-nv",
"date": "2026-07-01"
},
{
"user": "wujingyue",
"date": "2026-07-08"
}
]
182 changes: 179 additions & 3 deletions .github/workflows/claude_review.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,12 @@ on:
types: [created]

jobs:
review-on-comment:
name: Claude Review (comment trigger)
# ──────────────────────────────────────────────────────────────────
# Light review: quick pass for obvious bugs, typos, and test gaps
# Trigger: /claude review
# ──────────────────────────────────────────────────────────────────
light-review:
name: Claude Light Review
if: |
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
Expand Down Expand Up @@ -39,7 +43,7 @@ jobs:
--method POST \
-f content='eyes'

- name: Run Claude Code Review
- name: Run Claude Light Review
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
Expand Down Expand Up @@ -77,3 +81,175 @@ jobs:

It's perfectly acceptable to not have anything to comment on.
If you do not have anything to comment on, approve the PR with: gh pr review $PR_NUMBER --repo $REPO --approve --body "LGTM"

# ──────────────────────────────────────────────────────────────────
# Strict review: comprehensive Megatron-LM focused analysis
# covering precision, parallelism correctness, performance,
# backward compatibility, and code quality
# Trigger: /claude strict-review
# ──────────────────────────────────────────────────────────────────
strict-review:
name: Claude Strict Review
if: |
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
contains(github.event.comment.body, '/claude strict-review')
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
issues: write
id-token: write
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
PR_NUMBER: ${{ github.event.issue.number }}
steps:
- name: Get PR info
id: pr-info
run: |
PR_DATA=$(gh pr view $PR_NUMBER --repo $REPO --json headRefOid,baseRefName)
echo "sha=$(echo $PR_DATA | jq -r .headRefOid)" >> $GITHUB_OUTPUT
echo "base_ref=$(echo $PR_DATA | jq -r .baseRefName)" >> $GITHUB_OUTPUT

- name: Checkout repository
uses: actions/checkout@v6
with:
fetch-depth: 1
ref: ${{ steps.pr-info.outputs.sha }}

- name: Fetch base branch for diff analysis
run: git fetch origin ${{ steps.pr-info.outputs.base_ref }}

- name: React to trigger comment
run: |
gh api repos/$REPO/issues/comments/${{ github.event.comment.id }}/reactions \
--method POST \
-f content='eyes'

- name: Run Claude Strict Review
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
trigger_phrase: "/claude strict-review"
show_full_output: true
claude_args: |
--allowedTools "mcp__github_inline_comment__create_inline_comment,Bash(gh pr comment:*),Bash(gh pr diff:*),Bash(gh pr view:*),Bash(gh pr review:*),Bash(git diff:*),Bash(git show:*),Bash(git log:*)"
--model "claude-opus-4-6"
prompt: |
REPO: ${{ env.REPO }}
PR NUMBER: ${{ env.PR_NUMBER }}
BASE REF: origin/${{ steps.pr-info.outputs.base_ref }}

You are performing a strict, comprehensive code review on a **Megatron-LM** Pull Request.
Megatron-LM is NVIDIA's large-scale distributed training framework for LLMs.
Review the diff with a focus on **implementation correctness**, **training performance**, and **backward compatibility**.

## Review Procedure

1. Get PR metadata: `gh pr view $PR_NUMBER --repo $REPO --json title,body,baseRefName,headRefName,files,additions,deletions,changedFiles,author`
2. Get the full diff: `gh pr diff $PR_NUMBER --repo $REPO`
- For large PRs (>50 files), prioritize source code over config/lock/auto-generated files.
3. For each significant changed file, read the full file for surrounding context.
4. Trace data flow and dtype through computation paths to verify correctness.
5. For each newly introduced variable/argument/field, verify it has a meaningful runtime use path (see Mandatory Check below).
6. Post findings as inline comments with severity and category tags.

## Critical Issues (Must Fix)

### Implementation Correctness
- **dtype handling**: Verify operations use the correct dtype at each computation stage — explicit casts must be present at mixed-precision boundaries (e.g. fp16 compute → fp32 accumulation → fp16 output)
- **Loss scaling logic**: Verify DynamicLossScaler changes correctly detect inf/nan, adjust scale factor, and skip optimizer steps — incorrect logic causes training divergence or silent underflow
- **Reduction operations**: Verify reductions (sum, mean, allreduce) use correct dtype, reduction dimension, and normalization factor — wrong dimension or missing fp32 upcast produces silently wrong gradients
- **Normalization layers**: Verify LayerNorm/RMSNorm compute variance and mean on the correct dimension, with correct epsilon placement and upcast before rsqrt
- **Attention computation**: Verify QK^T scaling factor, softmax input dtype, causal mask application, and dropout placement match the intended algorithm
- **Residual connections**: Verify the correct tensor is added (pre-norm vs post-norm) with appropriate dtype for accumulation
- **Optimizer updates**: Verify state updates follow the correct formula — momentum/variance update order, bias correction, weight decay application
- **Gradient clipping**: Verify norm computation uses correct parameter set, norm type (L2 vs inf), and fp32 dtype
- **Embedding/output layer**: Verify weight tying is correctly wired, logit projection uses the right matrix, and output dtype matches expectation
- **MoE routing/aux loss**: Verify expert routing logic (top-k selection, capacity enforcement, token dropping) and auxiliary loss computation follow the intended algorithm

### Correctness
- **Tensor parallel**: Incorrect scatter/gather or allreduce placement — silent wrong results across TP ranks
- **Pipeline parallel**: Wrong microbatch scheduling, missing send/recv synchronization, incorrect grad accumulation across pipeline stages
- **Sequence parallel**: Incorrect sequence dimension partitioning or missing allgather/reduce-scatter in SP regions
- **Context parallel**: Incorrect KV cache partitioning or ring attention implementation errors
- **Expert parallel**: Token routing/dispatch errors across EP ranks, incorrect capacity factor handling
- **Gradient accumulation**: Missing no_sync() context or incorrect division factor when accumulating across microbatches
- **Checkpoint save/load**: State dict key mismatch, missing optimizer states, incorrect RNG state restoration — causes silent divergence after resume
- **RNG state management**: Incorrect random seed handling across TP/PP/DP ranks, causing correlated dropout masks or data sampling

## Important Issues (Should Fix)

### Training Performance
- **Unnecessary CPU-GPU sync**: .item(), .cpu(), torch.cuda.synchronize(), Python-side tensor value checks in training loop — kills throughput
- **Redundant communication**: Allreduce/allgather that could be fused, overlapped with compute, or eliminated
- **Memory inefficiency**: Missing activation checkpointing on memory-heavy layers, unnecessary tensor clones or .contiguous() calls
- **Communication-computation overlap**: Missed opportunities to overlap allreduce with backward, or allgather with forward
- **Kernel launch overhead**: Python loops over small ops that should be fused into a single kernel
- **CUDA graph compatibility**: Dynamic shapes, Python-side conditionals on tensor values, host-device sync inside captured region

### Backward Compatibility
- **Config/argument changes**: Renamed or removed arguments without deprecation path — breaks existing training scripts
- **Checkpoint format changes**: Modified state dict keys/structure without migration logic — makes existing checkpoints unloadable
- **Default value changes**: Changed defaults for training hyperparameters or parallelism settings — silently alters behavior for users relying on defaults
- **API contract changes**: Changed function signatures, return types, or side effects in megatron/core/ without backward-compat shim
- **Model architecture changes**: Altered layer ordering, initialization, or normalization placement — existing pretrained weights become incompatible

### Mandatory Check: Unused New Variables / Arguments
- For each changed file, list newly added identifiers (function args, config fields, locals).
- Verify each has a meaningful read/use path — not just declaration/docstring or discard assignment (_ = new_arg).
- Use Grep to search for usage beyond declaration sites.
- Treat placeholder discard patterns as findings unless explicitly documented as temporary migration shim.
- If usage is intentionally deferred, flag and request explicit TODO + migration note.

## Suggestions (Nice to Have)

### Naming
- Name must describe what the thing *is*, not what it's *used for*
- No abbreviations in parallel/distributed code — use full names (token_dispatcher, routing_map, comm_manager, world_size)
- Naming consistency within scope for variables serving the same role

### Function/Method Decomposition
- Functions over ~50 lines mixing data collection, reduction, computation, and I/O should be split
- Non-trivial logic blocks embedded in a method with different primary purpose should be extracted

### Simplification
- Redundant operations (e.g. .reshape(()) on 0-dim tensor, two-step constructions where one suffices)
- Setup constant across training should not run on every forward pass — move to __init__
- Dead complexity that doesn't achieve its stated purpose
- Unnecessary intermediate aliases adding indirection with no abstraction value

### Other
- Stale, imprecise, or misleading comments/docstrings — a wrong docstring is worse than none
- Missing shape/dtype assertions at parallelism boundaries

## What NOT to Comment On
- Style/formatting issues (leave to linters)
- Test code that is reasonably clear
- Clearly intentional design decisions by the author
- Pure refactoring that preserves identical behavior (verify via diff)
- Findings invalidated by deeper analysis — drop them entirely rather than hedging

## Comment Format

Prefix each comment with severity and category tag:
- `**[CRITICAL Implementation]**`, `**[CRITICAL Correctness]**`
- `**[IMPORTANT Performance]**`, `**[IMPORTANT Compatibility]**`
- `**[SUGGESTION Naming]**`, `**[SUGGESTION Simplification]**`

For each finding, explain: (1) what the issue is, (2) why it matters (impact/risk), (3) specific suggestion for fix.

Only use inline ```suggestion blocks for simple, self-contained line replacements (typos,
renames, single-line fixes). For structural changes that add, remove, or reorganize blocks
of code, use a top-level PR comment with a code block showing the proposed change instead.

## Completion

After posting all inline comments, post a summary PR comment:
- List total findings by severity (CRITICAL: N, IMPORTANT: N, SUGGESTION: N)
- Highlight the most impactful findings
- Overall assessment of the PR's risk level

If no significant issues are found, approve the PR:
gh pr review $PR_NUMBER --repo $REPO --approve --body "Strict review passed — no significant issues found. LGTM"
29 changes: 29 additions & 0 deletions .github/workflows/sync-skills.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
name: Sync skills → agent dirs

on:
workflow_dispatch:
push:
branches:
- main
paths:
- "skills/**"
- "AGENTS.md"

jobs:
sync:
uses: NVIDIA-NeMo/FW-CI-templates/.github/workflows/_sync_skills.yml@v0.91.0
secrets:
PAT: ${{ secrets.PAT }}
Loading
Loading