Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
49 commits
Select commit Hold shift + click to select a range
9a3c927
Fix nvtx_decorator to check _nvtx_enabled at call time (#4184)
minitu Apr 22, 2026
60f71e1
fix merges_file typo in megatron_hf_tokenizer (#4392)
chelseajohn Apr 22, 2026
c9dfe34
Enable NullTokenizer for pretraining to reduce I/O access (#4057)
asolergi-nv Apr 22, 2026
7073492
docs: Add SECURITY.md (#4431)
chtruong814 Apr 22, 2026
40627d0
Mamba inference opt (#4414)
wdykas Apr 22, 2026
55b8111
DDP refactoring: Extract parameter layout computation into optimizer …
deepakn94 Apr 22, 2026
90e09b6
Update PR template with explicit request for issue (#4409)
Phlip79 Apr 22, 2026
ab2b33d
Misc inference fixes (#4397)
sidsingh-nvidia Apr 23, 2026
60408d5
Rename Mamba to Hybrid outside megatron/core (#4159)
Phlip79 Apr 23, 2026
a52014c
Include mtp layers in token per expert logging (#4412)
Mellonta Apr 23, 2026
32275b2
fix: NVRx async compatibility and defer resiliency import (#4420)
sbak5 Apr 23, 2026
9bb35a8
ci: add base_sha to codecov/codecov-action upload step (#4445)
ko3n1g Apr 23, 2026
3034d86
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Apr 24, 2026
f78ed05
fix(checkpoint_inspector): allow empty --param-to-param-group-map-jso…
DAISY-gh Apr 24, 2026
4d6cdd5
Add the YARN support for hybrid_model (#4244)
guihong-nv Apr 24, 2026
41ffa83
[training migration] Add container class for config dataclasses (#4227)
maanug-nv Apr 24, 2026
a1165fa
Inference: Fix broken functional tests on gitlab (#4454)
sidsingh-nvidia Apr 24, 2026
d4cacef
SafeUnpickler class for safe pickle usage (#4319)
dimapihtar Apr 24, 2026
109feda
get rid of weights_only=False (#4434)
dimapihtar Apr 24, 2026
64870c1
Inference | Per-block MoE routing storage for prefix caching (#4301)
lmcafee-nvidia Apr 24, 2026
017e684
Add troubleshooting tip for 'access forbidden' (#4449)
balasaajay Apr 24, 2026
3d7bcd3
Fix checkpoint loading with rerun state machine (#4448)
YangFei1990 Apr 24, 2026
9b02206
Add misc CUDA graph sugar to CudaGraphManager (#4425)
tdene Apr 24, 2026
35f76df
Inference: Add the embedding and output layer in the full_iteration_i…
sidsingh-nvidia Apr 24, 2026
481efd0
Important bugfixes in local CG implementation that were leading to lo…
jiemingz Apr 24, 2026
e9abb6c
fix: Replace polynomial rolling hash with SHA-256 for prefix caching …
lmcafee-nvidia Apr 24, 2026
377af02
feat(ckpt): expose validate_access_integrity knob on dist-ckpt load (…
asolergi-nv Apr 24, 2026
241a5ca
Fix multivalidation (#3388)
RPrenger Apr 25, 2026
f2dcd42
Add missing knob for reduce_scatter_with_fp32_accumulation (#4410)
WanZzzzzz Apr 25, 2026
03f4111
Enable CUDA graphs for MTP inference (#4260)
santhnm2 Apr 26, 2026
1879dc2
chore(beep boop 🤖): Bump (main) (2026-04-27)
github-actions[bot] Apr 27, 2026
970c254
checkpoint integrity verification (#4305)
dimapihtar Apr 27, 2026
ebd70d3
Fix cache gating (#4455)
wdykas Apr 27, 2026
0447347
[Main] Fix FusedAdam.use_decoupled_grad mis-set for Megatron-FSDP. (#…
cspades Apr 27, 2026
8c5cf05
add permute fusion into hybrid ep (#4089)
Autumn1998 Apr 28, 2026
42e396e
Add ColocatedBridgeCommunicator for heterogeneous TP/DP MIMO training…
yashaswikarnati Apr 28, 2026
6fd6652
Fix incorrect bias display in extra_repr of Column/RowParallelLinear …
HelloWorldBeginner Apr 28, 2026
c8a4bfd
Fix assertion logic in combined_1f1b_schedule_for_interleaved_pipelin…
joapolarbear Apr 28, 2026
374fa85
ci: Fix event name reference in CI workflow condition for merge group…
balasaajay Apr 28, 2026
9c15290
Add manual sync workflow from main to dev (#4165)
Phlip79 Apr 28, 2026
9816140
fix: handle list-format quant_cfg from ModelOpt PR #1094 (#4187)
ChenhanYu Apr 28, 2026
9e98259
ci: also add Run MBridge tests label in nightly sync workflow (#4499)
Phlip79 Apr 28, 2026
702d4bf
Merge remote-tracking branch 'origin/main' into main2dev/28_04_2026
svcnvidia-nemo-ci Apr 28, 2026
cd2001d
chore: post-merge fixes for nightly sync main into dev (28_04_2026)
svcnvidia-nemo-ci Apr 28, 2026
54cbb38
fix: restore missing dev-only fields in TransformerConfig
svcnvidia-nemo-ci Apr 28, 2026
0e9f938
fix: restore num_sms_preprocessing_api in fused_a2a + golden config
svcnvidia-nemo-ci Apr 28, 2026
3d25b74
fix: HybridEPDispatch.backward must return 13 gradients, not 12
svcnvidia-nemo-ci Apr 29, 2026
a7f96d8
fix: restore Dynamic-CP collate_fn override in data_samplers
Phlip79 Apr 29, 2026
42f1108
fix: restore wrap_data_iterator calls for sequence packing scheduler
Phlip79 Apr 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/copy-pr-bot.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
enabled: true
auto_sync_draft: false
auto_sync_ready: true
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Mellonta", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "WanZzzzzz", "Wohox", "YangFei1990", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "aroshanghias-nvd", "asolergi-nv", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "frsun-nvda", "gautham-kollu", "gdengk", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kanz-nv", "keshavb96", "kevalmorabia97", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "parthmannan", "prajwal1210", "pthombre", "rhewett-nv", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "wdykas", "wplf", "wujingyue", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "Connor-XY", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Mellonta", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "WanZzzzzz", "Wohox", "YangFei1990", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "aroshanghias-nvd", "asolergi-nv", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "fitsumreda", "frsun-nvda", "gautham-kollu", "gdengk", "guihong-nv", "guyueh1", "hexinw-nvidia", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kajalj22", "kanz-nv", "kevalmorabia97", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "minitu", "mkhona-nvidia", "nanz-nv", "parthmannan", "prajwal1210", "pthombre", "rhewett-nv", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "sheliang-nv", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sraman-rgb", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "wdykas", "wplf", "wujingyue", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
9 changes: 9 additions & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,15 @@

:warning: For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact the @mcore-oncall.

## Issue tracking

For PRs from open-source community contributors:

- **New features**: a linked issue is **required**. Please open a [feature request](https://github.com/NVIDIA/Megatron-LM/issues/new?template=feature_request.md) and reference it here before submitting the PR.
- **Small updates (bug fixes, minor improvements)**: a linked issue is **recommended** and will accelerate the PR review process.

Linked issue: <!-- e.g. Fixes #1234 / Related to #1234 -->

## Contribution process

### Pre-checks
Expand Down
8 changes: 7 additions & 1 deletion .github/workflows/cicd-main.yml
Original file line number Diff line number Diff line change
Expand Up @@ -972,7 +972,7 @@ jobs:
(
needs.pre-flight.outputs.docs_only == 'true'
|| needs.pre-flight.outputs.is_deployment_workflow == 'true'
|| github.event == 'merge_group'
|| github.event_name == 'merge_group'
)
&& needs.pre-flight.outputs.is_ci_workload == 'false'
&& !cancelled()
Expand Down Expand Up @@ -1006,6 +1006,11 @@ jobs:
matrix:
flag: [unit-test]
steps:
- name: Get PR info
id: get-pr-info
if: startsWith(github.ref, 'refs/heads/pull-request/') && github.event_name == 'push'
uses: nv-gha-runners/get-pr-info@main

- name: Checkout
uses: actions/checkout@v6

Expand Down Expand Up @@ -1036,6 +1041,7 @@ jobs:
token: ${{ secrets.CODECOV_TOKEN }}
verbose: true
flags: ${{ matrix.flag }}
base_sha: ${{ fromJSON(steps.get-pr-info.outputs.pr-info || '{}').base.sha }}

- name: Upload artifacts
uses: actions/upload-artifact@v6
Expand Down
196 changes: 196 additions & 0 deletions .github/workflows/nightly-sync-main-to-dev.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
# Copyright (c) 2026, NVIDIA CORPORATION. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

name: Nightly Sync Main to Dev

on:
workflow_dispatch:
schedule:
# 21:00 UTC = 2 PM PDT (1 PM PST during winter — GitHub Actions cron
# is UTC-only and does not follow DST).
- cron: '0 21 * * *'

concurrency:
group: nightly-sync-main-to-dev
cancel-in-progress: false

permissions:
contents: write
pull-requests: write
issues: write
id-token: write

jobs:
sync-main-to-dev:
runs-on: ubuntu-latest
if: github.repository == 'NVIDIA/Megatron-LM'
timeout-minutes: 360
env:
GH_TOKEN: ${{ secrets.PAT }}
steps:
- name: Checkout repository
uses: actions/checkout@v6
with:
fetch-depth: 0
token: ${{ secrets.PAT }}

- name: Configure Git
run: |
git config user.name "svcnvidia-nemo-ci"
git config user.email "svcnvidia-nemo-ci@nvidia.com"

- name: Compute branch name
id: vars
run: |
DATE=$(date -u +%d_%m_%Y)
BRANCH="main2dev/${DATE}"
echo "branch=$BRANCH" >> "$GITHUB_OUTPUT"
echo "date=$DATE" >> "$GITHUB_OUTPUT"

- name: Close previous unmerged sync PRs
run: |
OPEN_PRS=$(gh pr list \
--repo "${{ github.repository }}" \
--base dev \
--state open \
--json number,headRefName \
--jq '.[] | select(.headRefName | startswith("main2dev/")) | .number')

for PR_NUM in $OPEN_PRS; do
echo "Closing stale sync PR #${PR_NUM}"
gh pr close "$PR_NUM" \
--repo "${{ github.repository }}" \
--comment "Superseded by today's nightly sync."
done

- name: Check if sync is needed
id: check-sync
run: |
git fetch origin main dev
AHEAD_COUNT=$(git rev-list --count origin/dev..origin/main)
echo "main is $AHEAD_COUNT commit(s) ahead of dev"
if [ "$AHEAD_COUNT" -eq 0 ]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "No changes to sync."
else
echo "skip=false" >> "$GITHUB_OUTPUT"
fi

- name: Run Claude Code to merge, fix, and iterate
if: steps.check-sync.outputs.skip != 'true'
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
github_token: ${{ secrets.PAT }}
prompt: |
You are an automated sync bot. Merge `main` into `dev`, create a
PR, ensure CI passes (fixing failures), and mark the PR ready.
There are 4 phases. You are NOT done until Phase 4 completes.

REPO: ${{ github.repository }}
BRANCH: ${{ steps.vars.outputs.branch }}
DATE: ${{ steps.vars.outputs.date }}

Read `.claude/skills/nightly-sync/SKILL.md` for the detailed
merge strategy, CI architecture, failure investigation procedures,
and known issues. Also read `.claude/skills/build-and-test/SKILL.md`
and `CLAUDE.md` for general CI and contribution guidelines.

## Hard Constraints

**Exit condition:** You MUST run `gh pr ready <PR_NUMBER>` before
exiting. That command is Phase 4. Do NOT exit after Phase 1, 2,
or 3 — not even if CI is "still running" or "stuck in queue."
Keep polling until it resolves, then act.

**NO background tasks. Ever.**
You are running inside a single GitHub Actions step. The step
process owns your shell. When you stop issuing tool calls, the
step ends and the runner container is DESTROYED — every
background process dies with it and cannot resume. There is no
"future session" to wake up into.

The following are strictly forbidden:
- `Bash` with `run_in_background: true`
- `Agent` with `run_in_background: true`
- `ScheduleWakeup` (nothing will ever wake up)
- Any shell command ending in `&`, or using `nohup`, `disown`,
or `setsid` to detach a process
- `tail -f` on a log produced by a backgrounded task

Required shape for every long wait: ONE foreground Bash tool
call containing an inline `while true; do ... sleep <N>; done`
or `until ...; do sleep <N>; done` loop that BLOCKS inside
that single tool call and only returns when the wait is
resolved (success, failure, or a clearly-classified terminal
state). Do NOT break a long wait into many short polls with
conversation in between — that wastes `--max-turns` and
creates windows where the agent could forget the loop.

**Source of truth for CI status:**
`gh pr view <PR_NUMBER> --repo $REPO --json statusCheckRollup`
This lists every required check — GitHub Actions jobs AND
external contexts (GitLab CI, `copy-pr-bot`, etc.). The
`gh api .../actions/runs/<RUN_ID>/jobs` endpoint alone is
NOT sufficient — it misses external contexts.

**Pre-existing failures:** MUST verify against recent dev CI
before classifying any failure as pre-existing. Run
`gh pr checks` on a recently merged dev PR. If the test passes
on dev, the failure is sync-caused and you must fix it. A
check that has never completed on your PR cannot be
pre-existing — wait for it to finish first.

**Phase 4 gate — strict "all terminal, all green":**
Do NOT run `gh pr ready` until every non-exempt required check
in `statusCheckRollup` satisfies BOTH:
- `status == "COMPLETED"` (NOT `QUEUED`, `IN_PROGRESS`,
`PENDING`, `WAITING`, or `REQUESTED`), AND
- `conclusion` ∈ {`SUCCESS`, `SKIPPED`, `NEUTRAL`}.
A check stuck in a runner queue is NOT complete. Never
classify queued/in-progress jobs as "infrastructure-blocked"
and ship anyway — wait for them to reach a terminal
conclusion, then act on that result. When a check fails,
loop: diagnose → fix → commit → push → `/ok to test <sha>` →
poll. Only exit the loop when the gate is satisfied on the
LATEST CI run against the current HEAD SHA.

**Exempt checks (may be ignored for the Phase 4 gate):**
These categories are pre-merge policy signals, not
correctness signals, so their failure must not block the
sync bot from marking the PR ready for human review.

- Approval / code-review: `codeowners-approval`,
`check-approval`, `multi-approval-bot-summary`,
`is-not-external-contributor`, any check whose name
contains `review` or `approval`.
- Code coverage: `Coverage (unit-test)`, `Coverage_Fake`,
any check whose name contains `codecov` or `coverage`
(case-insensitive).
- Docs: `build-docs / Build docs`, `build-docs-summary`,
any check whose name contains `build-docs`, `doc-build`,
`readthedocs`, or `sphinx`.

Everything else — unit tests (`tests/unit_tests/...`),
integration tests (`gpt/...`, `moe/...`, etc.), `linting`,
`cicd-container-build`, `cicd-mbridge-testing`,
`Nemo_CICD_Test`, `copyright-check`, `pre-flight`, wheel
builds, etc. — is NOT exempt and must reach a terminal
green conclusion.
show_full_output: true
claude_args: |
--allowedTools "Bash,Read,Edit,Write,Grep,Glob,Agent"
--model "opus[1m]"
--effort max
--max-turns 1500
25 changes: 25 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
## Security

NVIDIA is dedicated to the security and trust of our software products and services, including all source code repositories managed through our organization.

If you need to report a security issue, please use the appropriate contact points outlined below. **Please do not report security vulnerabilities through GitHub.** If a potential security issue is inadvertently reported via a public issue or pull request, NVIDIA maintainers may limit public discussion and redirect the reporter to the appropriate private disclosure channels.

## Reporting Potential Security Vulnerability in an NVIDIA Product

To report a potential security vulnerability in any NVIDIA product:

- Web: [Security Vulnerability Submission Form](https://www.nvidia.com/object/submit-security-vulnerability.html)
- E-Mail: psirt@nvidia.com
- We encourage you to use the following PGP key for secure email communication: [NVIDIA public PGP Key for communication](https://www.nvidia.com/en-us/security/pgp-key)
- Please include the following information:
- Product/Driver name and version/branch that contains the vulnerability
- Type of vulnerability (code execution, denial of service, buffer overflow, etc.)
- Instructions to reproduce the vulnerability
- Proof-of-concept or exploit code
- Potential impact of the vulnerability, including how an attacker could exploit the vulnerability

While NVIDIA currently does not have a bug bounty program, we do offer acknowledgement when an externally reported security issue is addressed under our coordinated vulnerability disclosure policy. Please visit our [Product Security Incident Response Team (PSIRT)](https://www.nvidia.com/en-us/security/psirt-policies/) policies page for more information.

## NVIDIA Product Security

For all security-related concerns, please visit NVIDIA's Product Security portal at https://www.nvidia.com/en-us/security
2 changes: 1 addition & 1 deletion docs/developer/contribute.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ This document outlines the processes and policies for issues and pull requests b

Everyone is welcome to contribute to the project! We recently migrated from using an internal repo to doing all development directly from the GitHub repository.

When contributing it is important to ensure that changes are in line with the project direction. Small changes to fix bugs are welcomed and appreciated. If proposing large architectural changes or changes for stylistic reasons open an issue first so we can discuss it.
When contributing it is important to ensure that changes are in line with the project direction. Small changes to fix bugs are welcomed and appreciated. **If proposing large architectural changes or changes for stylistic reasons open an issue first so we can discuss it.**

## Issue policy

Expand Down
33 changes: 30 additions & 3 deletions docs/user-guide/features/tokenizers.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,7 +149,24 @@ tokenizer = MegatronTokenizer.from_pretrained(

### Null Tokenizer

Use a null tokenizer for testing or non-text models:
The Null tokenizer is a lightweight, zero-I/O tokenizer that requires no model files.
It is useful in three scenarios:

1. **Performance benchmarking** with `--mock-data` where real tokenization is unnecessary.
2. **Testing** in functional tests and CI pipelines where tokenizer model files may not
be available. The Null tokenizer removes the dependency on external files, making
tests self-contained and portable.
3. **Pretraining with pretokenized data** where all data is already tokenized into
`.bin`/`.idx` files. In this case the tokenizer is only needed for metadata
(`vocab_size`, `eod`, `pad`) — not for actual tokenization. Using the Null tokenizer
avoids redundant filesystem access at scale, which is particularly beneficial on
shared filesystems like Lustre where thousands of ranks would otherwise all load the
same tokenizer files.

Properties derived from `--vocab-size N`:
- `vocab_size` = `N` (the exact value passed)
- `eod` = `N - 1` (last token in the vocabulary)
- `pad` = `0`

```python
tokenizer = MegatronTokenizer.from_pretrained(
Expand All @@ -165,10 +182,20 @@ tokenizer = MegatronTokenizer.from_pretrained(
The tokenizer system works with Megatron-LM training scripts:

```bash
# Null tokenizer for testing
# Null tokenizer for benchmarking with mock data
torchrun --nproc_per_node=8 pretrain_gpt.py \
--tokenizer-type NullTokenizer \
--vocab-size 131072 \
--mock-data \
...
```

```bash
# Null tokenizer for pretraining with pretokenized data (no tokenizer files needed)
torchrun --nproc_per_node=8 pretrain_gpt.py \
--tokenizer-type NullTokenizer \
--vocab-size 128256 \
--data-path /path/to/pretokenized_data \
...
```

Expand All @@ -195,7 +222,7 @@ The following table lists supported tokenizer backends:
| **SentencePiece** | Google's tokenizer | GPT-style models, custom vocabularies |
| **TikToken** | OpenAI's tokenizer | GPT-3.5/GPT-4 style tokenization |
| **Megatron** | Built-in tokenizers | Legacy GPT-2 BPE |
| **Null** | No-op tokenizer | Testing, non-text modalities |
| **Null** | Zero-I/O tokenizer | Benchmarking, pretokenized data |

## Common Tokenizer Types

Expand Down
4 changes: 2 additions & 2 deletions examples/mamba/run_text_gen_server_8b.sh
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ export NCCL_IB_QPS_PER_CONNECTION=4
export TRITON_CACHE_DIR="./triton-cache/"
export TRITON_CACHE_MANAGER="megatron.core.ssm.triton_cache_manager:ParallelFileCacheManager"

torchrun $DISTRIBUTED_ARGS ../../tools/run_mamba_text_generation_server.py \
torchrun $DISTRIBUTED_ARGS ../../tools/run_hybrid_text_generation_server.py \
--tensor-model-parallel-size 1 \
--pipeline-model-parallel-size 1 \
--untie-embeddings-and-output-weights \
Expand All @@ -46,5 +46,5 @@ torchrun $DISTRIBUTED_ARGS ../../tools/run_mamba_text_generation_server.py \
--bf16 \
--micro-batch-size 1 \
--use-mcore-models \
--spec megatron.core.models.mamba.mamba_layer_specs mamba_stack_spec \
--spec megatron.core.models.hybrid.hybrid_layer_specs hybrid_stack_spec \
--seed 42
4 changes: 2 additions & 2 deletions examples/mamba/train.sh
Original file line number Diff line number Diff line change
Expand Up @@ -96,8 +96,8 @@ options=" \
--eval-iters 32 \
--bf16 \
--use-mcore-models \
--spec megatron.core.models.mamba.mamba_layer_specs mamba_stack_spec \
--spec megatron.core.models.hybrid.hybrid_layer_specs hybrid_stack_spec \
--no-create-attention-mask-in-dataloader \
--tensorboard-dir ${TENSORBOARD_DIR}"

torchrun --nproc_per_node 8 ../../pretrain_mamba.py ${options}
torchrun --nproc_per_node 8 ../../pretrain_hybrid.py ${options}
Loading
Loading