sync: bump trainers-main to NVIDIA dev @ bfa33263 (DSv4 context-parallel, #5087) - #9
Closed
JackRao123 wants to merge 543 commits into
Closed
sync: bump trainers-main to NVIDIA dev @ bfa33263 (DSv4 context-parallel, #5087)#9JackRao123 wants to merge 543 commits into
JackRao123 wants to merge 543 commits into
Conversation
Signed-off-by: Charlie Truong <chtruong@nvidia.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
… arguments.py (NVIDIA#3266) Co-authored-by: Xin Yao <xiny@nvidia.com>
… state_dict (NVIDIA#3243) Co-authored-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com> Signed-off-by: Charlie Truong <chtruong@nvidia.com> Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com> Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Li Tao <lit@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: oliver könig <okoenig@nvidia.com> Signed-off-by: Charlie Truong <chtruong@nvidia.com> Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> Signed-off-by: Youngeun Kwon <youngeunk@nvidia.com> Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com> Signed-off-by: Jimmy Zhang <jiemingz@nvidia.com> Signed-off-by: Santosh Bhavani <santosh.bhavani@live.com> Signed-off-by: Deepak Narayanan <dnarayanan@nvidia.com> Signed-off-by: Hollow Man <hollowman@opensuse.org> Signed-off-by: Robin Zhang <robinz@nvidia.com> Signed-off-by: jinliangl <jinliangl@nvidia.com> Signed-off-by: Maanu Grover <maanug@nvidia.com> Signed-off-by: dimapihtar <dpihtar@gmail.com> Signed-off-by: xiaoxi-wangfj <690912414@qq.com> Signed-off-by: skydoorkai <htsantaclara@163.com> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Signed-off-by: meg miranda <mmiranda@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: sajadn <snorouzi@nvidia.com> Signed-off-by: lit <lit@nvidia.com> Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com> Signed-off-by: Cory Ye <cye@nvidia.com> Signed-off-by: adithyare <adithyare@nvidia.com> Signed-off-by: Soumye Singhal <soumyes@cw-dfw-cs-001-dc-01.cm.cluster> Signed-off-by: Ahmad Kiswani <kiswani.ahmad@gmail.com> Signed-off-by: mikail <mkhona@nvidia.com> Co-authored-by: HaochenYuan <106647990+HaochenYuan@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com> Co-authored-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Duncan Riach <33532941+duncanriach@users.noreply.github.com> Co-authored-by: yobi byte <yobibyte@users.noreply.github.com> Co-authored-by: Charlie Truong <chtruong@nvidia.com> Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com> Co-authored-by: root <root@gpu-h100-0348.cm.cluster> Co-authored-by: root <root@gpu-h100-0193.cm.cluster> Co-authored-by: root <root@gpu-h100-0082.cm.cluster> Co-authored-by: root <root@gpu-h100-0495.cm.cluster> Co-authored-by: William Dykas <wdykas@cw-pdx-cs-001-vscode-02.cm.cluster> Co-authored-by: root <root@gpu-h100-0213.cm.cluster> Co-authored-by: root <root@gpu-h100-0435.cm.cluster> Co-authored-by: root <root@gpu-h100-0188.cm.cluster> Co-authored-by: root <root@gpu-h100-0032.cm.cluster> Co-authored-by: root <root@gpu-h100-0023.cm.cluster> Co-authored-by: root <root@gpu-h100-0368.cm.cluster> Co-authored-by: root <root@gpu-h100-0203.cm.cluster> Co-authored-by: root <root@gpu-h100-0229.cm.cluster> Co-authored-by: root <root@gpu-h100-0123.cm.cluster> Co-authored-by: root <root@gpu-h100-0217.cm.cluster> Co-authored-by: root <root@gpu-h100-0496.cm.cluster> Co-authored-by: root <root@gpu-h100-0261.cm.cluster> Co-authored-by: GitHub Actions <github-actions[bot]@users.noreply.github.com> Co-authored-by: Jiayi Yan <66017932+1195343015@users.noreply.github.com> Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com> Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com> Co-authored-by: Youngeun Kwon <youngeunk@nvidia.com> Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com> Co-authored-by: Jimmy Zhang <133159885+jiemingz@users.noreply.github.com> Co-authored-by: tgkyrie <74066353+tgkyrie@users.noreply.github.com> Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com> Co-authored-by: Xin Yao <xiny@nvidia.com> Co-authored-by: rkarimimahab <rkarimimahab@nvidia.com> Co-authored-by: Rabeeh Mahabadi <rkarimimahab@nb-hel-cs-001-vscode-02.cm.cluster> Co-authored-by: Sanjeev Satheesh <sasatheesh@nvidia.com> Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com> Co-authored-by: Santosh Bhavani <santosh.bhavani@live.com> Co-authored-by: Ahmad Kiswani <kiswani.ahmad@gmail.com> Co-authored-by: Li Tao <lit@nvidia.com> Co-authored-by: Maanu Grover <109391026+maanug-nv@users.noreply.github.com> Co-authored-by: mvirts <mvirts@gmail.com> Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com> Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org> Co-authored-by: Robin Zhang <robinz@nvidia.com> Co-authored-by: Sheng Fu <shengf@nvidia.com> Co-authored-by: Venmugil Elango <498703+venmugil@users.noreply.github.com> Co-authored-by: mathemakitten <helenn@nvidia.com> Co-authored-by: Jared Casper <155158+jaredcasper@users.noreply.github.com> Co-authored-by: Parth Mannan <38387286+parthmannan@users.noreply.github.com> Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com> Co-authored-by: Tong Liu <tongliu@nvidia.com> Co-authored-by: Li Jinliang <jinliangl@nvidia.com> Co-authored-by: Jinliang Li <jinliangl@pool0-01676.cm.cluster> Co-authored-by: Jinliang Li <jinliangl@cw-dfw-cs-001-vscode-01.cm.cluster> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Nick Schank <nick@reflection.ai> Co-authored-by: Jeffrey Chen <jeffrey@reflection.ai> Co-authored-by: janEbert <janpabloe@nvidia.com> Co-authored-by: rj42 <lbkzman@gmail.com> Co-authored-by: Juntao Wang <juntaow@nvidia.com> Co-authored-by: Pingtian Li <158665726+Wohox@users.noreply.github.com> Co-authored-by: Chris Grimm <chris@reflection.ai> Co-authored-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com> Co-authored-by: Eric Harper <eharper@nvidia.com> Co-authored-by: xiaoxi-wangfj <690912414@qq.com> Co-authored-by: Jianbin Chang <shjwudp@gmail.com> Co-authored-by: c1lovez1 <141424951+c1lovez1@users.noreply.github.com> Co-authored-by: Zhang Haitao <htsantaclara@163.com> Co-authored-by: yeyu-nvidia <yeyu@nvidia.com> Co-authored-by: kwyss-nvidia <kwyss@nvidia.com> Co-authored-by: Jon Barker <jbarker@nvidia.com> Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Siddharth Singh <136645615+sidsingh-nvidia@users.noreply.github.com> Co-authored-by: megnvidia <mmiranda@nvidia.com> Co-authored-by: thecaptain789 <257642323+thecaptain789@users.noreply.github.com> Co-authored-by: thecaptain789 <thecaptain789@users.noreply.github.com> Co-authored-by: litianjian <litianjian@bytedance.com> Co-authored-by: Yan Bai <baiyan1996@icloud.com> Co-authored-by: xuwchen <xuwenc@nvidia.com> Co-authored-by: John St. John <jstjohn@users.noreply.github.com> Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> Co-authored-by: Robert Kirby <ArEsKay3@users.noreply.github.com> Co-authored-by: Siddharth Singh <sidsingh@nvidia.com> Co-authored-by: Robert Kirby <rkirby@cw-dfw-cs-001-vscode-01.cm.cluster> Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com> Co-authored-by: Dennis(Zhenhuan) Liu <denliu@nvidia.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Shanmugam Ramasamy <111910568+shanmugamr1992@users.noreply.github.com> Co-authored-by: vasunvidia <108759426+vasunvidia@users.noreply.github.com> Co-authored-by: Philip Petrakian <pgpetrak@gmail.com> Co-authored-by: Sajad Norouzi <sajad.n@gmail.com> Co-authored-by: Kunlun Li <94586211+kunlunl@users.noreply.github.com> Co-authored-by: xielaixin <xielx@shanghaitech.edu.cn> Co-authored-by: Robert Kirby <rkirby@nvidia.com> Co-authored-by: Ming <93323717+dndnda@users.noreply.github.com> Co-authored-by: liming127 <liming127@meituan.com> Co-authored-by: Jon Barker <jbarker@oci-hsg-cs-001-vscode-01.cm.cluster> Co-authored-by: helen ngo <helen.ngo14@gmail.com> Co-authored-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: yueshen2016 <39203804+yueshen2016@users.noreply.github.com> Co-authored-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com> Co-authored-by: Cory Ye <44509866+cspades@users.noreply.github.com> Co-authored-by: Adi Renduchintala <adithya.r@gmail.com> Co-authored-by: Soumye Singhal <soumyes@cw-dfw-cs-001-dc-01.cm.cluster> Co-authored-by: Seonjin Na <sna@nvidia.com> Co-authored-by: Seonmyeong Bak <sbak@nvidia.com> Co-authored-by: Mikail Khona (NVIDIA) <mkhona@nvidia.com>
…tp_size. (NVIDIA#3529) Co-authored-by: xiaotaoliu <xiaotaoliu@tencent.com> Co-authored-by: Yuzhong Wang <yuzhongw@nvidia.com> Co-authored-by: Zijie Yan <zijiey@nvidia.com>
…tOutput (NVIDIA#3641) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com> Signed-off-by: tailaim <tailaim@nvidia.com> Co-authored-by: kunlunl <kunlunl@nvidia.com>
…VIDIA#3668) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Hao Wu <skyw@nvidia.com> Co-authored-by: Hao Wu <skyw@nvidia.com>
Co-authored-by: Robin Zhang <robinz@nvidia.com>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Nightly sync of main into dev (22_06_2026). Resolves 16 conflicts preserving dev features (pre-push guard: 0 dropped dev lines); brings in main's inference shard-spec API additively. Supersedes NVIDIA#5429. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
…NVIDIA#5401) Signed-off-by: yangfan.bai <yangfan.bai@shopee.com> Co-authored-by: yangfan.bai <yangfan.bai@shopee.com>
…IDIA#5450) Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… flaky_in_dev (NVIDIA#5475) Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…DIA#5476) Signed-off-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1) test_optimizer.py: reverted to dev. The sync kept dev's multi_latent_attention.py (split q/kv down-proj, no _synthesize_fused_qkv_down_weight), but auto-merged main's test asserting the fused linear_qkv_down_proj.weight key. 2) training.py: guard the dev-only sequence_packing_scheduler config access with getattr (lines in train_step and train()). main's new MIMO schedule-plumbing test (NVIDIA#5333) passes an empty SimpleNamespace config; the reconciled training.py keeps dev's packing path, so the access must tolerate a config lacking the attribute. Real configs are unaffected (getattr returns the same value). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: xiaoyao0115 <1804647152@qq.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: hongbinl <hongbinl@nvidia.com> Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com> Signed-off-by: janEbert <janpabloe@nvidia.com> Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com> Signed-off-by: Helen Ngo <helenn@nvidia.com> Signed-off-by: ykarnati <ykarnati@nvidia.com> Signed-off-by: Shijie Wang <jaywan@nvidia.com> Signed-off-by: Ajay Balasa <abalasa@nvidia.com> Signed-off-by: oliver könig <okoenig@nvidia.com> Signed-off-by: Antoni-Joan Solergibert <asolergibert@nvidia.com> Signed-off-by: ilml <tolong@nvidia.com> Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com> Signed-off-by: sraman <sraman@nvidia.com> Signed-off-by: Jingyue Wu <wujingyue@gmail.com> Signed-off-by: Hollow Man <hollowman@opensuse.org> Signed-off-by: hongbinl <hongbinl@nvidia.com> Signed-off-by: Charlie Truong <chtruong@nvidia.com> Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com> Signed-off-by: wdykas <wdykas@nvidia.com> Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Signed-off-by: svcnvidia-nemo-ci <svc-nvidia-nemo-ci@nvidia.com> Co-authored-by: Teodor-Dumitru Ene <34819528+tdene@users.noreply.github.com> Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Jorge Albericio <jalbericiola@nvidia.com> Co-authored-by: Pranav Thombre <pthombre@nvidia.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: janEbert <janpabloe@nvidia.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com> Co-authored-by: mathemakitten <helenn@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Shijie <505749828@qq.com> Co-authored-by: Ajay <abalasa@nvidia.com> Co-authored-by: oliver könig <okoenig@nvidia.com> Co-authored-by: Antoni-Joan Solergibert <asolergibert@nvidia.com> Co-authored-by: Deepak Narayanan <dnarayanan@nvidia.com> Co-authored-by: Tom Long <tolong@nvidia.com> Co-authored-by: Keshav Santhanam <ksanthanam@nvidia.com> Co-authored-by: Teodor-Dumitru Ene <teodord.ene@gmail.com> Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com> Co-authored-by: Jingyue Wu <wujingyue@gmail.com> Co-authored-by: ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 <hollowman@opensuse.org> Co-authored-by: Hongbin Liu <lhb8125@users.noreply.github.com> Co-authored-by: Charlie Truong <chtruong@nvidia.com> Co-authored-by: Lawrence McAfee <85179052+lmcafee-nvidia@users.noreply.github.com> Co-authored-by: wdykas <73254672+wdykas@users.noreply.github.com> Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
…ention (NVIDIA#5011) Signed-off-by: Hongxiao Bai <hongxiaob@nvidia.com>
Signed-off-by: tailaim <tailaim@nvidia.com>
Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com> Co-authored-by: Kaixiang Lei <5780122+shyoshyo@users.noreply.github.com>
Signed-off-by: guihong-nv <guihongl@nvidia.com>
NVIDIA#5388) Signed-off-by: pingtianl <pingtianl@nvidia.com> Signed-off-by: Pingtian Li <pingtianl@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Yan Bai <bayan@nvidia.com>
Signed-off-by: HaochenYuan <haocheny@nvidia.com>
…IA#3282) Signed-off-by: Yuzhong Wang <yuzhongw@nvidia.com>
Signed-off-by: kunlunl <kunlunl@nvidia.com>
…it-recompute RoPE: cast pid_m (program_id) to int64 in all 4 fused_mla_yarn_rope_apply kernels so seq_index*stride (H*D=36864) doesn't overflow int32 at seq>58k (131k -> 4.83e9). Fixes cudaErrorIllegalAddress in the attention backward. split-recompute: recompute_split_attn_mlp staged recompute (attn and MoE backward never co-reside) to cut the first-backward memory peak.
* feat(dsa): GLM-5.2 IndexShare config + helpers on the new DSA base Add dsa_indexer_topk_freq / dsa_indexer_skip_topk_offset config (cross-layer top-k sharing) with "dsa"-variant validation, and the is_dsa_skip_topk_layer / source_dsa_compute_layer helpers. Leaves the new base's dsv4_hybrid path intact. Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * feat(dsa): fused absorbed-MLA DSA via dsa_kernels (DSAttentionFused) Add DSAttentionFused, a fused DSA core for the GLM-5.2 absorbed-MLA path: frozen-indexer top-k (indexer_topk) + FlashMLA sparse attention (dsa_sparse_attn), with cross-layer top-k sharing (IndexShare). Wire it as the apply_dsa_kernel_fusion branch of get_dsa_module_spec_for_backend alongside AbsorbedMLASelfAttention. Vendor the combined-kv AbsorbedMLASelfAttention from the trainers-main-dsv4-forward line (MQA matrix absorption: K up-proj folded into the query, V up-proj applied after core attention; combined linear_kv_up_proj split at runtime). On trainers-main this class was an unwired orphan with split k/v fields and no bridge mapping; the combined form maps 1:1 from HF kv_b_proj and is what the GLM-5.2 fused path requires. Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * feat(dsa): fold LoRA adapter into the absorbed kv up-projection AbsorbedMLASelfAttention consumes linear_kv_up_proj as a raw weight during matrix absorption (K's up-proj folded into the query, V's applied after core attention) instead of calling its forward, so a LoRA adapter on that module would otherwise be ignored. Add _effective_kv_up_weight(): when the module is LoRA-wrapped (duck-typed AdapterWrapper, to avoid a megatron.core -> megatron.bridge dependency) return W_base + scale * (B @ A) and feed that into the absorption einsums, so the adapter trains (grads flow to A/B) and serves consistently. Restricted to TP=1 (the fused DSA path already enforces it). Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * refactor(dsa): isolate GLM-5.2 fused DSA into additive files Keep GLM-5.2 support additive so it does not edit the actively-developed upstream DSA modules, minimizing rebase conflicts against NVIDIA dev. - glm_dsa_fused.py (new): DSAttentionFused + IndexShare helpers + the GLM fused-attention spec builder, importing shared primitives from dsa/dsa_kernels. - glm_absorbed_mla.py (new): GlmAbsorbedMLASelfAttention, folding the LoRA adapter into the kv up-projection effective weight via a _kv_up_proj_weight override. - dsa.py: DSAttentionFused + helpers removed (now byte-identical to base). - transformer_config.py: only the two IndexShare fields remain, declared so the GLM bridge's values survive the provider->config conversion. - absorbed_mla.py: _effective_kv_up_weight replaced by a minimal _kv_up_proj_weight seam. - module_specs.py: GLM fused branch -> build_glm_dsa_fused_attention_spec. Validated on 4xB200: 65k forward-backward PP16/EP2 loss=0.5747 (warmup 1.088 matches pre-refactor 1.089). Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * build(dsa): pin nvidia-cudnn-frontend[cutedsl]>=1.25.0 The fused DSA path imports the cuDNN-frontend DSA namespace (cudnn.DSA.*), which is only present in nvidia-cudnn-frontend>=1.25.0 with the cutedsl extra; 1.24.x ships NSA but not DSA. The previous unpinned spec resolved to 1.24.0, silently breaking the fused GLM-5.2 attention kernels at import time. Validated end-to-end on a 5-node B200 cluster: full forward/backward, optimizer steps, LoRA weight-sync, and sampling on the fused DSA path (apply_dsa_kernel_fusion=True). uv.lock regenerated with uv 0.8.22; dev extra now resolves cudnn-frontend 1.25.0 plus the cutedsl transitive deps (nvidia-cutlass-dsl 4.5.0, torch-c-dlpack-ext), lts extra unchanged (no DSA). Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * build(dsa): point flash_mla source at FlashMLA nv_dev (sparse fwd) The fused GLM-5.2 DSA path imports flash_mla_sparse_fwd, which only exists on FlashMLA's nv_dev branch. The previous rev (9edee0c, main) imports but lacks that symbol, so the trainer crashes at the first DSA sparse forward with ImportError. Point the source at nv_dev (b7643bd5). (This [tool.uv.sources] entry is the source of truth; the deployed trainer image vendors a prebuilt nv_dev wheel built with nvcc>=12.9 for sm100 — see basetenlabs/trainers server pyproject.) Signed-off-by: Paras Stefanopoulos <paras@parsed.com> * Use declared IndexShare config fields Signed-off-by: Paras Stefanopoulos <paras@parsed.com> --------- Signed-off-by: Paras Stefanopoulos <paras@parsed.com>
… linear dgrad
For a 3D grad_output, .matmul(weight) can be dispatched to a batched-GEMM
whose strideA argument is stored as int32 in the cuBLAS API. When
grad_output is a non-contiguous view (e.g. Megatron's standard [s, b, h]
layout on a [b, s, h]-contiguous storage), torch cannot collapse it to
2D without a copy and falls back to bmm. At long sequence and large
out-per-partition the resulting strideA = seq_len * out_per_partition
exceeds INT32_MAX and cuBLAS raises:
RuntimeError: at::cuda::blas::bgemm<at::BFloat16> argument ldb must
be positive and less than 2147483647 but got 2860646400
Repro: a frozen LM head under LoRA at seq=46080, vocab=248320, TP=4
(strideA = 46080 * 62080 = 2,860,646,400 > 2^31 - 1).
Flatten the leading dims into the M axis before the matmul so torch
routes through a single regular GEMM. The common Megatron-layout case
recovers the underlying [b, s, h]-contiguous view via a free
.transpose(0, 1) and the subsequent reshape becomes a pure view; for
any other 3D non-contiguous layout, fall back to an explicit reshape
that calls .contiguous() internally. The 2D path is unchanged.
Adds tests/unit_tests/tensor_parallel/test_layers.py
::test_LinearWithFrozenWeight_3d_non_contiguous_grad_output to defend
the dispatch path (the overflow itself only fires at sizes too large
for unit-test memory budgets; the test exercises the new code path at
small sizes against the same non-contiguous layout shape).
Signed-off-by: Kimbrian <kimbrian@parsed.com>
…fig (#6) GLM-5.2 declares indexer_rope_interleave: true (GPT-J interleaved rope in the DSA indexer; vLLM reads it as is_neox_style = not flag). The indexer's _apply_rope hardcoded the DeepSeek-V3.2 convention (non-interleaved), so the trainer scored its indexer on differently-rotated q/k than serving engines. Top-k selections still overlap at small candidate pools but diverge progressively with sequence length: measured trainer<->vLLM per-token logprob KL grows 0.010 -> 0.084 between 9k and 15k tokens, making long-context RL unusable. With the convention plumbed (dsa_indexer_rope_interleave, set by the GLM bridge from the checkpoint config), the same probe measures 0.009 at 15k — flat in length, zero systematic bias. Default False preserves DeepSeek-V3.2 behavior exactly.
…IA#5087) into trainers-main Brings in the dsv4-next stack the GLM-5.2 131k CP work builds on, most importantly kunlunl's DeepSeek-v4 Context Parallel support (NVIDIA#5087), plus THD packed-sequence support for DSv4 hybrid attention and related DSA fixes. Conflict resolution (pyproject.toml / uv.lock): take upstream's changes (flash-linear-attention>=0.4.2, flash_mla in no_pypi_wheels + no-build-isolation, cudnn-frontend git source pin, fusions coverage omit) but re-apply the two intentional Baseten deviations that a wholesale upstream take would have reverted: - nvidia-cudnn-frontend[cutedsl]>=1.25.0 version floor - FlashMLA pinned to commit b7643bd (not the floating nv_dev ref) with the comment explaining why the fused GLM-5.2 DSA path needs it. nv_dev already resolved to b7643bd in the lock, so the pinned commit is identical to what the branch was tested against; only the requested-rev label changed. Verified with 'uv lock --locked'. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
|
Superseded by the upstream rebase in |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Fast-forward
trainers-mainonto a newer NVIDIAdevsnapshot (upstreambfa33263) via a single merge commit, so the fork carries the dsv4-next stack the GLM-5.2 131k CP work is built on. Most importantly this brings in:build_flat_topk_idxs, etc.)devsince the June-23 baseAll 5 existing
[baseten]patches ontrainers-mainare preserved on top of the merge.Why split this out
This is the upstream-sync half of the GLM-5.2 131k work, separated from the actual GLM changes so each is reviewable on its own. The GLM adapter PR (
jackrao/glm-dsa-cp) has been rebased on top of this branch and now shows only the Baseten commits.Conflict resolution (
pyproject.toml/uv.lock)Took upstream's changes (
flash-linear-attention>=0.4.2,flash_mlainno_pypi_wheels+no-build-isolation, cudnn-frontend git source pin, fusions coverage omit) but re-applied the two intentional Baseten deviations a wholesale upstream take would have reverted:nvidia-cudnn-frontend[cutedsl]>=1.25.0version floorb7643bd(not the floatingnv_devref), with the comment explaining the fused GLM-5.2 DSA path needsflash_mla_sparse_fwdnv_devalready resolved tob7643bdin the lock, so the pinned dependency is byte-identical to what the branch was tested against — only the requested-rev label changed. Verified withuv lock --locked.🤖 Generated with Claude Code