Repository navigation
[Refactor] Let FusedMoE's weight-loading helpers read the layer's MoE-TP rank - #41814
Merged
Merged
Conversation
ch-wan
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Ying1123,
ispobock and
merrymercy
as code owners
September 30, 2026 03:07
This was referenced Sep 30, 2026
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 04:24
eea0f19 to
e5c3fef
Compare
ch-wan
requested review from
Qiaolin-Yu,
alexnails,
fzyzcjy,
hanming-lu,
hnyls2002,
kpham-sgl,
liusy58,
pyc96 and
xiezhq-hermann
as code owners
September 30, 2026 04:24
ch-wan
requested review from
1am9trash,
Alisehen,
ByronHsu,
Duyi-Wang,
OrangeRedeng,
ShangmingCai,
YAMY1234,
b8zhong,
hebiao064,
hubertlu-tw,
iforgetmyname,
kkHuang-amd,
mickqian,
mmangkad,
ping1jing2,
rainj-me,
sogalin,
whybeyoung,
yeahdongcn,
yhyang201 and
yuan-luo
as code owners
September 30, 2026 04:24
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 04:24
a89e8fe to
2917852
Compare
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 05:05
e5c3fef to
b996c34
Compare
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 05:05
2917852 to
73420b3
Compare
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 20:27
b996c34 to
aecc187
Compare
Base automatically changed from
cheng/refactor/runner-placement-snapshots
to
main
September 30, 2026 20:27
The seven loading helpers (`_load_w13`, `_load_w2`, the scale and g_idx loaders, the GGUF and FP8-shared-expert paths) each took a `tp_rank` argument. Every call is inside `FusedMoE`, and every value traces back to `tp_rank = self.moe_tp_rank` in the two weight loaders. Read `self.moe_tp_rank` in the helpers and drop the argument and the two locals. A layer that shards differently stores that in `self.moe_tp_rank` itself -- the Inkling shared experts and their loading helper keep the full-TP rank there -- so the helpers see the same value as before.
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 20:27
73420b3 to
9035759
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR is part of a stack (oldest at bottom):
Motivation
The seven
FusedMoEloading helpers (_load_w13,_load_w2, the scale andg_idxloaders, the GGUF and FP8 shared-expert paths) each take atp_rankargument. Every call is insideFusedMoE, and every value traces back totp_rank = self.moe_tp_rankin the two weight loaders.Modifications
Read
self.moe_tp_rankin the helpers, and drop the argument and the two locals.A layer that shards differently stores that choice in
self.moe_tp_rankitself: the Inkling shared experts and their loading helper keep the full-TP rank there. So the helpers see the same value as before.Accuracy Tests
H200, compared with
mainat the top of this stack, which contains this PR; 4 greedy prompts × 64 tokens and GSM8K (200 questions):--tp 2 --ep-size 2moe_wna16)--tp 2moe_wna16)--tp 2 --ep-size 2main/ 0.935 (the load fix is #41807, below in this stack)unit/models,unit/layers/moe,unit/layers/quantizationandtest_runtime_context: the same results as the parent commit.test/registered/unitat this PR's head, compared withmain: no new failures.Speed Tests and Profiling
Not applicable.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #36772813010
Latest PR Test (Extra): 🚫 Run #36772812682
Latest PR Test (AMD ROCm 10): ❌ Run #36772813054