Repository navigation
[Refactor] Stop passing models' layers the placement they already read - #41815
Merged
Merged
Conversation
This was referenced Sep 30, 2026
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 04:24
a89e8fe to
2917852
Compare
ch-wan
requested review from
Fridge003,
Qiaolin-Yu,
Ying1123,
alexnails,
fzyzcjy,
hanming-lu,
hnyls2002,
hzh0425,
ispobock,
kpham-sgl,
liusy58,
merrymercy,
pyc96,
xiezhq-hermann and
yizhang2077
as code owners
September 30, 2026 04:24
ch-wan
requested review from
1am9trash,
ByronHsu,
Duyi-Wang,
ShangmingCai,
YAMY1234,
hebiao064,
hubertlu-tw,
iforgetmyname,
kkHuang-amd,
mickqian,
ping1jing2,
rainj-me,
sogalin,
whybeyoung,
yeahdongcn,
yhyang201 and
yuan-luo
as code owners
September 30, 2026 04:24
ch-wan
force-pushed
the
cheng/refactor/models-drop-same-leaf-placement
branch
from
September 30, 2026 04:24
4b8856c to
6b0dfee
Compare
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 05:05
2917852 to
73420b3
Compare
ch-wan
force-pushed
the
cheng/refactor/models-drop-same-leaf-placement
branch
from
September 30, 2026 05:06
6b0dfee to
a5f7cf4
Compare
ch-wan
force-pushed
the
cheng/refactor/moe-loaders-read-own-rank
branch
from
September 30, 2026 20:27
73420b3 to
9035759
Compare
Base automatically changed from
cheng/refactor/moe-loaders-read-own-rank
to
main
September 30, 2026 20:28
Several model modules handed a layer exactly the placement the layer resolves by itself: - Grok's attention passed `tp_rank` / `tp_size` (in locals misnamed `attn_tp_*`) to its QKV and output projections, whose defaults are those values; keep only the TP width for the head arithmetic. - ZAYA's attention passed its stored TP rank and size to `CCA` and to `o_proj`, which default to the same values; its `tp_rank` copy is then unread and goes too. - IQuest-Q1 passed `kv_tp_rank` / `kv_tp_size` equal to the `tp_rank` / `tp_size` of the same call, which is what QKV defaults them to. - GLM-5 Next gave `sharded_weight_loader` a getter for the attention-TP rank, which is what the loader reads when no getter is given. - The DSpark draft attention passed `MqaAttentionBase` the attention-TP rank and size it defaults to, and no other caller sets them; drop the two parameters and keep the attributes, which decode-time attention TP still adjusts. No layer sees a different value.
ch-wan
force-pushed
the
cheng/refactor/models-drop-same-leaf-placement
branch
from
September 30, 2026 20:28
a5f7cf4 to
eccd9b9
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR is part of a stack (oldest at bottom):
Motivation
Several model modules hand a layer exactly the placement that the layer resolves by itself. That reads as a deliberate choice of group when there is none.
Modifications
tp_rank/tp_size(in locals misnamedattn_tp_*) to its QKV and output projections, whose defaults are those values. Keep only the TP width for the head arithmetic.CCAand too_proj, which default to the same values. Itstp_rankcopy is then unread and goes too.kv_tp_rank/kv_tp_sizeequal to thetp_rank/tp_sizeof the same call, which is what QKV defaults them to.sharded_weight_loadera getter for the attention-TP rank, which is what the loader reads when no getter is given.MqaAttentionBasethe attention-TP rank and size it defaults to, and no other caller sets them. Drop the two parameters and keep the attributes, which decode-time attention TP still adjusts.No layer sees a different value.
Accuracy Tests
Not applicable: no layer receives a different value.
test/registered/unitat this PR's head, compared withmain: no new failures.Speed Tests and Profiling
Not applicable.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #36772845494
Latest PR Test (Extra): 🚫 Run #36772845189
Latest PR Test (AMD ROCm 10): 🚫 Run #36772845538