Repository navigation
[Refactor] Drop placement values nothing reads - #41811
Merged
Merged
Conversation
ch-wan
requested review from
1am9trash,
Alisehen,
AniZpZ,
BBuf,
ByronHsu,
Duyi-Wang,
Edwardf0t1,
FlamingoPg,
Fridge003,
HaiShaw,
OrangeRedeng,
Qiaolin-Yu,
ShangmingCai,
YAMY1234,
Ying1123,
b8zhong,
hebiao064,
hnyls2002,
hubertlu-tw,
ispobock,
kkHuang-amd,
kpham-sgl,
merrymercy,
mickqian,
mmangkad,
pyc96,
rainj-me,
xiezhq-hermann,
yhyang201 and
yuan-luo
as code owners
September 30, 2026 03:07
This was referenced Sep 30, 2026
ch-wan
force-pushed
the
cheng/hot-fix/solar-model-construction
branch
from
September 30, 2026 04:24
3280536 to
4f20ec7
Compare
ch-wan
requested review from
alexnails,
alphabetc1,
hanming-lu,
huangtingwei9988,
hzh0425,
liusy58 and
yizhang2077
as code owners
September 30, 2026 04:24
ch-wan
force-pushed
the
cheng/refactor/drop-unread-placement
branch
from
September 30, 2026 04:24
52fc839 to
02fdd93
Compare
ch-wan
force-pushed
the
cheng/hot-fix/solar-model-construction
branch
from
September 30, 2026 05:05
4f20ec7 to
117af72
Compare
ch-wan
force-pushed
the
cheng/refactor/drop-unread-placement
branch
from
September 30, 2026 05:05
02fdd93 to
87b938b
Compare
ch-wan
force-pushed
the
cheng/hot-fix/solar-model-construction
branch
from
September 30, 2026 20:26
117af72 to
15cf773
Compare
Base automatically changed from
cheng/hot-fix/solar-model-construction
to
main
September 30, 2026 20:26
Remove parallel ranks, sizes and groups that are stored or passed but never read: - linear layers passed `tp_size` to `load_merged_column_weight`, which no parameter class reads; the W8A8 INT8 MoE method, the humming linear weight, the FlashMLA backend (`dcp_rank`), the DSA indexer (`cp_size`) and the Triton / Flash3 vision attention backends (`tp_size`) kept values with no reader; - `VocabParallelEmbedding.get_sharded_to_full_mapping` has no caller; - `BaseRunner` stored `tp_size` and `attn_tp_rank`, the decode and prefill graph runners re-stored `attn_tp_size`, `attn_tp_rank` and `dp_size` right after `BaseRunner.__init__` set them, and four draft graph runners copied `model_runner.tp_size`; none of these is read; - the n-gram worker kept an unread `tp_rank`; - `SchedulerRequestReceiver` carried `attn_tp_group` and `attn_cp_group` that it never reads (their CPU groups stay); - `PrefillBootstrapQueue` took and stored `gloo_group` and `tp_size`, and `DecodePreallocQueue` stored `tp_size` and `dp_size`, without reading them. No read value changes.
ch-wan
force-pushed
the
cheng/refactor/drop-unread-placement
branch
from
September 30, 2026 20:26
87b938b to
80e6449
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR is part of a stack (oldest at bottom):
Motivation
A number of classes store or pass parallel ranks, sizes and groups that nothing reads. Each one reads as if it mattered, and each is one more copy of the placement to keep consistent.
Modifications
Remove the unread values:
tp_sizetoload_merged_column_weight, which no parameter class reads; the W8A8 INT8 MoE method, the humming linear weight, the FlashMLA backend (dcp_rank), the DSA indexer (cp_size) and the Triton / Flash3 vision attention backends (tp_size) kept values with no reader;VocabParallelEmbedding.get_sharded_to_full_mappinghas no caller;BaseRunnerstoredtp_sizeandattn_tp_rank; the decode and prefill graph runners re-storedattn_tp_size,attn_tp_rankanddp_sizeright afterBaseRunner.__init__set them; four draft graph runners copiedmodel_runner.tp_size. None of these is read;tp_rank;SchedulerRequestReceivercarriedattn_tp_groupandattn_cp_groupthat it never reads (their CPU groups stay);PrefillBootstrapQueuetook and storedgloo_groupandtp_size, andDecodePreallocQueuestoredtp_sizeanddp_size, without reading them.No read value changes.
Accuracy Tests
Not applicable: nothing that is read changes.
test/registered/unitat this PR's head, compared withmain: no new failures.Speed Tests and Profiling
Not applicable.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #36772686757
Latest PR Test (Extra): 🚫 Run #36772686777
Latest PR Test (AMD ROCm 10): 🚫 Run #36772686680