Repository navigation
[Refactor] Keep only the TP and PP groups on the model runner - #41813
Merged
Merged
Conversation
ch-wan
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Qiaolin-Yu,
Ying1123,
hebiao064,
hnyls2002,
iforgetmyname,
ispobock,
merrymercy,
ping1jing2,
whybeyoung,
xiezhq-hermann and
yeahdongcn
as code owners
September 30, 2026 03:07
This was referenced Sep 30, 2026
ch-wan
requested review from
1am9trash,
Alisehen,
AniZpZ,
ByronHsu,
Duyi-Wang,
FlamingoPg,
OrangeRedeng,
ShangmingCai,
YAMY1234,
alphabetc1,
b8zhong,
huangtingwei9988,
hubertlu-tw,
hzh0425,
kkHuang-amd,
mickqian,
mmangkad,
rainj-me,
sogalin,
yhyang201 and
yuan-luo
as code owners
September 30, 2026 04:24
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 04:24
eea0f19 to
e5c3fef
Compare
ch-wan
force-pushed
the
cheng/refactor/drop-unread-model-placement
branch
from
September 30, 2026 05:05
791b0e2 to
134ee9d
Compare
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 05:05
e5c3fef to
b996c34
Compare
ch-wan
force-pushed
the
cheng/refactor/drop-unread-model-placement
branch
from
September 30, 2026 20:27
134ee9d to
afcf208
Compare
Base automatically changed from
cheng/refactor/drop-unread-model-placement
to
main
September 30, 2026 20:27
`ModelRunner.init_torch_distributed` copied fifteen placement values onto the runner. Only two of them have to outlive the scope the runner is built in: speculative workers re-enter a draft's scope through its TP group, and draft forwards read its PP group without the pipeline scope. Keep those two and drop the other thirteen (`attention_tp_group`, `tp_rank`, `tp_size`, `dp_size`, `attn_dp_size`, `pp_rank`, `pp_size`, `attn_cp_rank`, `attn_cp_size`, `attn_dcp_rank`, `attn_dcp_size`, `moe_ep_size`, `dp_rank`). Each reader now takes the value from where it runs: - readers that run while the runner is built, during a forward or a graph capture, or only on the target read `get_parallel()`, which in those scopes is what the runner saw when it copied the value; - readers that can run on a draft runner outside its scope read the runner's own groups: the KV cache configurator (built while pools are allocated) and the PP proxy hidden size (also reached when graphs are recaptured) use `pp_group.world_size` / `rank_in_group`, and the DFLASH / DSpark rank-0 log gates inside the draft scope read the target runner's `tp_group.rank_in_group`; - the HiSparse coordinator, built with the pools, reads `get_parallel().attn_tp_group`. Outside the scope that is the target's attention-TP group, which is the group an attention-owning draft runs on; a draft that does not own attention never narrows it. The runner-owned `tp_size` / `tp_rank` arguments to `build_load_config`, `dist_barrier_after_load`, the tensor-dump hook, `apply_torch_tp` and the quantized-MoE check now come from the context too. All of them run in the runner's constructor, except the barrier after an overlapped startup load, which the scheduler runs on the target's runners outside any draft scope. No value a reader sees changes. Tests that stood in for the retired fields follow: the MLX runner stub tests publish their attention-DP width, the DFLASH sampler and phase-1 tests give the fake a TP group and publish a topology, fake runners in the attention test kits and backend tests drop placement fields nothing reads, and the placement-freeze census asserts it finds the TP and PP groups.
ch-wan
force-pushed
the
cheng/refactor/runner-placement-snapshots
branch
from
September 30, 2026 20:27
b996c34 to
aecc187
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR is part of a stack (oldest at bottom):
Motivation
ModelRunner.init_torch_distributedcopies fifteen placement values onto the runner. Only two of them have to outlive the scope the runner is built in:The other thirteen are second copies of values that
get_parallel()already answers wherever they are read.Modifications
tp_groupandpp_group. Drop the other thirteen:attention_tp_group,tp_rank,tp_size,dp_size,attn_dp_size,pp_rank,pp_size,attn_cp_rank,attn_cp_size,attn_dcp_rank,attn_dcp_size,moe_ep_size,dp_rank.get_parallel(). In those scopes it answers what the runner copied.pp_group.world_size/rank_in_group. The DFLASH / DSpark rank-0 log gates inside the draft scope read the target runner'stp_group.rank_in_group.get_parallel().attn_tp_group. Outside the scope that is the target's attention-TP group, which is also the group an attention-owning draft runs on; a draft that does not own attention never narrows it.tp_size/tp_rankarguments tobuild_load_config,dist_barrier_after_load, the tensor-dump hook,apply_torch_tpand the quantized-MoE check come from the context too. All of these run in the runner's constructor, except the barrier after an overlapped startup load, which the scheduler runs on the target's runners outside any draft scope.No value a reader sees changes. Code outside the tree that reads, for example,
model_runner.tp_rankshould readget_parallel().tp_rank.Accuracy Tests
H200:
--random-seed 42,--tp-size 2and--tp-size 2 --dp-size 2 --enable-dp-attention: 4 greedy prompts × 32 tokens are identical to the parent commit.spec_verify_ctare identical tomain. Measured at the top of this stack, which contains this PR.mlxmodule, onmainand on this PR.test/registered/unitat this PR's head, compared withmain: no new failures.Speed Tests and Profiling
Not applicable.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #36772756644
Latest PR Test (Extra): 🚫 Run #36772756255
Latest PR Test (AMD ROCm 10): 🚫 Run #36772756706