Skip to content

[Refactor] Stop passing models' layers the placement they already read - #41815

Merged
ch-wan merged 1 commit into
mainfrom
cheng/refactor/models-drop-same-leaf-placement
Sep 30, 2026
Merged

ch-wan merged 1 commit into
mainfrom
cheng/refactor/models-drop-same-leaf-placement

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is part of a stack (oldest at bottom):

Motivation

Several model modules hand a layer exactly the placement that the layer resolves by itself. That reads as a deliberate choice of group when there is none.

Modifications

  • Grok's attention passed tp_rank / tp_size (in locals misnamed attn_tp_*) to its QKV and output projections, whose defaults are those values. Keep only the TP width for the head arithmetic.
  • ZAYA's attention passed its stored TP rank and size to CCA and to o_proj, which default to the same values. Its tp_rank copy is then unread and goes too.
  • IQuest-Q1 passed kv_tp_rank / kv_tp_size equal to the tp_rank / tp_size of the same call, which is what QKV defaults them to.
  • GLM-5 Next gave sharded_weight_loader a getter for the attention-TP rank, which is what the loader reads when no getter is given.
  • The DSpark draft attention passed MqaAttentionBase the attention-TP rank and size it defaults to, and no other caller sets them. Drop the two parameters and keep the attributes, which decode-time attention TP still adjusts.

No layer sees a different value.

Accuracy Tests

Not applicable: no layer receives a different value.

  • All modified modules import.
  • test/registered/unit at this PR's head, compared with main: no new failures.

Speed Tests and Profiling

Not applicable.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): 🚫 Run #36772845494
Latest PR Test (Extra): 🚫 Run #36772845189
Latest PR Test (AMD ROCm 10): 🚫 Run #36772845538

@ch-wan
ch-wan force-pushed the cheng/refactor/moe-loaders-read-own-rank branch from a89e8fe to 2917852 Compare September 30, 2026 04:24
@ch-wan
ch-wan force-pushed the cheng/refactor/models-drop-same-leaf-placement branch from 4b8856c to 6b0dfee Compare September 30, 2026 04:24
@ch-wan
ch-wan force-pushed the cheng/refactor/moe-loaders-read-own-rank branch from 2917852 to 73420b3 Compare September 30, 2026 05:05
@ch-wan
ch-wan force-pushed the cheng/refactor/models-drop-same-leaf-placement branch from 6b0dfee to a5f7cf4 Compare September 30, 2026 05:06
@ch-wan
ch-wan force-pushed the cheng/refactor/moe-loaders-read-own-rank branch from 73420b3 to 9035759 Compare September 30, 2026 20:27
Base automatically changed from cheng/refactor/moe-loaders-read-own-rank to main September 30, 2026 20:28
Several model modules handed a layer exactly the placement the layer
resolves by itself:

- Grok's attention passed `tp_rank` / `tp_size` (in locals misnamed
  `attn_tp_*`) to its QKV and output projections, whose defaults are
  those values; keep only the TP width for the head arithmetic.
- ZAYA's attention passed its stored TP rank and size to `CCA` and to
  `o_proj`, which default to the same values; its `tp_rank` copy is then
  unread and goes too.
- IQuest-Q1 passed `kv_tp_rank` / `kv_tp_size` equal to the `tp_rank` /
  `tp_size` of the same call, which is what QKV defaults them to.
- GLM-5 Next gave `sharded_weight_loader` a getter for the attention-TP
  rank, which is what the loader reads when no getter is given.
- The DSpark draft attention passed `MqaAttentionBase` the attention-TP
  rank and size it defaults to, and no other caller sets them; drop the
  two parameters and keep the attributes, which decode-time attention TP
  still adjusts.

No layer sees a different value.
@ch-wan
ch-wan force-pushed the cheng/refactor/models-drop-same-leaf-placement branch from a5f7cf4 to eccd9b9 Compare September 30, 2026 20:28
@ch-wan
ch-wan merged commit 00659ac into main Sep 30, 2026
9 of 18 checks passed
@ch-wan
ch-wan deleted the cheng/refactor/models-drop-same-leaf-placement branch September 30, 2026 20:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant