Skip to content

Check the topology identities where the layout is written, and build at the published widths - #40340

Merged
ch-wan merged 5 commits into
mainfrom
cheng/refactor/topology-identities
Sep 21, 2026
Merged

ch-wan merged 5 commits into
mainfrom
cheng/refactor/topology-identities

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

get_parallel().override(...) validated key names and nothing else. A
self-contradicting topology -- a rank at or past its width, a tp_size that
does not factor into the attention widths, a group whose width is not the one
configured -- went in quietly and surfaced much later as a hang or a wrong
answer in a collective, with nothing pointing back at the write.

At the other end, initialize_model_parallel took seven widths as arguments,
so every caller translated the published configuration into them again: the
scheduler from a per-runner record, the weight-cache daemon from its own
fields, the media encoder from neither. Three translations of one configuration
are three places for it to come out different, and the encoder's did -- it
built a layout the context never described.

Modifications

One set of identities, checked wherever the layout is written. Every rank
inside its width; tp_size factoring into the attention triple and into the
MoE triple; the attention rank layout the attention ranks are derived from; and
each built group's width equal to the configured one. They hold
unconditionally, on the four paths that write the namespace and at the group
build. The caller owns the arithmetic: stating one leaf without the quotients
that follow from it describes no real layout, so it is refused rather than
carried. Names that cannot be read are skipped -- a process that has published
nothing can still stamp a rank.

Two callers were stating a width the rest of the namespace disagreed with. The
shared-expert scope narrows moe_ep_size to one without sharding anything, so
it now says there is no expert-parallel group rather than leaving the wider one
installed. Publish stamped the placement in two calls, and the moment between
them described a half-placed process; it stamps once.

The group build reads the widths it builds at. The seven width parameters
are gone; the function reads them from the context. What is left in the
signature is not topology: the backend is decided by the device, two flags
belong to other namespaces, and the join parameters describe this particular
call. The encoder states the layout it has always built -- tp_size ranks
wide, no pipeline, no expert or MoE-DP dimension, no decode context parallelism
-- so what it answers and what it builds are the same thing again.

An elastic scale-up and the launch topology stop sharing two names. A
scale-up admits ranks into a WORLD that was pre-allocated to --max-ep-size;
it does not rebuild the process groups, which keep the width they were
constructed with. Writing the expanded replica count over attn_dp_size and
attn_dp_rank therefore left the launch names answering for groups of a
different size, and the identities say so: the rank lands outside its width and
tp_size stops factoring into the attention triple. elastic_dp_size and
elastic_dp_rank name the expanded replica set, each answering with its launch
counterpart until a scale-up moves it, and the readers that need the expanded
one -- the WORLD gather width, DP padding, KV routing, every place a gathered
list is indexed by replica, and KV-event sharding -- ask for it by name. That
includes the gather and scatter slices themselves: the list a scale-up gathers
spans WORLD, so indexing it by this process's place among the launch replicas
selects another rank's rows. The pair is range-checked like any other; it stays out of the identities
that describe the groups, because it does not describe them.

A draft on one context shard has no context-parallel communicator, and the
sampler asks for one only when there is more than one shard to reconcile --
otherwise the scope's answer, that there is no such group, is dereferenced.

The draft scope states a group for each width it narrows. It already said its
attention-TP width is the group it installs; it now hands over that group, and
says there is no attention-CP group, the way it already said there is no
expert-parallel one. The shared-expert scope states moe_dp_size for the same
reason -- narrowing two of the three MoE factors and leaving the third is a
triple that does not multiply out.

With the build checking what it made against what was configured,
get_tensor_model_parallel_world_size and its attention sibling answer the same
question as the configured widths rather than a second one, so their eight
readers move to the context.

Accuracy Tests

Not run for this revision. The build now reads the widths the previous code
passed it, and the encoder states the layout it already built, so no
configuration changes what is constructed.

Speed Tests and Profiling

Not applicable. The identity check is integer comparison on names already in
hand; it runs where a topology is written, not in a forward.

Checklist

Each identity is injected in the direction that breaks it and in the direction
that keeps it, because a guard that only ever fires is as uninformative as one
that never does. Twenty-one tests that built a topology by passing widths
publish it first, and ten benchmark and example entries do the same. The
scale-up test now publishes a configuration before it asserts: without one
every term was unreadable and the check it claimed to make was skipped.

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #35644382491
Latest PR Test (Extra): ❌ Run #35644381975
Latest PR Test (AMD ROCm 10): ❌ Run #35644382214

@github-actions github-actions Bot added amd Multi-modal multi-modal language model labels Sep 19, 2026
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch from 209d8fd to b433d52 Compare September 19, 2026 20:29
@github-actions github-actions Bot added the quant LLM Quantization label Sep 19, 2026
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch 2 times, most recently from 171a2a2 to 181497e Compare September 19, 2026 23:24
@ch-wan
ch-wan force-pushed the cheng/refactor/draft-scope-and-scheduler-reads branch from 2fe69c7 to dd50d6a Compare September 20, 2026 09:18
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch from b1c32c1 to 9097ca2 Compare September 20, 2026 09:18
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch from 9097ca2 to 9ab25ed Compare September 20, 2026 11:59
@ch-wan
ch-wan force-pushed the cheng/refactor/draft-scope-and-scheduler-reads branch from dd50d6a to 297fbc5 Compare September 21, 2026 11:29
@ch-wan
ch-wan requested a review from pyc96 as a code owner September 21, 2026 11:29
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch 3 times, most recently from 2c6baf4 to 45fe302 Compare September 21, 2026 12:56
@ch-wan
ch-wan force-pushed the cheng/refactor/draft-scope-and-scheduler-reads branch from 297fbc5 to 04e8f25 Compare September 21, 2026 19:05
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch from 45fe302 to 5a38ad3 Compare September 21, 2026 19:11
Base automatically changed from cheng/refactor/draft-scope-and-scheduler-reads to main September 21, 2026 19:19
`override()` validated key names and nothing else, so a self-contradicting
topology -- a rank at or past its width, a `tp_size` that does not factor into
the attention widths, a group whose width is not the one configured -- went in
quietly and surfaced much later as a hang or a wrong answer in a collective,
with nothing pointing back at the write.

The three identities now hold unconditionally on every path that writes the
namespace and at the group build. The caller owns the arithmetic: stating one
leaf without the quotients that follow from it describes no real layout, so it
is refused rather than carried. Names that cannot be read are skipped -- a
process that has published nothing can still stamp a rank, and a group that has
not been built answers nothing at all.

Two callers were stating a width the rest of the namespace disagreed with. The
draft's shared-expert scope narrows `moe_ep_size` to one without sharding
anything, so it now says there is no expert-parallel group rather than leaving
the wider one installed. Publish stamped the placement in two calls, and the
moment between them described a half-placed process; it stamps once.
`initialize_model_parallel` took seven widths as arguments, so every caller
translated the published configuration into them again: the scheduler from the
per-runner record, the weight-cache daemon from its own fields, the media
encoder from neither. Three translations of one configuration is three places
for it to come out different, and the encoder's did -- it built a layout the
context never described.

The widths now come from the context. What is left in the signature is not
topology: the backend is decided by the device, two flags belong to other
namespaces, and the join parameters describe this particular call rather than
the layout being joined.

The encoder states the layout it has always built -- `tp_size` ranks wide, no
pipeline, no expert or MoE-DP dimension, no decode context parallelism -- so
what it answers and what it builds are the same thing again.

Tests that built a topology by passing widths now publish it first, through the
same door production uses.
…sors

The widths are quotients of one another in two more ways than the attention
triple: `moe_tp_size` is what is left of `tp_size` after the expert and MoE-DP
dimensions, and the attention ranks are derived from `tp_rank` through a fixed
layout. Both relations already existed as one-directional arithmetic; stating
them as identities means a write that contradicts either is refused where it
happens.

Making them hold turned up one caller that was not saying what it meant: the
draft's tensor-parallel scope narrowed `tp_size` to the group it installs while
leaving the MoE widths on the target's answers. The draft runs the whole model
on that one group and has no expert dimension there -- the one worker that runs
a MoE draft declines this scope for exactly that reason -- so the scope says so.

With the group build now checking what it made against what was configured,
`get_tensor_model_parallel_world_size` and its attention sibling answer the
same question as the configured widths rather than a second one, so their eight
readers move to the context. `get_shared_experts_tp_group` gets a context name
and its one reader follows. A guard derives the accessor list from the source
and fails on any caller outside the package that defines them; the three that
remain are not topology and say why.
@ch-wan
ch-wan force-pushed the cheng/refactor/topology-identities branch from 5a38ad3 to 2a5ffc0 Compare September 21, 2026 19:22
@ch-wan
ch-wan merged commit 2d0e94e into main Sep 21, 2026
87 of 102 checks passed
@ch-wan
ch-wan deleted the cheng/refactor/topology-identities branch September 21, 2026 19:23
hdt98 added a commit to hdt98/sglang that referenced this pull request Sep 26, 2026
The topology check from sgl-project#40340 rejects moe_ep_size=4 with tp_size=1.
Same change as sgl-project#41002.

Co-authored-by: xinguozhu-2026 <xinguo.zhu@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd Multi-modal multi-modal language model quant LLM Quantization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant