Skip to content

[Feature] Enable GLM-5.3-flash PD disaggregation on Model Runner V1 - #16910

Merged
weijinqian0 merged 2 commits into
vllm-project:mainfrom
sunbaosong:feature/glm5next-pd-mrv1
Sep 22, 2026
Merged

weijinqian0 merged 2 commits into
vllm-project:mainfrom
sunbaosong:feature/glm5next-pd-mrv1

Conversation

@sunbaosong

@sunbaosong sunbaosong commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the kernel-block layout prefract landed for MRV2 (#16755). Without this, PD disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default runner rejects the layout at KV init and the cross-TP handshake fails.

Three adaptation points:

  • V1 reshape: two small-slot view branches mirrored from _reshape_kv_cache_v2 – the per-request kpool tail ring packs a contiguous suffix of the shared small slot, and the compressed indexer packs natural kernel blocks as a contiguous prefix, each bounded to half the slot.
  • GLM-Next cache layout: size the shared small slot by the unified page. _align_glm5_next_cache_specs derived the small page from the small candidates only, so with MRV1's unpadded worker specs the slot collapsed to the indexer's own page and the half-slot bound rejected the layout at KV init. MRV2 workers pre-pad their specs, so their behavior is unchanged.
  • V1 reshape: expose the NOPE main as natural 128-row kernel blocks instead of one page-sized block per scheduler block, so dim0 blocks stay byte-identical across engines with different TP. The page-sized form made the MooncakeConnectorV2 handshake reject cross-TP transfers with different local and remote kernel block size 1152 | 640.

Depends on #16755.

Does this PR introduce any user-facing change?

No new flags, config options, or API changes. For users running GLM-5.3-Flash with PD disaggregation on the default model runner, cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box. All other models and colocated workloads are unaffected – branch selection is identical to before (only GLM-Next publishes the stride marker).

How was this patch tested?

On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8, and P=TP8 → D=DP2xTP4):

  • Indexer and tail views bit-identical with the MRV2 reshape.

  • Equal-TP full-verify passes (Serial + concurrent, all match; per-rank ms-level pulls).

  • Unequal-TP chain: handshake clean, three factual questions all answered correctly, zero pull/transfer errors.

  • TextVQA authoritative rescore (per-row gold alignment, post- extraction): PD 92.0% identical to colocated 90.0% in the same batch.

  • vLLM main: vllm-project/vllm@84030bb

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enables support for GLM-5.3-Flash (GLM-5Next) KV cache geometry within the V1 model runner. It introduces mechanisms to handle compressed indexer caches and per-request tail ring caches, ensuring that memory layout remains consistent with the V2 implementation. These changes include updates to cache configuration, runner initialization, and KV transfer utilities to support the new model architecture.

Highlights

  • GLM-5.3-Flash Support: Implemented support for GLM-5.3-Flash (GLM-5Next) KV cache geometry within the V1 model runner.
  • Reshape Logic: Introduced specialized reshape logic for compressed indexer and tail ring caches to ensure consistent memory layout.
  • Cache Alignment: Aligned cache specifications to ensure consistent slot geometry across different runner implementations.
  • Testing: Added comprehensive unit tests to verify the new reshape logic and cache view geometry.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@sunbaosong sunbaosong changed the title Feature/glm5next pd mrv1 [WIP][Feature] Enable GLM-Next PD disaggregation on Model Runner V1 Sep 18, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support GLM-5.3-Flash kpool cache views in V1 and V2 model runners

Suggested PR Summary:

### What this PR does / why we need it?
This PR implements the GLM-5.3-Flash kpool cache views in both V1 and V2 model runners. It aligns the cache geometry contracts so that the compressed indexer packs natural kernel blocks as a contiguous prefix of the shared small slot, and the per-request tail ring packs a contiguous suffix. It also updates the KV transfer logic to support the new specs.

Several issues were identified in the review:
- Potential `IndexError` when accessing `raw_cache[0]` before checking if the tuple is empty in both `model_runner_v1.py` and `v2/attn_utils.py`.
- A `TypeError` in `v2/attn_utils.py` when attempting integer division on a list instead of its first element.
- Potential `AttributeError`s when directly accessing `model_version` and `non_causal_multi_token_decode` on standard `KVCacheSpec` objects.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with unit tests in `tests/ut/models/test_glm5next_kv_cache.py`, `tests/ut/worker/test_glm5next_pooled_cache.py`, and `tests/ut/worker/test_model_runner_v1_glm5next_reshape.py`.

Comment thread vllm_ascend/worker/model_runner_v1.py Outdated
Comment thread vllm_ascend/worker/v2/attn_utils.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
@sunbaosong
sunbaosong force-pushed the feature/glm5next-pd-mrv1 branch 2 times, most recently from e4465ee to c73b373 Compare September 20, 2026 02:36
@sunbaosong sunbaosong changed the title [WIP][Feature] Enable GLM-Next PD disaggregation on Model Runner V1 [Feature] Enable GLM-Next PD disaggregation on Model Runner V1 and Mooncake Connector V2 Sep 20, 2026
@sunbaosong
sunbaosong marked this pull request as ready for review September 20, 2026 02:52
@sunbaosong sunbaosong changed the title [Feature] Enable GLM-Next PD disaggregation on Model Runner V1 and Mooncake Connector V2 [Feature] Enable GLM-5.3-flash PD disaggregation on Model Runner V1 Sep 20, 2026
Comment thread vllm_ascend/attention/indexer_kpool.py
Comment thread vllm_ascend/attention/indexer_kpool.py
Comment thread vllm_ascend/attention/indexer_kpool.py
Comment thread vllm_ascend/attention/sfa_v1.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
Comment thread vllm_ascend/worker/v2/attn_utils.py
@sunbaosong
sunbaosong force-pushed the feature/glm5next-pd-mrv1 branch 2 times, most recently from 06797e3 to ae72306 Compare September 20, 2026 08:34
@freyfwt freyfwt added the ready-precise run selected e2e test for pr label Sep 20, 2026

@lijiahang226 lijiahang226 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Bring the GLM-5.3-Flash prefill/decode disaggregation to the default
model runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing
the kernel-block layout contract landed for MRV2:

- V1 reshape: mirror the two small-slot view branches from
  _reshape_kv_cache_v2 -- the per-request kpool tail ring packs a
  contiguous suffix of the shared small slot and the compressed indexer
  packs natural kernel blocks as a contiguous prefix, each bounded to
  half the slot.
- GLM-Next cache layout: size the shared small slot by the unified
  page. _align_glm5_next_cache_specs derived the small page from the
  small candidates only, so with MRV1's unpadded worker specs the slot
  collapsed to the indexer's own page and the half-slot bound rejected
  the layout at KV init. MRV2 workers pre-pad their specs, so their
  behavior is unchanged.
- V1 reshape: expose the NoPE main MLA as natural 128-row kernel
  blocks instead of one page-sized block per scheduler block, so dim0
  blocks stay byte-identical across engines with different TP; the
  page-sized form made the MooncakeConnectorV2 handshake reject
  cross-TP transfers with 'different local and remote kernel block
  size 1152 | 640'.

These view branches are model-owned layout logic and live in
vllm_ascend/models/glm5next/cache_views.py behind a single
view_glm5_next_cache entry. The runner calls it only when the layer
specs carry the glm5_next marker (is_glm5_next_cache_spec, mirroring
the is_deepseek_v41_cache gate in the same function) and falls through
to its generic reshape paths when the helper returns None, so
compressed specs owned by the existing V1 page-strided path (e.g.
DSV4) are untouched.

With these views the MooncakeConnectorV2 handshake (derived from the
actually registered caches) transfers at whole-block granularity with
zero connector changes on MRV1.

Verified on GLM-5.3-Flash w8a8 (TP8 single instance, P=TP8 <-> D=TP8,
and P=TP8 -> D=DP2xTP4): indexer and tail views bit-identical with the
MRV2 reshape, full-verify passes on equal TP, three factual questions
answer correctly through the unequal-TP chain with per-rank ms-level
pulls and zero errors, TextVQA authoritative rescore is identical
between PD (92.0%) and colocated (90.0%) in the same batch, and the
full UT suite matches the pre-change baseline (only the pre-existing
failures, re-verified after the view code moved into the model
directory).

Signed-off-by: sunbaosong <13793883820@163.com>
@sunbaosong
sunbaosong force-pushed the feature/glm5next-pd-mrv1 branch from 5c9040d to 6615b57 Compare September 21, 2026 11:48
@weijinqian0
weijinqian0 enabled auto-merge (squash) September 21, 2026 12:25
…el block

SparseMLAMetadataState sized its operator page from the spec's logical
block whenever it stayed under the operator limit, so the operator
consumed the cache at the scheduler's page granularity - 640 rows on a
TP8 GLM-Next hybrid pool, clamped to the kernel block only past 1024 -
and the block table was contracted back to that TP-dependent page size.
Cache views finer than the logical page, such as the 128-row kernel
blocks exposed for whole-block KV transfer, then tripped the split-only
divisibility guard.

Consume the sparse MLA cache at kernel-block granularity unconditionally:
the operator page is always the kernel block and the expanded block table
passes through as-is, so logical-page, kernel-block, and oversized hybrid
views all reach the operator through the same zero-copy re-view and the
metadata no longer depends on the unified page size.

Signed-off-by: sunbaosong <13793883820@163.com>
auto-merge was automatically disabled September 21, 2026 14:05

Head branch was pushed to by a user without write access

@sunbaosong
sunbaosong force-pushed the feature/glm5next-pd-mrv1 branch from 6615b57 to df31ceb Compare September 21, 2026 14:05
@AlfieJC

AlfieJC commented Sep 22, 2026 •

Copy link
Copy Markdown

/rerun
[Bot]: rerun command failed: you do not have permission. Only the PR author or users with triage+ permission can trigger /rerun.

@sunbaosong

sunbaosong commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@weijinqian0
weijinqian0 merged commit 18d161d into vllm-project:main Sep 22, 2026
55 of 57 checks passed
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…llm-project#16910)

### What this PR does / why we need it?

Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model
runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the
kernel-block layout prefract landed for MRV2 (vllm-project#16755). Without this, PD
disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default
runner rejects the layout at KV init and the cross-TP handshake fails.

Three adaptation points:
- V1 reshape: two small-slot view branches mirrored from
_reshape_kv_cache_v2 – the per-request kpool tail ring packs a
contiguous suffix of the shared small slot, and the compressed indexer
packs natural kernel blocks as a contiguous prefix, each bounded to half
the slot.
- GLM-Next cache layout: size the shared small slot by the unified page.
_align_glm5_next_cache_specs derived the small page from the small
candidates only, so with MRV1's unpadded worker specs the slot collapsed
to the indexer's own page and the half-slot bound rejected the layout at
KV init. MRV2 workers pre-pad their specs, so their behavior is
unchanged.
- V1 reshape: expose the NOPE main as natural 128-row kernel blocks
instead of one page-sized block per scheduler block, so dim0 blocks stay
byte-identical across engines with different TP. The page-sized form
made the MooncakeConnectorV2 handshake reject cross-TP transfers with
different local and remote kernel block size 1152 | 640.

Depends on vllm-project#16755.

### Does this PR introduce any user-facing change?

No new flags, config options, or API changes. For users running
GLM-5.3-Flash with PD disaggregation on the default model runner,
cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box.
All other models and colocated workloads are unaffected – branch
selection is identical to before (only GLM-Next publishes the stride
marker).

How was this patch tested?

On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8,
and P=TP8 → D=DP2xTP4):
- Indexer and tail views bit-identical with the MRV2 reshape.
- Equal-TP full-verify passes (Serial + concurrent, all match; per-rank
ms-level pulls).
- Unequal-TP chain: handshake clean, three factual questions all
answered correctly, zero pull/transfer errors.
- TextVQA authoritative rescore (per-row gold alignment, post-
extraction): PD 92.0% identical to colocated 90.0% in the same batch.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: sunbaosong <13793883820@163.com>
sunbaosong added a commit to sunbaosong/vllm-ascend that referenced this pull request Sep 26, 2026
The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Size the small page from the unpadded candidate pages instead --
the worker pre-pad that inflates page_size_bytes is deliberately
overridden, so the natural page wins on both runners -- alias the tail
ring at each small page head -- the same offset model upstream uses for
KpoolTailSpec -- and replace the half-slot bounds with an exact-fill
check on the indexer and a fit check on the tail. Slot soundness rests
on the global block-ID exclusivity the main MLA/KDA aliasing already
relies on. GLM-Next is exempted from the hybrid backing allocation:
the pool now spans two page classes, and its aliasing allocates per
descriptor as DeepSeek-V4 does. A zero-block pool builds empty views
into the raw slot storage, matching the other view paths.

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and
pull latency is unchanged.

Refs vllm-project#16910

Signed-off-by: sunbaosong <13793883820@163.com>
sunbaosong added a commit to sunbaosong/vllm-ascend that referenced this pull request Sep 26, 2026
The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Size the small page from the unpadded candidate pages instead --
the worker pre-pad that inflates page_size_bytes is deliberately
overridden, so the natural page wins on both runners -- alias the tail
ring at each small page head -- the same offset model upstream uses for
KpoolTailSpec -- and replace the half-slot bounds with an exact-fill
check on the indexer and a fit check on the tail. Slot soundness rests
on the global block-ID exclusivity the main MLA/KDA aliasing already
relies on. GLM-Next is exempted from the hybrid backing allocation:
the pool now spans two page classes, and its aliasing allocates per
descriptor as DeepSeek-V4 does. A zero-block pool builds empty views
into the raw slot storage, matching the other view paths.

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and
pull latency is unchanged.

Refs vllm-project#16910

Signed-off-by: sunbaosong <13793883820@163.com>
sunbaosong added a commit to sunbaosong/vllm-ascend that referenced this pull request Sep 27, 2026
The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Size the small page from the unpadded candidate pages instead,
and make the V2 spec collector honor the align_kv_cache_with_mamba
opt-out the auxiliary caches declare, so the worker pre-pad no longer
inflates the auxiliary specs and the natural page wins on both runners.
Alias the tail ring at each small page head -- the same offset model
upstream uses for KpoolTailSpec -- and replace the half-slot bounds
with an exact-fill check on the indexer and a fit check on the tail.
Slot soundness rests on the global block-ID exclusivity the main
MLA/KDA aliasing already relies on. GLM-Next is exempted from the
hybrid backing allocation: the pool now spans two page classes, and its
aliasing allocates per descriptor as DeepSeek-V4 does. A zero-block
pool builds empty views into the raw slot storage, matching the other
view paths.

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and
pull latency is unchanged.

Refs vllm-project#16910

Signed-off-by: sunbaosong <13793883820@163.com>
sunbaosong added a commit to sunbaosong/vllm-ascend that referenced this pull request Sep 27, 2026
The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Derive both page classes with one rule -- the max of the
candidates' claimed page_size_bytes and natural page sizes -- dropping
the unified main-page floor that inflated every small slot 16x. The V2
spec collector honors the align_kv_cache_with_mamba opt-out the
auxiliary caches declare, so nothing pre-pads the auxiliary specs and
the natural page wins on both runners. Alias the tail ring at each
small page head -- the same offset model upstream uses for
KpoolTailSpec -- and replace the half-slot bounds with an exact-fill
check on the indexer and a fit check on the tail. Slot soundness rests
on the global block-ID exclusivity the main MLA/KDA aliasing already
relies on. GLM-Next is exempted from the hybrid backing allocation:
the pool now spans two page classes, and its aliasing allocates per
descriptor as DeepSeek-V4 does. A zero-block pool builds empty views
into the raw slot storage, matching the other view paths.

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and
pull latency is unchanged.

Refs vllm-project#16910

Signed-off-by: sunbaosong <13793883820@163.com>
weijinqian0 pushed a commit that referenced this pull request Sep 28, 2026
### What this PR does / why we need it?

The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Size the small page from the unpadded candidate pages instead --
the worker pre-pad that inflates page_size_bytes is deliberately
overridden, so the natural page wins on both runners -- alias the tail
ring at each small page head -- the same offset model upstream uses for
KpoolTailSpec -- and replace the half-slot bounds with an exact-fill
check on the indexer and a fit check on the tail. Slot soundness rests
on the global block-ID exclusivity the main MLA/KDA aliasing already
relies on. GLM-Next is exempted from the hybrid backing allocation: the
pool now spans two page classes, and its aliasing allocates per
descriptor as DeepSeek-V4 does.

### Does this PR introduce _any_ user-facing change?

No new flags or config options. GLM-5.3-Flash KV capacity on the same
memory budget improves from ~630K to 1,044,480 tokens (+65%); all other
models and colocated workloads are unaffected.

### How was this patch tested?

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget, and pull latency is
unchanged.

Refs #16910

- vLLM main:
vllm-project/vllm@ced6857

Signed-off-by: sunbaosong <13793883820@163.com>
HackClawMxw pushed a commit to HackClawMxw/vllm-ascend-fork that referenced this pull request Sep 30, 2026
…ject#17545)

### What this PR does / why we need it?

The compressed indexer packs kernel rows across the full shared small
slot while the tail ring packs a suffix of the same slot, so the layout
inherited the unified main page and kept a 16x dead zone on every small
slot. Size the small page from the unpadded candidate pages instead --
the worker pre-pad that inflates page_size_bytes is deliberately
overridden, so the natural page wins on both runners -- alias the tail
ring at each small page head -- the same offset model upstream uses for
KpoolTailSpec -- and replace the half-slot bounds with an exact-fill
check on the indexer and a fit check on the tail. Slot soundness rests
on the global block-ID exclusivity the main MLA/KDA aliasing already
relies on. GLM-Next is exempted from the hybrid backing allocation: the
pool now spans two page classes, and its aliasing allocates per
descriptor as DeepSeek-V4 does.

### Does this PR introduce _any_ user-facing change?

No new flags or config options. GLM-5.3-Flash KV capacity on the same
memory budget improves from ~630K to 1,044,480 tokens (+65%); all other
models and colocated workloads are unaffected.

### How was this patch tested?

Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the
pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8
and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV
capacity 1,044,480 tokens on the same budget, and pull latency is
unchanged.

Refs vllm-project#16910

- vLLM main:
vllm-project/vllm@ced6857

Signed-off-by: sunbaosong <13793883820@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants