[Feature] Enable GLM-5.3-flash PD disaggregation on Model Runner V1 - #16910
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request enables support for GLM-5.3-Flash (GLM-5Next) KV cache geometry within the V1 model runner. It introduces mechanisms to handle compressed indexer caches and per-request tail ring caches, ensuring that memory layout remains consistent with the V2 implementation. These changes include updates to cache configuration, runner initialization, and KV transfer utilities to support the new model architecture. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Support GLM-5.3-Flash kpool cache views in V1 and V2 model runnersSuggested PR Summary:
### What this PR does / why we need it?
This PR implements the GLM-5.3-Flash kpool cache views in both V1 and V2 model runners. It aligns the cache geometry contracts so that the compressed indexer packs natural kernel blocks as a contiguous prefix of the shared small slot, and the per-request tail ring packs a contiguous suffix. It also updates the KV transfer logic to support the new specs.
Several issues were identified in the review:
- Potential `IndexError` when accessing `raw_cache[0]` before checking if the tuple is empty in both `model_runner_v1.py` and `v2/attn_utils.py`.
- A `TypeError` in `v2/attn_utils.py` when attempting integer division on a list instead of its first element.
- Potential `AttributeError`s when directly accessing `model_version` and `non_causal_multi_token_decode` on standard `KVCacheSpec` objects.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with unit tests in `tests/ut/models/test_glm5next_kv_cache.py`, `tests/ut/worker/test_glm5next_pooled_cache.py`, and `tests/ut/worker/test_model_runner_v1_glm5next_reshape.py`.e4465ee to
c73b373
Compare
06797e3 to
ae72306
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
a17fc1b to
5c9040d
Compare
Bring the GLM-5.3-Flash prefill/decode disaggregation to the default model runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the kernel-block layout contract landed for MRV2: - V1 reshape: mirror the two small-slot view branches from _reshape_kv_cache_v2 -- the per-request kpool tail ring packs a contiguous suffix of the shared small slot and the compressed indexer packs natural kernel blocks as a contiguous prefix, each bounded to half the slot. - GLM-Next cache layout: size the shared small slot by the unified page. _align_glm5_next_cache_specs derived the small page from the small candidates only, so with MRV1's unpadded worker specs the slot collapsed to the indexer's own page and the half-slot bound rejected the layout at KV init. MRV2 workers pre-pad their specs, so their behavior is unchanged. - V1 reshape: expose the NoPE main MLA as natural 128-row kernel blocks instead of one page-sized block per scheduler block, so dim0 blocks stay byte-identical across engines with different TP; the page-sized form made the MooncakeConnectorV2 handshake reject cross-TP transfers with 'different local and remote kernel block size 1152 | 640'. These view branches are model-owned layout logic and live in vllm_ascend/models/glm5next/cache_views.py behind a single view_glm5_next_cache entry. The runner calls it only when the layer specs carry the glm5_next marker (is_glm5_next_cache_spec, mirroring the is_deepseek_v41_cache gate in the same function) and falls through to its generic reshape paths when the helper returns None, so compressed specs owned by the existing V1 page-strided path (e.g. DSV4) are untouched. With these views the MooncakeConnectorV2 handshake (derived from the actually registered caches) transfers at whole-block granularity with zero connector changes on MRV1. Verified on GLM-5.3-Flash w8a8 (TP8 single instance, P=TP8 <-> D=TP8, and P=TP8 -> D=DP2xTP4): indexer and tail views bit-identical with the MRV2 reshape, full-verify passes on equal TP, three factual questions answer correctly through the unequal-TP chain with per-rank ms-level pulls and zero errors, TextVQA authoritative rescore is identical between PD (92.0%) and colocated (90.0%) in the same batch, and the full UT suite matches the pre-change baseline (only the pre-existing failures, re-verified after the view code moved into the model directory). Signed-off-by: sunbaosong <13793883820@163.com>
5c9040d to
6615b57
Compare
…el block SparseMLAMetadataState sized its operator page from the spec's logical block whenever it stayed under the operator limit, so the operator consumed the cache at the scheduler's page granularity - 640 rows on a TP8 GLM-Next hybrid pool, clamped to the kernel block only past 1024 - and the block table was contracted back to that TP-dependent page size. Cache views finer than the logical page, such as the 128-row kernel blocks exposed for whole-block KV transfer, then tripped the split-only divisibility guard. Consume the sparse MLA cache at kernel-block granularity unconditionally: the operator page is always the kernel block and the expanded block table passes through as-is, so logical-page, kernel-block, and oversized hybrid views all reach the operator through the same zero-copy re-view and the metadata no longer depends on the unified page size. Signed-off-by: sunbaosong <13793883820@163.com>
Head branch was pushed to by a user without write access
6615b57 to
df31ceb
Compare
|
/rerun |
|
/rerun Rerun (failed jobs only):
|
…llm-project#16910) ### What this PR does / why we need it? Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the kernel-block layout prefract landed for MRV2 (vllm-project#16755). Without this, PD disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default runner rejects the layout at KV init and the cross-TP handshake fails. Three adaptation points: - V1 reshape: two small-slot view branches mirrored from _reshape_kv_cache_v2 – the per-request kpool tail ring packs a contiguous suffix of the shared small slot, and the compressed indexer packs natural kernel blocks as a contiguous prefix, each bounded to half the slot. - GLM-Next cache layout: size the shared small slot by the unified page. _align_glm5_next_cache_specs derived the small page from the small candidates only, so with MRV1's unpadded worker specs the slot collapsed to the indexer's own page and the half-slot bound rejected the layout at KV init. MRV2 workers pre-pad their specs, so their behavior is unchanged. - V1 reshape: expose the NOPE main as natural 128-row kernel blocks instead of one page-sized block per scheduler block, so dim0 blocks stay byte-identical across engines with different TP. The page-sized form made the MooncakeConnectorV2 handshake reject cross-TP transfers with different local and remote kernel block size 1152 | 640. Depends on vllm-project#16755. ### Does this PR introduce any user-facing change? No new flags, config options, or API changes. For users running GLM-5.3-Flash with PD disaggregation on the default model runner, cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box. All other models and colocated workloads are unaffected – branch selection is identical to before (only GLM-Next publishes the stride marker). How was this patch tested? On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8, and P=TP8 → D=DP2xTP4): - Indexer and tail views bit-identical with the MRV2 reshape. - Equal-TP full-verify passes (Serial + concurrent, all match; per-rank ms-level pulls). - Unequal-TP chain: handshake clean, three factual questions all answered correctly, zero pull/transfer errors. - TextVQA authoritative rescore (per-row gold alignment, post- extraction): PD 92.0% identical to colocated 90.0% in the same batch. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: sunbaosong <13793883820@163.com>
The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Size the small page from the unpadded candidate pages instead -- the worker pre-pad that inflates page_size_bytes is deliberately overridden, so the natural page wins on both runners -- alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. A zero-block pool builds empty views into the raw slot storage, matching the other view paths. Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and pull latency is unchanged. Refs vllm-project#16910 Signed-off-by: sunbaosong <13793883820@163.com>
The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Size the small page from the unpadded candidate pages instead -- the worker pre-pad that inflates page_size_bytes is deliberately overridden, so the natural page wins on both runners -- alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. A zero-block pool builds empty views into the raw slot storage, matching the other view paths. Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and pull latency is unchanged. Refs vllm-project#16910 Signed-off-by: sunbaosong <13793883820@163.com>
The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Size the small page from the unpadded candidate pages instead, and make the V2 spec collector honor the align_kv_cache_with_mamba opt-out the auxiliary caches declare, so the worker pre-pad no longer inflates the auxiliary specs and the natural page wins on both runners. Alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. A zero-block pool builds empty views into the raw slot storage, matching the other view paths. Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and pull latency is unchanged. Refs vllm-project#16910 Signed-off-by: sunbaosong <13793883820@163.com>
The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Derive both page classes with one rule -- the max of the candidates' claimed page_size_bytes and natural page sizes -- dropping the unified main-page floor that inflated every small slot 16x. The V2 spec collector honors the align_kv_cache_with_mamba opt-out the auxiliary caches declare, so nothing pre-pads the auxiliary specs and the natural page wins on both runners. Alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. A zero-block pool builds empty views into the raw slot storage, matching the other view paths. Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget (was ~630K, +65%), and pull latency is unchanged. Refs vllm-project#16910 Signed-off-by: sunbaosong <13793883820@163.com>
### What this PR does / why we need it? The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Size the small page from the unpadded candidate pages instead -- the worker pre-pad that inflates page_size_bytes is deliberately overridden, so the natural page wins on both runners -- alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. ### Does this PR introduce _any_ user-facing change? No new flags or config options. GLM-5.3-Flash KV capacity on the same memory budget improves from ~630K to 1,044,480 tokens (+65%); all other models and colocated workloads are unaffected. ### How was this patch tested? Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget, and pull latency is unchanged. Refs #16910 - vLLM main: vllm-project/vllm@ced6857 Signed-off-by: sunbaosong <13793883820@163.com>
…ject#17545) ### What this PR does / why we need it? The compressed indexer packs kernel rows across the full shared small slot while the tail ring packs a suffix of the same slot, so the layout inherited the unified main page and kept a 16x dead zone on every small slot. Size the small page from the unpadded candidate pages instead -- the worker pre-pad that inflates page_size_bytes is deliberately overridden, so the natural page wins on both runners -- alias the tail ring at each small page head -- the same offset model upstream uses for KpoolTailSpec -- and replace the half-slot bounds with an exact-fill check on the indexer and a fit check on the tail. Slot soundness rests on the global block-ID exclusivity the main MLA/KDA aliasing already relies on. GLM-Next is exempted from the hybrid backing allocation: the pool now spans two page classes, and its aliasing allocates per descriptor as DeepSeek-V4 does. ### Does this PR introduce _any_ user-facing change? No new flags or config options. GLM-5.3-Flash KV capacity on the same memory budget improves from ~630K to 1,044,480 tokens (+65%); all other models and colocated workloads are unaffected. ### How was this patch tested? Verified on GLM-5.3-Flash w8a8 (.9/.7): UT full suite matches the pre-existing failure baseline, V1/V2 TP8 colocated and P=TP8 -> D=TP8 and P=TP8 -> D=DP2xTP4 disaggregation all come up healthy with KV capacity 1,044,480 tokens on the same budget, and pull latency is unchanged. Refs vllm-project#16910 - vLLM main: vllm-project/vllm@ced6857 Signed-off-by: sunbaosong <13793883820@163.com>
What this PR does / why we need it?
Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the kernel-block layout prefract landed for MRV2 (#16755). Without this, PD disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default runner rejects the layout at KV init and the cross-TP handshake fails.
Three adaptation points:
Depends on #16755.
Does this PR introduce any user-facing change?
No new flags, config options, or API changes. For users running GLM-5.3-Flash with PD disaggregation on the default model runner, cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box. All other models and colocated workloads are unaffected – branch selection is identical to before (only GLM-Next publishes the stride marker).
How was this patch tested?
On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8, and P=TP8 → D=DP2xTP4):
Indexer and tail views bit-identical with the MRV2 reshape.
Equal-TP full-verify passes (Serial + concurrent, all match; per-rank ms-level pulls).
Unequal-TP chain: handshake clean, three factual questions all answered correctly, zero pull/transfer errors.
TextVQA authoritative rescore (per-row gold alignment, post- extraction): PD 92.0% identical to colocated 90.0% in the same batch.
vLLM main: vllm-project/vllm@84030bb