Repository navigation
[Feature] Support layerwise KV pool for Qwen3.5 hybrid GDN+full_attention models - #12711
Conversation
…n models - Add wait_for_kv_layer_from_connector and record_attention_compute_start to GDN forward (ops/gdn.py) - Add has_connector_metadata guard to wait/save utils (attention/utils.py) - Skip touch_sending_mamba_blocks for layerwise to prevent block pool leak (pool_scheduler.py) Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for layerwise KV pooling in Qwen3.5 hybrid models that utilize both Gated Delta Net (GDN) and full attention layers. The changes ensure that GDN layers correctly participate in the KV transfer process, improve the reliability of connector metadata handling, and resolve potential memory leaks in the pool scheduler. Additionally, the update includes necessary adjustments to maintain compatibility with the latest vLLM framework versions. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Ops][Misc] Integrate connector metadata checks and GDN attention synchronizationSuggested PR Summary:
### What this PR does / why we need it?
This pull request introduces several improvements and synchronization mechanisms for the Ascend backend:
1. Adds checks for `connector.has_connector_metadata()` in KV layer wait and save utilities to prevent operations when metadata is missing.
2. Bypasses `touch_sending_mamba_blocks` when `use_layerwise` is enabled.
3. Integrates `wait_for_kv_layer_from_connector` and `record_attention_compute_start` into the GDN attention forward pass to ensure proper synchronization.
Feedback:
- In `vllm_ascend/attention/utils.py`, defensive checks should be added to ensure `connector` is not `None` and implements `has_connector_metadata` before invocation to avoid potential `AttributeError` or `TypeError`.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Not specified in the PR. Please ensure existing tests pass and add unit tests for the new synchronization logic.| forward_context: ForwardContext = get_forward_context() | ||
| attn_metadata = forward_context.attn_metadata | ||
| if attn_metadata is None: | ||
| if attn_metadata is None or not connector.has_connector_metadata(): |
There was a problem hiding this comment.
To prevent potential AttributeError or TypeError if connector is None or does not implement has_connector_metadata, we should defensively check if connector is not None and has the attribute before calling it.
| if attn_metadata is None or not connector.has_connector_metadata(): | |
| if attn_metadata is None or connector is None or not getattr(connector, "has_connector_metadata", lambda: False)(): |
| forward_context: ForwardContext = get_forward_context() | ||
| attn_metadata = forward_context.attn_metadata | ||
| if attn_metadata is None: | ||
| if attn_metadata is None or not connector.has_connector_metadata(): |
There was a problem hiding this comment.
To prevent potential AttributeError or TypeError if connector is None or does not implement has_connector_metadata, we should defensively check if connector is not None and has the attribute before calling it.
| if attn_metadata is None or not connector.has_connector_metadata(): | |
| if attn_metadata is None or connector is None or not getattr(connector, "has_connector_metadata", lambda: False)(): |
|
/rerun Rerun:
|
2ae7efb to
42e4049
Compare
|
/cancel https://github.com/vllm-project/vllm-ascend/actions/runs/32341607564 |
…2711 Signed-off-by: F.Liu <1661888967@qq.com>
The layerwise KV pool feature added wait_for_kv_layer_from_connector(self.prefix) and record_attention_compute_start() to AscendGatedDeltaNetAttention.forward. Both are NPU-only side effects invoked from the traced (compiled) forward, so torch.compile(fullgraph=True) graph-broke on the Mock connector call and the global/lock access, failing test_connector_observes_updated_gdn_state_for_each_compiled_call on the CPU and A2 runners. Stub them (consistent with the existing clear_ssm_states/get_pcp_group stubs) so the inductor graph stays fullgraph. Signed-off-by: tyy0829 <1455207791@qq.com>
ba8c68c to
c9708cb
Compare
|
/rerun Rerun (failed jobs only):
|
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
|
/rerun Rerun (failed jobs only):
|
Add a comment block on the deferred Mamba state copy hook explaining why the copy must run after the layerwise KV load finishes and how the per-layer scheduling overlaps the copy with the remaining layer loads. Signed-off-by: tyy0829 <1455207791@qq.com>
|
/cancel |
|
/rerun Rerun (failed jobs only):
|
|
Superseded by #15479, which re-implements this feature on the V2 model runner ( The V1-runner implementation from this PR is preserved on branch Closing this one in favor of #15479; please review there. Thanks! |
|
(Reopened per author request — keeping both PRs open for now: this V1-runner implementation and #15479, which re-implements the same feature on the V2 model runner.) |
|
/rerun Rerun (failed jobs only):
|
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
|
/rerun Rerun (failed jobs only):
|
|
@Wangbei25 PR related to qwen3.5, PTAL |
|
/cancel https://github.com/vllm-project/vllm-ascend/actions/runs/33503476206 |
…tion models (vllm-project#12711) ## Description This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next) hybrid models that combine GDN (Gated Delta Net) linear attention layers with standard full attention layers. ## Changes 1. **GDN forward (`ops/gdn.py`)**: Add `wait_for_kv_layer_from_connector(self.prefix)` at the start of `forward()` and `record_attention_compute_start()` before the custom op call. GDN layers do not go through the `@maybe_transfer_kv_layer` decorator, so they must explicitly call these functions to participate in layerwise KV transfer. 2. **Attention utils (`attention/utils.py`)**: Add `connector.has_connector_metadata()` guard to `wait_for_kv_layer_from_connector` and `maybe_save_kv_layer_to_connector`, matching the guard in the `@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer` counter increments during non-save steps (e.g. profile run). 3. **Pool scheduler (`pool_scheduler.py`)**: Skip `touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise send thread (`KVCacheStoreLayerSendingThread`) does not use the `completed_events` mechanism, so touched mamba blocks would never be freed, leaking the block pool. 4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update `postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding `state_dim_row_count_ptr`, `state_dim_row_stride_ptr`, `idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and `PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend pointer-type-cast optimization. ## Testing Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B: - Service starts successfully with `use_layerwise=true` and MemCache backend - KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix cache hit rate 22.9% - 5 concurrent requests all succeed with cache hits - MTP speculative decoding (`num_speculative_tokens=3, method=qwen3_5_mtp`) works correctly --- (Replaces vllm-project#12526, which was locked-closed after a force-push to the head branch; this PR is based on the clean single-commit feature.) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com> Signed-off-by: tyy0829 <1455207791@qq.com> Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
…tion models (vllm-project#12711) ## Description This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next) hybrid models that combine GDN (Gated Delta Net) linear attention layers with standard full attention layers. ## Changes 1. **GDN forward (`ops/gdn.py`)**: Add `wait_for_kv_layer_from_connector(self.prefix)` at the start of `forward()` and `record_attention_compute_start()` before the custom op call. GDN layers do not go through the `@maybe_transfer_kv_layer` decorator, so they must explicitly call these functions to participate in layerwise KV transfer. 2. **Attention utils (`attention/utils.py`)**: Add `connector.has_connector_metadata()` guard to `wait_for_kv_layer_from_connector` and `maybe_save_kv_layer_to_connector`, matching the guard in the `@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer` counter increments during non-save steps (e.g. profile run). 3. **Pool scheduler (`pool_scheduler.py`)**: Skip `touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise send thread (`KVCacheStoreLayerSendingThread`) does not use the `completed_events` mechanism, so touched mamba blocks would never be freed, leaking the block pool. 4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update `postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding `state_dim_row_count_ptr`, `state_dim_row_stride_ptr`, `idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and `PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend pointer-type-cast optimization. ## Testing Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B: - Service starts successfully with `use_layerwise=true` and MemCache backend - KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix cache hit rate 22.9% - 5 concurrent requests all succeed with cache hits - MTP speculative decoding (`num_speculative_tokens=3, method=qwen3_5_mtp`) works correctly --- (Replaces vllm-project#12526, which was locked-closed after a force-push to the head branch; this PR is based on the clean single-commit feature.) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com> Signed-off-by: tyy0829 <1455207791@qq.com> Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
…tion models (vllm-project#12711) ## Description This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next) hybrid models that combine GDN (Gated Delta Net) linear attention layers with standard full attention layers. ## Changes 1. **GDN forward (`ops/gdn.py`)**: Add `wait_for_kv_layer_from_connector(self.prefix)` at the start of `forward()` and `record_attention_compute_start()` before the custom op call. GDN layers do not go through the `@maybe_transfer_kv_layer` decorator, so they must explicitly call these functions to participate in layerwise KV transfer. 2. **Attention utils (`attention/utils.py`)**: Add `connector.has_connector_metadata()` guard to `wait_for_kv_layer_from_connector` and `maybe_save_kv_layer_to_connector`, matching the guard in the `@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer` counter increments during non-save steps (e.g. profile run). 3. **Pool scheduler (`pool_scheduler.py`)**: Skip `touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise send thread (`KVCacheStoreLayerSendingThread`) does not use the `completed_events` mechanism, so touched mamba blocks would never be freed, leaking the block pool. 4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update `postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding `state_dim_row_count_ptr`, `state_dim_row_stride_ptr`, `idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and `PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend pointer-type-cast optimization. ## Testing Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B: - Service starts successfully with `use_layerwise=true` and MemCache backend - KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix cache hit rate 22.9% - 5 concurrent requests all succeed with cache hits - MTP speculative decoding (`num_speculative_tokens=3, method=qwen3_5_mtp`) works correctly --- (Replaces vllm-project#12526, which was locked-closed after a force-push to the head branch; this PR is based on the clean single-commit feature.) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com> Signed-off-by: tyy0829 <1455207791@qq.com> Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
…n3.5 hybrid GDN+full-attention models on Model Runner V2 (#15479) ## Summary Supersedes #12711 (V1-runner version, preserved on branch [`backup/pr-12711-layerwise-v1-runner`](https://github.com/tyy0829/vllm-ascend/tree/backup/pr-12711-layerwise-v1-runner)). This PR re-implements layerwise KV pool support for **Qwen3.5 hybrid GDN+full-attention models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`), which is where new hybrid/mamba features are expected to land. ### What's inside 1. **Keep hybrid-aligned block size against draft-config re-verification** (`utils.py`): the V2 MTP speculator builds the draft `VllmConfig` via `dataclasses.replace`, which re-runs the config hooks on the *shared* `cache_config`. Because the draft model is not hybrid, `refresh_block_size` downgraded the mamba-page-aligned attention block size (1536 for Qwen3.5) back to 128, corrupting the layerwise pool's address layout (batch_copy `-3102`, vector-core exceptions) and shrinking usable KV capacity ~5.4x. The fix skips the generic downgrade when the shared cache_config was already sized for a hybrid model (`mamba_page_size_padded` set). 2. **NPU-safe mamba align pre-copy kernel** (`ops/triton/mamba/precopy.py`): upstream's `precopy_mamba_align_fused_kernel` vectorizes the temporal-state copy through uint64 loads/stores, which the Ascend vector core does not support (vector-core exception as soon as the V2 mamba align path runs). Adds a byte-wise uint8 variant (same semantics, pointer cast hoisted out of the loop, mirroring the existing NPU `postprocess_mamba_fused_kernel` fix) and installs it from `patch_mamba_utils`. 3. **Per-layer mamba state copy deferred behind layerwise KV loads** (`worker/v2/model_states/mamba_hybrid.py` + connectors): `AscendMambaHybridModelState.preprocess_state` keeps the upstream GPU-resident decision kernel but hands the actual align pre-copy to a layerwise-capable connector via `prepare_mamba_state_copy`. Each mamba layer's copy is launched from `do_mamba_copy_for_layer` right after that layer's KV load (conv/ssm state included) completes — sliced to the layer's own rows of the GPU-resident metadata, no CPU-GPU sync. The next step's `preprocess_state` validates every deferred layer executed. `AscendStoreConnector`/`AscendMultiConnector` expose this V2 interface; `use_layerwise` now requires the V2 model runner. 4. **Layerwise pool scheduler fixes carried over from #12711** (runner-agnostic): conditional eagle/MTP tail trim in `pool_scheduler` (fixes external hit 61.4%→76.7%), full-hit tail re-store in `pool_worker`, GDN layerwise wait inside the custom-op body, `has_connector_metadata` guard for the wait/save hooks. ## Verification (Atlas 800T A2, Qwen3.5-27B, TP=4, MTP num_speculative_tokens=3, memcache backend) | Scenario | Requests | External pool hit rate | Notes | |---|---|---|---| | PD-mixed (`kv_both`, in10000/out100/rr0.9, cc200) | 200/200 OK | **76.3%** (V1 baseline 76.7%) | no crashes | | PD-disaggregated (P: cards 12-15, D: cards 8-11, in15360/out200/rr0.9, cc200) | 400/400 OK | **P 89.6% / D 100.0%** (V1 baseline 89.95%/99.99%) | no crashes | | Control: unlayerwise + FULL_DECODE_ONLY | 200/200 OK | 76.3% | MTP acceptance 69.3% | | Control: unlayerwise + PIECEWISE | 200/200 OK | — | MTP acceptance 13.7% | Layer protocol counters verified healthy under load: 4200 consecutive steps each saw exactly 67 wait/save hook calls with `current_layer` reaching `num_layers=65` (no interleaving/race under async scheduling). ### Known limitation (not introduced by this PR) MTP acceptance length drops under **PIECEWISE cudagraph mode + MTP on the V2 runner** (1.0-1.3 vs 3.0+ tokens/step under FULL_DECODE_ONLY). Layerwise KV pool requires PIECEWISE (same as V1), so it surfaces there, but the control experiment above shows unlayerwise+PIECEWISE is equally affected — it is an upstream V2 speculator/PIECEWISE interaction that deserves a separate issue. This PR introduces no MTP regression relative to other PIECEWISE configurations (9.5% vs 13.7% within the same degraded mode). ## Unit tests - `tests/ut/worker/test_model_runner_v2_mamba.py`: per-layer deferral scheduling, sliced kernel launch, missing-layer validation, connector takeover/release - `tests/ut/distributed/ascend_store/test_ascend_store_connector.py`: copy-after-load ordering, V2 gate, non-layerwise keeps bulk path - `tests/ut/kv_offload/test_ascend_multi_connector.py`: copy runs after all child connectors finish the layer load All 384 UTs in the touched modules pass locally; `ruff format --check` and `ruff check` clean. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: tyy0829 <1455207791@qq.com>
…n V2 runner Enable the layerwise KV pool (pooling + use_layerwise) for Kimi-K3 hybrid KDA + MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1), the same scenario previously enabled for Qwen3.5 GDN hybrids (vllm-project#12711/vllm-project#15479). - ops/kimi_kda.py: KDA layers bypass the standard attention layer, so like the GDN path they must drive the per-layer layerwise protocol themselves. Add wait_for_kv_layer_from_connector + record_attention_compute_start at the top of the eager-break _forward body (before the conv/recurrent kernels touch mamba state) and maybe_save_kv_layer_to_connector on every exit path. Without these hooks the pool worker's current_layer counter desyncs (only the MLA layer advanced it) and the first multi-block save trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. - worker/v2/model_states/mamba_hybrid.py: port the per-layer mamba align pre-copy deferral from the Qwen3.5 V2 work. preprocess_state keeps the GPU-resident decision kernel but hands the actual copy to a layerwise connector (prepare_mamba_state_copy); each KDA layer's copy is launched from do_mamba_copy_for_layer right after that layer's KV load (conv/ssm state included) completes, so the copy never reads half-loaded state. - AscendStoreConnector/AscendMultiConnector: prepare_mamba_state_copy now accepts either the V1 mamba copy buffers or the V2 mamba hybrid model state; wait_for_layer_load triggers the per-layer copy after the load finishes; finish_mamba_state_copy releases the handle. - UT: per-layer deferral scheduling (decision kernel still runs, bulk pre-copy skipped, sliced kernel launch, missing-layer raises), copy-after- load ordering and finish-release semantics for both connectors; V1 copy-bufs mock is now spec'd so duck-typing keeps it on the bulk path. Verified on Atlas 800I A3 (Kimi-K3-w4a8-4layer, TP=8, EP, eager, AscendStoreConnector memcache device_sdma backend, use_layerwise=true): - Service starts under VLLM_USE_V2_MODEL_RUNNER=1; /health 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (kvpool hit tokens: 768), repeat request latency 3.4s -> 0.57s - Stress 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - tests/ut: 44/44 pass (mamba_hybrid deferral + connector duck-typing) Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: tyy0829 <1455207791@qq.com>
### What this PR does / why we need it? Mooncake hybrid layerwise currently rejects layouts containing recurrent Mamba state. Removing that check alone is unsafe: GDN, Kimi KDA, and other Mamba-like layers update conv/recurrent state inside attention, so an attention-entry PUT can publish the previous state while the following request reports a cache hit. This change: - identifies physical layers containing `MambaSpec`, including per-layer `UniformTypeKVCacheSpecs`; - keeps full/sliding-window attention PUTs at attention entry; - preserves the post-compute save hook for recurrent layers so Mooncake publishes updated state; - reuses the stage-local layer mapping and PP-scoped key namespace from #17183, mapping global recurrent layer names to the correct local layer on each pipeline stage; - applies the existing send-backlog policy to recurrent and record-only save paths; - drains queued PUTs at step completion, including when the final layer has no ranges to save. The generic GDN/Mamba load, deferred state-copy, and connector hooks from #12711 are already present on main. This PR only adds the Mooncake block-key publication semantics. The implementation and lifecycle coverage are based on #17194 and credit its author in the commit. | Model | Cards | Dtype| layerwise | kv pool | input | output | External prefix cache | Prefix cache | P | D | Batchsize | Request | TTFT/s | TPOT/ms | TPS | |------|------|----------|------------------|----------|------|------|------------------------|--------------|-----------|-----------|--------|----------|--------|---------|------------------| | Qwen3.5-35B | 2 | w8a8 | × | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 | 128 | 6.07 | / | 41606 | | Qwen3.5-35B | 2 | w8a8 | √ | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 | 128 | 4.93 | / | 53418 | ### Does this PR introduce _any_ user-facing change? Yes. `AscendStoreConnector` with `backend="mooncake"` and `use_layerwise=true` now accepts hybrid recurrent-state groups using `mamba_cache_mode="align"`. This also composes with Mooncake layerwise pipeline parallel support from #17183. Existing topology and prefill/decode TP-mismatch restrictions remain unchanged. ### How was this patch tested? - Targeted Mooncake layerwise/Mamba mock tests: 26 passed. - PP+Mamba worker regression: 5 passed, 2 subtests; the new case covers PP rank 1, global-to-local layer mapping, regular attention-entry PUT, and Mamba post-compute PUT for raw and uniform-wrapped specs. - CI PP layer-count regression: 2 passed; non-layerwise partial worker initialization no longer reads layerwise-only cache specs. - `test_pool_worker.py` excluding unrelated KVPP tests: 96 passed, 2 deselected, 48 subtests passed. - Broader AscendStore mock suite: 440 passed, with the same 3 environment/mock failures reproduced on unmodified main (435 passed, same failures). - `ruff check` passed for all changed Python files. - `ruff format --check` passed for all changed Python files. - `python -m py_compile` passed for all changed Python files. - Added a raw-token lifecycle regression covering Mamba2, GDN, generic linear attention, Full/SWA, merged/per-layer specs, chunked prefill, decode, cache hits, shared prefixes, and missing-state fallback. This file requires a full vLLM/PyTorch CPU test environment and was not executed in the local Windows mock environment. - The equivalent patch on commit `ca91dfa600373c56cfffc44be1d77028b0dd621d` was validated successfully by the contributor on the target deployment. This latest-main rebase preserves the same transfer semantics; real NPU/Mooncake validation was not rerun in the local development environment. `bash format.sh ci` could not run locally because `pre-commit` is not installed. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: fangrongcan <17343701736@163.com> Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
…17287) ### What this PR does / why we need it? Mooncake hybrid layerwise currently rejects layouts containing recurrent Mamba state. Removing that check alone is unsafe: GDN, Kimi KDA, and other Mamba-like layers update conv/recurrent state inside attention, so an attention-entry PUT can publish the previous state while the following request reports a cache hit. This change: - identifies physical layers containing `MambaSpec`, including per-layer `UniformTypeKVCacheSpecs`; - keeps full/sliding-window attention PUTs at attention entry; - preserves the post-compute save hook for recurrent layers so Mooncake publishes updated state; - reuses the stage-local layer mapping and PP-scoped key namespace from vllm-project#17183, mapping global recurrent layer names to the correct local layer on each pipeline stage; - applies the existing send-backlog policy to recurrent and record-only save paths; - drains queued PUTs at step completion, including when the final layer has no ranges to save. The generic GDN/Mamba load, deferred state-copy, and connector hooks from vllm-project#12711 are already present on main. This PR only adds the Mooncake block-key publication semantics. The implementation and lifecycle coverage are based on vllm-project#17194 and credit its author in the commit. | Model | Cards | Dtype| layerwise | kv pool | input | output | External prefix cache | Prefix cache | P | D | Batchsize | Request | TTFT/s | TPOT/ms | TPS | |------|------|----------|------------------|----------|------|------|------------------------|--------------|-----------|-----------|--------|----------|--------|---------|------------------| | Qwen3.5-35B | 2 | w8a8 | × | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 | 128 | 6.07 | / | 41606 | | Qwen3.5-35B | 2 | w8a8 | √ | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 | 128 | 4.93 | / | 53418 | ### Does this PR introduce _any_ user-facing change? Yes. `AscendStoreConnector` with `backend="mooncake"` and `use_layerwise=true` now accepts hybrid recurrent-state groups using `mamba_cache_mode="align"`. This also composes with Mooncake layerwise pipeline parallel support from vllm-project#17183. Existing topology and prefill/decode TP-mismatch restrictions remain unchanged. ### How was this patch tested? - Targeted Mooncake layerwise/Mamba mock tests: 26 passed. - PP+Mamba worker regression: 5 passed, 2 subtests; the new case covers PP rank 1, global-to-local layer mapping, regular attention-entry PUT, and Mamba post-compute PUT for raw and uniform-wrapped specs. - CI PP layer-count regression: 2 passed; non-layerwise partial worker initialization no longer reads layerwise-only cache specs. - `test_pool_worker.py` excluding unrelated KVPP tests: 96 passed, 2 deselected, 48 subtests passed. - Broader AscendStore mock suite: 440 passed, with the same 3 environment/mock failures reproduced on unmodified main (435 passed, same failures). - `ruff check` passed for all changed Python files. - `ruff format --check` passed for all changed Python files. - `python -m py_compile` passed for all changed Python files. - Added a raw-token lifecycle regression covering Mamba2, GDN, generic linear attention, Full/SWA, merged/per-layer specs, chunked prefill, decode, cache hits, shared prefixes, and missing-state fallback. This file requires a full vLLM/PyTorch CPU test environment and was not executed in the local Windows mock environment. - The equivalent patch on commit `ca91dfa600373c56cfffc44be1d77028b0dd621d` was validated successfully by the contributor on the target deployment. This latest-main rebase preserves the same transfer semantics; real NPU/Mooncake validation was not rerun in the local development environment. `bash format.sh ci` could not run locally because `pre-commit` is not installed. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: fangrongcan <17343701736@163.com> Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
Description
This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next) hybrid models that combine GDN (Gated Delta Net) linear attention layers with standard full attention layers.
Changes
GDN forward (
ops/gdn.py): Addwait_for_kv_layer_from_connector(self.prefix)at the start offorward()andrecord_attention_compute_start()before the custom op call. GDN layers do not go through the@maybe_transfer_kv_layerdecorator, so they must explicitly call these functions to participate in layerwise KV transfer.Attention utils (
attention/utils.py): Addconnector.has_connector_metadata()guard towait_for_kv_layer_from_connectorandmaybe_save_kv_layer_to_connector, matching the guard in the@maybe_transfer_kv_layerdecorator. Prevents spuriouscurrent_layercounter increments during non-save steps (e.g. profile run).Pool scheduler (
pool_scheduler.py): Skiptouch_sending_mamba_blockswhenuse_layerwise=True. The layerwise send thread (KVCacheStoreLayerSendingThread) does not use thecompleted_eventsmechanism, so touched mamba blocks would never be freed, leaking the block pool.Postprocess kernel (
ops/triton/mamba/postprocess.py): Updatepostprocess_mamba_fused_kernelsignature to match vLLM v0.25.0, addingstate_dim_row_count_ptr,state_dim_row_stride_ptr,idx_mapping_ptr,CONV_STATE_DIM_FIRST,HAS_IDX_MAPPING, andPRECOMPUTED_NEW_COMPUTEDparameters. Preserves the triton-ascend pointer-type-cast optimization.Testing
Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B:
use_layerwise=trueand MemCache backendnum_speculative_tokens=3, method=qwen3_5_mtp) works correctly(Replaces #12526, which was locked-closed after a force-push to the head branch; this PR is based on the clean single-commit feature.)