Repository navigation
[Feature] [KV Pool][model_runner_v2]Support layerwise KV pool for Qwen3.5 hybrid GDN+full-attention models on Model Runner V2 - #15479
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces layerwise KV pool support for Qwen3.5 hybrid models on the V2 model runner. The implementation focuses on deferring mamba state copies to align with remote layerwise KV loads, preventing race conditions and ensuring compatibility with Ascend hardware. Additionally, the PR includes necessary kernel adjustments and scheduler improvements to maintain performance and stability. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Defer Mamba state copy behind layerwise KV pool loads on AscendSuggested PR Summary:
### What this PR does / why we need it?
This PR implements deferred per-layer Mamba state copy behind layerwise KV pool loads for the V2 model runner on Ascend. When a layerwise KV transfer connector is active, the actual state copies are deferred and executed one layer at a time right after that layer's KV load finishes, preventing race conditions where the copy reads half-loaded state.
### Does this PR introduce _any_ user-facing change?
No user-facing changes. This is an internal optimization and bug fix for Mamba hybrid models running with layerwise KV transfer on Ascend.
### How was this patch tested?
Added new unit tests in:
- tests/ut/distributed/ascend_store/test_ascend_store_connector.py
- tests/ut/kv_offload/test_ascend_multi_connector.py
- tests/ut/worker/test_model_runner_v2_mamba.py|
CI gate is now only waiting on the (Note: this PR and #12711 - the V1-runner implementation of the same feature - are both kept open on purpose; reviewers can pick whichever fits the merge strategy. The V1 code is also preserved on the |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
71dd500 to
898350c
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
898350c to
a08ddff
Compare
… hooks The V2 MTP speculator builds the draft VllmConfig with dataclasses.replace, which re-runs VllmConfig.__post_init__ on the shared cache_config. Because the draft model is not hybrid, refresh_block_size downgraded the mamba-page-aligned attention block size (e.g. 1536 for Qwen3.5 hybrid) back to the generic 128. The mis-sized pages then broke the layerwise KV pool's address layout (batch_copy -3102 / vector core exceptions) and shrank usable KV capacity ~5.4x. Skip the generic downgrade when the shared cache_config has already been sized for a hybrid attention+mamba model (mamba_page_size_padded set). Signed-off-by: tyy0829 <1455207791@qq.com>
The upstream V2 preprocess_mamba_align_fused_kernel path launches precopy_mamba_align_fused_kernel, whose temporal-state copy is vectorized through uint64 loads/stores. The Ascend vector core does not support 8-byte vector accesses and aborts with a vector core exception once the V2 mamba align path runs. Add a byte-wise uint8 variant (same grid, signature and copy semantics, with the pointer cast hoisted out of the copy loop, mirroring the existing NPU postprocess_mamba_fused_kernel fix) and install it from patch_mamba_utils so MambaSpecDecodeGPUContext.run_fused_precopy and the per-layer deferral resolve the kernel through the module attribute. Signed-off-by: tyy0829 <1455207791@qq.com>
…d models - pool_scheduler: only trim the external hit when it reaches into the prompt's final granularity block, so the MTP draft's recomputation zone stays intact (full-hit tail was over-trimmed, dropping the external hit rate from 76.7% to 61.4% on in10000/rr0.9). Layerwise transfer frees mamba blocks on its own completion path; skip the bulk touch that would double-reference the block pool. - pool_worker: never skip the trailing block store on the full-hit recompute path with eagle/MTP; it is recomputed and re-stored by the normal save path. - gdn.py: wait for the layerwise KV load inside the custom-op body (the Dynamo-traced caller region must stay free of host-side side effects) and record the attention-compute start fence before the kernels touch mamba state. - attention/utils.py: also check has_connector_metadata before the layerwise wait/save hooks so non-transfer steps are skipped. Signed-off-by: tyy0829 <1455207791@qq.com>
…oads on V2 Port the layerwise KV pool for Qwen3.5 hybrid GDN+full-attention models to the V2 model runner: - AscendMambaHybridModelState.preprocess_state keeps running the upstream GPU-resident decision kernel, but hands the actual align pre-copy to a layerwise-capable connector (prepare_mamba_state_copy). Each mamba layer's copy is then launched from do_mamba_copy_for_layer right after that layer's KV load (conv/ssm state included) completes, sliced to the layer's own rows of the GPU-resident metadata, so the copy never reads half-loaded state and no CPU-GPU sync is needed. - The next step's preprocess_state validates every deferred layer executed and releases the connector handle. - AscendStoreConnector/AscendMultiConnector expose the V2 interface: prepare_mamba_state_copy takes the model state, wait_for_layer_load triggers the per-layer copy after the layer loads, finish drops the handle. use_layerwise now requires the V2 model runner (the per-layer copy is only implemented for vllm_ascend/worker/v2). Verified on Atlas 800T A2 (Qwen3.5-27B, TP=4, MTP): - PD-mixed: 200/200 requests, external pool hit 76.3% (V1 baseline 76.7%) - PD-disaggregated: 400/400 requests, P 89.6% / D 100.0% hit (V1 baseline 89.95% / 99.99%) Signed-off-by: tyy0829 <1455207791@qq.com>
The order-recording lists mixed str markers with (name, layer) tuples, which mypy rejects as incompatible list item types. Signed-off-by: tyy0829 <1455207791@qq.com>
a08ddff to
242da78
Compare
…n V2 runner Enable the layerwise KV pool (pooling + use_layerwise) for Kimi-K3 hybrid KDA + MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1), the same scenario previously enabled for Qwen3.5 GDN hybrids (vllm-project#12711/vllm-project#15479). - ops/kimi_kda.py: KDA layers bypass the standard attention layer, so like the GDN path they must drive the per-layer layerwise protocol themselves. Add wait_for_kv_layer_from_connector + record_attention_compute_start at the top of the eager-break _forward body (before the conv/recurrent kernels touch mamba state) and maybe_save_kv_layer_to_connector on every exit path. Without these hooks the pool worker's current_layer counter desyncs (only the MLA layer advanced it) and the first multi-block save trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. - worker/v2/model_states/mamba_hybrid.py: port the per-layer mamba align pre-copy deferral from the Qwen3.5 V2 work. preprocess_state keeps the GPU-resident decision kernel but hands the actual copy to a layerwise connector (prepare_mamba_state_copy); each KDA layer's copy is launched from do_mamba_copy_for_layer right after that layer's KV load (conv/ssm state included) completes, so the copy never reads half-loaded state. - AscendStoreConnector/AscendMultiConnector: prepare_mamba_state_copy now accepts either the V1 mamba copy buffers or the V2 mamba hybrid model state; wait_for_layer_load triggers the per-layer copy after the load finishes; finish_mamba_state_copy releases the handle. - UT: per-layer deferral scheduling (decision kernel still runs, bulk pre-copy skipped, sliced kernel launch, missing-layer raises), copy-after- load ordering and finish-release semantics for both connectors; V1 copy-bufs mock is now spec'd so duck-typing keeps it on the bulk path. Verified on Atlas 800I A3 (Kimi-K3-w4a8-4layer, TP=8, EP, eager, AscendStoreConnector memcache device_sdma backend, use_layerwise=true): - Service starts under VLLM_USE_V2_MODEL_RUNNER=1; /health 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (kvpool hit tokens: 768), repeat request latency 3.4s -> 0.57s - Stress 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - tests/ut: 44/44 pass (mamba_hybrid deferral + connector duck-typing) Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: tyy0829 <1455207791@qq.com>
…n V2 runner Extend the layerwise KV pool (AscendStoreConnector + use_layerwise, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to Kimi-K3 hybrid KDA+MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the @maybe_transfer_kv_layer decorator. Without per-layer hooks the pool worker's current_layer counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. Add the same hooks as ops/gdn.py to the KDA eager-break _forward body: - wait_for_kv_layer_from_connector + record_attention_compute_start after the attn_metadata None check (before conv/recurrent kernels touch mamba state, ordering the deferred per-layer mamba state copy and the layer load) - maybe_save_kv_layer_to_connector on the idle early-exit path and the normal exit path The mamba-hybrid deferral, connector duck-typing and UTs needed by this scenario were merged upstream in vllm-project#15479; this PR only adds the KDA-side hooks. A separate mla_v1 get_kv_cache_shape cache_dtype_str fix that was carried here earlier has been solved upstream by vllm-project#15514. Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer debug checkpoint, TP=8, EP, memcache backend (device_sdma), use_layerwise=true, VLLM_USE_V2_MODEL_RUNNER=1: 104/104 stress requests OK, external prefix cache hit rate 64.6% (90% repeat workload). AIS-Bench A/B vs vLLM HBM prefix caching (input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%) at equal token-level hit rate (85.93%) - the gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover. Signed-off-by: tyy0829 <1455207791@qq.com>
…n V2 runner Extend the layerwise KV pool (AscendStoreConnector + use_layerwise, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to Kimi-K3 hybrid KDA+MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the @maybe_transfer_kv_layer decorator. Without per-layer hooks the pool worker's current_layer counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread. Add the same hooks as ops/gdn.py to the KDA eager-break _forward body: - wait_for_kv_layer_from_connector + record_attention_compute_start after the attn_metadata None check (before conv/recurrent kernels touch mamba state, ordering the deferred per-layer mamba state copy and the layer load) - maybe_save_kv_layer_to_connector on the idle early-exit path and the normal exit path The mamba-hybrid deferral, connector duck-typing and UTs needed by this scenario were merged upstream in vllm-project#15479; this PR only adds the KDA-side hooks. A separate mla_v1 get_kv_cache_shape cache_dtype_str fix that was carried here earlier has been solved upstream by vllm-project#15514. Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer debug checkpoint, TP=8, EP, memcache backend (device_sdma), use_layerwise=true, VLLM_USE_V2_MODEL_RUNNER=1: 104/104 stress requests OK, external prefix cache hit rate 64.6% (90% repeat workload). AIS-Bench A/B vs vLLM HBM prefix caching (input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%) at equal token-level hit rate (85.93%) - the gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover. Signed-off-by: tyy0829 <1455207791@qq.com>
…hecks Three host-side optimizations for the layerwise KV pool transfer path that cut per-step RPC overhead under prefix-sharing workloads: 1. Defer the last layer's save drain out of the forward critical path (pool_worker). The host-side wait + event reset for the final layer's save now runs at the start of the NEXT step's start_load_kv, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via VLLM_ASCEND_LW_DEFER_LAST_SAVE (default on). 2. Per-step RPC result caches (pool_worker). Concurrent requests sharing a prefix issue identical memcache queries; keyinfo / lease / exist results are cached per step (reset at the start of process_layer_data) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with valid_gva_indices to avoid mis-zeroing GVAs. 3. Scheduler-side block hit cache (pool_scheduler). The layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per (group, block_hash) with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors save_kv_layer's num_local computation (layerwise_key_layers under pp_size > 1) so only per-stage local layers are drained. Measured on Atlas 800T A3 with the memcache backend (aisbench prefix_cache, 512 requests, 16384 input / 1 output, concurrency 40, prefix repeat rate 0.9): - Qwen3.5-35B-A3B (TP8): layerwise TTFT-avg degradation vs HBM-only narrows from +18.1% to +1.3% (2775.5ms -> 2380.4ms, 3-run mean) and TTFT P90 overtakes HBM by 11.8%. - Qwen3.5-397B-A17B w8a8 (TP16): degradation narrows from +16.4% to +7.7% and TTFT P90 stays below HBM across repeated runs. Verification was performed on PR vllm-project#15479 as of commit a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and all ascend_store unit tests pass locally (418 passed). Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
…LA on the V2 model runner (#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in #15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in #15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by #15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
…hecks Three host-side optimizations for the layerwise KV pool transfer path that cut per-step RPC overhead under prefix-sharing workloads: 1. Defer the last layer's save drain out of the forward critical path (pool_worker). The host-side wait + event reset for the final layer's save now runs at the start of the NEXT step's start_load_kv, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via VLLM_ASCEND_LW_DEFER_LAST_SAVE (default on). 2. Per-step RPC result caches (pool_worker). Concurrent requests sharing a prefix issue identical memcache queries; keyinfo / lease / exist results are cached per step (reset at the start of process_layer_data) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with valid_gva_indices to avoid mis-zeroing GVAs. 3. Scheduler-side block hit cache (pool_scheduler). The layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per (group, block_hash) with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors save_kv_layer's num_local computation (layerwise_key_layers under pp_size > 1) so only per-stage local layers are drained. Measured on Atlas 800T A3 with the memcache backend (aisbench prefix_cache, 512 requests, 16384 input / 1 output, concurrency 40, prefix repeat rate 0.9): - Qwen3.5-35B-A3B (TP8): layerwise TTFT-avg degradation vs HBM-only narrows from +18.1% to +1.3% (2775.5ms -> 2380.4ms, 3-run mean) and TTFT P90 overtakes HBM by 11.8%. - Qwen3.5-397B-A17B w8a8 (TP16): degradation narrows from +16.4% to +7.7% and TTFT P90 stays below HBM across repeated runs. Verification was performed on PR vllm-project#15479 as of commit a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and all ascend_store unit tests pass locally (418 passed). Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
…hecks Three host-side optimizations for the layerwise KV pool transfer path that cut per-step RPC overhead under prefix-sharing workloads: 1. Defer the last layer's save drain out of the forward critical path (pool_worker). The host-side wait + event reset for the final layer's save now runs at the start of the NEXT step's start_load_kv, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via VLLM_ASCEND_LW_DEFER_LAST_SAVE (default on). 2. Per-step RPC result caches (pool_worker). Concurrent requests sharing a prefix issue identical memcache queries; keyinfo / lease / exist results are cached per step (reset at the start of process_layer_data) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with valid_gva_indices to avoid mis-zeroing GVAs. 3. Scheduler-side block hit cache (pool_scheduler). The layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per (group, block_hash) with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors save_kv_layer's num_local computation (layerwise_key_layers under pp_size > 1) so only per-stage local layers are drained. Measured on Atlas 800T A3 with the memcache backend (aisbench prefix_cache, 512 requests, 16384 input / 1 output, concurrency 40, prefix repeat rate 0.9): - Qwen3.5-35B-A3B (TP8): layerwise TTFT-avg degradation vs HBM-only narrows from +18.1% to +1.3% (2775.5ms -> 2380.4ms, 3-run mean) and TTFT P90 overtakes HBM by 11.8%. - Qwen3.5-397B-A17B w8a8 (TP16): degradation narrows from +16.4% to +7.7% and TTFT P90 stays below HBM across repeated runs. Verification was performed on PR vllm-project#15479 as of commit a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and all ascend_store unit tests pass locally (418 passed). Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
…hecks Three host-side optimizations for the layerwise KV pool transfer path that cut per-step RPC overhead under prefix-sharing workloads: 1. Defer the last layer's save drain out of the forward critical path (pool_worker). The host-side wait + event reset for the final layer's save now runs at the start of the NEXT step's start_load_kv, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via VLLM_ASCEND_LW_DEFER_LAST_SAVE (default on). 2. Per-step RPC result caches (pool_worker). Concurrent requests sharing a prefix issue identical memcache queries; keyinfo / lease / exist results are cached per step (reset at the start of process_layer_data) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with valid_gva_indices to avoid mis-zeroing GVAs. 3. Scheduler-side block hit cache (pool_scheduler). The layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per (group, block_hash) with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors save_kv_layer's num_local computation (layerwise_key_layers under pp_size > 1) so only per-stage local layers are drained. Measured on Atlas 800T A3 with the memcache backend (aisbench prefix_cache, 512 requests, 16384 input / 1 output, concurrency 40, prefix repeat rate 0.9): - Qwen3.5-35B-A3B (TP8): layerwise TTFT-avg degradation vs HBM-only narrows from +18.1% to +1.3% (2775.5ms -> 2380.4ms, 3-run mean) and TTFT P90 overtakes HBM by 11.8%. - Qwen3.5-397B-A17B w8a8 (TP16): degradation narrows from +16.4% to +7.7% and TTFT P90 stays below HBM across repeated runs. Verification was performed on PR vllm-project#15479 as of commit a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and all ascend_store unit tests pass locally (418 passed). Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
…LA on the V2 model runner (vllm-project#16320) ### What this PR does / why we need it? Extends the layerwise KV pool (`AscendStoreConnector` + `use_layerwise`, enabled for Qwen3.5 GDN hybrids in vllm-project#15479) to **Kimi-K3 hybrid KDA + MLA models on the V2 model runner** (`VLLM_USE_V2_MODEL_RUNNER=1`). Kimi-K3's KDA (delta attention) layers reuse the GDN state layout but bypass the standard attention layer, so they never went through the `@maybe_transfer_kv_layer` decorator. Without per-layer hooks the pool worker's `current_layer` counter desyncs (only the MLA layer advances it) and the first multi-block prefill trips `thread: 0 save failed` in `KVCacheStoreLayerSendingThread`. This PR adds the same hooks that `ops/gdn.py` already has to the KDA eager-break `_forward` body (`ops/kimi_kda.py`, +17 lines): - `wait_for_kv_layer_from_connector` + `record_attention_compute_start` right after the `attn_metadata` None check, before the conv/recurrent kernels touch mamba state (this also orders the deferred per-layer mamba state copy and the layer load) - `maybe_save_kv_layer_to_connector` on the idle early-exit path and on the normal exit path **Scope note**: the mamba-hybrid per-layer copy deferral, the connector V1/V2 duck-typing and the accompanying UTs required by this scenario were **merged upstream in vllm-project#15479** and are not duplicated here. An `mla_v1.get_kv_cache_shape` `cache_dtype_str` signature fix that was briefly carried in this PR was solved upstream by vllm-project#15514 and has been dropped. ### Does this PR introduce _any_ user-facing change? No API/config change. The layerwise KV pool now works for Kimi-K3 (KDA+MLA hybrid) under the V2 model runner; previously it crashed on the first multi-block prefill. ### How was this patch tested? Verified on Atlas 800I A3 (8 NPU) with the Kimi-K3-w4a8-4layer checkpoint (reduced 4-layer debug build of the ModelSlim W4A8 quantized model), TP=8, EP, eager, memcache backend (`device_sdma`), `use_layerwise=true`, `VLLM_USE_V2_MODEL_RUNNER=1`: - Service starts; `/health` 200; text requests 200 - Long-prefix repeats hit the external pool: 768-token block restored per repeat request (`kvpool hit tokens: 768`); repeat latency 3.4s -> 0.57s - Stress: 104/104 requests OK (8-way concurrency, 90% repeat rate), external prefix cache hit rate 64.6%, zero worker errors - AIS-Bench A/B vs vLLM HBM prefix caching (prefix dataset: input 16088 / output 50, 400 requests, concurrency 8, repeat_rate 0.9): **TTFT avg 1515ms -> 608ms (-60%), TTFT P90 5966ms -> 312ms (-95%), input throughput 42.2k -> 65.1k tok/s (+54%)** at equal token-level hit rate (85.93% both). The gain comes from restoring KDA states on hits, which HBM prefix caching cannot cover (it only caches the MLA layer KV, the KDA layers still recompute). Remaining: MTP path needs validation once a Kimi-K3 checkpoint with `num_nextn_predict_layers > 0` MTP weights is available (the 4-layer debug checkpoint ships none). - vLLM main: vllm-project/vllm@84030bb Signed-off-by: tyy0829 <1455207791@qq.com>
…hecks Three host-side optimizations for the layerwise KV pool transfer path that cut per-step RPC overhead under prefix-sharing workloads: 1. Defer the last layer's save drain out of the forward critical path (pool_worker). The host-side wait + event reset for the final layer's save now runs at the start of the NEXT step's start_load_kv, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via VLLM_ASCEND_LW_DEFER_LAST_SAVE (default on). 2. Per-step RPC result caches (pool_worker). Concurrent requests sharing a prefix issue identical memcache queries; keyinfo / lease / exist results are cached per step (reset at the start of process_layer_data) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with valid_gva_indices to avoid mis-zeroing GVAs. 3. Scheduler-side block hit cache (pool_scheduler). The layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per (group, block_hash) with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors save_kv_layer's num_local computation (layerwise_key_layers under pp_size > 1) so only per-stage local layers are drained. Measured on Atlas 800T A3 with the memcache backend (aisbench prefix_cache, 512 requests, 16384 input / 1 output, concurrency 40, prefix repeat rate 0.9): - Qwen3.5-35B-A3B (TP8): layerwise TTFT-avg degradation vs HBM-only narrows from +18.1% to +1.3% (2775.5ms -> 2380.4ms, 3-run mean) and TTFT P90 overtakes HBM by 11.8%. - Qwen3.5-397B-A17B w8a8 (TP16): degradation narrows from +16.4% to +7.7% and TTFT P90 stays below HBM across repeated runs. Verification was performed on PR vllm-project#15479 as of commit a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and all ascend_store unit tests pass locally (418 passed). Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
…se hit checks (#16947) ## Summary Three host-side optimizations for the layerwise KV pool transfer path (`pool_worker.py` / `pool_scheduler.py`) that cut per-step memcache RPC overhead under prefix-sharing workloads, building on the layerwise KV pool for the V2 model runner (#15479). ## What's inside 1. **Deferred last-layer save drain** (`pool_worker.py`, **opt-in**): the host-side wait + event reset for the final layer's save can run at the start of the NEXT step's `start_load_kv`, absorbing the send thread's tail latency into the inter-step scheduling gap. Switchable via `VLLM_ASCEND_LW_DEFER_LAST_SAVE` (default **off** — the long-standing synchronous flush contract is preserved by default; upstream tests and downstream readers rely on it). Note: the deferral relaxes the synchronous final-layer flush contract (`_wait_for_final_layer_save` used to run inline at the end of `save_kv_layer`); the flush is instead confirmed at the next step's `start_load_kv`. `test_mooncake_hybrid.py` asserts post-save visibility, so it gains explicit drain sync points mirroring its existing intermediate-layer event-wait pattern. 2. **Per-step RPC result caches** (`pool_worker.py`): concurrent requests sharing a prefix issue identical memcache queries; `batch_get_key_info` / `batch_add_lease` / `batch_is_exist` results are cached per step (reset at the start of `process_layer_data`) so each distinct key is queried at most once per step. Lease result codes are cached per key and the result list stays aligned with `valid_gva_indices` to avoid mis-zeroing GVAs. 3. **Scheduler-side block hit cache** (`pool_scheduler.py`): the layerwise hit check is a sequential prefix scan; a block's all-rank presence is cached per `(group, block_hash)` with a 30s TTL and a 200k-entry cap. A cached miss short-circuits the whole group with zero RPCs; uncached segments are fetched in one batch and the result (hit or miss) is recorded for later requests. The deferred drain is pipeline-parallel aware: it mirrors `save_kv_layer`'s `num_local` computation (`layerwise_key_layers` under `pp_size > 1`) so only per-stage local layers are drained. ## Verification (Atlas 800T A3, memcache backend, aisbench prefix_cache) 512 requests / input 16384 / output 1 / concurrency 40 / prefix repeat rate 0.9. | Model | Metric | HBM-only | Layerwise before | Layerwise after | |---|---|---:|---:|---:| | Qwen3.5-35B-A3B (TP8) | TTFT avg | 2350.2 ms | 2775.5 ms (+18.1%) | **2380.4 ms (+1.3%, 3-run mean)** | | Qwen3.5-35B-A3B (TP8) | TTFT P90 | 2752.1 ms | 2857.3 ms | **2427.6 ms (-11.8%, overtakes HBM)** | | Qwen3.5-397B-A17B w8a8 (TP16) | TTFT avg | 6127.5 ms | 7127.0 ms (+16.4%) | **6597.4 ms (+7.7%)** | | Qwen3.5-397B-A17B w8a8 (TP16) | TTFT P90 | 7216.6 ms | 7342.7 ms | **6799.7 ms (-5.8%, overtakes HBM)** | Prefix hit-rate equivalence is maintained: the external pool matches the internal HBM prefix cache (87-88% across all models and runs). Note: performance numbers were measured on #15479 @ a08ddff in a vllm 0.27.1-based environment; this patch is rebased onto current main and validated via the unit tests below. ## Unit tests All 440 tests in `tests/ut/distributed/ascend_store/` pass locally against vLLM 84030bbe (current main baseline, including the Mooncake hybrid suite); `ruff format --check` and `ruff check` clean. ## vLLM dependency - vLLM main: vllm-project/vllm@ced6857 --------- Signed-off-by: zss-hp <321708722+zss-hp@users.noreply.github.com> Co-authored-by: zss-hp <321708722+zss-hp@users.noreply.github.com>
Summary
Supersedes #12711 (V1-runner version, preserved on branch
backup/pr-12711-layerwise-v1-runner). This PR re-implements layerwise KV pool support for Qwen3.5 hybrid GDN+full-attention models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1), which is where new hybrid/mamba features are expected to land.What's inside
Keep hybrid-aligned block size against draft-config re-verification (
utils.py): the V2 MTP speculator builds the draftVllmConfigviadataclasses.replace, which re-runs the config hooks on the sharedcache_config. Because the draft model is not hybrid,refresh_block_sizedowngraded the mamba-page-aligned attention block size (1536 for Qwen3.5) back to 128, corrupting the layerwise pool's address layout (batch_copy-3102, vector-core exceptions) and shrinking usable KV capacity ~5.4x. The fix skips the generic downgrade when the shared cache_config was already sized for a hybrid model (mamba_page_size_paddedset).NPU-safe mamba align pre-copy kernel (
ops/triton/mamba/precopy.py): upstream'sprecopy_mamba_align_fused_kernelvectorizes the temporal-state copy through uint64 loads/stores, which the Ascend vector core does not support (vector-core exception as soon as the V2 mamba align path runs). Adds a byte-wise uint8 variant (same semantics, pointer cast hoisted out of the loop, mirroring the existing NPUpostprocess_mamba_fused_kernelfix) and installs it frompatch_mamba_utils.Per-layer mamba state copy deferred behind layerwise KV loads (
worker/v2/model_states/mamba_hybrid.py+ connectors):AscendMambaHybridModelState.preprocess_statekeeps the upstream GPU-resident decision kernel but hands the actual align pre-copy to a layerwise-capable connector viaprepare_mamba_state_copy. Each mamba layer's copy is launched fromdo_mamba_copy_for_layerright after that layer's KV load (conv/ssm state included) completes — sliced to the layer's own rows of the GPU-resident metadata, no CPU-GPU sync. The next step'spreprocess_statevalidates every deferred layer executed.AscendStoreConnector/AscendMultiConnectorexpose this V2 interface;use_layerwisenow requires the V2 model runner.Layerwise pool scheduler fixes carried over from [Feature] Support layerwise KV pool for Qwen3.5 hybrid GDN+full_attention models #12711 (runner-agnostic): conditional eagle/MTP tail trim in
pool_scheduler(fixes external hit 61.4%→76.7%), full-hit tail re-store inpool_worker, GDN layerwise wait inside the custom-op body,has_connector_metadataguard for the wait/save hooks.Verification (Atlas 800T A2, Qwen3.5-27B, TP=4, MTP num_speculative_tokens=3, memcache backend)
kv_both, in10000/out100/rr0.9, cc200)Layer protocol counters verified healthy under load: 4200 consecutive steps each saw exactly 67 wait/save hook calls with
current_layerreachingnum_layers=65(no interleaving/race under async scheduling).Known limitation (not introduced by this PR)
MTP acceptance length drops under PIECEWISE cudagraph mode + MTP on the V2 runner (1.0-1.3 vs 3.0+ tokens/step under FULL_DECODE_ONLY). Layerwise KV pool requires PIECEWISE (same as V1), so it surfaces there, but the control experiment above shows unlayerwise+PIECEWISE is equally affected — it is an upstream V2 speculator/PIECEWISE interaction that deserves a separate issue. This PR introduces no MTP regression relative to other PIECEWISE configurations (9.5% vs 13.7% within the same degraded mode).
Unit tests
tests/ut/worker/test_model_runner_v2_mamba.py: per-layer deferral scheduling, sliced kernel launch, missing-layer validation, connector takeover/releasetests/ut/distributed/ascend_store/test_ascend_store_connector.py: copy-after-load ordering, V2 gate, non-layerwise keeps bulk pathtests/ut/kv_offload/test_ascend_multi_connector.py: copy runs after all child connectors finish the layer loadAll 384 UTs in the touched modules pass locally;
ruff format --checkandruff checkclean.