Skip to content

[Feature] Support layerwise KV pool for Qwen3.5 hybrid GDN+full_attention models - #12711

Merged
kunpengW-code merged 5 commits into
vllm-project:mainfrom
tyy0829:main
Sep 2, 2026
Merged

kunpengW-code merged 5 commits into
vllm-project:mainfrom
tyy0829:main

Conversation

@tyy0829

@tyy0829 tyy0829 commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Description

This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next) hybrid models that combine GDN (Gated Delta Net) linear attention layers with standard full attention layers.

Changes

  1. GDN forward (ops/gdn.py): Add wait_for_kv_layer_from_connector(self.prefix) at the start of forward() and record_attention_compute_start() before the custom op call. GDN layers do not go through the @maybe_transfer_kv_layer decorator, so they must explicitly call these functions to participate in layerwise KV transfer.

  2. Attention utils (attention/utils.py): Add connector.has_connector_metadata() guard to wait_for_kv_layer_from_connector and maybe_save_kv_layer_to_connector, matching the guard in the @maybe_transfer_kv_layer decorator. Prevents spurious current_layer counter increments during non-save steps (e.g. profile run).

  3. Pool scheduler (pool_scheduler.py): Skip touch_sending_mamba_blocks when use_layerwise=True. The layerwise send thread (KVCacheStoreLayerSendingThread) does not use the completed_events mechanism, so touched mamba blocks would never be freed, leaking the block pool.

  4. Postprocess kernel (ops/triton/mamba/postprocess.py): Update postprocess_mamba_fused_kernel signature to match vLLM v0.25.0, adding state_dim_row_count_ptr, state_dim_row_stride_ptr, idx_mapping_ptr, CONV_STATE_DIM_FIRST, HAS_IDX_MAPPING, and PRECOMPUTED_NEW_COMPUTED parameters. Preserves the triton-ascend pointer-type-cast optimization.

Testing

Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B:

  • Service starts successfully with use_layerwise=true and MemCache backend
  • KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix cache hit rate 22.9%
  • 5 concurrent requests all succeed with cache hits
  • MTP speculative decoding (num_speculative_tokens=3, method=qwen3_5_mtp) works correctly

(Replaces #12526, which was locked-closed after a force-push to the head branch; this PR is based on the clean single-commit feature.)

…n models

- Add wait_for_kv_layer_from_connector and record_attention_compute_start to GDN forward (ops/gdn.py)
- Add has_connector_metadata guard to wait/save utils (attention/utils.py)
- Skip touch_sending_mamba_blocks for layerwise to prevent block pool leak (pool_scheduler.py)

Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for layerwise KV pooling in Qwen3.5 hybrid models that utilize both Gated Delta Net (GDN) and full attention layers. The changes ensure that GDN layers correctly participate in the KV transfer process, improve the reliability of connector metadata handling, and resolve potential memory leaks in the pool scheduler. Additionally, the update includes necessary adjustments to maintain compatibility with the latest vLLM framework versions.

Highlights

  • GDN Layer Integration: Updated GDN forward operations to explicitly call KV layer synchronization and attention compute recording, enabling participation in layerwise KV transfer.
  • KV Connector Robustness: Added metadata checks to KV connector utilities to prevent spurious counter increments during non-save operations.
  • Memory Leak Prevention: Modified the pool scheduler to skip block touching when layerwise transfer is enabled, preventing memory leaks in the block pool.
  • Triton Kernel Compatibility: Updated the mamba postprocess kernel signature to maintain compatibility with vLLM v0.25.0 requirements.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Ops][Misc] Integrate connector metadata checks and GDN attention synchronization

Suggested PR Summary:

### What this PR does / why we need it?
This pull request introduces several improvements and synchronization mechanisms for the Ascend backend:
1. Adds checks for `connector.has_connector_metadata()` in KV layer wait and save utilities to prevent operations when metadata is missing.
2. Bypasses `touch_sending_mamba_blocks` when `use_layerwise` is enabled.
3. Integrates `wait_for_kv_layer_from_connector` and `record_attention_compute_start` into the GDN attention forward pass to ensure proper synchronization.

Feedback:
- In `vllm_ascend/attention/utils.py`, defensive checks should be added to ensure `connector` is not `None` and implements `has_connector_metadata` before invocation to avoid potential `AttributeError` or `TypeError`.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Not specified in the PR. Please ensure existing tests pass and add unit tests for the new synchronization logic.

forward_context: ForwardContext = get_forward_context()
attn_metadata = forward_context.attn_metadata
if attn_metadata is None:
if attn_metadata is None or not connector.has_connector_metadata():

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

To prevent potential AttributeError or TypeError if connector is None or does not implement has_connector_metadata, we should defensively check if connector is not None and has the attribute before calling it.

Suggested change
if attn_metadata is None or not connector.has_connector_metadata():
if attn_metadata is None or connector is None or not getattr(connector, "has_connector_metadata", lambda: False)():

forward_context: ForwardContext = get_forward_context()
attn_metadata = forward_context.attn_metadata
if attn_metadata is None:
if attn_metadata is None or not connector.has_connector_metadata():

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

To prevent potential AttributeError or TypeError if connector is None or does not implement has_connector_metadata, we should defensively check if connector is not None and has the attribute before calling it.

Suggested change
if attn_metadata is None or not connector.has_connector_metadata():
if attn_metadata is None or connector is None or not getattr(connector, "has_connector_metadata", lambda: False)():

@tyy0829

tyy0829 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@tyy0829
tyy0829 force-pushed the main branch 2 times, most recently from 2ae7efb to 42e4049 Compare August 20, 2026 06:34
@Pz1116 Pz1116 added the ready label Aug 20, 2026
@tyy0829

tyy0829 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor Author

/cancel
[Bot]: successfully force-cancelled the following runs:

https://github.com/vllm-project/vllm-ascend/actions/runs/32341607564

ader47 added a commit to ader47/vllm-ascend that referenced this pull request Aug 20, 2026
@wangxiyuan wangxiyuan removed the ready label Aug 20, 2026
The layerwise KV pool feature added wait_for_kv_layer_from_connector(self.prefix) and record_attention_compute_start() to AscendGatedDeltaNetAttention.forward. Both are NPU-only side effects invoked from the traced (compiled) forward, so torch.compile(fullgraph=True) graph-broke on the Mock connector call and the global/lock access, failing test_connector_observes_updated_gdn_state_for_each_compiled_call on the CPU and A2 runners. Stub them (consistent with the existing clear_ssm_states/get_pcp_group stubs) so the inductor graph stays fullgraph.

Signed-off-by: tyy0829 <1455207791@qq.com>
@Pz1116 Pz1116 added the ready-precise run selected e2e test for pr label Aug 21, 2026
@tyy0829
tyy0829 force-pushed the main branch 2 times, most recently from ba8c68c to c9708cb Compare August 26, 2026 09:31
@tyy0829

tyy0829 commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@tyy0829

tyy0829 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

Add a comment block on the deferred Mamba state copy hook explaining why
the copy must run after the layerwise KV load finishes and how the
per-layer scheduling overlaps the copy with the remaining layer loads.

Signed-off-by: tyy0829 <1455207791@qq.com>
@tyy0829

tyy0829 commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

/cancel
[Bot]: cancel completed. No running workflow runs found for this commit.

@tyy0829

tyy0829 commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@tyy0829

tyy0829 commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #15479, which re-implements this feature on the V2 model runner (vllm_ascend/worker/v2), where hybrid/mamba feature development is now expected to land.

The V1-runner implementation from this PR is preserved on branch backup/pr-12711-layerwise-v1-runner for reference.

Closing this one in favor of #15479; please review there. Thanks!

@tyy0829 tyy0829 closed this Sep 1, 2026
@tyy0829 tyy0829 reopened this Sep 1, 2026
@tyy0829

tyy0829 commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

(Reopened per author request — keeping both PRs open for now: this V1-runner implementation and #15479, which re-implements the same feature on the V2 model runner.)

@tyy0829

tyy0829 commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@tyy0829

tyy0829 commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@kunpengW-code

Copy link
Copy Markdown
Collaborator

@Wangbei25 PR related to qwen3.5, PTAL

@tyy0829

tyy0829 commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor Author

/cancel
[Bot]: successfully force-cancelled the following runs:

https://github.com/vllm-project/vllm-ascend/actions/runs/33503476206

@kunpengW-code
kunpengW-code merged commit 9915508 into vllm-project:main Sep 2, 2026
32 checks passed
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 3, 2026
…tion models (vllm-project#12711)

## Description

This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next)
hybrid models that combine GDN (Gated Delta Net) linear attention layers
with standard full attention layers.

## Changes

1. **GDN forward (`ops/gdn.py`)**: Add
`wait_for_kv_layer_from_connector(self.prefix)` at the start of
`forward()` and `record_attention_compute_start()` before the custom op
call. GDN layers do not go through the `@maybe_transfer_kv_layer`
decorator, so they must explicitly call these functions to participate
in layerwise KV transfer.

2. **Attention utils (`attention/utils.py`)**: Add
`connector.has_connector_metadata()` guard to
`wait_for_kv_layer_from_connector` and
`maybe_save_kv_layer_to_connector`, matching the guard in the
`@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer`
counter increments during non-save steps (e.g. profile run).

3. **Pool scheduler (`pool_scheduler.py`)**: Skip
`touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise
send thread (`KVCacheStoreLayerSendingThread`) does not use the
`completed_events` mechanism, so touched mamba blocks would never be
freed, leaking the block pool.

4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update
`postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding
`state_dim_row_count_ptr`, `state_dim_row_stride_ptr`,
`idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and
`PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend
pointer-type-cast optimization.

## Testing

Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B:
- Service starts successfully with `use_layerwise=true` and MemCache
backend
- KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix
cache hit rate 22.9%
- 5 concurrent requests all succeed with cache hits
- MTP speculative decoding (`num_speculative_tokens=3,
method=qwen3_5_mtp`) works correctly

---
(Replaces vllm-project#12526, which was locked-closed after a force-push to the head
branch; this PR is based on the clean single-commit feature.)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
…tion models (vllm-project#12711)

## Description

This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next)
hybrid models that combine GDN (Gated Delta Net) linear attention layers
with standard full attention layers.

## Changes

1. **GDN forward (`ops/gdn.py`)**: Add
`wait_for_kv_layer_from_connector(self.prefix)` at the start of
`forward()` and `record_attention_compute_start()` before the custom op
call. GDN layers do not go through the `@maybe_transfer_kv_layer`
decorator, so they must explicitly call these functions to participate
in layerwise KV transfer.

2. **Attention utils (`attention/utils.py`)**: Add
`connector.has_connector_metadata()` guard to
`wait_for_kv_layer_from_connector` and
`maybe_save_kv_layer_to_connector`, matching the guard in the
`@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer`
counter increments during non-save steps (e.g. profile run).

3. **Pool scheduler (`pool_scheduler.py`)**: Skip
`touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise
send thread (`KVCacheStoreLayerSendingThread`) does not use the
`completed_events` mechanism, so touched mamba blocks would never be
freed, leaking the block pool.

4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update
`postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding
`state_dim_row_count_ptr`, `state_dim_row_stride_ptr`,
`idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and
`PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend
pointer-type-cast optimization.

## Testing

Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B:
- Service starts successfully with `use_layerwise=true` and MemCache
backend
- KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix
cache hit rate 22.9%
- 5 concurrent requests all succeed with cache hits
- MTP speculative decoding (`num_speculative_tokens=3,
method=qwen3_5_mtp`) works correctly

---
(Replaces vllm-project#12526, which was locked-closed after a force-push to the head
branch; this PR is based on the clean single-commit feature.)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…tion models (vllm-project#12711)

## Description

This PR adds layerwise KV pool support for Qwen3.5 (and Qwen3-Next)
hybrid models that combine GDN (Gated Delta Net) linear attention layers
with standard full attention layers.

## Changes

1. **GDN forward (`ops/gdn.py`)**: Add
`wait_for_kv_layer_from_connector(self.prefix)` at the start of
`forward()` and `record_attention_compute_start()` before the custom op
call. GDN layers do not go through the `@maybe_transfer_kv_layer`
decorator, so they must explicitly call these functions to participate
in layerwise KV transfer.

2. **Attention utils (`attention/utils.py`)**: Add
`connector.has_connector_metadata()` guard to
`wait_for_kv_layer_from_connector` and
`maybe_save_kv_layer_to_connector`, matching the guard in the
`@maybe_transfer_kv_layer` decorator. Prevents spurious `current_layer`
counter increments during non-save steps (e.g. profile run).

3. **Pool scheduler (`pool_scheduler.py`)**: Skip
`touch_sending_mamba_blocks` when `use_layerwise=True`. The layerwise
send thread (`KVCacheStoreLayerSendingThread`) does not use the
`completed_events` mechanism, so touched mamba blocks would never be
freed, leaking the block pool.

4. **Postprocess kernel (`ops/triton/mamba/postprocess.py`)**: Update
`postprocess_mamba_fused_kernel` signature to match vLLM v0.25.0, adding
`state_dim_row_count_ptr`, `state_dim_row_stride_ptr`,
`idx_mapping_ptr`, `CONV_STATE_DIM_FIRST`, `HAS_IDX_MAPPING`, and
`PRECOMPUTED_NEW_COMPUTED` parameters. Preserves the triton-ascend
pointer-type-cast optimization.

## Testing

Verified on Atlas 800T A2 (Ascend 910B3, 8 NPU) with Qwen3.5-9B:
- Service starts successfully with `use_layerwise=true` and MemCache
backend
- KV cache hit confirmed: 4 groups all hit 512 tokens, External prefix
cache hit rate 22.9%
- 5 concurrent requests all succeed with cache hits
- MTP speculative decoding (`num_speculative_tokens=3,
method=qwen3_5_mtp`) works correctly

---
(Replaces vllm-project#12526, which was locked-closed after a force-push to the head
branch; this PR is based on the clean single-commit feature.)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: tyy0829 <tyy0829@users.noreply.github.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
Co-authored-by: tyy0829 <tyy0829@users.noreply.github.com>
yiz-liu pushed a commit that referenced this pull request Sep 18, 2026
…n3.5 hybrid GDN+full-attention models on Model Runner V2 (#15479)

## Summary

Supersedes #12711 (V1-runner version, preserved on branch
[`backup/pr-12711-layerwise-v1-runner`](https://github.com/tyy0829/vllm-ascend/tree/backup/pr-12711-layerwise-v1-runner)).
This PR re-implements layerwise KV pool support for **Qwen3.5 hybrid
GDN+full-attention models on the V2 model runner**
(`VLLM_USE_V2_MODEL_RUNNER=1`), which is where new hybrid/mamba features
are expected to land.

### What's inside

1. **Keep hybrid-aligned block size against draft-config
re-verification** (`utils.py`): the V2 MTP speculator builds the draft
`VllmConfig` via `dataclasses.replace`, which re-runs the config hooks
on the *shared* `cache_config`. Because the draft model is not hybrid,
`refresh_block_size` downgraded the mamba-page-aligned attention block
size (1536 for Qwen3.5) back to 128, corrupting the layerwise pool's
address layout (batch_copy `-3102`, vector-core exceptions) and
shrinking usable KV capacity ~5.4x. The fix skips the generic downgrade
when the shared cache_config was already sized for a hybrid model
(`mamba_page_size_padded` set).

2. **NPU-safe mamba align pre-copy kernel**
(`ops/triton/mamba/precopy.py`): upstream's
`precopy_mamba_align_fused_kernel` vectorizes the temporal-state copy
through uint64 loads/stores, which the Ascend vector core does not
support (vector-core exception as soon as the V2 mamba align path runs).
Adds a byte-wise uint8 variant (same semantics, pointer cast hoisted out
of the loop, mirroring the existing NPU `postprocess_mamba_fused_kernel`
fix) and installs it from `patch_mamba_utils`.

3. **Per-layer mamba state copy deferred behind layerwise KV loads**
(`worker/v2/model_states/mamba_hybrid.py` + connectors):
`AscendMambaHybridModelState.preprocess_state` keeps the upstream
GPU-resident decision kernel but hands the actual align pre-copy to a
layerwise-capable connector via `prepare_mamba_state_copy`. Each mamba
layer's copy is launched from `do_mamba_copy_for_layer` right after that
layer's KV load (conv/ssm state included) completes — sliced to the
layer's own rows of the GPU-resident metadata, no CPU-GPU sync. The next
step's `preprocess_state` validates every deferred layer executed.
`AscendStoreConnector`/`AscendMultiConnector` expose this V2 interface;
`use_layerwise` now requires the V2 model runner.

4. **Layerwise pool scheduler fixes carried over from #12711**
(runner-agnostic): conditional eagle/MTP tail trim in `pool_scheduler`
(fixes external hit 61.4%→76.7%), full-hit tail re-store in
`pool_worker`, GDN layerwise wait inside the custom-op body,
`has_connector_metadata` guard for the wait/save hooks.

## Verification (Atlas 800T A2, Qwen3.5-27B, TP=4, MTP
num_speculative_tokens=3, memcache backend)

| Scenario | Requests | External pool hit rate | Notes |
|---|---|---|---|
| PD-mixed (`kv_both`, in10000/out100/rr0.9, cc200) | 200/200 OK |
**76.3%** (V1 baseline 76.7%) | no crashes |
| PD-disaggregated (P: cards 12-15, D: cards 8-11, in15360/out200/rr0.9,
cc200) | 400/400 OK | **P 89.6% / D 100.0%** (V1 baseline 89.95%/99.99%)
| no crashes |
| Control: unlayerwise + FULL_DECODE_ONLY | 200/200 OK | 76.3% | MTP
acceptance 69.3% |
| Control: unlayerwise + PIECEWISE | 200/200 OK | — | MTP acceptance
13.7% |

Layer protocol counters verified healthy under load: 4200 consecutive
steps each saw exactly 67 wait/save hook calls with `current_layer`
reaching `num_layers=65` (no interleaving/race under async scheduling).

### Known limitation (not introduced by this PR)

MTP acceptance length drops under **PIECEWISE cudagraph mode + MTP on
the V2 runner** (1.0-1.3 vs 3.0+ tokens/step under FULL_DECODE_ONLY).
Layerwise KV pool requires PIECEWISE (same as V1), so it surfaces there,
but the control experiment above shows unlayerwise+PIECEWISE is equally
affected — it is an upstream V2 speculator/PIECEWISE interaction that
deserves a separate issue. This PR introduces no MTP regression relative
to other PIECEWISE configurations (9.5% vs 13.7% within the same
degraded mode).

## Unit tests

- `tests/ut/worker/test_model_runner_v2_mamba.py`: per-layer deferral
scheduling, sliced kernel launch, missing-layer validation, connector
takeover/release
- `tests/ut/distributed/ascend_store/test_ascend_store_connector.py`:
copy-after-load ordering, V2 gate, non-layerwise keeps bulk path
- `tests/ut/kv_offload/test_ascend_multi_connector.py`: copy runs after
all child connectors finish the layer load

All 384 UTs in the touched modules pass locally; `ruff format --check`
and `ruff check` clean.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: tyy0829 <1455207791@qq.com>
tyy0829 added a commit to tyy0829/vllm-ascend that referenced this pull request Sep 20, 2026
…n V2 runner

Enable the layerwise KV pool (pooling + use_layerwise) for Kimi-K3 hybrid
KDA + MLA models on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1), the
same scenario previously enabled for Qwen3.5 GDN hybrids (vllm-project#12711/vllm-project#15479).

- ops/kimi_kda.py: KDA layers bypass the standard attention layer, so like
  the GDN path they must drive the per-layer layerwise protocol themselves.
  Add wait_for_kv_layer_from_connector + record_attention_compute_start at
  the top of the eager-break _forward body (before the conv/recurrent
  kernels touch mamba state) and maybe_save_kv_layer_to_connector on every
  exit path. Without these hooks the pool worker's current_layer counter
  desyncs (only the MLA layer advanced it) and the first multi-block save
  trips 'thread: 0 save failed' in KVCacheStoreLayerSendingThread.
- worker/v2/model_states/mamba_hybrid.py: port the per-layer mamba align
  pre-copy deferral from the Qwen3.5 V2 work. preprocess_state keeps the
  GPU-resident decision kernel but hands the actual copy to a layerwise
  connector (prepare_mamba_state_copy); each KDA layer's copy is launched
  from do_mamba_copy_for_layer right after that layer's KV load (conv/ssm
  state included) completes, so the copy never reads half-loaded state.
- AscendStoreConnector/AscendMultiConnector: prepare_mamba_state_copy now
  accepts either the V1 mamba copy buffers or the V2 mamba hybrid model
  state; wait_for_layer_load triggers the per-layer copy after the load
  finishes; finish_mamba_state_copy releases the handle.
- UT: per-layer deferral scheduling (decision kernel still runs, bulk
  pre-copy skipped, sliced kernel launch, missing-layer raises), copy-after-
  load ordering and finish-release semantics for both connectors; V1
  copy-bufs mock is now spec'd so duck-typing keeps it on the bulk path.

Verified on Atlas 800I A3 (Kimi-K3-w4a8-4layer, TP=8, EP, eager,
AscendStoreConnector memcache device_sdma backend, use_layerwise=true):
- Service starts under VLLM_USE_V2_MODEL_RUNNER=1; /health 200
- Long-prefix repeats hit the external pool: 768-token block restored per
  repeat request (kvpool hit tokens: 768), repeat request latency 3.4s ->
  0.57s
- Stress 104/104 requests OK (8-way concurrency, 90% repeat rate),
  external prefix cache hit rate 64.6%, zero worker errors
- tests/ut: 44/44 pass (mamba_hybrid deferral + connector duck-typing)

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
weijinqian0 pushed a commit that referenced this pull request Sep 23, 2026
### What this PR does / why we need it?

Mooncake hybrid layerwise currently rejects layouts containing recurrent
Mamba state. Removing that check alone is unsafe: GDN, Kimi KDA, and
other Mamba-like layers update conv/recurrent state inside attention, so
an attention-entry PUT can publish the previous state while the
following request reports a cache hit.

This change:

- identifies physical layers containing `MambaSpec`, including per-layer
`UniformTypeKVCacheSpecs`;
- keeps full/sliding-window attention PUTs at attention entry;
- preserves the post-compute save hook for recurrent layers so Mooncake
publishes updated state;
- reuses the stage-local layer mapping and PP-scoped key namespace from
#17183, mapping global recurrent layer names to the correct local layer
on each pipeline stage;
- applies the existing send-backlog policy to recurrent and record-only
save paths;
- drains queued PUTs at step completion, including when the final layer
has no ranges to save.

The generic GDN/Mamba load, deferred state-copy, and connector hooks
from #12711 are already present on main. This PR only adds the Mooncake
block-key publication semantics. The implementation and lifecycle
coverage are based on #17194 and credit its author in the commit.

| Model | Cards | Dtype| layerwise | kv pool | input | output | External
prefix cache | Prefix cache | P | D | Batchsize | Request | TTFT/s |
TPOT/ms | TPS |

|------|------|----------|------------------|----------|------|------|------------------------|--------------|-----------|-----------|--------|----------|--------|---------|------------------|
| Qwen3.5-35B | 2 | w8a8 | × | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 |
128 | 6.07 | / | 41606 |
| Qwen3.5-35B | 2 | w8a8 | √ | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 |
128 | 4.93 | / | 53418 |

### Does this PR introduce _any_ user-facing change?

Yes. `AscendStoreConnector` with `backend="mooncake"` and
`use_layerwise=true` now accepts hybrid recurrent-state groups using
`mamba_cache_mode="align"`. This also composes with Mooncake layerwise
pipeline parallel support from #17183. Existing topology and
prefill/decode TP-mismatch restrictions remain unchanged.

### How was this patch tested?

- Targeted Mooncake layerwise/Mamba mock tests: 26 passed.
- PP+Mamba worker regression: 5 passed, 2 subtests; the new case covers
PP rank 1, global-to-local layer mapping, regular attention-entry PUT,
and Mamba post-compute PUT for raw and uniform-wrapped specs.
- CI PP layer-count regression: 2 passed; non-layerwise partial worker
initialization no longer reads layerwise-only cache specs.
- `test_pool_worker.py` excluding unrelated KVPP tests: 96 passed, 2
deselected, 48 subtests passed.
- Broader AscendStore mock suite: 440 passed, with the same 3
environment/mock failures reproduced on unmodified main (435 passed,
same failures).
- `ruff check` passed for all changed Python files.
- `ruff format --check` passed for all changed Python files.
- `python -m py_compile` passed for all changed Python files.
- Added a raw-token lifecycle regression covering Mamba2, GDN, generic
linear attention, Full/SWA, merged/per-layer specs, chunked prefill,
decode, cache hits, shared prefixes, and missing-state fallback. This
file requires a full vLLM/PyTorch CPU test environment and was not
executed in the local Windows mock environment.
- The equivalent patch on commit
`ca91dfa600373c56cfffc44be1d77028b0dd621d` was validated successfully by
the contributor on the target deployment. This latest-main rebase
preserves the same transfer semantics; real NPU/Mooncake validation was
not rerun in the local development environment.

`bash format.sh ci` could not run locally because `pre-commit` is not
installed.


- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: fangrongcan <17343701736@163.com>
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
hw-qianlingfeng pushed a commit to hw-qianlingfeng/vllm-ascend that referenced this pull request Sep 27, 2026
…17287)

### What this PR does / why we need it?

Mooncake hybrid layerwise currently rejects layouts containing recurrent
Mamba state. Removing that check alone is unsafe: GDN, Kimi KDA, and
other Mamba-like layers update conv/recurrent state inside attention, so
an attention-entry PUT can publish the previous state while the
following request reports a cache hit.

This change:

- identifies physical layers containing `MambaSpec`, including per-layer
`UniformTypeKVCacheSpecs`;
- keeps full/sliding-window attention PUTs at attention entry;
- preserves the post-compute save hook for recurrent layers so Mooncake
publishes updated state;
- reuses the stage-local layer mapping and PP-scoped key namespace from
vllm-project#17183, mapping global recurrent layer names to the correct local layer
on each pipeline stage;
- applies the existing send-backlog policy to recurrent and record-only
save paths;
- drains queued PUTs at step completion, including when the final layer
has no ranges to save.

The generic GDN/Mamba load, deferred state-copy, and connector hooks
from vllm-project#12711 are already present on main. This PR only adds the Mooncake
block-key publication semantics. The implementation and lifecycle
coverage are based on vllm-project#17194 and credit its author in the commit.

| Model | Cards | Dtype| layerwise | kv pool | input | output | External
prefix cache | Prefix cache | P | D | Batchsize | Request | TTFT/s |
TPOT/ms | TPS |

|------|------|----------|------------------|----------|------|------|------------------------|--------------|-----------|-----------|--------|----------|--------|---------|------------------|
| Qwen3.5-35B | 2 | w8a8 | × | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 |
128 | 6.07 | / | 41606 |
| Qwen3.5-35B | 2 | w8a8 | √ | √ | 16k | 1 | 90% | / | DP1TP1 | / | 32 |
128 | 4.93 | / | 53418 |

### Does this PR introduce _any_ user-facing change?

Yes. `AscendStoreConnector` with `backend="mooncake"` and
`use_layerwise=true` now accepts hybrid recurrent-state groups using
`mamba_cache_mode="align"`. This also composes with Mooncake layerwise
pipeline parallel support from vllm-project#17183. Existing topology and
prefill/decode TP-mismatch restrictions remain unchanged.

### How was this patch tested?

- Targeted Mooncake layerwise/Mamba mock tests: 26 passed.
- PP+Mamba worker regression: 5 passed, 2 subtests; the new case covers
PP rank 1, global-to-local layer mapping, regular attention-entry PUT,
and Mamba post-compute PUT for raw and uniform-wrapped specs.
- CI PP layer-count regression: 2 passed; non-layerwise partial worker
initialization no longer reads layerwise-only cache specs.
- `test_pool_worker.py` excluding unrelated KVPP tests: 96 passed, 2
deselected, 48 subtests passed.
- Broader AscendStore mock suite: 440 passed, with the same 3
environment/mock failures reproduced on unmodified main (435 passed,
same failures).
- `ruff check` passed for all changed Python files.
- `ruff format --check` passed for all changed Python files.
- `python -m py_compile` passed for all changed Python files.
- Added a raw-token lifecycle regression covering Mamba2, GDN, generic
linear attention, Full/SWA, merged/per-layer specs, chunked prefill,
decode, cache hits, shared prefixes, and missing-state fallback. This
file requires a full vLLM/PyTorch CPU test environment and was not
executed in the local Windows mock environment.
- The equivalent patch on commit
`ca91dfa600373c56cfffc44be1d77028b0dd621d` was validated successfully by
the contributor on the target deployment. This latest-main rebase
preserves the same transfer semantics; real NPU/Mooncake validation was
not rerun in the local development environment.

`bash format.sh ci` could not run locally because `pre-commit` is not
installed.


- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: fangrongcan <17343701736@163.com>
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants