Repository navigation
[Refactor][Model] Decouple indexer from SFA with unified forward pipeline - #15669
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request refactors the indexer implementation to improve modularity by decoupling it from the AscendSFAImpl class. By introducing a standalone AscendSFAIndexerBackend that handles its own compute and metadata, the architecture becomes more maintainable and better suited for NPU-specific optimizations. The changes also streamline the interaction between the indexer and the SFA layer, replacing hardcoded logic with a cleaner delegation pattern. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request refactors the SFA indexer cache implementation by introducing AscendSFAIndexerBackend and AscendSFAIndexerMetadata to handle compute and cache persistence, moving logic out of the IndexerWrapper and AscendSFAImpl. It adds support for NPU-based kernels to replace hardcoded CUDA fp8 paths, introduces metadata builders for efficient cache management, and updates tests to reflect these architectural changes. A critical issue was identified regarding a potential AttributeError when accessing k_cache.prefix if k_cache is None.
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
9edf15b to
8fd7046
Compare
a45803e to
a5dec30
Compare
137267a to
ec00247
Compare
a99c3ac to
989cfd1
Compare
- Consolidate AscendLightningIndexer into AscendSFAIndexerBackend and drop the AscendIndexerBase indirection (ops/indexer base.py/lightning.py removed) - Unify forward_k / write_cache / forward into a single backend forward: k path -> parallel-layout gather (PCP / DSA-CP) -> cache write -> top-k selection - Indexer builds its own attn metadata via AscendSFAIndexerMetadataBuilder (slot mapping, block table, seq lens, group_*, cos/sin); SFA fetches it by k_cache prefix - IndexerWrapper in ops/mla.py instantiates the backend directly and keeps weight module paths for load/quant compatibility - skip_topk layers still run the k path and cache write (compute_topk=False) so the shared-index cache stays up to date - Adapt to upstream c8_reshape_optim_enabled config rename (vllm-project#15752) - Update UTs: indexer gather/metadata build, wrapper delegation, SFA metadata fetch Signed-off-by: Biuapha <1731372716@qq.com>
The indexer k path now lives inside the indexer backend, so k_li/k_li_scale are always None at the _prepare_kv_for_parallel / _store_parallel_kv call sites. Remove the dead parameters, the unreachable k_li branches in the PCP overrides, and fold the identical two-way splits; shrink the return tuples accordingly. Addresses review comments from weijinqian0. Signed-off-by: Biuapha <1731372716@qq.com>
During MTP draft propose the draft forward runs under the main model runner's forward context, whose metadata dict carries only main-layer keys - neither the draft's own indexer cache prefix nor its attn layer name is present. Restore the kv_sharing_target_layer_name resolution so a KV-sharing layer looks up the sharing target's indexer cache prefix first, keeping the layer-name fallback for the proposer-built context. Fixes a RuntimeError crash with MTP + FULL_DECODE_ONLY graph replay. Signed-off-by: Biuapha <1731372716@qq.com>
AscendSFAImpl instances created via __new__ in unit tests (and any legacy construction path that skips __init__) have no kv_sharing_target_layer_name attribute; use getattr with a None default, and add a UT covering the kv-sharing target prefix resolution. Signed-off-by: Biuapha <1731372716@qq.com>
…line (vllm-project#15669) ### What this PR does / why we need it? Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA: - Introduces `AscendSFAIndexerBackend` (`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in one unified `forward` - k-path (`forward_k`) -> parallel-layout gather (`_gather_cache_inputs`) -> `write_cache` -> top-k selection. `compute_topk=False` keeps the skip_topk layers' cache up to date while reusing shared top-k indices. - The indexer builds its own attention metadata via `AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA's `attn_metadata`, so a variant indexer with a compressed cache layout gets its own slot mapping. - Parallel-layout transforms are internal to the backend: PCP all-gathers the prefill region (reordering slot mapping to the gathered layout), DSA-CP all-gathers the indexer k across the TP group. SFA no longer carries indexer-specific gather/write hooks. - `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (`...indexer.<name>`) that weight loading and quant name mapping key off, and delegates the unified `forward` to the backend. - Removes the now-dead `ops/indexer/base.py` / `ops/indexer/lightning.py`. ### Does this PR introduce _any_ user-facing change? No. Pure internal refactor; no API, config, or behavior change. ### How was this patch tested? - Unit tests updated/added: `tests/ut/attention/test_indexer.py` (metadata builder, unified forward, PCP/DSA-CP gather), `tests/ut/attention/test_sfa_v1.py`, `tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`. - Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts). - Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`, `enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture passes with DSA-CP enabled). - GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`, num_speculative_tokens=5) and `enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on: - pre-refactor baseline: 1283/1319 = 97.27% - with this PR: 1286/1319 = 97.50% Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Biuapha <1731372716@qq.com>
…line (vllm-project#15669) ### What this PR does / why we need it? Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA: - Introduces `AscendSFAIndexerBackend` (`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in one unified `forward` - k-path (`forward_k`) -> parallel-layout gather (`_gather_cache_inputs`) -> `write_cache` -> top-k selection. `compute_topk=False` keeps the skip_topk layers' cache up to date while reusing shared top-k indices. - The indexer builds its own attention metadata via `AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA's `attn_metadata`, so a variant indexer with a compressed cache layout gets its own slot mapping. - Parallel-layout transforms are internal to the backend: PCP all-gathers the prefill region (reordering slot mapping to the gathered layout), DSA-CP all-gathers the indexer k across the TP group. SFA no longer carries indexer-specific gather/write hooks. - `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (`...indexer.<name>`) that weight loading and quant name mapping key off, and delegates the unified `forward` to the backend. - Removes the now-dead `ops/indexer/base.py` / `ops/indexer/lightning.py`. ### Does this PR introduce _any_ user-facing change? No. Pure internal refactor; no API, config, or behavior change. ### How was this patch tested? - Unit tests updated/added: `tests/ut/attention/test_indexer.py` (metadata builder, unified forward, PCP/DSA-CP gather), `tests/ut/attention/test_sfa_v1.py`, `tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`. - Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts). - Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`, `enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture passes with DSA-CP enabled). - GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`, num_speculative_tokens=5) and `enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on: - pre-refactor baseline: 1283/1319 = 97.27% - with this PR: 1286/1319 = 97.50% Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Biuapha <1731372716@qq.com>
While this branch was open, main grew its own implementations of both features it was proposing: - vllm-project#15913 GLM-Next KV cache management, which keeps the incomplete pool in an absolute-position FP32 Glm5NextStateCache rather than in this branch's paged Glm5NextTailCache, and stores completed pools as unquantized BF16 - vllm-project#15669 decoupled the indexer from SFA, restructuring the very plumbing this branch's kpool backend was wired into - AscendDflash2Proposer, selected by is_dflash2_draft() under method dflash Those designs are mutually exclusive with this branch's, not textually conflicting with it, so every conflict is resolved in favour of main and the branch's own kpool backend, tail cache and DFlash2 patch are dropped. The follow-up commit re-adds only what main still lacks. Also drops CI-config churn this branch had picked up from a stale tree: it reverted its own parent vllm-project#14683 by deleting GLM-5.2-W4A8C8-SFA-DCP.yaml and moving nightly entries into weekly, and added an unreferenced DeepSeek-V3.2-W8A8-DCP.yaml. Signed-off-by: yiminghub2024 <482890@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
…line (vllm-project#15669) ### What this PR does / why we need it? Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA: - Introduces `AscendSFAIndexerBackend` (`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in one unified `forward` - k-path (`forward_k`) -> parallel-layout gather (`_gather_cache_inputs`) -> `write_cache` -> top-k selection. `compute_topk=False` keeps the skip_topk layers' cache up to date while reusing shared top-k indices. - The indexer builds its own attention metadata via `AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA's `attn_metadata`, so a variant indexer with a compressed cache layout gets its own slot mapping. - Parallel-layout transforms are internal to the backend: PCP all-gathers the prefill region (reordering slot mapping to the gathered layout), DSA-CP all-gathers the indexer k across the TP group. SFA no longer carries indexer-specific gather/write hooks. - `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (`...indexer.<name>`) that weight loading and quant name mapping key off, and delegates the unified `forward` to the backend. - Removes the now-dead `ops/indexer/base.py` / `ops/indexer/lightning.py`. ### Does this PR introduce _any_ user-facing change? No. Pure internal refactor; no API, config, or behavior change. ### How was this patch tested? - Unit tests updated/added: `tests/ut/attention/test_indexer.py` (metadata builder, unified forward, PCP/DSA-CP gather), `tests/ut/attention/test_sfa_v1.py`, `tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`. - Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts). - Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`, `enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture passes with DSA-CP enabled). - GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`, num_speculative_tokens=5) and `enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on: - pre-refactor baseline: 1283/1319 = 97.27% - with this PR: 1286/1319 = 97.50% Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Biuapha <1731372716@qq.com>
…ard pipeline (vllm-project#15669)" This reverts commit 44e3160. Signed-off-by: chenzeyu <2978509328@qq.com> # Conflicts: # vllm_ascend/attention/indexer.py
…line (vllm-project#15669) ### What this PR does / why we need it? Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA: - Introduces `AscendSFAIndexerBackend` (`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in one unified `forward` - k-path (`forward_k`) -> parallel-layout gather (`_gather_cache_inputs`) -> `write_cache` -> top-k selection. `compute_topk=False` keeps the skip_topk layers' cache up to date while reusing shared top-k indices. - The indexer builds its own attention metadata via `AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA's `attn_metadata`, so a variant indexer with a compressed cache layout gets its own slot mapping. - Parallel-layout transforms are internal to the backend: PCP all-gathers the prefill region (reordering slot mapping to the gathered layout), DSA-CP all-gathers the indexer k across the TP group. SFA no longer carries indexer-specific gather/write hooks. - `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (`...indexer.<name>`) that weight loading and quant name mapping key off, and delegates the unified `forward` to the backend. - Removes the now-dead `ops/indexer/base.py` / `ops/indexer/lightning.py`. ### Does this PR introduce _any_ user-facing change? No. Pure internal refactor; no API, config, or behavior change. ### How was this patch tested? - Unit tests updated/added: `tests/ut/attention/test_indexer.py` (metadata builder, unified forward, PCP/DSA-CP gather), `tests/ut/attention/test_sfa_v1.py`, `tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`. - Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts). - Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`, `enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture passes with DSA-CP enabled). - GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`, num_speculative_tokens=5) and `enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on: - pre-refactor baseline: 1283/1319 = 97.27% - with this PR: 1286/1319 = 97.50% Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Biuapha <1731372716@qq.com> Signed-off-by: like-0517 <ithwlike@126.com>
) ### What this PR does / why we need it? After #15669 separated the SFA indexer cache and metadata builder, replicated DCP can expose an address mismatch: SFA uses a rank-local KV-cache layout, while the indexer cache is replicated across DCP ranks. Building indexer metadata from the SFA-local block table or slot mapping can therefore read or write the wrong indexer cache locations. This PR makes the indexer builder derive its complete cache view from common attention metadata and the explicitly supplied CP context: `CommonAttentionMetadata + CP context -> indexer builder -> AscendSFAIndexerMetadata` The builder owns the final block table and write-ready slots, RoPE tables, query/key sequence lengths, PCP decode boundary, and LI-C8 cache-write grouping metadata. Mutable derived buffers are keyed by the persistent common slot-mapping buffer so graph capture, replay, and MTP draft steps use the appropriate storage. Existing parallel calculations are shared through stateless helpers: `get_cp_local_query_key_lens` and `build_pcp_ordered_slot_mapping` in `context_parallel/common_cp.py`, and replicated-DCP address / PCP global-view helpers in `context_parallel/sfa_dcp_utils.py`. SFA and indexer builders call those helpers and retain their own metadata buffers. #### Why `sfa_v1.py` changes | Change | Reason | | --- | --- | | Move LI-C8 group metadata construction to the indexer builder; remove the old SFA-side construction and field assignments. | The grouping must describe the indexer's final slots and block size. Removing the unused SFA-side call also avoids duplicate auxiliary-metadata work; SFA's own C8 KV-write path remains intact. | | Look up the sharing target's indexer metadata, then the draft indexer's own metadata; remove fallback to the SFA layer's metadata. | KV-sharing draft layers can have their own metadata even when physical cache storage is shared. SFA metadata is not a valid substitute for the indexer's replicated cache view. | | Remove forward-time assignments of `actual_seq_lengths_query`, `actual_seq_lengths_key`, and `num_decode_tokens`. | These fields are already constructed by the indexer builder. Copying them from SFA would preserve a dependency on SFA's construction and buffer ownership. | | Read `cos` / `sin` inside indexer forward from indexer metadata instead of passing SFA's tensors. | This completes the indexer's metadata ownership, including CP slicing and graph/MTP storage. It reuses the existing RoPE calculation. | The direct DCP bug fix is the replicated cache-address construction. The RoPE interface change and removal of unused SFA auxiliary fields complete the independent-metadata design; they do not introduce a different SFA attention algorithm. The proposer retains metadata builders for cache-only layers. It identifies these groups through the existing backend contract (`get_impl_cls() is None`), not a concrete KV-cache-spec type, so an independent tail-cache spec is not mistaken for the primary attention group. The backend query runs in the proposer's vLLM configuration context. During MTP, common sequence/position/slot state advances once through the primary attention group, then each cache-only builder constructs its own metadata from that updated common view. Cache-only groups are excluded from primary graph/backend selection, not from metadata construction. This remains the existing one-main-attention-group path; it does not add generic multi-KV-cache scheduling or heterogeneous full-graph support. The Step3.5 and Gemma4 proposers are unchanged. ### Does this PR introduce _any_ user-facing change? It fixes indexer cache addressing and removes reliance on SFA metadata being used as an indexer fallback or patched into indexer metadata during forward. There is no public API or configuration-format change. The separate compiled MoE reduction issue is addressed by #16495. The hardware comparison below includes that patch only in an isolated test overlay; it is not included in this PR. ### How was this patch tested? #### Revision and source checks - **Latest-main conflict resolution (`ec6e301d`, 2026-09-15):** rebased the single PR commit onto main `0e0910b0`. The only textual conflict was the existing `_make_runner()` helper in `tests/ut/worker/test_model_runner_v2.py`: main #16574 already supplies `adaptive_verification` and `use_fia`, so this PR now retains only the still-missing `attn_groups = []` initialization. `git range-diff` shows no production-logic change from the rebase. Ruff check passed for all 19 changed files, the conflict file passes Ruff format check, `git diff --check` passes, and the paired vLLM pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. - **DSpark adaptive graph CI fix (`ec6e301d`, 2026-09-15):** the failed A3 eight-card case was `test_glm5_2_dspark_eager[adaptive]`; MTP and fixed DSpark passed in the same job. Adaptive FULL graph records its graph-shaped token width in `positions.shape[0]`, while `common_attn_metadata.num_input_tokens` can retain the compact count. Before the split, indexer RoPE came from SFA, whose builder already handles this contract. The independent indexer builder now applies the same width when constructing its own positions/RoPE/slot view. One focused unit regression was added; Ruff check, targeted format check, and `git diff --check` pass. Local full pytest was not claimed because the available Windows helper environment lacks the paired vLLM dependency set; CI is the authoritative rerun. No DSpark/model-runner/CI test behavior or assertion was changed. The first CI revision exposed four existing test doubles that omit the optional `method` field; the final predicate uses `getattr(..., None)` for that field as well, without changing those fixtures. That CPU-UT run reported 4624 passed / 4 failed before this one-line compatibility correction. - **Current CI follow-up (`f0945801`, 2026-09-15):** compared with `242e9751`, this adds only three initialization lines to the existing `_make_runner()` test helper in `tests/ut/worker/test_model_runner_v2.py`: `attn_groups = []`, `adaptive_verification = None`, and `use_fia = False`. No production code or PR base changed; the PR remains one signed-off commit. The [previous CPU-UT job](https://github.com/vllm-project/vllm-ascend/actions/runs/34923604678/job/104237578682) reported 4611 passed / 3 failed after CI rebased onto main `a643db116`. Both the tested runner code and newly added tests came from main, not this PR: their lightweight `__new__` fixture omitted fields normally initialized by the real runner. The fix preserves all assertions and changes neither production defaults nor execution paths. Ruff check, format check and `git diff --check` pass; the new [pre-commit/mypy job](https://github.com/vllm-project/vllm-ascend/actions/runs/34925651210/job/104243025168) and [CPU-UT job](https://github.com/vllm-project/vllm-ascend/actions/runs/34925651210/job/104243993813) both completed successfully. This does not claim that every NPU matrix job has finished. An independent CPU reproduction used main `a643db116` and its paired vLLM `84030bbe3d74d99bad477a3d2e37a973ccd8865c`: the complete `test_model_runner_v2.py` reproduced **33 passed / 3 failed** before the fix, then **36 passed / 0 failed** after the same three-line patch. The complete `test_attn_utils_v2.py` also passed **23/23**. No assertions were removed, dependencies installed, or NPU probes performed. Logs: `/mnt/share/g00672350/refactor/pr16325-ci-242e9751-20260915/112_cpu_model_runner_v2_unfixed.log`, `112_cpu_model_runner_v2_fixed.log`, and `112_cpu_attn_utils_v2_fixed.log`. The archive identities were verified: before `cd13ae104`, after `cbcd83bc3`; these reproduce the CI main + PR combination without changing this PR's base. Targeted Ruff checks pass. The repository-wide auto-formatting command was not run because it could modify unrelated files; full pre-commit/mypy remains covered by CI. - **Earlier mypy fix (`242e9751`):** declared `consumes_pcp_context = False` on the existing `_PrefillStateBuilder` test double; no production changes. The [pre-commit job including mypy](https://github.com/vllm-project/vllm-ascend/actions/runs/34923604678/job/104236813190) passed. This is previous-run evidence, not a claim that all checks for the new head have passed. - **Current CPU regression (112, 2026-09-15):** the complete 165-item focused Eagle/SFA/indexer suite passed under the repository's standard CPU conftest profile: **155 passed, 10 skipped, 0 failed**. It ran on exact source `242e9751` with paired vLLM `a97dacb`; current production code is identical. Log: `/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/112_cpu_eagle_sfa_indexer_standard_profile.log`. An earlier A5-profile CPU attempt had two custom-op-dispatch expectation failures; the authoritative result is the independent full rerun under the standard CPU profile, not combined partial results. The priority cache-only group/context set also passed 10/10. - **Current NPU precision/acceptance (111, 2026-09-15):** MTP1, TP8/DP1, DSA-CP off, target `FULL_DECODE_ONLY` with compilation enabled and breakable disabled: **24/24 passed**, including 9/9 original >=2K exact retrieval answers and 4/4 additional exact answers at **12,761 / 25,801 / 38,051 / 39,441 input tokens**. The full suite also contains short sanity cases. Raw Prometheus counter deltas were 64 drafts, 64 drafted tokens, 58 accepted tokens: **90.625% draft-token acceptance**, mean acceptance length **1.90625** (including the bonus token). This is a small long-input/short-output regression workload, not a matched pre-refactor acceptance-rate baseline or a throughput benchmark. Both SFA-C8 and LI-C8 were enabled; the speculative configuration retained `enforce_eager=true` for the drafter. No GSM8K. - **Current MTP5/DSA-CP-off result (111, 2026-09-15):** the same TP8/DP1, W4A4C8, `FULL_DECODE_ONLY` + #16495 configuration passed **24/24**, including all 9 original >=2K exact retrieval cases and all 4 extra 12,761 / 25,801 / 38,051 / 39,441-token cases. Counter deltas: **41 drafts, 205 drafted tokens, 103 accepted tokens**; aggregate draft acceptance **50.2439%**, mean acceptance length **3.5122**. Acceptance by draft position 1-5 was **92.6829% / 58.5366% / 51.2195% / 39.0244% / 9.7561%** (denominator 41 drafts for each position). These small, mostly short-answer requests do not establish equality to the historical interval metrics or to a matched pre-refactor/eager baseline. Evidence directory: `/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/111_mtp5_graph_dsa_off_retry1`. The initial off attempt failed before readiness because memory from the preceding stopped service had not yet been released; retry began only after all eight cards were healthy and free. No memory threshold was lowered, no device/container was reset, and no unrelated process was stopped. - **NPU source identity and remaining cases:** the tested archive is exact PR revision `242e9751` plus authorized #16495 patch `d9bae5c4614df6959caf93d9b6f6ad05778fd52f` (isolated overlay `da1e17c53d7362766162ab0583ea267ba0880f02`), paired with vLLM `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`. Current PR `ec6e301d` additionally contains the focused DSpark-adaptive indexer graph-width fix above; the prior NPU results do not validate that new line-level change. Evidence: `/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/111_mtp1_graph_dsa_off_retry2/{status.json,results/summary.json,metrics_before.txt,metrics_after.txt,service.log}`. MTP5 + DSA-CP-on at TP4/DP2 was rejected before readiness by the paired vLLM's spec-decode/SP graph-shape validation (query length 6 versus TP4); **no precision pass or failure is claimed for that configuration**. Its failed service was stopped. The MTP5/DSA-CP-off case subsequently passed as reported above. The author approved TP2/DP4 for the MTP5 + DSA-CP-on graph case. Its validation supervisor was externally terminated with exit 137 before readiness; no requests or precision/acceptance result were obtained. Container/host checks did not show OOM. The author confirmed external process termination and requested that validation wait, so further NPU launches/retries are now paused. Only task-owned test processes are being cleaned up; no device/container reset or unrelated process cleanup is performed. The TP2/DP4 result remains unvalidated, and the topology difference must be retained in any later result. 112 cards 1/2 report UB/RAS `81AF8000` alarms, so 112 is used only for CPU tests. NPU validation is now paused again at the author's request; the old recurring task remains paused and will not automatically restart tests. - **Cache-only group follow-up (`061cf2ab4`, 2026-09-15):** relative to `249c63d3`, only `llm_base_proposer.py` (+20/-14) and its existing Eagle proposer test file (+38/-5) changed. Removed the concrete SFA-indexer-spec check, used the existing backend capability for primary selection and the MTP cache-only group list, and retained the existing single-main-group restriction. No broad proposer/CP/Gemma4 refactor was restored. - Extended the existing focused test to cover all six permutations of main attention + SFA indexer + an independent `SlidingWindowSpec`-shaped cache-only tail group, configuration scope, per-layer metadata ownership, graph backend selection and the no-executable-group error. The test's six cases passed locally against the actual production method ASTs with runtime imports/spec classes/config context isolated. The old HEAD was also checked and does select an independent tail spec as primary when it comes first. **This is isolated method-level testing, not a full runtime pytest result.** - Ruff 0.14.0 `check`, `format --check` and `git diff --check` passed on both files. `format.sh ci` could not proceed because local `pre-commit` is not installed. NPU/model validation and the recurring task remain paused; no new precision or acceptance-rate result is claimed. - The author's 111 `/mnt/share/g00672350/refactor/prci` registers `KpoolTailSpec` from `models/glm5next/kv_cache.py`, where it derives from `SlidingWindowSpec`; its layer refers to `vllm.v1.attention.backends.mla.indexer.KpoolTailBackend`. That backend class was not present in the vLLM source mapped by the specified `xhg_refactor_sfa` container at inspection time. The new selection path supports such a tail backend when it follows the cache-only `get_impl_cls() is None` contract; complete KPool runtime/cache-addressing compatibility has not been validated. - **Previous test-only cleanup (`249c63d3`, 2026-09-15):** production code is unchanged from restored revision `199a459c`. Test changes relative to the PR base were reduced from 11 files (+793/-47) to 9 files (+502/-49). - Removed the extra standalone common-CP helper tests, KPool wrapper argument-binding test and large proposer dummy-capture scaffold. Merged duplicate PCP decode-boundary, DSA padding and capture-buffer cases through parameterization; reused the existing Model Runner V2 context-propagation test. Direct indexer cache-address/C8/buffer and split-MTP metadata regressions, plus required existing-test signature/mock adaptations, remain. - Ruff 0.14.0 `check`, `format --check` and `git diff --check` passed for that test-only cleanup. Runtime pytest and NPU validation were **not rerun**; validation and the recurring task remain paused at the author's request. CI/UT and hardware results below are historical evidence, not new results for this test-only revision. - **2026-09-15 restoration, before the test-only cleanup:** restored the exact pre-generalization commit `199a459c7e3b2ac749e05e6163b5dc7bcdc401bc`. The broader proposer/Gemma4/MLA-DCP generalization is no longer included. Validation and the recurring task are paused at the author's request; the results below are previously recorded evidence, not newly rerun tests. - Single commit: `ec6e301d183c5da1023652744b4e9d6395172b86`, based on current main `0e0910b060f58a8400397331462c306040dc34e4`. - At `199a459c`, the CPU-UT follow-up changed only two test files relative to `e889fc608bf1b60854816d7d9f6bbdd09429a39a`; its production files were identical to the hardware-tested pre-squash head `c30f4f421e0eea29c043f3ee8131c0333eb5100e`. `git diff --check` passed. - The previous CPU-UT run had six failures caused by stale test doubles: four SFA forward cases used the old indexer call signature, and two Eagle cases omitted the indexer cache spec's `block_size`. The focused recheck passed: SFA forward 4/4; Eagle proposer 8 passed, 10 skipped in the existing 111 container with batch-invariant mode to avoid its missing custom-op registration. Ruff `check` and `format --check` passed on both touched tests. The historical [PR CI run](https://github.com/vllm-project/vllm-ascend/actions/runs/34845930945) passed pre-commit and the full CPU-UT job; the six previous test-double failures are gone. Device-test matrix jobs were still running when that evidence was recorded; this is not a current CI status report. - Pre-generalization CP-helper checks: six dependency-free common-CP tests and **1,500 exact before/after DSA-CP comparisons passed**. They cover all ranks at TP1/2/4/8, main/MTP steps, padding, stable buffer reuse, independent buffers, and SFA/indexer calculations. The PCP orchestration regression passed against both revisions. These checks execute repository method bodies with distributed/NPU imports excluded. - Ruff 0.14.0 `check`, `format --check`, and AST parsing passed on the eight Python files touched by the final helper cleanup. The full `format.sh ci` command was unavailable locally because `pre-commit` was not installed. - An earlier 11-file focused pytest run at `c9916556c`, using vLLM `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`, reported **204 passed, 13 skipped, 0 failed**. It covered attention/indexer, proposer, worker metadata plumbing and GLM/DeepSeek indexers. This earlier run is not a full pytest result for the final tree. #### A5 long-input precision checks (2026-09-14) Setup: A5 node 112, Ascend 950DT, GLM-5.2 W4A4C8, vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`, EP enabled, `max_model_len=40960`, `max_num_batched_tokens=4096`, `max_num_seqs=8`, and breakable graph disabled. The following settings apply to **every row in this NPU matrix**, both with and without the #16495 overlay, and to the four additional 12K-39K retrieval checks below: - **MTP: disabled.** No speculative-decoding configuration was supplied. These are non-MTP precision results; they do not validate MTP-enabled execution or the MTP + C8 combination. - **C8: enabled for both SFA and the indexer.** The model is W4A4C8, with `additional_config.enable_sparse_sfa_c8=True` and `additional_config.enable_sparse_li_c8=True`. - **The tested Ascend/vLLM source commits are paired correctly.** Both Ascend test commits (`c30f4f421e0eea29c043f3ee8131c0333eb5100e` and overlay `3bc93d4dae7a8fe397d3262880ac9da78ca83375`) pin vLLM `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` in `.github/vllm-main-verified.commit`. The NPU checkout has that exact vLLM HEAD, and the service logs confirm imports from `/mnt/share/g00672350/refactor/vllm-a97dacb/vllm/`. The bot-managed vLLM footer reflects the base branch and is not the version used for these hardware results. The reference script was used unchanged: eight serial requests, four concurrent requests, then eight concurrent requests. Prompt lengths were `17/14/32/752/1832/2192/2677/3700` tokens. The nine requests at 2,192 / 2,677 / 3,700 tokens were also checked for exact expected-code output. | Code | Topology | DSA-CP requested / effective | Execution | Script passes | >=2K exact | | --- | --- | --- | --- | ---: | ---: | | This PR's tree | TP8/DP1 | off / off | eager | 20/20 | 9/9 | | This PR's tree | TP8/DP1 | off / off | `FULL_DECODE_ONLY` | **2/20** | **0/9** | | This PR's tree | TP8/DP1 | on / **off** | eager | 20/20 | 9/9 | | This PR's tree | TP8/DP1 | on / **off** | `FULL_DECODE_ONLY` | **3/20** | **0/9** | | This PR's tree | TP4/DP2 | on / on | eager | 20/20 | 9/9 | | This PR's tree | TP4/DP2 | on / on | `FULL_DECODE_ONLY` | 20/20 | 9/9 | | This PR + #16495 | TP8/DP1 | off / off | `FULL_DECODE_ONLY` | **20/20** | **9/9** | | This PR + #16495 | TP4/DP2 | on / on | `FULL_DECODE_ONLY` | 20/20 | 9/9 | All requests returned HTTP 200. The two failing graph rows produced corrupted text; the script's substring check can count a corrupted short answer as a pass, so the exact long-input result is the stronger signal. With this vLLM configuration, TP8/DP1 makes `use_sequence_parallel_moe` false and Ascend explicitly disables requested DSA-CP. TP4/DP2 enables it. The table therefore records different topologies rather than implying a controlled DSA-CP toggle at one TP/DP shape. The overlay was #16495 head `d9bae5c4614df6959caf93d9b6f6ad05778fd52f` cherry-picked onto the tested pre-squash head, producing test commit `3bc93d4dae7a8fe397d3262880ac9da78ca83375`. In its TP8/DP1 graph mode, four additional original-style sequential retrieval prompts at **12,761 / 25,801 / 35,271 / 39,441 tokens** all returned the exact expected code with `finish_reason=stop`. Evidence on the shared A5 mount: `/mnt/share/g00672350/refactor/pr16325-longinput-testkit-112/results_*/summary.json`. Overlay service log: `/mnt/share/g00672350/refactor/longinput_112_tp8dp1_graph_dsa_off_plus16495_3bc93d.log`. #### A5 MTP1/MTP5 long-input precision checks (2026-09-14) A5 node 111, Ascend 950DT, GLM-5.2 W4A4C8, TP8/DP1, EP enabled, DSA-CP off, `FULL_DECODE_ONLY` graph mode, breakable graph disabled. Both `enable_sparse_sfa_c8` and `enable_sparse_li_c8` were `True`. The server used Ascend test overlay `3bc93d4dae7a8fe397d3262880ac9da78ca83375` (this PR + #16495) with its matching vLLM `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`; `max_model_len=40960`, `max_num_batched_tokens=4096`, and `max_num_seqs=8`. The engine logs confirmed effective MTP speculative settings of `num_spec_tokens=1` and `num_spec_tokens=5`, respectively, and successful decode-graph capture. | MTP draft tokens | Original serial/concurrent cases | Original >=2K exact answers | Additional exact long inputs | | --- | ---: | ---: | ---: | | 1 | 20/20 | 9/9 | 4/4 | | 5 | 20/20 | 9/9 | 4/4 | The additional prompt lengths were **12,761 / 25,801 / 38,051 / 39,441 tokens**. Each returned HTTP 200, `finish_reason=stop`, and content exactly equal to the expected retrieval code; both full 24-case suites passed. The original 20 cases used the same eight serial, four concurrent, and eight concurrent requests as the non-MTP matrix. MTP-enabled hardware evidence covers this TP8/DP1, DSA-CP-off configuration with #16495 applied; it does not establish MTP precision for the PR alone or for DSA-CP-on configurations. #### MTP draft acceptance observed during the long-input runs The server's `SpecDecoding metrics` log lines show that speculative decoding was active. During the sustained request windows, MTP1 had roughly **90–95% average draft acceptance** and **1.9–2.0 mean acceptance length**. MTP5 had roughly **57–66% average draft acceptance** and **3.9–4.3 mean acceptance length**; its per-position acceptance declined from about **90–94% at draft position 1** to about **31–44% at position 5**. These are interval metrics from the service logs, not an aggregate over the 24 requests. Alongside the 24/24 exact-output checks for each setting, they show no obvious acceptance collapse. No matched pre-refactor acceptance-rate baseline was collected, so these numbers do **not** establish acceptance-rate equivalence with the old implementation. Evidence on the shared A5 mount: `/mnt/share/g00672350/refactor/mtp1_graph_111_results_v2/summary.json` and `/mnt/share/g00672350/refactor/mtp5_graph_111_results/summary.json`. Service logs: `/mnt/share/g00672350/refactor/mtp1_graph_111.log` and `/mnt/share/g00672350/refactor/mtp5_graph_111.log`. #### A5 Model Runner V2 long-input precision checks (2026-09-14) A5 node 111, the existing `xhg_refactor_sfa` container, GLM-5.2 W4A4C8, both sparse SFA-C8 and LI-C8 enabled, `VLLM_USE_V2_MODEL_RUNNER=1`, `FULL_DECODE_ONLY` graph mode, breakable graph disabled, EP enabled, and the same Ascend #16325 + #16495 test overlay `3bc93d4dae7a8fe397d3262880ac9da78ca83375` paired with vLLM `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`. Worker logs explicitly confirmed the V2 model-runner path and successful decode-graph capture. The overlay contains the same PR production code as restored revision `199a459c`; #16495 is not part of this PR. | Topology | DSA-CP | MTP draft tokens | Long-input exact-output cases | | --- | --- | ---: | ---: | | TP8/DP1 | off | 0 | **24/24** | | TP8/DP1 | off | 1 | **24/24** | | TP8/DP1 | off | 5 | **24/24** | | TP4/DP2 | on | 0 | **24/24** | Each suite comprised the prior 20 serial/concurrent requests plus four exact-code retrieval prompts of **12,761 / 25,801 / 38,051 / 39,441 tokens**. All reported cases returned HTTP 200, `finish_reason=stop`, and the expected output. These results validate the tested Model Runner V2 graph configurations, not eager mode or this PR without #16495. The TP4/DP2 + DSA-CP-on + MTP1/MTP5 combinations were **not run** because the A5 nodes became occupied by other services; no result is claimed for them. Evidence on the shared A5 mount: `/mnt/share/g00672350/refactor/mrv2_mtp0_graph_111_results/summary.json`, `mrv2_mtp1_graph_111_results/summary.json`, `mrv2_mtp5_graph_111_results/summary.json`, and `mrv2_tp4dp2_dsa_mtp0_graph_111_results/summary.json`. Corresponding service logs use the same names without `_results/summary.json` and with `.log` suffix. Validation limits: the A5 profile rejects replicated SFA DCP before startup, so replicated-DCP address transforms are covered by unit tests rather than this hardware matrix. MTP was disabled in the preceding non-MTP matrix; the separate MTP1/MTP5 checks are reported above. GSM8K and full repository pytest have not been rerun on the final tree; the long-input checks are not a throughput benchmark. <details> <summary>Earlier combined-build GSM8K evidence</summary> These results belong to combined validation commit `5f0326bb3368df11a3067b9fe8c8bc88f9e5463a`, base `fb00c6be102f31e356bea4ab0f5a64e5091608c6`, and vLLM `b2f685834a6456197e7033966fdef52a23f1abcd`, before the indexer and MoE fixes were separated. They are not final-tree GSM8K results. Single-node W4A4C8, TP4/DP2, sparse SFA/LI-C8 enabled, MTP and MLAPO disabled, `max_num_batched_tokens=4096`, `max_num_seqs=8`, breakable graph disabled: | DSA-CP | Execution | Long-input script | GSM8K | | --- | --- | ---: | ---: | | off | eager | 20/20 | 178/200 (89.0%) | | off | `FULL_DECODE_ONLY` | 20/20 | 178/200 (89.0%) | | on | eager | 20/20 | 183/200 (91.5%) | | on | `FULL_DECODE_ONLY` | 20/20 | 180/200 (90.0%) | </details> - vLLM main: vllm-project/vllm@84030bb Signed-off-by: Biuapha <1731372716@qq.com>
What this PR does / why we need it?
Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA:
AscendSFAIndexerBackend(vllm_ascend/attention/indexer.py): owns the whole indexer pipeline in one unifiedforward- k-path (forward_k) -> parallel-layout gather (_gather_cache_inputs) ->write_cache-> top-k selection.compute_topk=Falsekeeps the skip_topk layers' cache up to date while reusing shared top-k indices.AscendSFAIndexerMetadataBuilder(slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA'sattn_metadata, so a variant indexer with a compressed cache layout gets its own slot mapping.IndexerWrapper(vllm_ascend/ops/mla.py) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (...indexer.<name>) that weight loading and quant name mapping key off, and delegates the unifiedforwardto the backend.ops/indexer/base.py/ops/indexer/lightning.py.Does this PR introduce any user-facing change?
No. Pure internal refactor; no API, config, or behavior change.
How was this patch tested?
Unit tests updated/added:
tests/ut/attention/test_indexer.py(metadata builder, unified forward, PCP/DSA-CP gather),tests/ut/attention/test_sfa_v1.py,tests/ut/attention/test_sfa_cp.py,tests/ut/ops/test_mla.py.Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts).
Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with
enable_sparse_sfa_c8,enable_sparse_li_c8,enable_dsa_cp, flashcomm1, multistream-overlap; smoke prompts correct in both eager andFULL_DECODE_ONLYACLGraph modes (graph capture passes with DSA-CP enabled).GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (
deepseek_mtp, num_speculative_tokens=5) andenable_sparse_sfa_c8/enable_sparse_li_c8/enable_dsa_cpon:Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5.
vLLM main: vllm-project/vllm@b2f6858