Skip to content

[Refactor][Model] Decouple indexer from SFA with unified forward pipeline - #15669

Merged
zzzzwwjj merged 4 commits into
vllm-project:mainfrom
Biuapha:main
Sep 8, 2026
Merged

zzzzwwjj merged 4 commits into
vllm-project:mainfrom
Biuapha:main

Conversation

@Biuapha

@Biuapha Biuapha commented Sep 3, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Decouples the sparse-attention indexer from the SFA attention impl so indexer variants (different cache layout / compression / top-k kernel) can be plugged in without touching SFA:

  • Introduces AscendSFAIndexerBackend (vllm_ascend/attention/indexer.py): owns the whole indexer pipeline in one unified forward - k-path (forward_k) -> parallel-layout gather (_gather_cache_inputs) -> write_cache -> top-k selection. compute_topk=False keeps the skip_topk layers' cache up to date while reusing shared top-k indices.
  • The indexer builds its own attention metadata via AscendSFAIndexerMetadataBuilder (slot_mapping / block_table / seq lens / rope tables), instead of borrowing SFA's attn_metadata, so a variant indexer with a compressed cache layout gets its own slot mapping.
  • Parallel-layout transforms are internal to the backend: PCP all-gathers the prefill region (reordering slot mapping to the gathered layout), DSA-CP all-gathers the indexer k across the TP group. SFA no longer carries indexer-specific gather/write hooks.
  • IndexerWrapper (vllm_ascend/ops/mla.py) instantiates the backend directly, registers the indexer weights on the wrapper so module-tree paths keep the pre-refactor layout (...indexer.<name>) that weight loading and quant name mapping key off, and delegates the unified forward to the backend.
  • Removes the now-dead ops/indexer/base.py / ops/indexer/lightning.py.

Does this PR introduce any user-facing change?

No. Pure internal refactor; no API, config, or behavior change.

How was this patch tested?

  • Unit tests updated/added: tests/ut/attention/test_indexer.py (metadata builder, unified forward, PCP/DSA-CP gather), tests/ut/attention/test_sfa_v1.py, tests/ut/attention/test_sfa_cp.py, tests/ut/ops/test_mla.py.

  • Equivalence review against pre-refactor flow (k-path ordering, cache write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather layouts).

  • Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8): service up with enable_sparse_sfa_c8, enable_sparse_li_c8, enable_dsa_cp, flashcomm1, multistream-overlap; smoke prompts correct in both eager and FULL_DECODE_ONLY ACLGraph modes (graph capture passes with DSA-CP enabled).

  • GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240) on the dual-node PP2TP8EP8 deployment with MTP5 (deepseek_mtp, num_speculative_tokens=5) and enable_sparse_sfa_c8/enable_sparse_li_c8/enable_dsa_cp on:

    • pre-refactor baseline: 1283/1319 = 97.27%
    • with this PR: 1286/1319 = 97.50%
      Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP mean acceptance length 4.00/5.
  • vLLM main: vllm-project/vllm@b2f6858

@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request refactors the indexer implementation to improve modularity by decoupling it from the AscendSFAImpl class. By introducing a standalone AscendSFAIndexerBackend that handles its own compute and metadata, the architecture becomes more maintainable and better suited for NPU-specific optimizations. The changes also streamline the interaction between the indexer and the SFA layer, replacing hardcoded logic with a cleaner delegation pattern.

Highlights

  • Architecture Refactoring: Decoupled the indexer logic from the main SFA implementation by introducing a standalone AscendSFAIndexerBackend class.
  • Indexer Backend Implementation: Implemented AscendSFAIndexerBackend as an nn.Module to manage its own compute path, cache persistence, and metadata, replacing previous hardcoded logic.
  • NPU Kernel Integration: Replaced hardcoded CUDA-specific paths with NPU-specific kernels for the indexer forward path to improve performance and compatibility.
  • Wrapper Update: Updated IndexerWrapper to act as a clean interface that delegates indexer-related operations to the new backend instance.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the SFA indexer cache implementation by introducing AscendSFAIndexerBackend and AscendSFAIndexerMetadata to handle compute and cache persistence, moving logic out of the IndexerWrapper and AscendSFAImpl. It adds support for NPU-based kernels to replace hardcoded CUDA fp8 paths, introduces metadata builders for efficient cache management, and updates tests to reflect these architectural changes. A critical issue was identified regarding a potential AttributeError when accessing k_cache.prefix if k_cache is None.

Comment thread vllm_ascend/attention/indexer.py
Comment thread vllm_ascend/attention/indexer.py
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@Biuapha Biuapha changed the title [Refactor]Decouple indexer from SFA [Refactor][Model] Decouple indexer from SFA with unified forward pipeline Sep 7, 2026
@Biuapha
Biuapha force-pushed the main branch 3 times, most recently from 9edf15b to 8fd7046 Compare September 7, 2026 11:22
@Biuapha Biuapha added the ready label Sep 7, 2026
@Biuapha
Biuapha force-pushed the main branch 3 times, most recently from a45803e to a5dec30 Compare September 7, 2026 12:21
@Biuapha Biuapha added the ready-all run all e2e test for pr label Sep 7, 2026
@Biuapha
Biuapha force-pushed the main branch 2 times, most recently from 137267a to ec00247 Compare September 7, 2026 15:31
@Biuapha Biuapha added glm-5 and removed glm-5 labels Sep 7, 2026
@Biuapha
Biuapha force-pushed the main branch 2 times, most recently from a99c3ac to 989cfd1 Compare September 8, 2026 01:53
@Biuapha Biuapha removed the ready label Sep 8, 2026
@Biuapha Biuapha closed this Sep 8, 2026
@Biuapha Biuapha reopened this Sep 8, 2026
weijinqian0
weijinqian0 previously approved these changes Sep 8, 2026
Comment thread vllm_ascend/attention/sfa_v1.py Outdated
Comment thread vllm_ascend/attention/sfa_v1.py
weijinqian0
weijinqian0 previously approved these changes Sep 8, 2026
- Consolidate AscendLightningIndexer into AscendSFAIndexerBackend and drop the AscendIndexerBase indirection (ops/indexer base.py/lightning.py removed)
- Unify forward_k / write_cache / forward into a single backend forward: k path -> parallel-layout gather (PCP / DSA-CP) -> cache write -> top-k selection
- Indexer builds its own attn metadata via AscendSFAIndexerMetadataBuilder (slot mapping, block table, seq lens, group_*, cos/sin); SFA fetches it by k_cache prefix
- IndexerWrapper in ops/mla.py instantiates the backend directly and keeps weight module paths for load/quant compatibility
- skip_topk layers still run the k path and cache write (compute_topk=False) so the shared-index cache stays up to date
- Adapt to upstream c8_reshape_optim_enabled config rename (vllm-project#15752)
- Update UTs: indexer gather/metadata build, wrapper delegation, SFA metadata fetch

Signed-off-by: Biuapha <1731372716@qq.com>
The indexer k path now lives inside the indexer backend, so k_li/k_li_scale are always None at the _prepare_kv_for_parallel / _store_parallel_kv call sites. Remove the dead parameters, the unreachable k_li branches in the PCP overrides, and fold the identical two-way splits; shrink the return tuples accordingly. Addresses review comments from weijinqian0.

Signed-off-by: Biuapha <1731372716@qq.com>
During MTP draft propose the draft forward runs under the main model runner's forward context, whose metadata dict carries only main-layer keys - neither the draft's own indexer cache prefix nor its attn layer name is present. Restore the kv_sharing_target_layer_name resolution so a KV-sharing layer looks up the sharing target's indexer cache prefix first, keeping the layer-name fallback for the proposer-built context. Fixes a RuntimeError crash with MTP + FULL_DECODE_ONLY graph replay.

Signed-off-by: Biuapha <1731372716@qq.com>
AscendSFAImpl instances created via __new__ in unit tests (and any legacy construction path that skips __init__) have no kv_sharing_target_layer_name attribute; use getattr with a None default, and add a UT covering the kv-sharing target prefix resolution.

Signed-off-by: Biuapha <1731372716@qq.com>
@zzzzwwjj
zzzzwwjj merged commit 44e3160 into vllm-project:main Sep 8, 2026
38 of 52 checks passed
jiangli221 pushed a commit to jiangli221/vllm-ascend that referenced this pull request Sep 9, 2026
…line (vllm-project#15669)

### What this PR does / why we need it?

Decouples the sparse-attention indexer from the SFA attention impl so
indexer variants (different cache layout / compression / top-k kernel)
can be plugged in without touching SFA:

- Introduces `AscendSFAIndexerBackend`
(`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in
one unified `forward` - k-path (`forward_k`) -> parallel-layout gather
(`_gather_cache_inputs`) -> `write_cache` -> top-k selection.
`compute_topk=False` keeps the skip_topk layers' cache up to date while
reusing shared top-k indices.
- The indexer builds its own attention metadata via
`AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens
/ rope tables), instead of borrowing SFA's `attn_metadata`, so a variant
indexer with a compressed cache layout gets its own slot mapping.
- Parallel-layout transforms are internal to the backend: PCP
all-gathers the prefill region (reordering slot mapping to the gathered
layout), DSA-CP all-gathers the indexer k across the TP group. SFA no
longer carries indexer-specific gather/write hooks.
- `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend
directly, registers the indexer weights on the wrapper so module-tree
paths keep the pre-refactor layout (`...indexer.<name>`) that weight
loading and quant name mapping key off, and delegates the unified
`forward` to the backend.
- Removes the now-dead `ops/indexer/base.py` /
`ops/indexer/lightning.py`.

### Does this PR introduce _any_ user-facing change?

No. Pure internal refactor; no API, config, or behavior change.

### How was this patch tested?

- Unit tests updated/added: `tests/ut/attention/test_indexer.py`
(metadata builder, unified forward, PCP/DSA-CP gather),
`tests/ut/attention/test_sfa_v1.py`,
`tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`.
- Equivalence review against pre-refactor flow (k-path ordering, cache
write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather
layouts).
- Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8):
service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`,
`enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct
in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture
passes with DSA-CP enabled).
- GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240)
on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`,
num_speculative_tokens=5) and
`enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on:
  - pre-refactor baseline: 1283/1319 = 97.27%
  - with this PR: 1286/1319 = 97.50%
Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP
mean acceptance length 4.00/5.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Biuapha <1731372716@qq.com>
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
…line (vllm-project#15669)

### What this PR does / why we need it?

Decouples the sparse-attention indexer from the SFA attention impl so
indexer variants (different cache layout / compression / top-k kernel)
can be plugged in without touching SFA:

- Introduces `AscendSFAIndexerBackend`
(`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in
one unified `forward` - k-path (`forward_k`) -> parallel-layout gather
(`_gather_cache_inputs`) -> `write_cache` -> top-k selection.
`compute_topk=False` keeps the skip_topk layers' cache up to date while
reusing shared top-k indices.
- The indexer builds its own attention metadata via
`AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens
/ rope tables), instead of borrowing SFA's `attn_metadata`, so a variant
indexer with a compressed cache layout gets its own slot mapping.
- Parallel-layout transforms are internal to the backend: PCP
all-gathers the prefill region (reordering slot mapping to the gathered
layout), DSA-CP all-gathers the indexer k across the TP group. SFA no
longer carries indexer-specific gather/write hooks.
- `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend
directly, registers the indexer weights on the wrapper so module-tree
paths keep the pre-refactor layout (`...indexer.<name>`) that weight
loading and quant name mapping key off, and delegates the unified
`forward` to the backend.
- Removes the now-dead `ops/indexer/base.py` /
`ops/indexer/lightning.py`.

### Does this PR introduce _any_ user-facing change?

No. Pure internal refactor; no API, config, or behavior change.

### How was this patch tested?

- Unit tests updated/added: `tests/ut/attention/test_indexer.py`
(metadata builder, unified forward, PCP/DSA-CP gather),
`tests/ut/attention/test_sfa_v1.py`,
`tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`.
- Equivalence review against pre-refactor flow (k-path ordering, cache
write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather
layouts).
- Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8):
service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`,
`enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct
in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture
passes with DSA-CP enabled).
- GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240)
on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`,
num_speculative_tokens=5) and
`enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on:
  - pre-refactor baseline: 1283/1319 = 97.27%
  - with this PR: 1286/1319 = 97.50%
Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP
mean acceptance length 4.00/5.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Biuapha <1731372716@qq.com>
yiminghub2024 added a commit to yiminghub2024/vllm-ascend that referenced this pull request Sep 10, 2026
While this branch was open, main grew its own implementations of both features
it was proposing:

  - vllm-project#15913 GLM-Next KV cache management, which keeps the incomplete pool in an
    absolute-position FP32 Glm5NextStateCache rather than in this branch's
    paged Glm5NextTailCache, and stores completed pools as unquantized BF16
  - vllm-project#15669 decoupled the indexer from SFA, restructuring the very plumbing this
    branch's kpool backend was wired into
  - AscendDflash2Proposer, selected by is_dflash2_draft() under method dflash

Those designs are mutually exclusive with this branch's, not textually
conflicting with it, so every conflict is resolved in favour of main and the
branch's own kpool backend, tail cache and DFlash2 patch are dropped. The
follow-up commit re-adds only what main still lacks.

Also drops CI-config churn this branch had picked up from a stale tree: it
reverted its own parent vllm-project#14683 by deleting GLM-5.2-W4A8C8-SFA-DCP.yaml and
moving nightly entries into weekly, and added an unreferenced
DeepSeek-V3.2-W8A8-DCP.yaml.

Signed-off-by: yiminghub2024 <482890@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…line (vllm-project#15669)

### What this PR does / why we need it?

Decouples the sparse-attention indexer from the SFA attention impl so
indexer variants (different cache layout / compression / top-k kernel)
can be plugged in without touching SFA:

- Introduces `AscendSFAIndexerBackend`
(`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in
one unified `forward` - k-path (`forward_k`) -> parallel-layout gather
(`_gather_cache_inputs`) -> `write_cache` -> top-k selection.
`compute_topk=False` keeps the skip_topk layers' cache up to date while
reusing shared top-k indices.
- The indexer builds its own attention metadata via
`AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens
/ rope tables), instead of borrowing SFA's `attn_metadata`, so a variant
indexer with a compressed cache layout gets its own slot mapping.
- Parallel-layout transforms are internal to the backend: PCP
all-gathers the prefill region (reordering slot mapping to the gathered
layout), DSA-CP all-gathers the indexer k across the TP group. SFA no
longer carries indexer-specific gather/write hooks.
- `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend
directly, registers the indexer weights on the wrapper so module-tree
paths keep the pre-refactor layout (`...indexer.<name>`) that weight
loading and quant name mapping key off, and delegates the unified
`forward` to the backend.
- Removes the now-dead `ops/indexer/base.py` /
`ops/indexer/lightning.py`.

### Does this PR introduce _any_ user-facing change?

No. Pure internal refactor; no API, config, or behavior change.

### How was this patch tested?

- Unit tests updated/added: `tests/ut/attention/test_indexer.py`
(metadata builder, unified forward, PCP/DSA-CP gather),
`tests/ut/attention/test_sfa_v1.py`,
`tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`.
- Equivalence review against pre-refactor flow (k-path ordering, cache
write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather
layouts).
- Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8):
service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`,
`enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct
in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture
passes with DSA-CP enabled).
- GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240)
on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`,
num_speculative_tokens=5) and
`enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on:
  - pre-refactor baseline: 1283/1319 = 97.27%
  - with this PR: 1286/1319 = 97.50%
Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP
mean acceptance length 4.00/5.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Biuapha <1731372716@qq.com>
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 12, 2026
…ard pipeline (vllm-project#15669)"

This reverts commit 44e3160.

Signed-off-by: chenzeyu <2978509328@qq.com>

# Conflicts:
#	vllm_ascend/attention/indexer.py
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…line (vllm-project#15669)

### What this PR does / why we need it?

Decouples the sparse-attention indexer from the SFA attention impl so
indexer variants (different cache layout / compression / top-k kernel)
can be plugged in without touching SFA:

- Introduces `AscendSFAIndexerBackend`
(`vllm_ascend/attention/indexer.py`): owns the whole indexer pipeline in
one unified `forward` - k-path (`forward_k`) -> parallel-layout gather
(`_gather_cache_inputs`) -> `write_cache` -> top-k selection.
`compute_topk=False` keeps the skip_topk layers' cache up to date while
reusing shared top-k indices.
- The indexer builds its own attention metadata via
`AscendSFAIndexerMetadataBuilder` (slot_mapping / block_table / seq lens
/ rope tables), instead of borrowing SFA's `attn_metadata`, so a variant
indexer with a compressed cache layout gets its own slot mapping.
- Parallel-layout transforms are internal to the backend: PCP
all-gathers the prefill region (reordering slot mapping to the gathered
layout), DSA-CP all-gathers the indexer k across the TP group. SFA no
longer carries indexer-specific gather/write hooks.
- `IndexerWrapper` (`vllm_ascend/ops/mla.py`) instantiates the backend
directly, registers the indexer weights on the wrapper so module-tree
paths keep the pre-refactor layout (`...indexer.<name>`) that weight
loading and quant name mapping key off, and delegates the unified
`forward` to the backend.
- Removes the now-dead `ops/indexer/base.py` /
`ops/indexer/lightning.py`.

### Does this PR introduce _any_ user-facing change?

No. Pure internal refactor; no API, config, or behavior change.

### How was this patch tested?

- Unit tests updated/added: `tests/ut/attention/test_indexer.py`
(metadata builder, unified forward, PCP/DSA-CP gather),
`tests/ut/attention/test_sfa_v1.py`,
`tests/ut/attention/test_sfa_cp.py`, `tests/ut/ops/test_mla.py`.
- Equivalence review against pre-refactor flow (k-path ordering, cache
write timing, skip_topk/index-cache reuse, CP/PCP/DSA-CP gather
layouts).
- Two-node deployment on Ascend 950 (PP2TP8EP8, GLM-5.3 W8A8C8 MXFP8):
service up with `enable_sparse_sfa_c8`, `enable_sparse_li_c8`,
`enable_dsa_cp`, flashcomm1, multistream-overlap; smoke prompts correct
in both eager and `FULL_DECODE_ONLY` ACLGraph modes (graph capture
passes with DSA-CP enabled).
- GSM8K zero-shot CoT full-set eval (1319 questions, max_tokens=10240)
on the dual-node PP2TP8EP8 deployment with MTP5 (`deepseek_mtp`,
num_speculative_tokens=5) and
`enable_sparse_sfa_c8`/`enable_sparse_li_c8`/`enable_dsa_cp` on:
  - pre-refactor baseline: 1283/1319 = 97.27%
  - with this PR: 1286/1319 = 97.50%
Accuracy is consistent (delta +0.23%, within run-to-run variance); MTP
mean acceptance length 4.00/5.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Biuapha <1731372716@qq.com>
Signed-off-by: like-0517 <ithwlike@126.com>
weijinqian0 pushed a commit that referenced this pull request Sep 17, 2026
)

### What this PR does / why we need it?

After #15669 separated the SFA indexer cache and metadata builder,
replicated DCP can expose an address mismatch: SFA uses a rank-local
KV-cache layout, while the indexer cache is replicated across DCP ranks.
Building indexer metadata from the SFA-local block table or slot mapping
can therefore read or write the wrong indexer cache locations.

This PR makes the indexer builder derive its complete cache view from
common attention metadata and the explicitly supplied CP context:

`CommonAttentionMetadata + CP context -> indexer builder ->
AscendSFAIndexerMetadata`

The builder owns the final block table and write-ready slots, RoPE
tables, query/key sequence lengths, PCP decode boundary, and LI-C8
cache-write grouping metadata. Mutable derived buffers are keyed by the
persistent common slot-mapping buffer so graph capture, replay, and MTP
draft steps use the appropriate storage.

Existing parallel calculations are shared through stateless helpers:
`get_cp_local_query_key_lens` and `build_pcp_ordered_slot_mapping` in
`context_parallel/common_cp.py`, and replicated-DCP address / PCP
global-view helpers in `context_parallel/sfa_dcp_utils.py`. SFA and
indexer builders call those helpers and retain their own metadata
buffers.

#### Why `sfa_v1.py` changes

| Change | Reason |
| --- | --- |
| Move LI-C8 group metadata construction to the indexer builder; remove
the old SFA-side construction and field assignments. | The grouping must
describe the indexer's final slots and block size. Removing the unused
SFA-side call also avoids duplicate auxiliary-metadata work; SFA's own
C8 KV-write path remains intact. |
| Look up the sharing target's indexer metadata, then the draft
indexer's own metadata; remove fallback to the SFA layer's metadata. |
KV-sharing draft layers can have their own metadata even when physical
cache storage is shared. SFA metadata is not a valid substitute for the
indexer's replicated cache view. |
| Remove forward-time assignments of `actual_seq_lengths_query`,
`actual_seq_lengths_key`, and `num_decode_tokens`. | These fields are
already constructed by the indexer builder. Copying them from SFA would
preserve a dependency on SFA's construction and buffer ownership. |
| Read `cos` / `sin` inside indexer forward from indexer metadata
instead of passing SFA's tensors. | This completes the indexer's
metadata ownership, including CP slicing and graph/MTP storage. It
reuses the existing RoPE calculation. |

The direct DCP bug fix is the replicated cache-address construction. The
RoPE interface change and removal of unused SFA auxiliary fields
complete the independent-metadata design; they do not introduce a
different SFA attention algorithm.

The proposer retains metadata builders for cache-only layers. It
identifies these groups through the existing backend contract
(`get_impl_cls() is None`), not a concrete KV-cache-spec type, so an
independent tail-cache spec is not mistaken for the primary attention
group. The backend query runs in the proposer's vLLM configuration
context. During MTP, common sequence/position/slot state advances once
through the primary attention group, then each cache-only builder
constructs its own metadata from that updated common view. Cache-only
groups are excluded from primary graph/backend selection, not from
metadata construction. This remains the existing
one-main-attention-group path; it does not add generic multi-KV-cache
scheduling or heterogeneous full-graph support. The Step3.5 and Gemma4
proposers are unchanged.

### Does this PR introduce _any_ user-facing change?

It fixes indexer cache addressing and removes reliance on SFA metadata
being used as an indexer fallback or patched into indexer metadata
during forward. There is no public API or configuration-format change.

The separate compiled MoE reduction issue is addressed by #16495. The
hardware comparison below includes that patch only in an isolated test
overlay; it is not included in this PR.

### How was this patch tested?

#### Revision and source checks
- **Latest-main conflict resolution (`ec6e301d`, 2026-09-15):** rebased
the single PR commit onto main `0e0910b0`. The only textual conflict was
the existing `_make_runner()` helper in
`tests/ut/worker/test_model_runner_v2.py`: main #16574 already supplies
`adaptive_verification` and `use_fia`, so this PR now retains only the
still-missing `attn_groups = []` initialization. `git range-diff` shows
no production-logic change from the rebase. Ruff check passed for all 19
changed files, the conflict file passes Ruff format check, `git diff
--check` passes, and the paired vLLM pin remains
`84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- **DSpark adaptive graph CI fix (`ec6e301d`, 2026-09-15):** the failed
A3 eight-card case was `test_glm5_2_dspark_eager[adaptive]`; MTP and
fixed DSpark passed in the same job. Adaptive FULL graph records its
graph-shaped token width in `positions.shape[0]`, while
`common_attn_metadata.num_input_tokens` can retain the compact count.
Before the split, indexer RoPE came from SFA, whose builder already
handles this contract. The independent indexer builder now applies the
same width when constructing its own positions/RoPE/slot view. One
focused unit regression was added; Ruff check, targeted format check,
and `git diff --check` pass. Local full pytest was not claimed because
the available Windows helper environment lacks the paired vLLM
dependency set; CI is the authoritative rerun. No DSpark/model-runner/CI
test behavior or assertion was changed. The first CI revision exposed
four existing test doubles that omit the optional `method` field; the
final predicate uses `getattr(..., None)` for that field as well,
without changing those fixtures. That CPU-UT run reported 4624 passed /
4 failed before this one-line compatibility correction.

- **Current CI follow-up (`f0945801`, 2026-09-15):** compared with
`242e9751`, this adds only three initialization lines to the existing
`_make_runner()` test helper in
`tests/ut/worker/test_model_runner_v2.py`: `attn_groups = []`,
`adaptive_verification = None`, and `use_fia = False`. No production
code or PR base changed; the PR remains one signed-off commit. The
[previous CPU-UT
job](https://github.com/vllm-project/vllm-ascend/actions/runs/34923604678/job/104237578682)
reported 4611 passed / 3 failed after CI rebased onto main `a643db116`.
Both the tested runner code and newly added tests came from main, not
this PR: their lightweight `__new__` fixture omitted fields normally
initialized by the real runner. The fix preserves all assertions and
changes neither production defaults nor execution paths. Ruff check,
format check and `git diff --check` pass; the new [pre-commit/mypy
job](https://github.com/vllm-project/vllm-ascend/actions/runs/34925651210/job/104243025168)
and [CPU-UT
job](https://github.com/vllm-project/vllm-ascend/actions/runs/34925651210/job/104243993813)
both completed successfully. This does not claim that every NPU matrix
job has finished. An independent CPU reproduction used main `a643db116`
and its paired vLLM `84030bbe3d74d99bad477a3d2e37a973ccd8865c`: the
complete `test_model_runner_v2.py` reproduced **33 passed / 3 failed**
before the fix, then **36 passed / 0 failed** after the same three-line
patch. The complete `test_attn_utils_v2.py` also passed **23/23**. No
assertions were removed, dependencies installed, or NPU probes
performed. Logs:
`/mnt/share/g00672350/refactor/pr16325-ci-242e9751-20260915/112_cpu_model_runner_v2_unfixed.log`,
`112_cpu_model_runner_v2_fixed.log`, and
`112_cpu_attn_utils_v2_fixed.log`. The archive identities were verified:
before `cd13ae104`, after `cbcd83bc3`; these reproduce the CI main + PR
combination without changing this PR's base. Targeted Ruff checks pass.
The repository-wide auto-formatting command was not run because it could
modify unrelated files; full pre-commit/mypy remains covered by CI.
- **Earlier mypy fix (`242e9751`):** declared `consumes_pcp_context =
False` on the existing `_PrefillStateBuilder` test double; no production
changes. The [pre-commit job including
mypy](https://github.com/vllm-project/vllm-ascend/actions/runs/34923604678/job/104236813190)
passed. This is previous-run evidence, not a claim that all checks for
the new head have passed.
- **Current CPU regression (112, 2026-09-15):** the complete 165-item
focused Eagle/SFA/indexer suite passed under the repository's standard
CPU conftest profile: **155 passed, 10 skipped, 0 failed**. It ran on
exact source `242e9751` with paired vLLM `a97dacb`; current production
code is identical. Log:
`/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/112_cpu_eagle_sfa_indexer_standard_profile.log`.
An earlier A5-profile CPU attempt had two custom-op-dispatch expectation
failures; the authoritative result is the independent full rerun under
the standard CPU profile, not combined partial results. The priority
cache-only group/context set also passed 10/10.
- **Current NPU precision/acceptance (111, 2026-09-15):** MTP1, TP8/DP1,
DSA-CP off, target `FULL_DECODE_ONLY` with compilation enabled and
breakable disabled: **24/24 passed**, including 9/9 original >=2K exact
retrieval answers and 4/4 additional exact answers at **12,761 / 25,801
/ 38,051 / 39,441 input tokens**. The full suite also contains short
sanity cases. Raw Prometheus counter deltas were 64 drafts, 64 drafted
tokens, 58 accepted tokens: **90.625% draft-token acceptance**, mean
acceptance length **1.90625** (including the bonus token). This is a
small long-input/short-output regression workload, not a matched
pre-refactor acceptance-rate baseline or a throughput benchmark. Both
SFA-C8 and LI-C8 were enabled; the speculative configuration retained
`enforce_eager=true` for the drafter. No GSM8K.
- **Current MTP5/DSA-CP-off result (111, 2026-09-15):** the same
TP8/DP1, W4A4C8, `FULL_DECODE_ONLY` + #16495 configuration passed
**24/24**, including all 9 original >=2K exact retrieval cases and all 4
extra 12,761 / 25,801 / 38,051 / 39,441-token cases. Counter deltas:
**41 drafts, 205 drafted tokens, 103 accepted tokens**; aggregate draft
acceptance **50.2439%**, mean acceptance length **3.5122**. Acceptance
by draft position 1-5 was **92.6829% / 58.5366% / 51.2195% / 39.0244% /
9.7561%** (denominator 41 drafts for each position). These small, mostly
short-answer requests do not establish equality to the historical
interval metrics or to a matched pre-refactor/eager baseline. Evidence
directory:
`/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/111_mtp5_graph_dsa_off_retry1`.
The initial off attempt failed before readiness because memory from the
preceding stopped service had not yet been released; retry began only
after all eight cards were healthy and free. No memory threshold was
lowered, no device/container was reset, and no unrelated process was
stopped.

- **NPU source identity and remaining cases:** the tested archive is
exact PR revision `242e9751` plus authorized #16495 patch
`d9bae5c4614df6959caf93d9b6f6ad05778fd52f` (isolated overlay
`da1e17c53d7362766162ab0583ea267ba0880f02`), paired with vLLM
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`. Current PR `ec6e301d`
additionally contains the focused DSpark-adaptive indexer graph-width
fix above; the prior NPU results do not validate that new line-level
change. Evidence:
`/mnt/share/g00672350/refactor/pr16325-validation-242e9751-20260915/111_mtp1_graph_dsa_off_retry2/{status.json,results/summary.json,metrics_before.txt,metrics_after.txt,service.log}`.
MTP5 + DSA-CP-on at TP4/DP2 was rejected before readiness by the paired
vLLM's spec-decode/SP graph-shape validation (query length 6 versus
TP4); **no precision pass or failure is claimed for that
configuration**. Its failed service was stopped. The MTP5/DSA-CP-off
case subsequently passed as reported above. The author approved TP2/DP4
for the MTP5 + DSA-CP-on graph case. Its validation supervisor was
externally terminated with exit 137 before readiness; no requests or
precision/acceptance result were obtained. Container/host checks did not
show OOM. The author confirmed external process termination and
requested that validation wait, so further NPU launches/retries are now
paused. Only task-owned test processes are being cleaned up; no
device/container reset or unrelated process cleanup is performed. The
TP2/DP4 result remains unvalidated, and the topology difference must be
retained in any later result. 112 cards 1/2 report UB/RAS `81AF8000`
alarms, so 112 is used only for CPU tests. NPU validation is now paused
again at the author's request; the old recurring task remains paused and
will not automatically restart tests.

- **Cache-only group follow-up (`061cf2ab4`, 2026-09-15):** relative to
`249c63d3`, only `llm_base_proposer.py` (+20/-14) and its existing Eagle
proposer test file (+38/-5) changed. Removed the concrete
SFA-indexer-spec check, used the existing backend capability for primary
selection and the MTP cache-only group list, and retained the existing
single-main-group restriction. No broad proposer/CP/Gemma4 refactor was
restored.
- Extended the existing focused test to cover all six permutations of
main attention + SFA indexer + an independent `SlidingWindowSpec`-shaped
cache-only tail group, configuration scope, per-layer metadata
ownership, graph backend selection and the no-executable-group error.
The test's six cases passed locally against the actual production method
ASTs with runtime imports/spec classes/config context isolated. The old
HEAD was also checked and does select an independent tail spec as
primary when it comes first. **This is isolated method-level testing,
not a full runtime pytest result.**
- Ruff 0.14.0 `check`, `format --check` and `git diff --check` passed on
both files. `format.sh ci` could not proceed because local `pre-commit`
is not installed. NPU/model validation and the recurring task remain
paused; no new precision or acceptance-rate result is claimed.
- The author's 111 `/mnt/share/g00672350/refactor/prci` registers
`KpoolTailSpec` from `models/glm5next/kv_cache.py`, where it derives
from `SlidingWindowSpec`; its layer refers to
`vllm.v1.attention.backends.mla.indexer.KpoolTailBackend`. That backend
class was not present in the vLLM source mapped by the specified
`xhg_refactor_sfa` container at inspection time. The new selection path
supports such a tail backend when it follows the cache-only
`get_impl_cls() is None` contract; complete KPool
runtime/cache-addressing compatibility has not been validated.

- **Previous test-only cleanup (`249c63d3`, 2026-09-15):** production
code is unchanged from restored revision `199a459c`. Test changes
relative to the PR base were reduced from 11 files (+793/-47) to 9 files
(+502/-49).
- Removed the extra standalone common-CP helper tests, KPool wrapper
argument-binding test and large proposer dummy-capture scaffold. Merged
duplicate PCP decode-boundary, DSA padding and capture-buffer cases
through parameterization; reused the existing Model Runner V2
context-propagation test. Direct indexer cache-address/C8/buffer and
split-MTP metadata regressions, plus required existing-test
signature/mock adaptations, remain.
- Ruff 0.14.0 `check`, `format --check` and `git diff --check` passed
for that test-only cleanup. Runtime pytest and NPU validation were **not
rerun**; validation and the recurring task remain paused at the author's
request. CI/UT and hardware results below are historical evidence, not
new results for this test-only revision.

- **2026-09-15 restoration, before the test-only cleanup:** restored the
exact pre-generalization commit
`199a459c7e3b2ac749e05e6163b5dc7bcdc401bc`. The broader
proposer/Gemma4/MLA-DCP generalization is no longer included. Validation
and the recurring task are paused at the author's request; the results
below are previously recorded evidence, not newly rerun tests.

- Single commit: `ec6e301d183c5da1023652744b4e9d6395172b86`, based on
current main `0e0910b060f58a8400397331462c306040dc34e4`.
- At `199a459c`, the CPU-UT follow-up changed only two test files
relative to `e889fc608bf1b60854816d7d9f6bbdd09429a39a`; its production
files were identical to the hardware-tested pre-squash head
`c30f4f421e0eea29c043f3ee8131c0333eb5100e`. `git diff --check` passed.
- The previous CPU-UT run had six failures caused by stale test doubles:
four SFA forward cases used the old indexer call signature, and two
Eagle cases omitted the indexer cache spec's `block_size`. The focused
recheck passed: SFA forward 4/4; Eagle proposer 8 passed, 10 skipped in
the existing 111 container with batch-invariant mode to avoid its
missing custom-op registration. Ruff `check` and `format --check` passed
on both touched tests. The historical [PR CI
run](https://github.com/vllm-project/vllm-ascend/actions/runs/34845930945)
passed pre-commit and the full CPU-UT job; the six previous test-double
failures are gone. Device-test matrix jobs were still running when that
evidence was recorded; this is not a current CI status report.
- Pre-generalization CP-helper checks: six dependency-free common-CP
tests and **1,500 exact before/after DSA-CP comparisons passed**. They
cover all ranks at TP1/2/4/8, main/MTP steps, padding, stable buffer
reuse, independent buffers, and SFA/indexer calculations. The PCP
orchestration regression passed against both revisions. These checks
execute repository method bodies with distributed/NPU imports excluded.
- Ruff 0.14.0 `check`, `format --check`, and AST parsing passed on the
eight Python files touched by the final helper cleanup. The full
`format.sh ci` command was unavailable locally because `pre-commit` was
not installed.
- An earlier 11-file focused pytest run at `c9916556c`, using vLLM
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`, reported **204 passed, 13
skipped, 0 failed**. It covered attention/indexer, proposer, worker
metadata plumbing and GLM/DeepSeek indexers. This earlier run is not a
full pytest result for the final tree.

#### A5 long-input precision checks (2026-09-14)

Setup: A5 node 112, Ascend 950DT, GLM-5.2 W4A4C8, vLLM main
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`, EP enabled,
`max_model_len=40960`, `max_num_batched_tokens=4096`, `max_num_seqs=8`,
and breakable graph disabled.

The following settings apply to **every row in this NPU matrix**, both
with and without the #16495 overlay, and to the four additional 12K-39K
retrieval checks below:

- **MTP: disabled.** No speculative-decoding configuration was supplied.
These are non-MTP precision results; they do not validate MTP-enabled
execution or the MTP + C8 combination.
- **C8: enabled for both SFA and the indexer.** The model is W4A4C8,
with `additional_config.enable_sparse_sfa_c8=True` and
`additional_config.enable_sparse_li_c8=True`.
- **The tested Ascend/vLLM source commits are paired correctly.** Both
Ascend test commits (`c30f4f421e0eea29c043f3ee8131c0333eb5100e` and
overlay `3bc93d4dae7a8fe397d3262880ac9da78ca83375`) pin vLLM
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` in
`.github/vllm-main-verified.commit`. The NPU checkout has that exact
vLLM HEAD, and the service logs confirm imports from
`/mnt/share/g00672350/refactor/vllm-a97dacb/vllm/`. The bot-managed vLLM
footer reflects the base branch and is not the version used for these
hardware results.

The reference script was used unchanged: eight serial requests, four
concurrent requests, then eight concurrent requests. Prompt lengths were
`17/14/32/752/1832/2192/2677/3700` tokens. The nine requests at 2,192 /
2,677 / 3,700 tokens were also checked for exact expected-code output.

| Code | Topology | DSA-CP requested / effective | Execution | Script
passes | >=2K exact |
| --- | --- | --- | --- | ---: | ---: |
| This PR's tree | TP8/DP1 | off / off | eager | 20/20 | 9/9 |
| This PR's tree | TP8/DP1 | off / off | `FULL_DECODE_ONLY` | **2/20** |
**0/9** |
| This PR's tree | TP8/DP1 | on / **off** | eager | 20/20 | 9/9 |
| This PR's tree | TP8/DP1 | on / **off** | `FULL_DECODE_ONLY` |
**3/20** | **0/9** |
| This PR's tree | TP4/DP2 | on / on | eager | 20/20 | 9/9 |
| This PR's tree | TP4/DP2 | on / on | `FULL_DECODE_ONLY` | 20/20 | 9/9
|
| This PR + #16495 | TP8/DP1 | off / off | `FULL_DECODE_ONLY` |
**20/20** | **9/9** |
| This PR + #16495 | TP4/DP2 | on / on | `FULL_DECODE_ONLY` | 20/20 |
9/9 |

All requests returned HTTP 200. The two failing graph rows produced
corrupted text; the script's substring check can count a corrupted short
answer as a pass, so the exact long-input result is the stronger signal.

With this vLLM configuration, TP8/DP1 makes `use_sequence_parallel_moe`
false and Ascend explicitly disables requested DSA-CP. TP4/DP2 enables
it. The table therefore records different topologies rather than
implying a controlled DSA-CP toggle at one TP/DP shape.

The overlay was #16495 head `d9bae5c4614df6959caf93d9b6f6ad05778fd52f`
cherry-picked onto the tested pre-squash head, producing test commit
`3bc93d4dae7a8fe397d3262880ac9da78ca83375`. In its TP8/DP1 graph mode,
four additional original-style sequential retrieval prompts at **12,761
/ 25,801 / 35,271 / 39,441 tokens** all returned the exact expected code
with `finish_reason=stop`.

Evidence on the shared A5 mount:
`/mnt/share/g00672350/refactor/pr16325-longinput-testkit-112/results_*/summary.json`.
Overlay service log:
`/mnt/share/g00672350/refactor/longinput_112_tp8dp1_graph_dsa_off_plus16495_3bc93d.log`.

#### A5 MTP1/MTP5 long-input precision checks (2026-09-14)

A5 node 111, Ascend 950DT, GLM-5.2 W4A4C8, TP8/DP1, EP enabled, DSA-CP
off, `FULL_DECODE_ONLY` graph mode, breakable graph disabled. Both
`enable_sparse_sfa_c8` and `enable_sparse_li_c8` were `True`. The server
used Ascend test overlay `3bc93d4dae7a8fe397d3262880ac9da78ca83375`
(this PR + #16495) with its matching vLLM
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`; `max_model_len=40960`,
`max_num_batched_tokens=4096`, and `max_num_seqs=8`. The engine logs
confirmed effective MTP speculative settings of `num_spec_tokens=1` and
`num_spec_tokens=5`, respectively, and successful decode-graph capture.

| MTP draft tokens | Original serial/concurrent cases | Original >=2K
exact answers | Additional exact long inputs |
| --- | ---: | ---: | ---: |
| 1 | 20/20 | 9/9 | 4/4 |
| 5 | 20/20 | 9/9 | 4/4 |

The additional prompt lengths were **12,761 / 25,801 / 38,051 / 39,441
tokens**. Each returned HTTP 200, `finish_reason=stop`, and content
exactly equal to the expected retrieval code; both full 24-case suites
passed. The original 20 cases used the same eight serial, four
concurrent, and eight concurrent requests as the non-MTP matrix.
MTP-enabled hardware evidence covers this TP8/DP1, DSA-CP-off
configuration with #16495 applied; it does not establish MTP precision
for the PR alone or for DSA-CP-on configurations.

#### MTP draft acceptance observed during the long-input runs

The server's `SpecDecoding metrics` log lines show that speculative
decoding was active. During the sustained request windows, MTP1 had
roughly **90–95% average draft acceptance** and **1.9–2.0 mean
acceptance length**. MTP5 had roughly **57–66% average draft
acceptance** and **3.9–4.3 mean acceptance length**; its per-position
acceptance declined from about **90–94% at draft position 1** to about
**31–44% at position 5**. These are interval metrics from the service
logs, not an aggregate over the 24 requests. Alongside the 24/24
exact-output checks for each setting, they show no obvious acceptance
collapse. No matched pre-refactor acceptance-rate baseline was
collected, so these numbers do **not** establish acceptance-rate
equivalence with the old implementation.

Evidence on the shared A5 mount:
`/mnt/share/g00672350/refactor/mtp1_graph_111_results_v2/summary.json`
and `/mnt/share/g00672350/refactor/mtp5_graph_111_results/summary.json`.
Service logs: `/mnt/share/g00672350/refactor/mtp1_graph_111.log` and
`/mnt/share/g00672350/refactor/mtp5_graph_111.log`.

#### A5 Model Runner V2 long-input precision checks (2026-09-14)

A5 node 111, the existing `xhg_refactor_sfa` container, GLM-5.2 W4A4C8,
both sparse SFA-C8 and LI-C8 enabled, `VLLM_USE_V2_MODEL_RUNNER=1`,
`FULL_DECODE_ONLY` graph mode, breakable graph disabled, EP enabled, and
the same Ascend #16325 + #16495 test overlay
`3bc93d4dae7a8fe397d3262880ac9da78ca83375` paired with vLLM
`a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`. Worker logs explicitly
confirmed the V2 model-runner path and successful decode-graph capture.
The overlay contains the same PR production code as restored revision
`199a459c`; #16495 is not part of this PR.

| Topology | DSA-CP | MTP draft tokens | Long-input exact-output cases |
| --- | --- | ---: | ---: |
| TP8/DP1 | off | 0 | **24/24** |
| TP8/DP1 | off | 1 | **24/24** |
| TP8/DP1 | off | 5 | **24/24** |
| TP4/DP2 | on | 0 | **24/24** |

Each suite comprised the prior 20 serial/concurrent requests plus four
exact-code retrieval prompts of **12,761 / 25,801 / 38,051 / 39,441
tokens**. All reported cases returned HTTP 200, `finish_reason=stop`,
and the expected output. These results validate the tested Model Runner
V2 graph configurations, not eager mode or this PR without #16495. The
TP4/DP2 + DSA-CP-on + MTP1/MTP5 combinations were **not run** because
the A5 nodes became occupied by other services; no result is claimed for
them.

Evidence on the shared A5 mount:
`/mnt/share/g00672350/refactor/mrv2_mtp0_graph_111_results/summary.json`,
`mrv2_mtp1_graph_111_results/summary.json`,
`mrv2_mtp5_graph_111_results/summary.json`, and
`mrv2_tp4dp2_dsa_mtp0_graph_111_results/summary.json`. Corresponding
service logs use the same names without `_results/summary.json` and with
`.log` suffix.

Validation limits: the A5 profile rejects replicated SFA DCP before
startup, so replicated-DCP address transforms are covered by unit tests
rather than this hardware matrix. MTP was disabled in the preceding
non-MTP matrix; the separate MTP1/MTP5 checks are reported above. GSM8K
and full repository pytest have not been rerun on the final tree; the
long-input checks are not a throughput benchmark.

<details>
<summary>Earlier combined-build GSM8K evidence</summary>

These results belong to combined validation commit
`5f0326bb3368df11a3067b9fe8c8bc88f9e5463a`, base
`fb00c6be102f31e356bea4ab0f5a64e5091608c6`, and vLLM
`b2f685834a6456197e7033966fdef52a23f1abcd`, before the indexer and MoE
fixes were separated. They are not final-tree GSM8K results.

Single-node W4A4C8, TP4/DP2, sparse SFA/LI-C8 enabled, MTP and MLAPO
disabled, `max_num_batched_tokens=4096`, `max_num_seqs=8`, breakable
graph disabled:

| DSA-CP | Execution | Long-input script | GSM8K |
| --- | --- | ---: | ---: |
| off | eager | 20/20 | 178/200 (89.0%) |
| off | `FULL_DECODE_ONLY` | 20/20 | 178/200 (89.0%) |
| on | eager | 20/20 | 183/200 (91.5%) |
| on | `FULL_DECODE_ONLY` | 20/20 | 180/200 (90.0%) |

</details>

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: Biuapha <1731372716@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants