Skip to content

[Performance][Attention] Fuse SFA K-path ops and enable PROLOG_V3 fused decode for non-PD serving - #16328

Merged
ZT-AIA merged 2 commits into
vllm-project:mainfrom
tanjiangshan:perf/glm-kpath-fusion
Sep 18, 2026
Merged

ZT-AIA merged 2 commits into
vllm-project:mainfrom
tanjiangshan:perf/glm-kpath-fusion

Conversation

@tanjiangshan

@tanjiangshan tanjiangshan commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

GLM5.2 SFA profiling shows several removable ops in the K processing path of every layer. This PR removes them and makes PROLOG_V3 the default fused preprocessing for quantized SFA layers in every deployment:

  1. Per-step int64 slot cast: exec_kv re-cast the shared slot mapping to int64 for npu_kv_rmsnorm_rope_cache in every layer, although all layers of a scheduling step receive the same int32 slot tensor. The conversion is now cached on the attention metadata (_int64_kv_slots), so one Cast kernel runs per step instead of one per layer (~5us x num_layers per step). The PROLOG_V3 fused preprocess reuses the same cached conversion for its int64 cache indices.

  2. Duplicate indexer GEMM: the indexer's k path (forward_k) and top-k stage (forward) both ran the same wk_weights_proj GEMM ([tokens, hidden] x [160, hidden]) on the same hidden states, once for the indexer K and once for the lightning-indexer weights. forward_k now returns the non-K tail of the GEMM output and forward reuses it, removing one GEMM plus its slice copy per indexer layer per step (falling back to the GEMM only when the two stages are handed different tensors).

  3. Redundant .contiguous() copies: npu_rms_norm returns a contiguous tensor so the copy before the C8 block-quant view was discarded, and torch.cat already allocates contiguous outputs for the sparse-attention query concat.

  4. PROLOG_V3 by default, no new switch: PROLOG_V3 (the npu_mla_prolog_v3 single fused op covering qkv proj + norm + rope + q up-proj + C8 quantize/pack + direct cache write) was previously gated on is_kv_consumer, i.e. only PD-disaggregated decode workers could take it. It is now the default fused preprocessing for quantized SFA layers in every deployment (plain serving, PD KV producers and KV consumers) and serves every attention state: prefill and decode steps both take the fused path (the per-step attention-state fallback to NATIVE is gone; only MLAPO keeps its token-count limit). Switch convergence:

    • enable_dsa_cp is the prefill/P-node route selector: it routes to AscendSFADSACPImpl, which unconditionally disables fused preprocessing, so the two are mutually exclusive by construction (dsa_cp on => prolog off, dsa_cp off => prolog on);
    • the C8 switches (enable_sparse_sfa_c8 / enable_sparse_li_c8) only select the KV cache layout and are orthogonal to this choice; W8A8Dynamic layers no longer require enable_sparse_sfa_c8 to take the fused path;
    • unquantized layers keep the NATIVE chain outside KV consumers because the unquantized weight preparation transposes fused_qkv_a_proj.weight in place, which the NATIVE fallback still consumes;
    • dispose_layer stays gated on is_kv_consumer so producers and plain-serving workers keep the fallback weights (the cost is the extra PROLOG_V3 weight copies: memory, not correctness).

Does this PR introduce any user-facing change?

Yes, a default behavior change: quantized (W8A8Dynamic / W8A8MXFP8) SFA deployments now take the PROLOG_V3 fused preprocessing for both prefill and decode steps by default, without any additional-config option. Deployments on enable_dsa_cp (prefill/P-node CP route) and unquantized (bf16) layers are unaffected. The default trades extra NPU weight memory (the retained qkv_a/q_b fallback weights on producers and plain-serving workers) for kernel savings.

How was this patch tested?

  • Unit tests added/updated in tests/ut/attention/test_sfa_v1.py:

    • per-step int64 slot caching (passthrough / convert-and-cache / re-convert on new step) and exec_kv reusing the cached slots across layers;
    • single wk_weights_proj invocation across the indexer k path and top-k stage (plus the fallback recomputation when the tensors differ);
    • PROLOG_V3 routing matrix for the default gate (non-PD W8A8Dynamic with and without C8 -> PROLOG_V3; non-PD MXFP8 -> PROLOG_V3; non-PD unquantized -> NATIVE; KV producer quantized -> PROLOG_V3, unquantized -> NATIVE);
    • weight-disposal guard confirming non-consumer workers keep the fallback weights.
  • uvx ruff==0.14.0 check and format --check on all touched Python files.

  • CI cpu-ut green (4431+ tests).

End-to-end A/B benchmark: DSA-CP route (main) vs PROLOG_V3 route (this PR) (Atlas 800 A3, 16x 910B, CANN 9.1.0, vllm 0.28.0, GLM-5.2-w4a8c8, DP2xTP8 + EP):

The comparison is between the two decode preprocessing routes, each in its best usable configuration. enable_dsa_cp requires SP-MoE and unconditionally disables the fused preprocessing path, so the two options are mutually exclusive by construction; everything else is identical on both sides:

item base (main 125924bb2) PR (cb525a869)
decode preprocessing route enable_dsa_cp=true PROLOG_V3 route (now the default; measured with the earlier opt-in build of this PR)
SP-MoE / sequence parallelism (VLLM_ASCEND_ENABLE_FLASHCOMM1=1; auto-enabled by DSA-CP on the base side) on on
enable_sparse_sfa_c8 + enable_sparse_li_c8 on on
enable_balance_scheduling, enable_fused_mc2=0 on on
--enable-expert-parallel (EP), DP2xTP8 on on
MTP speculative decoding (deepseek_mtp, num_speculative_tokens=3, enforce_eager) on on
cudagraph FULL_DECODE_ONLY, --quantization ascend on on
multistream_overlap_shared_expert off off (incompatible with FlashComm1 on DP>1, see #16446)

Both sides verified via serve logs: no "Disabling DSA-CP" / sp-MoE active on the base side, MlaPrologV3 kernels present in the PR-side profile only. Workload: GSM8K test full 1319 prompts, ais-bench stream mode, concurrency 8, temperature 0.

metric base (DSA-CP route) PR (PROLOG_V3 route) delta
TPOT avg 27.9 ms 22.9 ms -17.9%
TPOT median 27.8 ms 22.8 ms -18.0%
E2EL avg 7,181.6 ms 5,885.0 ms -18.1%
TTFT avg 459.9 ms 395.8 ms -13.9%
per-request output throughput 33.73 tok/s 40.91 tok/s +21.3%
request throughput (aggregate) 1.1117 req/s 1.3564 req/s +22.0%
failed requests 0/1319 0/1319 =

Kernel-level verification (rank0 profile of a 500-token decode request):

kernel base (DSA-CP) PR (PROLOG_V3) note
MlaPrologV3 0 13,369 fused decode path taken over
InterleaveRope 28,980 1,156 -96%: NATIVE rope chain leaves decode
DynamicBlockQuant 14,490 578 -96%: c8 quant chain leaves decode
ScatterNdUpdate 21,035 6,967 -67%
Slice 31,556 12,211 -61% count
QuantBatchMatmulV3 73,652 (1,629 ms) 44,207 (835 ms) -40% count / -49% time
KvQuantSparseFlashAttention 14,190 13,703 ~same (attention body)
total decode kernels 711,070 479,326 -32.6%

E2E coverage for the PROLOG_V3 default route:

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests module:core labels Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request implements several performance optimizations for the Sparse Flash Attention (SFA) K-path and expands the availability of the PROLOG_V3 fused decode preprocessing. By caching metadata conversions, reusing GEMM outputs, and eliminating redundant memory copies, the changes reduce kernel overhead per layer. Additionally, a new configuration option allows non-PD serving workers to utilize the PROLOG_V3 fused path, trading a small amount of memory for significant decode kernel savings.

Highlights

  • Performance Optimization: Cached int64 slot mapping on attention metadata to eliminate redundant Cast kernel operations per layer.
  • GEMM Reuse: Optimized the indexer forward path by reusing the wk_weights_proj GEMM output across both forward_k and forward stages.
  • Redundancy Removal: Removed unnecessary .contiguous() calls in SFA and device operators where tensor operations already guarantee contiguous output.
  • Feature Enablement: Enabled PROLOG_V3 fused decode for non-PD serving via the new enable_sfa_prolog_v3 configuration option.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces SFA PROLOG_V3 fused decode preprocessing support for plain serving via a new configuration option, alongside several K-path optimizations such as caching int64 slot conversions and reusing indexer weights to avoid redundant GEMM operations. Redundant .contiguous() calls have also been cleaned up.

I have no feedback to provide as there are no review comments.

Suggested PR Title:

[Attention][Feature] Enable SFA PROLOG_V3 fused decode preprocessing in plain serving and optimize K-path fusions

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces several optimizations and features for the Ascend SFA (Sparse Flash Attention) attention backend:
1. **SFA PROLOG_V3 Opt-in for Plain Serving**: Enables the `PROLOG_V3` fused decode preprocessing path (using a single `npu_mla_prolog_v3` op) outside of PD-disaggregated KV-consumer workers (i.e., in plain serving) via the new `enable_sfa_prolog_v3` configuration option.
2. **K-Path Fusions & Optimizations**:
   - Caches the `int64` converted KV slot mapping once per scheduling step (`_int64_kv_slots`) to avoid redundant Cast kernels across layers.
   - Reuses the non-K tail of the `wk_weights_proj` GEMM output (`indexer_weights`) in the indexer's top-k stage when the hidden states match, avoiding duplicate GEMM computations.
3. **Redundancy Cleanup**: Removes redundant `.contiguous()` calls after `torch.cat` and `npu_rms_norm` since these operations already return contiguous tensors.

### Does this PR introduce _any_ user-facing change?
Yes, it introduces a new configuration option `enable_sfa_prolog_v3` (boolean, default `False`) to allow plain serving to opt into the SFA PROLOG_V3 fused decode preprocessing path.

### How was this patch tested?
Added new unit tests in `tests/ut/attention/test_sfa_v1.py` covering:
- `_int64_kv_slots` caching and passthrough behavior.
- Reuse of `int64` slots across layers in `exec_kv`.
- Reuse of `wk_weights_proj` weights in the indexer.
- Path resolution logic for non-PD opt-in configurations.

@lijiahang226

Copy link
Copy Markdown
Collaborator

Thanks for the contribution. Could you provide a performance comparison between the MLA Prolog v3 path and DSACP under prefill / mixed-deployment (hybrid) scenarios? These two features currently conflict with each other.

tanjiangshan added a commit to tanjiangshan/vllm-ascend that referenced this pull request Sep 14, 2026
The cpu-ut CI on PR vllm-project#16328 failed on three tests:

1. tests/ut/ops/test_mla.py (2 tests): the forward_k mocks still
   return 2-tuples while forward now unpacks 3 values (k, scale,
   weights tail) - update the mocks to return 3-tuples.

2. tests/ut/attention/test_sfa_v1.py::test_indexer_forward_reuses_
   wk_weights_proj: the test emulated the pre-squash calling pattern
   (forward_k called externally before forward). forward now calls
   forward_k internally, so the same sequence counts one extra GEMM.
   Rework the assertions for the current architecture: reset the GEMM
   mock before forward, expect exactly one call when both stages share
   the same hidden-states tensor (reuse), expect three calls (one in
   forward_k + one fallback) for a distinct top-k input, and check the
   weights tail passed to indexer_select_post_process by value instead
   of identity.

Signed-off-by: huamus <1943805462@qq.com>
@lijiahang226 lijiahang226 added the ready-precise run selected e2e test for pr label Sep 14, 2026
tanjiangshan added a commit to tanjiangshan/vllm-ascend that referenced this pull request Sep 14, 2026
The cpu-ut CI on PR vllm-project#16328 failed on three tests:

1. tests/ut/ops/test_mla.py (2 tests): the forward_k mocks still
   return 2-tuples while forward now unpacks 3 values (k, scale,
   weights tail) - update the mocks to return 3-tuples.

2. tests/ut/attention/test_sfa_v1.py::test_indexer_forward_reuses_
   wk_weights_proj: the test emulated the pre-squash calling pattern
   (forward_k called externally before forward). forward now calls
   forward_k internally, so the same sequence counts one extra GEMM.
   Rework the assertions for the current architecture: reset the GEMM
   mock before forward, expect exactly one call when both stages share
   the same hidden-states tensor (reuse), expect three calls (one in
   forward_k + one fallback) for a distinct top-k input, and check the
   weights tail passed to indexer_select_post_process by value instead
   of identity.

Signed-off-by: huamus <1943805462@qq.com>
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

tanjiangshan added a commit to tanjiangshan/vllm-ascend that referenced this pull request Sep 15, 2026
The cpu-ut CI on PR vllm-project#16328 failed on three tests:

1. tests/ut/ops/test_mla.py (2 tests): the forward_k mocks still
   return 2-tuples while forward now unpacks 3 values (k, scale,
   weights tail) - update the mocks to return 3-tuples.

2. tests/ut/attention/test_sfa_v1.py::test_indexer_forward_reuses_
   wk_weights_proj: the test emulated the pre-squash calling pattern
   (forward_k called externally before forward). forward now calls
   forward_k internally, so the same sequence counts one extra GEMM.
   Rework the assertions for the current architecture: reset the GEMM
   mock before forward, expect exactly one call when both stages share
   the same hidden-states tensor (reuse), expect three calls (one in
   forward_k + one fallback) for a distinct top-k input, and check the
   weights tail passed to indexer_select_post_process by value instead
   of identity.

Signed-off-by: huamus <1943805462@qq.com>
@tanjiangshan

tanjiangshan commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

tanjiangshan added a commit to tanjiangshan/vllm-ascend that referenced this pull request Sep 15, 2026
The cpu-ut CI on PR vllm-project#16328 failed on three tests:

1. tests/ut/ops/test_mla.py (2 tests): the forward_k mocks still
   return 2-tuples while forward now unpacks 3 values (k, scale,
   weights tail) - update the mocks to return 3-tuples.

2. tests/ut/attention/test_sfa_v1.py::test_indexer_forward_reuses_
   wk_weights_proj: the test emulated the pre-squash calling pattern
   (forward_k called externally before forward). forward now calls
   forward_k internally, so the same sequence counts one extra GEMM.
   Rework the assertions for the current architecture: reset the GEMM
   mock before forward, expect exactly one call when both stages share
   the same hidden-states tensor (reuse), expect three calls (one in
   forward_k + one fallback) for a distinct top-k input, and check the
   weights tail passed to indexer_select_post_process by value instead
   of identity.

Signed-off-by: huamus <1943805462@qq.com>
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

tanjiangshan added a commit to tanjiangshan/vllm-ascend that referenced this pull request Sep 16, 2026
The cpu-ut CI on PR vllm-project#16328 failed on three tests:

1. tests/ut/ops/test_mla.py (2 tests): the forward_k mocks still
   return 2-tuples while forward now unpacks 3 values (k, scale,
   weights tail) - update the mocks to return 3-tuples.

2. tests/ut/attention/test_sfa_v1.py::test_indexer_forward_reuses_
   wk_weights_proj: the test emulated the pre-squash calling pattern
   (forward_k called externally before forward). forward now calls
   forward_k internally, so the same sequence counts one extra GEMM.
   Rework the assertions for the current architecture: reset the GEMM
   mock before forward, expect exactly one call when both stages share
   the same hidden-states tensor (reuse), expect three calls (one in
   forward_k + one fallback) for a distinct top-k input, and check the
   weights tail passed to indexer_select_post_process by value instead
   of identity.

Signed-off-by: huamus <1943805462@qq.com>
ZT-AIA
ZT-AIA previously approved these changes Sep 16, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions github-actions Bot removed documentation Improvements or additions to documentation module:core merge-conflicts labels Sep 17, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

GLM5.2 SFA profiling shows removable ops in the K processing path of
every layer; PROLOG_V3 (the npu_mla_prolog_v3 single fused op covering
qkv proj + norm + rope + q up-proj + C8 quantize/pack + direct cache
write) was previously gated on is_kv_consumer, i.e. only
PD-disaggregated decode workers could take it.

1. exec_kv re-casts the shared slot mapping to int64 for
   npu_kv_rmsnorm_rope_cache in every layer, although all layers of a
   scheduling step receive the same int32 slot tensor. The conversion is
   now cached on the attention metadata, so one Cast kernel runs per
   step instead of one per layer; the PROLOG_V3 fused preprocess reuses
   the same cached conversion for its int64 cache indices.

2. The indexer k path (forward_k) and top-k stage (forward) both run the
   same wk_weights_proj GEMM on the same hidden states, once for the
   indexer K and once for the lightning-indexer weights. forward_k now
   returns the non-K tail of the GEMM output and forward reuses it,
   removing one GEMM plus its slice copy per indexer layer per step
   (falling back to the GEMM only when the two stages are handed
   different tensors).

3. Redundant .contiguous() copies removed: npu_rms_norm returns a
   contiguous tensor so the copy before the C8 block-quant view is
   discarded, and torch.cat already allocates contiguous outputs for the
   sparse-attention query concat.

4. PROLOG_V3 becomes the default fused preprocessing for quantized SFA
   layers in every deployment (plain serving, PD KV producers and KV
   consumers) and serves every attention state: prefill and decode steps
   both take the fused path, and the per-step attention-state fallback
   to NATIVE is gone (only MLAPO keeps its token-count limit).
   enable_dsa_cp remains the prefill/P-node route selector: it routes to
   AscendSFADSACPImpl, which unconditionally disables fused
   preprocessing, so the two are mutually exclusive by construction.
   The C8 switches only select the KV cache layout and are orthogonal to
   this choice; W8A8Dynamic layers no longer require enable_sparse_sfa_c8
   to take the fused path. Unquantized layers keep the NATIVE chain
   outside KV consumers because the unquantized weight preparation
   transposes fused_qkv_a_proj.weight in place, which the NATIVE fallback
   still consumes; dispose_layer stays gated on is_kv_consumer so
   producers and plain-serving workers keep the fallback weights (the
   cost is the extra PROLOG_V3 weight copies: memory, not correctness).

Signed-off-by: huamus <1943805462@qq.com>
@lijiahang226

Copy link
Copy Markdown
Collaborator

Could you please double-check whether we really need this many UTs for this change?

- routing matrix for the default gate: non-PD W8A8Dynamic (with and
  without C8) and MXFP8 resolve to PROLOG_V3, unquantized stays NATIVE,
  KV producers take the fused path too (quantized) or stay NATIVE
  (unquantized); the KV-consumer cases and the weight-disposal guard
  are unchanged;
- per-step int64 slot-conversion caching (passthrough / convert-and-cache /
  re-convert on new step) and exec_kv reusing the cached slots across layers;
- a single wk_weights_proj invocation across the indexer k path and the
  top-k stage, plus the fallback recomputation when the two stages are
  handed different tensors;
- adapt the forward_k mocks in test_mla.py to the 3-tuple return.

Signed-off-by: huamus <1943805462@qq.com>
@ZT-AIA
ZT-AIA merged commit bdd53a2 into vllm-project:main Sep 18, 2026
22 checks passed
ningjingbengxiaohai pushed a commit that referenced this pull request Sep 19, 2026
### What this PR does / why we need it?

The A5 `GLM5_1_W4A4_A5.yaml` nightly fails during graph warmup with
`aclnnMlaPrologV3WeightNz` error 561002: `Parameter KvQuantMode of
MlaPrologV3 has incorrect value 3. It should be 1, 0.` See [the failing
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35381103773/job/105718954194).

The K3 custom operator introduced in #14454 shares its ACLNN symbols and
OPP operator type with CANN's MLAPO v3. Its restricted template set
(#16516) rejects the per-tile C8 configuration used by SFA's
`torch_npu.npu_mla_prolog_v3` call. #16328 exposed that fused path to
ordinary quantized serving.

Give the K3 implementation its own identity throughout the build and
runtime:

- Rename its Torch schema to `npu_mla_prolog_v3_k3`, OPP type to
`MlaPrologV3K3`, kernel to `mla_prolog_v3_k3`, and public/inner ACLNN
symbols to their `K3` variants.
- Route K3's no-RoPE MLA path to the renamed custom operator. Keep
ordinary MLA and SFA C8 per-tile mode 3 on
`torch_npu.npu_mla_prolog_v3`, backed by CANN.
- Update build metadata, bindings, documentation and tests without
changing the kernel math, supported K3 modes or cache layouts.

### Does this PR introduce _any_ user-facing change?

GLM5 C8 serving can use CANN's per-tile MLAPO v3 while K3 custom
operators are installed. Direct users of the internal
`_C_ascend.npu_mla_prolog_v3` entry point must use
`_C_ascend.npu_mla_prolog_v3_k3`. Rebuild/reinstall the complete custom
operator package to remove the old unsuffixed registration; no
compatibility alias is retained because it would recreate the collision.

### How was this patch tested?

Validated on Ascend950DT with CANN 9.1.0 and torch_npu 2.10.0.post4:

- `pytest -q tests/ut/attention/test_mla_v1.py
tests/ut/attention/test_sfa_v1.py`: **182 passed**, including
native/custom dispatch and MXFP8 C8 mode 3 routing regressions.
- `cd csrc && bash build.sh --pkg --ops=mla_prolog_v3_k3 --soc=ascend950
-j256`: passed. Full `vllm_ascend_C` Torch extension CMake build also
passed with `SOC_VERSION=ascend950dt_9582` and parallelism 256.
- Ran all **8 existing K3 NPU cases** with the freshly built extension
and OPP package, covering head96 BF16, no-RoPE, query norm, BSND,
PA_BSND, PA_NZ and noncontiguous cache layouts: passed.
- Added and ran
`test_cann_mla_prolog_v3_mxfp8_per_tile_with_k3_registered`: the K3
no-RoPE operator executes first, followed by native MXFP8 weight mode 3
/ KV mode 3; outputs are finite and only the selected KV slots are
written.
- Compared a separate pure-CANN process with native mode 3 after
executing K3: query, query RoPE and packed KV cache are **bitwise
identical**. Native mode 3 graph capture/replay with K3 loaded also
matches that baseline bitwise.
- Audited the fresh `libcust_opapi.so`: only
`aclnnMlaPrologV3K3WeightNz*` and `aclnnInnerMlaPrologV3K3*` are
exported for this operator; there are no unsuffixed MLAPO v3 exports.
Generated OPP metadata uses `MlaPrologV3K3`.

- `bash format.sh ci`: all hooks passed (Ruff, codespell, typos,
Markdown, actionlint, secret scan, shellcheck and repository checks).

The NPU checks load the fresh extension and operator package directly to
avoid importing the image's older vLLM-Ascend binaries. The
maintainer-triggered [GLM5_1_W4A4_A5 nightly
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35426937149/job/105855093279)
subsequently passed on PR commit
`a49f62d1ec61b80b729f8f09925605ceff1b889f`: `1 passed, 14 warnings in
405.38s`. The log confirms successful server startup, ACL graph replay
and a completion request under TP8/DP1/EP, C8, prefix caching, async
scheduling and MTP3. This is the configured functional test; no
accuracy/performance benchmark artifacts were produced. The workflow
aggregate reports failure and its artifact-merge check reports `No
artifacts found matching pattern nightly-test-benchmark-results-*`; the
GLM5 test job itself is successful.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: MQ <maomaoyu870@gmail.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…t#16924)

### What this PR does / why we need it?

The A5 `GLM5_1_W4A4_A5.yaml` nightly fails during graph warmup with
`aclnnMlaPrologV3WeightNz` error 561002: `Parameter KvQuantMode of
MlaPrologV3 has incorrect value 3. It should be 1, 0.` See [the failing
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35381103773/job/105718954194).

The K3 custom operator introduced in vllm-project#14454 shares its ACLNN symbols and
OPP operator type with CANN's MLAPO v3. Its restricted template set
(vllm-project#16516) rejects the per-tile C8 configuration used by SFA's
`torch_npu.npu_mla_prolog_v3` call. vllm-project#16328 exposed that fused path to
ordinary quantized serving.

Give the K3 implementation its own identity throughout the build and
runtime:

- Rename its Torch schema to `npu_mla_prolog_v3_k3`, OPP type to
`MlaPrologV3K3`, kernel to `mla_prolog_v3_k3`, and public/inner ACLNN
symbols to their `K3` variants.
- Route K3's no-RoPE MLA path to the renamed custom operator. Keep
ordinary MLA and SFA C8 per-tile mode 3 on
`torch_npu.npu_mla_prolog_v3`, backed by CANN.
- Update build metadata, bindings, documentation and tests without
changing the kernel math, supported K3 modes or cache layouts.

### Does this PR introduce _any_ user-facing change?

GLM5 C8 serving can use CANN's per-tile MLAPO v3 while K3 custom
operators are installed. Direct users of the internal
`_C_ascend.npu_mla_prolog_v3` entry point must use
`_C_ascend.npu_mla_prolog_v3_k3`. Rebuild/reinstall the complete custom
operator package to remove the old unsuffixed registration; no
compatibility alias is retained because it would recreate the collision.

### How was this patch tested?

Validated on Ascend950DT with CANN 9.1.0 and torch_npu 2.10.0.post4:

- `pytest -q tests/ut/attention/test_mla_v1.py
tests/ut/attention/test_sfa_v1.py`: **182 passed**, including
native/custom dispatch and MXFP8 C8 mode 3 routing regressions.
- `cd csrc && bash build.sh --pkg --ops=mla_prolog_v3_k3 --soc=ascend950
-j256`: passed. Full `vllm_ascend_C` Torch extension CMake build also
passed with `SOC_VERSION=ascend950dt_9582` and parallelism 256.
- Ran all **8 existing K3 NPU cases** with the freshly built extension
and OPP package, covering head96 BF16, no-RoPE, query norm, BSND,
PA_BSND, PA_NZ and noncontiguous cache layouts: passed.
- Added and ran
`test_cann_mla_prolog_v3_mxfp8_per_tile_with_k3_registered`: the K3
no-RoPE operator executes first, followed by native MXFP8 weight mode 3
/ KV mode 3; outputs are finite and only the selected KV slots are
written.
- Compared a separate pure-CANN process with native mode 3 after
executing K3: query, query RoPE and packed KV cache are **bitwise
identical**. Native mode 3 graph capture/replay with K3 loaded also
matches that baseline bitwise.
- Audited the fresh `libcust_opapi.so`: only
`aclnnMlaPrologV3K3WeightNz*` and `aclnnInnerMlaPrologV3K3*` are
exported for this operator; there are no unsuffixed MLAPO v3 exports.
Generated OPP metadata uses `MlaPrologV3K3`.

- `bash format.sh ci`: all hooks passed (Ruff, codespell, typos,
Markdown, actionlint, secret scan, shellcheck and repository checks).

The NPU checks load the fresh extension and operator package directly to
avoid importing the image's older vLLM-Ascend binaries. The
maintainer-triggered [GLM5_1_W4A4_A5 nightly
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35426937149/job/105855093279)
subsequently passed on PR commit
`a49f62d1ec61b80b729f8f09925605ceff1b889f`: `1 passed, 14 warnings in
405.38s`. The log confirms successful server startup, ACL graph replay
and a completion request under TP8/DP1/EP, C8, prefix caching, async
scheduling and MTP3. This is the configured functional test; no
accuracy/performance benchmark artifacts were produced. The workflow
aggregate reports failure and its artifact-merge check reports `No
artifacts found matching pattern nightly-test-benchmark-results-*`; the
GLM5 test job itself is successful.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: MQ <maomaoyu870@gmail.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
…t#16924)

### What this PR does / why we need it?

The A5 `GLM5_1_W4A4_A5.yaml` nightly fails during graph warmup with
`aclnnMlaPrologV3WeightNz` error 561002: `Parameter KvQuantMode of
MlaPrologV3 has incorrect value 3. It should be 1, 0.` See [the failing
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35381103773/job/105718954194).

The K3 custom operator introduced in vllm-project#14454 shares its ACLNN symbols and
OPP operator type with CANN's MLAPO v3. Its restricted template set
(vllm-project#16516) rejects the per-tile C8 configuration used by SFA's
`torch_npu.npu_mla_prolog_v3` call. vllm-project#16328 exposed that fused path to
ordinary quantized serving.

Give the K3 implementation its own identity throughout the build and
runtime:

- Rename its Torch schema to `npu_mla_prolog_v3_k3`, OPP type to
`MlaPrologV3K3`, kernel to `mla_prolog_v3_k3`, and public/inner ACLNN
symbols to their `K3` variants.
- Route K3's no-RoPE MLA path to the renamed custom operator. Keep
ordinary MLA and SFA C8 per-tile mode 3 on
`torch_npu.npu_mla_prolog_v3`, backed by CANN.
- Update build metadata, bindings, documentation and tests without
changing the kernel math, supported K3 modes or cache layouts.

### Does this PR introduce _any_ user-facing change?

GLM5 C8 serving can use CANN's per-tile MLAPO v3 while K3 custom
operators are installed. Direct users of the internal
`_C_ascend.npu_mla_prolog_v3` entry point must use
`_C_ascend.npu_mla_prolog_v3_k3`. Rebuild/reinstall the complete custom
operator package to remove the old unsuffixed registration; no
compatibility alias is retained because it would recreate the collision.

### How was this patch tested?

Validated on Ascend950DT with CANN 9.1.0 and torch_npu 2.10.0.post4:

- `pytest -q tests/ut/attention/test_mla_v1.py
tests/ut/attention/test_sfa_v1.py`: **182 passed**, including
native/custom dispatch and MXFP8 C8 mode 3 routing regressions.
- `cd csrc && bash build.sh --pkg --ops=mla_prolog_v3_k3 --soc=ascend950
-j256`: passed. Full `vllm_ascend_C` Torch extension CMake build also
passed with `SOC_VERSION=ascend950dt_9582` and parallelism 256.
- Ran all **8 existing K3 NPU cases** with the freshly built extension
and OPP package, covering head96 BF16, no-RoPE, query norm, BSND,
PA_BSND, PA_NZ and noncontiguous cache layouts: passed.
- Added and ran
`test_cann_mla_prolog_v3_mxfp8_per_tile_with_k3_registered`: the K3
no-RoPE operator executes first, followed by native MXFP8 weight mode 3
/ KV mode 3; outputs are finite and only the selected KV slots are
written.
- Compared a separate pure-CANN process with native mode 3 after
executing K3: query, query RoPE and packed KV cache are **bitwise
identical**. Native mode 3 graph capture/replay with K3 loaded also
matches that baseline bitwise.
- Audited the fresh `libcust_opapi.so`: only
`aclnnMlaPrologV3K3WeightNz*` and `aclnnInnerMlaPrologV3K3*` are
exported for this operator; there are no unsuffixed MLAPO v3 exports.
Generated OPP metadata uses `MlaPrologV3K3`.

- `bash format.sh ci`: all hooks passed (Ruff, codespell, typos,
Markdown, actionlint, secret scan, shellcheck and repository checks).

The NPU checks load the fresh extension and operator package directly to
avoid importing the image's older vLLM-Ascend binaries. The
maintainer-triggered [GLM5_1_W4A4_A5 nightly
job](https://github.com/vllm-project/vllm-ascend/actions/runs/35426937149/job/105855093279)
subsequently passed on PR commit
`a49f62d1ec61b80b729f8f09925605ceff1b889f`: `1 passed, 14 warnings in
405.38s`. The log confirms successful server startup, ACL graph replay
and a completion request under TP8/DP1/EP, C8, prefix caching, async
scheduling and MTP3. This is the configured functional test; no
accuracy/performance benchmark artifacts were produced. The workflow
aggregate reports failure and its artifact-merge check reports `No
artifacts found matching pattern nightly-test-benchmark-results-*`; the
GLM5 test job itself is successful.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: MQ <maomaoyu870@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants