Skip to content

[BugFix][Model] Fix the GLM-5.3-Flash NoPE and KDA convolution paths on Ascend - #15885

Open
yiminghub2024 wants to merge 3 commits into
vllm-project:mainfrom
yiminghub2024:pr/glm53-kda-conv1d
Open

yiminghub2024 wants to merge 3 commits into
vllm-project:mainfrom
yiminghub2024:pr/glm53-kda-conv1d

Conversation

@yiminghub2024

@yiminghub2024 yiminghub2024 commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

[BugFix][Model] Fix the GLM-5.3-Flash NoPE and KDA convolution paths on Ascend

What this PR does / why we need it?

Four things GLM-5.3-Flash needs before it runs correctly on Ascend, and one bug
that reaches well past it.

GLM-5.3-Flash applies no rope, so two paths that assume one had to be taught to
step aside: the QKNorm/rope fusion pass, which fused a rope that is not there,
and the MLA rope-cache operators, which a NoPE layer must not call.

Its KDA layers convolve through causal_conv1d_update. The per-request PyTorch
path read query_start_loc with .item() once per request, which is rejected
outright while an ACL graph is being captured, so decode-FULL capture aborted;
it also looped in Python once per KDA layer on every decode step. Route to the
fused npu_causal_conv1d_custom operator where the hardware registers it, and
to a batched torch expression where it does not -- A5 withholds
RUNTIME_CUSTOM_OPS (#7157), so custom ops are off there entirely.

The last commit is the one worth reading. Both torch paths read
num_accepted_tokens as how many of this step's tokens to convolve. It
describes the step before: which drafts the sampler keeps is only known once
the model has run. A steady decode accepts one token, so seven of every eight
draft-verify rows fell out and reached the recurrent layers as raw projections.
Their logits could not match a draft, acceptance sat at zero past the first
position, and generation came out garbled.

Upstream's kernel uses the count for one thing, the column a request's history
starts at:

conv_state_token_offset = tl.load(num_accepted_tokens_ptr + idx_seq) - 1

and widens the state it writes back to width - 1 + (seqlen - 1), which is the
room that rewind reads from. Both paths now do the same: convolve every row,
take the count as the read offset, and store the history shifted by one, so
that the window ending at this step's a-th token sits at column a - 1.

The per-request path in vllm_ascend/ops/causal_conv1d.py is the fallback for
every model whose hardware withholds the fused operator, Qwen3-Next GDN and
Kimi KDA among them, and it had no rewind at all: it only ever read and wrote
the first width - 1 columns of a row that is allocated num_spec wider.
Speculative decoding on those models is affected the same way.

Does this PR introduce any user-facing change?

No new options or configuration. Speculative decoding on models that fall back
to the torch conv1d path now produces correct output and a usable acceptance
rate where it previously produced neither.

How was this patch tested?

Unit tests, 35 across three files:

pytest -sv tests/ut/models/test_glm5next_causal_conv1d.py \
           tests/ut/attention/test_mla_nope_cache_paths.py \
           tests/ut/compilation/test_qknorm_rope_fusion_pass.py

The conv1d tests previously checked the batched expression only against the
per-request one, which is what let a mistake shared by both go unseen. They now
also check against the operator's definition written out directly: that every
row of a verify step is convolved, and that reading the stored row at offset
a - 1 lands on the window ending with this step's a-th token, for every a.

On hardware, GLM-5.3-Flash TP8 on Ascend A5 with DFlash2 and seven draft
tokens, per-position acceptance:

  • before: 0.000 at every position
  • after: 0.583 0.403 0.278 0.181 0.167 0.111 0.069

Output went from garbled to coherent on the same prompt.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses critical compatibility and performance issues for the GLM-5.3-Flash model on Ascend hardware. It introduces robust handling for NoPE (No RoPE) layers, ensures proper graph capture by eliminating host-syncing operations in the KDA convolution path, and corrects logic errors in speculative decoding history management. These changes significantly improve the reliability and correctness of speculative decoding on models falling back to the torch convolution path.

Highlights

  • NoPE Layer Support: Updated MLA and QKNorm/RoPE fusion passes to correctly handle models without rotary embeddings (NoPE), preventing invalid rope cache access.
  • Batched Convolution Update: Implemented a batched causal convolution update to eliminate host-syncing .item() calls, enabling successful ACL graph capture on Ascend hardware.
  • Speculative Decoding Logic: Corrected the num_accepted_tokens logic in the convolution path to ensure accurate history rewinding during speculative decoding.
  • Test Coverage: Added 35 unit tests across three files to validate GLM-5.3-Flash convolution paths, MLA NoPE branches, and fusion pass gating.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][BugFix] Support host-sync-free causal conv1d update and fix NoPE layer gating

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the causal convolution 1D operations for GLM-5.3-Flash to ensure they are free of host syncs (such as `.item()` calls) during ACL graph capture. It introduces a batched PyTorch fallback expression when the fused `npu_causal_conv1d_custom` operator is unavailable (e.g., on A5 hardware). Additionally, it fixes gating for MLA NoPE layers to prevent them from entering paths that expect a non-empty RoPE cache, and disables QKNorm-RoPE fusion when `rope_dim <= 0`.

Review feedback suggests replacing `index_copy_` with standard PyTorch tensor assignment in `causal_conv1d.py` to ensure safer and more idiomatic execution on NPU backends.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
New unit tests have been added under `tests/ut/` covering the MLA NoPE cache paths, QKNorm-RoPE fusion pass gating, and GLM-5.3-Flash causal conv1d updates.

Comment thread vllm_ascend/models/glm5next/ops/causal_conv1d.py Outdated
Comment thread vllm_ascend/models/glm5next/ops/causal_conv1d.py Outdated
@zhangxinyuehfad zhangxinyuehfad added the ready-precise run selected e2e test for pr label Sep 7, 2026
Comment thread vllm_ascend/ops/causal_conv1d.py Outdated
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@yiminghub2024
yiminghub2024 force-pushed the pr/glm53-kda-conv1d branch 2 times, most recently from 61f79b2 to 6487bf5 Compare September 10, 2026 11:09
…on Ascend

Linearize onto current main so CI's rebase no longer conflicts in
tests/ut/conftest.py. Keep only the npugraph_ex/torchair CPU-UT mocks
on top of main's conftest.

Signed-off-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com>
The linearized push dropped the whole module while __init__.py still
imports it, so mypy failed with import-untyped. Keep the intended
19-line KDA bind removal and put the rest of the file back.

Signed-off-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com>
The previous E2E run sat on Initialize containers for >70m with no
logs available. A new SHA cancels that stuck run via workflow
concurrency (cancel-in-progress).

Signed-off-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com>
@yiminghub2024

Copy link
Copy Markdown
Contributor Author

Retriggered CI with an empty commit (8743a2b).

The previous E2E run (https://github.com/vllm-project/vllm-ascend/actions/runs/34471051271) assigned runners for the three prepare-csrc-cache jobs at 11:40 UTC, then sat on Initialize containers for 70+ minutes with no logs (BlobNotFound). /rerun only covers failed jobs, and I cannot cancel Actions on this repo.

pr_test.yaml concurrency is cancel-in-progress, so this SHA should cancel the hung run and start a fresh E2E pass.

@yiminghub2024

Copy link
Copy Markdown
Contributor Author

a3-4 card-(part 2-2) failed on test_deepseek_v4_dsa_pcp_dspark (job):

Acceptance rate at draft position 2 is 0.4933, below minimum 0.55
(tolerance 0.03, so need >= 0.52)

The sibling test in the same file passed. This PR does not touch DeepSeek-V4; the miss is a noisy DSpark acceptance-rate check, not the GLM-5.3 NoPE/KDA change. Remaining selected tests are still running — I will /rerun the failed job after this E2E run finishes.

@yiminghub2024

yiminghub2024 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@Ruiqiu-Zheng

Ruiqiu-Zheng commented Sep 11, 2026 •

Copy link
Copy Markdown

I reviewed the GLM-5.3 KDA causal-conv ordinary non-spec prefill portion of this PR with official GLM-5.3-Flash layer0 BF16 q/k/v conv weights. This is component/wrapper evidence only; it is not a full-model, E2E, MTP/decode/speculative, production, all-rank/all-layer, or whole-PR readiness claim.

Evidence summary:

  • Official weights: q/k/v layer0 tensors were range-extracted from zai-org/GLM-5.3-Flash, each BF16 [8192,1,4]. Raw tensor hashes: q 714fbd54ce53f7ef926b4c26bb67b22bee38758862625a46f4af0db05f9f4381, k 81a666b5a7f4a6211101de40cab8dd69841d967143e0426fa7fcd1d092a4f528, v c3d2ce7063c7bfddc66bbcd4707ab0e1884f5c48a863534f8ebd10fa6496746c.
  • Correctness: using current-main-compatible source plus the focused [BugFix][Model] Fix the GLM-5.3-Flash NoPE and KDA convolution paths on Ascend #15885 causal-conv subset at head 8743a2b5e4be853d12d7e141e32f6101394959f5, the harness imported the actual patched wrapper and actual _pack_conv_weight. Rank0 and rank7 over A/B/C/D deterministic cases were output/state bitwise exact, max abs 0.0, with no NaN/Inf.
  • Performance: actual-wrapper timing on server15 910B2, physical /dev/davinci6, included the wrapper guard, transposes, output allocation, native dispatch, and fallback reference path. With 5 warmups and 20 steady observations per arm, rank0 B/D speedups were 4.244x and 10.389x; rank7 B/D speedups were 4.276x and 10.753x. A/C were also strong.

Coordination note: PR #16251 also touches the ordinary non-spec GLM KDA causal-conv path and calls the same custom op through a different metadata/staging integration surface. The overlap appears partial rather than cleanly duplicate: #15885 carries fallback/native availability behavior, load-time packing, and explicit pack-order/layout validation, while #16251 routes through GDN metadata and non-contiguous state staging/writeback. I am not recommending which whole PR to merge here; this note is limited to the causal-conv ordinary-prefill evidence for #15885.

Remaining non-scientific/review items observed outside the numerical result include CRLF line endings in the causal-conv-relevant files at the tested PR head, plus normal PR-wide review/merge-state handling.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants