Skip to content

[Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash - #16251

Merged
ZT-AIA merged 1 commit into
vllm-project:mainfrom
lijiahang226:glm53-v5-kda-conv
Sep 14, 2026
Merged

ZT-AIA merged 1 commit into
vllm-project:mainfrom
lijiahang226:glm53-v5-kda-conv

Conversation

@lijiahang226

@lijiahang226 lijiahang226 commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Refs #15665

Integrate the existing AscendC recurrent/chunk KDA and causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share the native KDA call helpers in vllm_ascend/ops/kda.py, while retaining their model-specific projections, weight loading, beta preparation, cache updates, and metadata handling.

Preserve physical-page state views and cache-spec capability detection without changing the cache allocation, grouping, or shared MLA/Mamba tensor ownership from #15913. Mask inactive convolution slots before pointer arithmetic and preserve writeback to noncontiguous state views.

The native operators and GDN interfaces are already in main. Full-model validation combines the GLM Flash integration PRs and their prerequisites; prerequisite commits are excluded from this branch.

Does this PR introduce any user-facing change?

Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported operator configurations are rejected during initialization. Kimi K3 retains its existing beta and gate semantics through the shared helpers.

How was this patch tested?

A3 integration validation included PR revision 2fc40a144:

  • 320 unit tests passed, 1 skipped, covering GLM/Kimi KDA, cache pages, KeyPool routing, GDN metadata, MTP, ModelSlim, and shared SFA. Updated coverage includes tests/ut/ops/test_kimi_kda.py and tests/ut/models/test_glm5next_kda_contracts.py.
  • Four native NPU tests passed, including tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.py for inactive slots and strided state layouts, plus existing Kimi recurrent/prefill tests.
  • Twelve comparisons against the prior implementations produced exactly equal outputs and backing-state tensors for GLM/Kimi gate variants, decode, MTP verification, and 65/131-token prefill.
  • GLM-5.3-Flash-w8a8 passed 8/8 inference cases, including a 5264-token prompt, with TP8/EP8, FULL_DECODE_ONLY, and MTP=3.

The current revision rebases on #16067 and preserves its direct Q/K/V view handling and removal of redundant Kimi output masks. The shared recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases assert that contiguous and separately strided views reach the native operator without Python-side materialization.

Eight focused GLM/Kimi contract cases passed in an isolated CPU harness executing the actual function bodies and test cases. Ruff, formatting, git diff --check, and GitHub pre-commit passed. The hardware results above apply to the earlier integration; this rebased revision has not had a new hardware run.

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests labels Sep 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates AscendC-based KDA and causal convolution operators into the GLM-5.3-Flash model. It enhances cache management by ensuring shared physical-page strides are correctly handled for Mamba cache views, preventing potential data corruption between scheduler blocks. Additionally, it introduces robust testing coverage for operator dispatch and cache isolation to ensure stability.

Highlights

  • AscendC Integration: Integrated AscendC-based KDA and causal convolution operators for the GLM-5.3-Flash model to improve performance and operator support.
  • Cache Management: Implemented shared physical-page stride logic for GLM Mamba cache views to ensure block isolation and prevent scheduler block ID overwrites.
  • Testing Infrastructure: Added comprehensive unit and regression tests covering operator contracts, dispatch logic, and cache storage preservation.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Support GLM5-Next short convolution and KDA operators on Ascend

Suggested PR Summary:

### What this PR does / why we need it?
This PR integrates the AscendC custom operators (`npu_causal_conv1d_custom`, `recurrent_kda`, and `chunk_kda_fwd`) into the GLM5-Next model implementation for Ascend. It replaces previous Triton/fallback implementations with optimized AscendC operators for prefill, decode, and MTP verification. It also introduces a Triton-based state copy kernel for non-contiguous views and updates the model runner to support shared MLA and Mamba physical pages.

Feedback has been provided to use `tl.where` to clamp the slot index to a safe value (e.g., `0`) when inactive in the Triton state copy kernel (`_copy_conv_state`) to prevent undefined behavior or compilation errors from negative pointer arithmetic on Ascend hardware.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with new end-to-end and unit tests:
- `test_glm5next_causal_conv1d.py`
- `test_glm5next_kda_contracts.py`
- `test_glm5next_kda_dispatch.py`
- `test_glm5next_pooled_cache.py`

Comment thread vllm_ascend/models/glm5next/ops/causal_conv1d.py Outdated
@github-actions github-actions Bot removed the documentation Improvements or additions to documentation label Sep 10, 2026
@lijiahang226
lijiahang226 force-pushed the glm53-v5-kda-conv branch 2 times, most recently from 37920cd to 4b11596 Compare September 10, 2026 12:05
@github-actions github-actions Bot added documentation Improvements or additions to documentation module:ops labels Sep 10, 2026
@github-actions github-actions Bot removed documentation Improvements or additions to documentation module:ops labels Sep 10, 2026
@lijiahang226 lijiahang226 added the ready-precise run selected e2e test for pr label Sep 10, 2026 — with ChatGPT Codex Connector
@lijiahang226
lijiahang226 marked this pull request as ready for review September 11, 2026 00:38
@Ruiqiu-Zheng

Ruiqiu-Zheng commented Sep 11, 2026 •

Copy link
Copy Markdown

Current-head applicability update (2026-09-11): #16251 is now at 2a889f62374619dc8aeda08c0aad4e627d55e600. The diff from the validated c40b9e5d918b61fc6a40ff631f7d94b1c6f631bb to 2a889f6... changes only tests/ut/worker/test_glm5next_pooled_cache.py, vllm_ascend/core/kv_cache_interface.py, and vllm_ascend/worker/model_runner_v1.py; it does not change vllm_ascend/models/glm5next/kda.py or vllm_ascend/models/glm5next/ops/causal_conv1d.py. Therefore the causal-conv/KDA wrapper evidence below remains source-applicable to the current head, while it still does not cover the newly changed pooled-cache/model-runner behavior.

I ran a bounded standalone validation of the exact #16251 GLM-5.3 KDA causal-conv ordinary non-spec prefill path at head c40b9e5d918b61fc6a40ff631f7d94b1c6f631bb. Scope was the #16251 wrapper, GDN metadata, and non-contiguous conv-state staging/writeback on server15 physical NPU7, rank0, using exact R13 fixtures and official GLM-5.3-Flash layer0 BF16 q/k/v conv weights. This does not satisfy complete-model KDA/MTP validation and is not an E2E, production, all-rank/all-layer, decode/speculative, or whole-PR readiness claim.

Correctness:

  • A/B/C/D fixtures all matched the R13 fixture hashes.
  • Output, touched state, and untouched state were bitwise exact for all four cases.
  • Max abs error was 0.0 for all tested outputs/states.
  • _copy_conv_state staging and writeback were observed.
  • Pre/post correctness checks for timed B and D were true.

Performance protocol: 5 warmups, 20 observations per arm, balanced AB/BA order.

  • B seq_8192: reference median 1.2604128569364548 ms, candidate median 0.5797962658107281 ms, speedup 2.173889228441342x.
  • D 256/768/2048/5120 total8192 mixed: reference median 3.1065656803548336 ms, candidate median 0.6571072153747082 ms, speedup 4.727638972254092x.

On the negative-slot review comment: the reviewed #16251 wrapper computes cache_offsets before applying active, and clamping invalid or negative slots is a reasonable broader hardening request, especially for spec/padding paths. In the ordinary non-spec prefill source reachability inspected here, current upstream semantics use PAD_SLOT_ID=-1 and NULL_BLOCK_ID=0; #16251 gdn_attn_builder.py uses PAD_SLOT_ID only for spec graph reset in the inspected paths, ordinary non-spec decode graph padding is filled with NULL_BLOCK_ID, and the current vLLM GPU runner fills unused block-table request rows with NULL_BLOCK_ID=0. mamba_get_block_table_tensor preserves the block table for modes all/none and clamps gather start index for zero-length align mode; it does not convert ordinary prefill null rows to -1. Exact #16251 tests show mixed non-spec prefill can include zero-length rows, but the helper gives those rows positive block ids, not negative ones.

So the source-reachability evidence in this ordinary-prefill scope did not establish a negative slot reaching the tested path, and it does not downgrade the R16e valid-row prefill result. It also does not prove the broader wrapper is safe; the suggested clamp remains a legitimate review/hardening item.

Coordination note: PR #15885 also covers the ordinary non-spec GLM KDA causal-conv path through a narrower integration surface. I independently validated its actual wrapper as bitwise exact on the same official-weight family, with strong component speedups, while #16251 additionally owns GDN metadata and non-contiguous state staging/writeback. The source overlap is therefore partial rather than a clean duplicate; this note is intended to provide evidence for maintainer coordination, not to recommend which whole PR should win ownership.

Comment thread vllm_ascend/models/glm5next/kda.py
@Ruiqiu-Zheng

Copy link
Copy Markdown

The current ready-precise failure is tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py::test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy (3 passed, 1 failed in that job), while this PR changes GLM-5.3 KDA/causal-conv and pooled-cache integration. Retrying the failed selected-test job once to distinguish unrelated/noisy DSV3 SFA PCP accuracy from a reproducible regression.

/rerun

@Ruiqiu-Zheng

Ruiqiu-Zheng commented Sep 11, 2026 •

Copy link
Copy Markdown

/rerun
[Bot]: rerun command failed: you do not have permission. Only the PR author or users with triage+ permission can trigger /rerun.

lijiahang226 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@lijiahang226

lijiahang226 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@lijiahang226

lijiahang226 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed. No failed jobs found.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

…olution state

Integrate AscendC KDA and convolution for GLM using the same kernel helpers as Kimi K3. Preserve each model's beta processing, state updates and projections. Mask inactive convolution slots before address arithmetic. Retain cache-spec capability detection and physical-page views without changing shared cache allocation or grouping.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>

Copy link
Copy Markdown

Current-head hardware validation for 6102f76e95e1b24b2daa3ba62c23f99c40cba9b5:

  • Server15 / Ascend 910B2, quay.io/ascend/vllm-ascend:v0.23.0rc1, with the previously coherent causal-conv extension/kernel/custom-op bundle re-read at point of use (image, four runtime hashes, and registered operator schema all matched).
  • Scope: ordinary non-spec prefill causal-conv boundary only. The recurrent-KDA view-stride change is downstream of this boundary and was not re-benchmarked/revalidated here.
  • dim=3072, conv width 4, state length 3, BF16 activations/state.
  • Both fallback and native arms consumed the same BF16-narrowed packed q|k|v weight values, constructed from FP32 checkpoint-style q/k/v weights using the current model-path packing (cat -> transpose -> BF16 -> contiguous). Weight SHA256 was identical in both arms: abdf3496c8b60dbe2e4f6b36036f35a1e035538f3a8ea277d73e2e2ebd4a51cd.
  • Frozen comparison tolerance: rtol=1e-2, atol=1e-2.
  • Four deterministic cases passed 4/4: contiguous/no-initial, contiguous/initial, non-contiguous physical-page/no-initial, and non-contiguous/mixed-initial. Output max-abs deltas were 0.03125 / 0.0625 / 0.03125 / 0.03125; all four satisfy the fixed allclose criterion. Touched final state/cache rows were bit-exact in all four cases; no NaN/Inf.

So the current 6102f76 revision now has direct NPU evidence for the matched BF16 ordinary-prefill causal-conv path. This does not by itself claim recurrent-KDA, full-layer/full-model, performance, or production validation. A separate earlier diagnostic that compared original FP32 weight values directly against the BF16-packed native path remains a different test and is not overwritten by this result.

Comment thread vllm_ascend/models/glm5next/ops/causal_conv1d.py

@ZT-AIA ZT-AIA left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There are too many duplicate file names. I will restructure them later.

@ZT-AIA
ZT-AIA merged commit 5bba97c into vllm-project:main Sep 14, 2026
28 checks passed
windshado added a commit to windshado/vllm-ascend that referenced this pull request Sep 14, 2026
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba

* 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits)
  Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py
  [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344)
  [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300)
  [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253)
  [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415)
  [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251)
  [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343)
  [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422)
  [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232)
  [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319)
  [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189)
  [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043)
  [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)
  [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243)
  [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252)
  [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321)
  [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918)
  [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067)
  [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214)
  [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259)
  ...
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…-5.3-Flash (vllm-project#16251)

### What this PR does / why we need it?

Refs vllm-project#15665

Integrate the existing AscendC recurrent/chunk KDA and
causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share
the native KDA call helpers in `vllm_ascend/ops/kda.py`, while retaining
their model-specific projections, weight loading, beta preparation,
cache updates, and metadata handling.

Preserve physical-page state views and cache-spec capability detection
without changing the cache allocation, grouping, or shared MLA/Mamba
tensor ownership from vllm-project#15913. Mask inactive convolution slots before
pointer arithmetic and preserve writeback to noncontiguous state views.

The native operators and GDN interfaces are already in main. Full-model
validation combines the GLM Flash integration PRs and their
prerequisites; prerequisite commits are excluded from this branch.

### Does this PR introduce _any_ user-facing change?

Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported
operator configurations are rejected during initialization. Kimi K3
retains its existing beta and gate semantics through the shared helpers.

### How was this patch tested?

A3 integration validation included PR revision `2fc40a144`:

- 320 unit tests passed, 1 skipped, covering GLM/Kimi KDA, cache pages,
KeyPool routing, GDN metadata, MTP, ModelSlim, and shared SFA. Updated
coverage includes `tests/ut/ops/test_kimi_kda.py` and
`tests/ut/models/test_glm5next_kda_contracts.py`.
- Four native NPU tests passed, including
`tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.py`
for inactive slots and strided state layouts, plus existing Kimi
recurrent/prefill tests.
- Twelve comparisons against the prior implementations produced exactly
equal outputs and backing-state tensors for GLM/Kimi gate variants,
decode, MTP verification, and 65/131-token prefill.
- GLM-5.3-Flash-w8a8 passed 8/8 inference cases, including a 5264-token
prompt, with TP8/EP8, FULL_DECODE_ONLY, and MTP=3.

The current revision rebases on vllm-project#16067 and preserves its direct Q/K/V
view handling and removal of redundant Kimi output masks. The shared
recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases
assert that contiguous and separately strided views reach the native
operator without Python-side materialization.

Eight focused GLM/Kimi contract cases passed in an isolated CPU harness
executing the actual function bodies and test cases. Ruff, formatting,
`git diff --check`, and GitHub pre-commit passed. The hardware results
above apply to the earlier integration; this rebased revision has not
had a new hardware run.

- vLLM main:
vllm-project/vllm@a97dacb

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…-5.3-Flash (vllm-project#16251)

### What this PR does / why we need it?

Refs vllm-project#15665

Integrate the existing AscendC recurrent/chunk KDA and
causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share
the native KDA call helpers in `vllm_ascend/ops/kda.py`, while retaining
their model-specific projections, weight loading, beta preparation,
cache updates, and metadata handling.

Preserve physical-page state views and cache-spec capability detection
without changing the cache allocation, grouping, or shared MLA/Mamba
tensor ownership from vllm-project#15913. Mask inactive convolution slots before
pointer arithmetic and preserve writeback to noncontiguous state views.

The native operators and GDN interfaces are already in main. Full-model
validation combines the GLM Flash integration PRs and their
prerequisites; prerequisite commits are excluded from this branch.

### Does this PR introduce _any_ user-facing change?

Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported
operator configurations are rejected during initialization. Kimi K3
retains its existing beta and gate semantics through the shared helpers.

### How was this patch tested?

A3 integration validation included PR revision `2fc40a144`:

- 320 unit tests passed, 1 skipped, covering GLM/Kimi KDA, cache pages,
KeyPool routing, GDN metadata, MTP, ModelSlim, and shared SFA. Updated
coverage includes `tests/ut/ops/test_kimi_kda.py` and
`tests/ut/models/test_glm5next_kda_contracts.py`.
- Four native NPU tests passed, including
`tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.py`
for inactive slots and strided state layouts, plus existing Kimi
recurrent/prefill tests.
- Twelve comparisons against the prior implementations produced exactly
equal outputs and backing-state tensors for GLM/Kimi gate variants,
decode, MTP verification, and 65/131-token prefill.
- GLM-5.3-Flash-w8a8 passed 8/8 inference cases, including a 5264-token
prompt, with TP8/EP8, FULL_DECODE_ONLY, and MTP=3.

The current revision rebases on vllm-project#16067 and preserves its direct Q/K/V
view handling and removal of redundant Kimi output masks. The shared
recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases
assert that contiguous and separately strided views reach the native
operator without Python-side materialization.

Eight focused GLM/Kimi contract cases passed in an isolated CPU harness
executing the actual function bodies and test cases. Ruff, formatting,
`git diff --check`, and GitHub pre-commit passed. The hardware results
above apply to the earlier integration; this rebased revision has not
had a new hardware run.

- vLLM main:
vllm-project/vllm@a97dacb

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: like-0517 <ithwlike@126.com>
tracellex pushed a commit to tracellex/vllm-ascend that referenced this pull request Sep 18, 2026
- xr-conv-native: decode 因果卷积原生化(main vllm-project#16251 第 1 步)
- pr15-mooncake-connector-glm5.3-flash-a3: mooncake connector 适配
- patch02_topk_fix: lightning indexer 规避 CANN 9.1 TopKV2 宽归约故障(宽度<=4096)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants