Skip to content

[Feature][Operator] Add DeepSeek V4.1 sparse attention operators - #16422

Merged
weijinqian0 merged 7 commits into
vllm-project:mainfrom
GDzhu01:codex/deepseek-v4.1-operators
Sep 14, 2026
Merged

weijinqian0 merged 7 commits into
vllm-project:mainfrom
GDzhu01:codex/deepseek-v4.1-operators

Conversation

@GDzhu01

@GDzhu01 GDzhu01 commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Adds the native and Triton operator layer required by DeepSeek V4.1 on Ascend:

  • Quant Lightning Indexer V2 and metadata operators
  • Sparse Flash MLA and metadata operators
  • compressor, index preparation, query quantization, Engram INT8, and required MoE operator updates
  • operator documentation, native examples, static checks, and operator-level unit/e2e tests

This is the operator-only first part of the DeepSeek V4.1 enablement. Framework/model integration is intentionally split into dependent PR #16423.

Does this PR introduce any user-facing change?

Yes. It exposes the native operator capabilities required by DeepSeek V4.1 serving, but does not register or enable the model integration by itself.

How was this patch tested?

  • operator static checks: 8 passed
  • focused related unit tests: 342 passed, 1 skipped
  • broad related unit tests: 855 passed, 12 skipped
  • clean native rebuild on Ascend A3: passed
  • final single-card NPU operator regression: 359 passed
  • pre-commit scoped checks: Ruff, Ruff format, codespell, typos, clang-format, markdownlint, and forbidden-import check passed

End-to-end model validation is recorded in dependent PR #16423.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request adds the necessary operator infrastructure to support the DeepSeek V4.1 model on Ascend NPU hardware. It focuses on implementing high-performance native and Triton kernels for sparse attention and quantization operations. This is the first phase of the model enablement, providing the required operator layer while deferring full framework integration to a subsequent pull request.

Highlights

  • DeepSeek V4.1 Operator Support: Introduced native and Triton operator layers for DeepSeek V4.1 on Ascend, including Quant Lightning Indexer V2, Sparse Flash MLA, and various MoE-related operators.
  • Operator Testing and Validation: Added comprehensive operator-level unit and end-to-end tests, along with documentation and static checks to ensure operator stability.
  • Performance Benchmarking: Included new benchmark scripts for indexer index preparation and query quantization to optimize operator performance on NPU.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch from 4a8723a to 1ca0e0c Compare September 12, 2026 12:30
@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests module:ops labels Sep 12, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new two-level TopK candidate selection mechanism for the QuantLightningIndexerV2 operator, specifically optimized for the Ascend910B (arch22) architecture. The changes include adding support for TND layout, non-contiguous key strides, and an output index offset feature. The PR also includes comprehensive test assets and infrastructure updates to support batch consistency testing and performance profiling. My review identified several critical issues, including incorrect API usage for DataCopyPad, a compilation error due to a variable naming mismatch in the invocation macro, missing type-fixing wrappers for special quantization modes, and a potential crash in the InferShape implementation when handling optional outputs.

Comment on lines +69 to +76
if (tailSize > 0) {
AscendC::DataCopyExtParams dataCopyParams;
dataCopyParams.blockCount = 1;
dataCopyParams.blockLen = tailSize * sizeof(T);
dataCopyParams.srcStride = 0;
dataCopyParams.dstStride = 0;
AscendC::DataCopyPad(outGm[gmOffset], popBuffer, dataCopyParams);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

In Ascend C, DataCopyPad expects DataCopyPadParams (or DataCopyPadExtParams), whereas DataCopy expects DataCopyExtParams. Since dataCopyParams is declared as DataCopyExtParams, calling DataCopyPad with it will result in a compilation error or undefined behavior due to struct layout mismatch. Please use AscendC::DataCopy instead of AscendC::DataCopyPad when copying the tail.

            if (tailSize > 0) {
                AscendC::DataCopyExtParams dataCopyParams;
                dataCopyParams.blockCount = 1;
                dataCopyParams.blockLen = tailSize * sizeof(T);
                dataCopyParams.srcStride = 0;
                dataCopyParams.dstStride = 0;
                AscendC::DataCopy(outGm[gmOffset], popBuffer, dataCopyParams);
            }

Comment on lines +35 to +42
// arch22: 含 candidate_topk_index 输入/输出 (两级TopK)
#define INVOKE_LI_CANDIDATE_OP_IMPL(templateClass, ...) \
do { \
templateClass<QLIV2Type<__VA_ARGS__>> op; \
op.Init(query, key, weights, queryScale, keyScale, cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, \
blockTable, outputIdxOffset, metadata, candidateTopkIndex, sparseIndices, sparseValues, \
candidateTopkIndexOut, user, tiling_data, &tPipe); \
op.Process(); \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The parameter in the quant_lightning_indexer_v2 function signature is named workspace, but the macro INVOKE_LI_CANDIDATE_OP_IMPL passes user to op.Init. This will cause a compilation error because user is not declared in this scope. Please change user to workspace.

#define INVOKE_LI_CANDIDATE_OP_IMPL(templateClass, ...) \
    do { \
        templateClass<QLIV2Type<__VA_ARGS__>> op; \
        op.Init(query, key, weights, queryScale, keyScale, cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, \
                blockTable, outputIdxOffset, metadata, candidateTopkIndex, sparseIndices, sparseValues, \
                candidateTopkIndexOut, workspace, tiling_data, &tPipe); \
        op.Process(); \
    } while (0)

Comment on lines +152 to +156
int64_t keyStride0 = key.stride(0);
int64_t keyScaleStride0 = keyDequantScale.stride(0);

EXEC_NPU_CMD(aclnnQuantLightningIndexerV2, query, key, weights, queryDequantScale, keyDequantScale,
cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, blockTable, outputIdxOffset,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

In quant_lightning_indexer_v2_torch_adpt.h, query, key, queryDequantScale, and keyDequantScale are passed directly to EXEC_NPU_CMD. However, for special quantization modes (like HIFLOAT8, FLOAT4_E2M1, and MXFP8), their ACL data types must be fixed using FixQLIV2AclDtypes and MakeWrapper (as done in quant_lightning_indexer.cpp), otherwise EXEC_NPU_CMD will use default PyTorch scalar type mappings (e.g., ACL_UINT8 instead of ACL_HIFLOAT8), causing runtime errors or incorrect execution. Please apply the type fixing wrappers before calling EXEC_NPU_CMD.

    auto queryWrapper = at_npu::native::TensorWrapper{query, at_npu::native::ConvertToAclDataType(query.scalar_type())};
    auto keyWrapper = at_npu::native::TensorWrapper{key, at_npu::native::ConvertToAclDataType(key.scalar_type())};
    auto queryScaleWrapper = at_npu::native::TensorWrapper{queryDequantScale, at_npu::native::ConvertToAclDataType(queryDequantScale.scalar_type())};
    auto keyScaleWrapper = at_npu::native::TensorWrapper{keyDequantScale, at_npu::native::ConvertToAclDataType(keyDequantScale.scalar_type())};

    if (quantMode == 4) {
        queryWrapper.dtype = ACL_HIFLOAT8;
        keyWrapper.dtype = ACL_HIFLOAT8;
    } else if (quantMode == 5) {
        queryWrapper.dtype = ACL_FLOAT4_E2M1;
        keyWrapper.dtype = ACL_FLOAT4_E2M1;
    } else if (quantMode == 3) {
        queryScaleWrapper.dtype = ACL_FLOAT8_E8M0;
        keyScaleWrapper.dtype = ACL_FLOAT8_E8M0;
    }

    EXEC_NPU_CMD(aclnnQuantLightningIndexerV2, queryWrapper, keyWrapper, weights, queryScaleWrapper, keyScaleWrapper,
              cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, blockTable, outputIdxOffset,
              metadata, candidateTopkIndexIn, topk, quantMode, maxSeqlenQ, queryLayoutPtr, keyLayoutPtr, maskMode,
              cmpRatio, returnValue, candidateMode, candidateTopkBlocks, candidateBlockSize, keyStride0,
              keyScaleStride0, sparseIndicesOut, sparseValuesOut, candidateTopkIndexOut);

Comment on lines +93 to +109
OP_CHECK_NULL_WITH_CONTEXT(context, candidateTopkIndexShape);
const int32_t *candidate_mode = attrs->GetAttrPointer<int32_t>(ATTR_CANDIDATE_MODE_INDEX);
uint32_t candidateMode = (candidate_mode != nullptr) ? static_cast<uint32_t>(*candidate_mode) : 3U;
if (candidateMode == CANDIDATE_MODE_SOURCE) {
OP_CHECK_IF(inputLayoutQueryPtrStr != "BSND",
OP_LOGE("QuantLightningIndexerV2",
"candidate_mode=1 only supports layout_q=BSND, but got %s.",
inputLayoutQueryPtrStr.c_str()),
return GRAPH_FAILED);
const int64_t *candidate_topk_blocks = attrs->GetAttrPointer<int64_t>(ATTR_CANDIDATE_TOPK_BLOCKS_INDEX);
int64_t candBlocks = (candidate_topk_blocks != nullptr) ? *candidate_topk_blocks : CANDIDATE_TOPK_BLOCKS_FIX;
*candidateTopkIndexShape = *sparseIndicesShape;
candidateTopkIndexShape->SetDim(sparseIndicesShape->GetDimNum() - 1, candBlocks);
} else {
candidateTopkIndexShape->SetDimNum(1);
candidateTopkIndexShape->SetDim(0, 0);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Since candidate_topk_index_out is an optional output, context->GetOutputShape(CANDIDATE_TOPK_INDEX_OUTPUT_INDEX) can return nullptr if the output is not bound in the graph. Using OP_CHECK_NULL_WITH_CONTEXT on it will cause the entire InferShape process to fail with GRAPH_FAILED when the optional output is not bound, which defeats the purpose of it being optional. Please wrap the shape setting in a null check instead of failing.

    if (candidateTopkIndexShape != nullptr) {
        const int32_t *candidate_mode = attrs->GetAttrPointer<int32_t>(ATTR_CANDIDATE_MODE_INDEX);
        uint32_t candidateMode = (candidate_mode != nullptr) ? static_cast<uint32_t>(*candidate_mode) : 3U;
        if (candidateMode == CANDIDATE_MODE_SOURCE) {
            OP_CHECK_IF(inputLayoutQueryPtrStr != "BSND",
                        OP_LOGE("QuantLightningIndexerV2",
                                "candidate_mode=1 only supports layout_q=BSND, but got %s.",
                                inputLayoutQueryPtrStr.c_str()),
                        return GRAPH_FAILED);
            const int64_t *candidate_topk_blocks = attrs->GetAttrPointer<int64_t>(ATTR_CANDIDATE_TOPK_BLOCKS_INDEX);
            int64_t candBlocks = (candidate_topk_blocks != nullptr) ? *candidate_topk_blocks : CANDIDATE_TOPK_BLOCKS_FIX;
            *candidateTopkIndexShape = *sparseIndicesShape;
            candidateTopkIndexShape->SetDim(sparseIndicesShape->GetDimNum() - 1, candBlocks);
        } else {
            candidateTopkIndexShape->SetDimNum(1);
            candidateTopkIndexShape->SetDim(0, 0);
        }
    }

@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch from 1ca0e0c to 0963d88 Compare September 12, 2026 12:55
@GDzhu01
GDzhu01 marked this pull request as ready for review September 12, 2026 13:01
@GDzhu01 GDzhu01 added the ready-precise run selected e2e test for pr label Sep 12, 2026
@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch 2 times, most recently from 85e70ee to 90971c5 Compare September 12, 2026 13:06
@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch 2 times, most recently from 563b456 to 17a523e Compare September 12, 2026 13:10
Add Quant Lightning Indexer v2 enhancements, Sparse Flash MLA and metadata operators, Triton compressor helpers, and focused operator tests required by DeepSeek V4.1.

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch from 17a523e to caa11e6 Compare September 12, 2026 13:15
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@GDzhu01
GDzhu01 force-pushed the codex/deepseek-v4.1-operators branch from aa48a0c to fd87bf4 Compare September 12, 2026 13:40
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@weijinqian0
weijinqian0 merged commit bd69bad into vllm-project:main Sep 14, 2026
33 checks passed
windshado added a commit to windshado/vllm-ascend that referenced this pull request Sep 14, 2026
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba

* 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits)
  Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py
  [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344)
  [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300)
  [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253)
  [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415)
  [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251)
  [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343)
  [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422)
  [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232)
  [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319)
  [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189)
  [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043)
  [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)
  [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243)
  [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252)
  [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321)
  [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918)
  [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067)
  [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214)
  [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259)
  ...
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…m-project#16422)

### What this PR does / why we need it?

Adds the native and Triton operator layer required by DeepSeek V4.1 on
Ascend:

- Quant Lightning Indexer V2 and metadata operators
- Sparse Flash MLA and metadata operators
- compressor, index preparation, query quantization, Engram INT8, and
required MoE operator updates
- operator documentation, native examples, static checks, and
operator-level unit/e2e tests

This is the operator-only first part of the DeepSeek V4.1 enablement.
Framework/model integration is intentionally split into dependent PR
vllm-project#16423.

### Does this PR introduce _any_ user-facing change?

Yes. It exposes the native operator capabilities required by DeepSeek
V4.1 serving, but does not register or enable the model integration by
itself.

### How was this patch tested?

- operator static checks: 8 passed
- focused related unit tests: 342 passed, 1 skipped
- broad related unit tests: 855 passed, 12 skipped
- clean native rebuild on Ascend A3: passed
- final single-card NPU operator regression: 359 passed
- pre-commit scoped checks: Ruff, Ruff format, codespell, typos,
clang-format, markdownlint, and forbidden-import check passed

End-to-end model validation is recorded in dependent PR vllm-project#16423.

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…m-project#16422)

### What this PR does / why we need it?

Adds the native and Triton operator layer required by DeepSeek V4.1 on
Ascend:

- Quant Lightning Indexer V2 and metadata operators
- Sparse Flash MLA and metadata operators
- compressor, index preparation, query quantization, Engram INT8, and
required MoE operator updates
- operator documentation, native examples, static checks, and
operator-level unit/e2e tests

This is the operator-only first part of the DeepSeek V4.1 enablement.
Framework/model integration is intentionally split into dependent PR
vllm-project#16423.

### Does this PR introduce _any_ user-facing change?

Yes. It exposes the native operator capabilities required by DeepSeek
V4.1 serving, but does not register or enable the model integration by
itself.

### How was this patch tested?

- operator static checks: 8 passed
- focused related unit tests: 342 passed, 1 skipped
- broad related unit tests: 855 passed, 12 skipped
- clean native rebuild on Ascend A3: passed
- final single-card NPU operator regression: 359 passed
- pre-commit scoped checks: Ruff, Ruff format, codespell, typos,
clang-format, markdownlint, and forbidden-import check passed

End-to-end model validation is recorded in dependent PR vllm-project#16423.

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: like-0517 <ithwlike@126.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 15, 2026
vllm-project#16422 reintroduced SDK Shape ToString in QLIV2 and SparseFlashMLA
tiling error logs. CANN 9.1.0 does not export that symbol, so A5 OPC
fails to load liboptiling.so. Restore the vllm-project#16016 local formatter;
tiling checks and kernel math are unchanged.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
ZT-AIA pushed a commit that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

Fix two A5 QLIV2 build regressions introduced by #16422:

- Restore the local shape formatter from #16016. The [A5 Ubuntu amd64
image
build](https://github.com/vllm-project/vllm-ascend/actions/runs/34829507531/job/103929405987)
cannot load `liboptiling.so` because the 53 restored
`Ops::Base::ToString(const gert::Shape&)` calls leave
`_ZN3Ops4Base8ToStringERKN4gert5ShapeE` unresolved. Use the
bracket-preserving `ShapeToStringForLog` wrapper around the existing
`ToStringRaw`.
- Match the A5 dependency with `"ascend950" IN_LIST
ASCEND_COMPUTE_UNIT`. The standard build passes `ascend950;`; the
string-equality check skips the sibling dependency and causes
`lightning_indexer_v2_vector1_base.h` to be missing from the staged
kernel sources. List membership also handles multi-target configurations
while keeping the dependency A5-specific.

### Does this PR introduce _any_ user-facing change?

Yes: affected A5 source builds can resolve these QLIV2 Host and
kernel-source dependencies. Operator interfaces, numerical computation,
and diagnostic formatting are unchanged.

### How was this patch tested?

Rebased on main `fbb75a47436901616849da155828f755ec58b619` (including
the CPU UT fix #16574). Current HEAD is
`5793e3135057281e08f36ab66f807d2c229843f3`; `git range-diff` confirms
both patches are unchanged. All 18 native lint hooks passed on this
HEAD. A new [CI
run](https://github.com/vllm-project/vllm-ascend/actions/runs/34960944860)
was triggered. The full build evidence below belongs to the pre-rebase
HEAD; the full build was not repeated after rebase.

In the A5-107 aarch64 development container with CANN 9.1.0:

- Compiled the complete QLIV2 Host translation unit before and after the
formatter fix into separate objects. `nm -u` confirms the exact external
Shape formatter reference in the baseline object and its absence after
the fix.
- Compared the actual local formatting helpers against the SDK
formatter: all seven shape cases match.
- Built the shared Host tiling library with the formatter fix and loaded
it in a fresh process using `RTLD_NOW`; the library has neither the
unresolved Shape formatter nor a direct `libops_base.so` dependency.
- Reproduced the missing sibling header with an actual QLIV2 kernel
compilation. After the CMake fix, standard reconfiguration records the
dependency and stages the missing header; all five generated QLIV2
configurations (0–4) compile successfully in isolated output directories
and produce their final object and JSON artifacts.
- On pre-rebase HEAD `5837fe671ec6ef180e48a350e9ecc33d71547b5f`, `bash
format.sh ci` passed all 18 hooks with a clean worktree; Gitleaks passed
for the two-commit baseline-to-HEAD range.

- On that pre-rebase HEAD, the standard `python3 setup.py build`
completed with exit 0 (`SOC_VERSION=ascend950dt_9582`, `MAX_JOBS=8`,
`OPS_CPU_NUMBER=8`, `VLLM_BATCH_INVARIANT=0`). All 212 A5 kernel targets
completed, including all five QLIV2 targets; the OPP installer and
Python C++ extension were generated successfully.
- Verified the packaged output: 212 kernel objects, the sibling header
byte-identical to its source, and `liboptiling.so` loading successfully
in a fresh `RTLD_NOW` process without `LD_PRELOAD` or
`ASCEND_CUSTOM_OPP_PATH`. The installed library has no unresolved Shape
formatter reference or direct `libops_base.so` dependency.

The final build reused completed targets from the initial formatter-only
attempt and the compiler cache; it is not an all-cache-disabled rebuild.
This is a native aarch64 container build, not a rerun of the amd64
Docker image CI. NPU numerical/performance tests were not run. No new
Python unit test is added: the native compilation, symbol,
library-loading, and package checks directly exercise these build
failures.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: Foriv <2293567056@qq.com>
weijinqian0 added a commit that referenced this pull request Sep 18, 2026
…#16544)

### What this PR does / why we need it?

Add DeepSeek V4.1 framework support on top of the native operators
introduced in #16422, porting the framework integration from #16423 and
adapting it to main.

- Register the V4.1 text and vision models and integrate model
configuration, weight loading, ModelSlim quantization, and frontend
encoding/reasoning/tool-call handling.
- Add V4.1 DSA attention and context parallelism, compressed/shared KV
cache allocation, ring-buffer state, and DSpark integration while
preserving main's model-runner and padded-page cache behavior.
- Support Engram HBM storage and optional host offload. Host offload
keeps tables on CPU and stages selected BF16 rows through bounded pinned
buffers, with a fused INT8 CPU lookup operator and fixed-address graph
inputs.
- Add regression coverage, the model tutorial, and A3 launch examples.
Fix fresh-process model import ordering and the API-node/headless DP
launch arguments.

The branch contains two commits: framework integration (`4e07c7f`) and
Engram host offload (`9038b37`). The integration baseline is main
`b8db6e985`, paired with vLLM
`84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

### Does this PR introduce _any_ user-facing change?

Yes. Users can select the V4.1 model/frontend and configure Engram
storage/offload. The tutorial and launch examples cover A3 deployment
with DSpark and full decode graphs. Host offload adds a native CPU
extension build dependency through the existing CMake build.

### How was this patch tested?

Added tests cover V4.1 cache allocation and metadata, compressor/indexer
and attention operator contracts, model configuration, frontend
encoding, DSpark, fresh-process imports, Engram storage/offload, and
integration with the worker/model runner.

Earlier integration state `2ac02e6` with vLLM `a97dacb71` passed 400
framework unit tests, 3 fresh-process import cases, and 44 targeted
regressions. An A3 single-node TP8/DP2/EP16 deployment with Engram
disabled also exercised ordinary generation, thinking mode, and graph
replay. These results predate the current main rebase and host-offload
commit and are not claimed as full validation of this PR head.

Current-head CI and complete end-to-end/GPQA accuracy results remain to
be confirmed. No passing two-node Engram or host-offload accuracy result
is claimed here.

Reproduction entry points:
- `examples/deepseek_v41/serve_a3.sh`
- `examples/deepseek_v41/serve_a3_offload.sh`
- `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: chenmenglong <chenmenglong1@huawei.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: drslark <slarksblood@qq.com>
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: Yizhou Liu <liu_yizhou@outlook.com>
Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com>
Signed-off-by: Angazenn <supperccell@163.com>
Co-authored-by: Angazenn <supperccell@163.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: jack <QwertyJack@users.noreply.github.com>
Co-authored-by: nwpu-zxr <zhouxuerong2@huawei.com>
Co-authored-by: pgzddxx <184603735+pgzddxx@users.noreply.github.com>
Co-authored-by: Qi Mao <maomaoyu870@gmail.com>
Co-authored-by: Qiu Chunshuo <qiuchunshuo@huawei.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: weijinqian0 <1184188277@qq.com>
Co-authored-by: Zhu Yi Lin <GDzhu01@users.noreply.github.com>
Co-authored-by: zhuyilin (A) <z00911894@china.huawei.com>
Co-authored-by: ZT-AIA <1028681969@qq.com>
Co-authored-by: drslark <slarksblood@qq.com>
Co-authored-by: GPT-5.6 Codex <codex@openai.com>
Co-authored-by: Yizhou Liu <liu_yizhou@outlook.com>
Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
linfeng-yuan pushed a commit that referenced this pull request Sep 21, 2026
…#16925)

### What this PR does / why we need it?

Re-submit the complete DeepSeek V4.1 framework integration from #16544
after its revert in #16905, updated for the current vllm-ascend main.
This builds on the native operators from #16422.

- Register the V4.1 text and vision models and integrate model
configuration, weight loading, ModelSlim quantization, and frontend
encoding, reasoning, and tool-call handling.
- Add V4.1 DSA attention and context parallelism, compressed/shared KV
cache allocation, ring-buffer state from #16890, and DSpark integration
while preserving the current model-runner and padded-page cache
behavior.
- Support Engram HBM storage and optional CPU host offload with bounded
pinned buffers, fused INT8 lookup, and fixed-address graph inputs.
- Restore the regression tests and model tutorial, including its A3
launch commands. Resolve integration conflicts against current main,
including its KV-cache dtype API and drafter typing contract.

### Does this PR introduce _any_ user-facing change?

Yes. Users can select the V4.1 model/frontend and configure Engram
storage/offload. The tutorial's launch commands cover A3 deployment with
DSpark and full decode graphs. Host offload uses the native CPU
extension provided by the operator prerequisite.

### How was this patch tested?

The original #16544 description records framework-unit-test,
fresh-process import, targeted-regression, and A3 single-node results
from earlier integration states. Those results predate this resubmission
and are not claimed as validation of this PR head.

For this head, all changed Python files parse and `git diff --check`
passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks
pass. Selected A2/A3/310P hardware tests are pending. No end-to-end
accuracy or two-node Engram/host-offload result is claimed.

Reproduction entry point:
`docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`.

Based on #16544. cc @dragondream-chen @GDzhu01 for review and
attribution.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: pgzddxx <1697817735@qq.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…vllm-project#16925)

### What this PR does / why we need it?

Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544
after its revert in vllm-project#16905, updated for the current vllm-ascend main.
This builds on the native operators from vllm-project#16422.

- Register the V4.1 text and vision models and integrate model
configuration, weight loading, ModelSlim quantization, and frontend
encoding, reasoning, and tool-call handling.
- Add V4.1 DSA attention and context parallelism, compressed/shared KV
cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration
while preserving the current model-runner and padded-page cache
behavior.
- Support Engram HBM storage and optional CPU host offload with bounded
pinned buffers, fused INT8 lookup, and fixed-address graph inputs.
- Restore the regression tests and model tutorial, including its A3
launch commands. Resolve integration conflicts against current main,
including its KV-cache dtype API and drafter typing contract.

### Does this PR introduce _any_ user-facing change?

Yes. Users can select the V4.1 model/frontend and configure Engram
storage/offload. The tutorial's launch commands cover A3 deployment with
DSpark and full decode graphs. Host offload uses the native CPU
extension provided by the operator prerequisite.

### How was this patch tested?

The original vllm-project#16544 description records framework-unit-test,
fresh-process import, targeted-regression, and A3 single-node results
from earlier integration states. Those results predate this resubmission
and are not claimed as validation of this PR head.

For this head, all changed Python files parse and `git diff --check`
passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks
pass. Selected A2/A3/310P hardware tests are pending. No end-to-end
accuracy or two-node Engram/host-offload result is claimed.

Reproduction entry point:
`docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`.

Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and
attribution.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: pgzddxx <1697817735@qq.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
…vllm-project#16925)

Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544
after its revert in vllm-project#16905, updated for the current vllm-ascend main.
This builds on the native operators from vllm-project#16422.

- Register the V4.1 text and vision models and integrate model
configuration, weight loading, ModelSlim quantization, and frontend
encoding, reasoning, and tool-call handling.
- Add V4.1 DSA attention and context parallelism, compressed/shared KV
cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration
while preserving the current model-runner and padded-page cache
behavior.
- Support Engram HBM storage and optional CPU host offload with bounded
pinned buffers, fused INT8 lookup, and fixed-address graph inputs.
- Restore the regression tests and model tutorial, including its A3
launch commands. Resolve integration conflicts against current main,
including its KV-cache dtype API and drafter typing contract.

Yes. Users can select the V4.1 model/frontend and configure Engram
storage/offload. The tutorial's launch commands cover A3 deployment with
DSpark and full decode graphs. Host offload uses the native CPU
extension provided by the operator prerequisite.

The original vllm-project#16544 description records framework-unit-test,
fresh-process import, targeted-regression, and A3 single-node results
from earlier integration states. Those results predate this resubmission
and are not claimed as validation of this PR head.

For this head, all changed Python files parse and `git diff --check`
passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks
pass. Selected A2/A3/310P hardware tests are pending. No end-to-end
accuracy or two-node Engram/host-offload result is claimed.

Reproduction entry point:
`docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`.

Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and
attribution.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: pgzddxx <1697817735@qq.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…vllm-project#16925)

### What this PR does / why we need it?

Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544
after its revert in vllm-project#16905, updated for the current vllm-ascend main.
This builds on the native operators from vllm-project#16422.

- Register the V4.1 text and vision models and integrate model
configuration, weight loading, ModelSlim quantization, and frontend
encoding, reasoning, and tool-call handling.
- Add V4.1 DSA attention and context parallelism, compressed/shared KV
cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration
while preserving the current model-runner and padded-page cache
behavior.
- Support Engram HBM storage and optional CPU host offload with bounded
pinned buffers, fused INT8 lookup, and fixed-address graph inputs.
- Restore the regression tests and model tutorial, including its A3
launch commands. Resolve integration conflicts against current main,
including its KV-cache dtype API and drafter typing contract.

### Does this PR introduce _any_ user-facing change?

Yes. Users can select the V4.1 model/frontend and configure Engram
storage/offload. The tutorial's launch commands cover A3 deployment with
DSpark and full decode graphs. Host offload uses the native CPU
extension provided by the operator prerequisite.

### How was this patch tested?

The original vllm-project#16544 description records framework-unit-test,
fresh-process import, targeted-regression, and A3 single-node results
from earlier integration states. Those results predate this resubmission
and are not claimed as validation of this PR head.

For this head, all changed Python files parse and `git diff --check`
passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks
pass. Selected A2/A3/310P hardware tests are pending. No end-to-end
accuracy or two-node Engram/host-offload result is claimed.

Reproduction entry point:
`docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`.

Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and
attribution.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: pgzddxx <1697817735@qq.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:ops module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants