[Feature][Operator] Add DeepSeek V4.1 sparse attention operators - #16422
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request adds the necessary operator infrastructure to support the DeepSeek V4.1 model on Ascend NPU hardware. It focuses on implementing high-performance native and Triton kernels for sparse attention and quantization operations. This is the first phase of the model enablement, providing the required operator layer while deferring full framework integration to a subsequent pull request. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
4a8723a to
1ca0e0c
Compare
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
This pull request introduces a new two-level TopK candidate selection mechanism for the QuantLightningIndexerV2 operator, specifically optimized for the Ascend910B (arch22) architecture. The changes include adding support for TND layout, non-contiguous key strides, and an output index offset feature. The PR also includes comprehensive test assets and infrastructure updates to support batch consistency testing and performance profiling. My review identified several critical issues, including incorrect API usage for DataCopyPad, a compilation error due to a variable naming mismatch in the invocation macro, missing type-fixing wrappers for special quantization modes, and a potential crash in the InferShape implementation when handling optional outputs.
| if (tailSize > 0) { | ||
| AscendC::DataCopyExtParams dataCopyParams; | ||
| dataCopyParams.blockCount = 1; | ||
| dataCopyParams.blockLen = tailSize * sizeof(T); | ||
| dataCopyParams.srcStride = 0; | ||
| dataCopyParams.dstStride = 0; | ||
| AscendC::DataCopyPad(outGm[gmOffset], popBuffer, dataCopyParams); | ||
| } |
There was a problem hiding this comment.
In Ascend C, DataCopyPad expects DataCopyPadParams (or DataCopyPadExtParams), whereas DataCopy expects DataCopyExtParams. Since dataCopyParams is declared as DataCopyExtParams, calling DataCopyPad with it will result in a compilation error or undefined behavior due to struct layout mismatch. Please use AscendC::DataCopy instead of AscendC::DataCopyPad when copying the tail.
if (tailSize > 0) {
AscendC::DataCopyExtParams dataCopyParams;
dataCopyParams.blockCount = 1;
dataCopyParams.blockLen = tailSize * sizeof(T);
dataCopyParams.srcStride = 0;
dataCopyParams.dstStride = 0;
AscendC::DataCopy(outGm[gmOffset], popBuffer, dataCopyParams);
}| // arch22: 含 candidate_topk_index 输入/输出 (两级TopK) | ||
| #define INVOKE_LI_CANDIDATE_OP_IMPL(templateClass, ...) \ | ||
| do { \ | ||
| templateClass<QLIV2Type<__VA_ARGS__>> op; \ | ||
| op.Init(query, key, weights, queryScale, keyScale, cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, \ | ||
| blockTable, outputIdxOffset, metadata, candidateTopkIndex, sparseIndices, sparseValues, \ | ||
| candidateTopkIndexOut, user, tiling_data, &tPipe); \ | ||
| op.Process(); \ |
There was a problem hiding this comment.
The parameter in the quant_lightning_indexer_v2 function signature is named workspace, but the macro INVOKE_LI_CANDIDATE_OP_IMPL passes user to op.Init. This will cause a compilation error because user is not declared in this scope. Please change user to workspace.
#define INVOKE_LI_CANDIDATE_OP_IMPL(templateClass, ...) \
do { \
templateClass<QLIV2Type<__VA_ARGS__>> op; \
op.Init(query, key, weights, queryScale, keyScale, cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, \
blockTable, outputIdxOffset, metadata, candidateTopkIndex, sparseIndices, sparseValues, \
candidateTopkIndexOut, workspace, tiling_data, &tPipe); \
op.Process(); \
} while (0)| int64_t keyStride0 = key.stride(0); | ||
| int64_t keyScaleStride0 = keyDequantScale.stride(0); | ||
|
|
||
| EXEC_NPU_CMD(aclnnQuantLightningIndexerV2, query, key, weights, queryDequantScale, keyDequantScale, | ||
| cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, blockTable, outputIdxOffset, |
There was a problem hiding this comment.
In quant_lightning_indexer_v2_torch_adpt.h, query, key, queryDequantScale, and keyDequantScale are passed directly to EXEC_NPU_CMD. However, for special quantization modes (like HIFLOAT8, FLOAT4_E2M1, and MXFP8), their ACL data types must be fixed using FixQLIV2AclDtypes and MakeWrapper (as done in quant_lightning_indexer.cpp), otherwise EXEC_NPU_CMD will use default PyTorch scalar type mappings (e.g., ACL_UINT8 instead of ACL_HIFLOAT8), causing runtime errors or incorrect execution. Please apply the type fixing wrappers before calling EXEC_NPU_CMD.
auto queryWrapper = at_npu::native::TensorWrapper{query, at_npu::native::ConvertToAclDataType(query.scalar_type())};
auto keyWrapper = at_npu::native::TensorWrapper{key, at_npu::native::ConvertToAclDataType(key.scalar_type())};
auto queryScaleWrapper = at_npu::native::TensorWrapper{queryDequantScale, at_npu::native::ConvertToAclDataType(queryDequantScale.scalar_type())};
auto keyScaleWrapper = at_npu::native::TensorWrapper{keyDequantScale, at_npu::native::ConvertToAclDataType(keyDequantScale.scalar_type())};
if (quantMode == 4) {
queryWrapper.dtype = ACL_HIFLOAT8;
keyWrapper.dtype = ACL_HIFLOAT8;
} else if (quantMode == 5) {
queryWrapper.dtype = ACL_FLOAT4_E2M1;
keyWrapper.dtype = ACL_FLOAT4_E2M1;
} else if (quantMode == 3) {
queryScaleWrapper.dtype = ACL_FLOAT8_E8M0;
keyScaleWrapper.dtype = ACL_FLOAT8_E8M0;
}
EXEC_NPU_CMD(aclnnQuantLightningIndexerV2, queryWrapper, keyWrapper, weights, queryScaleWrapper, keyScaleWrapper,
cuSeqlensQ, cuSeqlensK, sequsedQ, sequsedK, cmpResidualK, blockTable, outputIdxOffset,
metadata, candidateTopkIndexIn, topk, quantMode, maxSeqlenQ, queryLayoutPtr, keyLayoutPtr, maskMode,
cmpRatio, returnValue, candidateMode, candidateTopkBlocks, candidateBlockSize, keyStride0,
keyScaleStride0, sparseIndicesOut, sparseValuesOut, candidateTopkIndexOut);| OP_CHECK_NULL_WITH_CONTEXT(context, candidateTopkIndexShape); | ||
| const int32_t *candidate_mode = attrs->GetAttrPointer<int32_t>(ATTR_CANDIDATE_MODE_INDEX); | ||
| uint32_t candidateMode = (candidate_mode != nullptr) ? static_cast<uint32_t>(*candidate_mode) : 3U; | ||
| if (candidateMode == CANDIDATE_MODE_SOURCE) { | ||
| OP_CHECK_IF(inputLayoutQueryPtrStr != "BSND", | ||
| OP_LOGE("QuantLightningIndexerV2", | ||
| "candidate_mode=1 only supports layout_q=BSND, but got %s.", | ||
| inputLayoutQueryPtrStr.c_str()), | ||
| return GRAPH_FAILED); | ||
| const int64_t *candidate_topk_blocks = attrs->GetAttrPointer<int64_t>(ATTR_CANDIDATE_TOPK_BLOCKS_INDEX); | ||
| int64_t candBlocks = (candidate_topk_blocks != nullptr) ? *candidate_topk_blocks : CANDIDATE_TOPK_BLOCKS_FIX; | ||
| *candidateTopkIndexShape = *sparseIndicesShape; | ||
| candidateTopkIndexShape->SetDim(sparseIndicesShape->GetDimNum() - 1, candBlocks); | ||
| } else { | ||
| candidateTopkIndexShape->SetDimNum(1); | ||
| candidateTopkIndexShape->SetDim(0, 0); | ||
| } |
There was a problem hiding this comment.
Since candidate_topk_index_out is an optional output, context->GetOutputShape(CANDIDATE_TOPK_INDEX_OUTPUT_INDEX) can return nullptr if the output is not bound in the graph. Using OP_CHECK_NULL_WITH_CONTEXT on it will cause the entire InferShape process to fail with GRAPH_FAILED when the optional output is not bound, which defeats the purpose of it being optional. Please wrap the shape setting in a null check instead of failing.
if (candidateTopkIndexShape != nullptr) {
const int32_t *candidate_mode = attrs->GetAttrPointer<int32_t>(ATTR_CANDIDATE_MODE_INDEX);
uint32_t candidateMode = (candidate_mode != nullptr) ? static_cast<uint32_t>(*candidate_mode) : 3U;
if (candidateMode == CANDIDATE_MODE_SOURCE) {
OP_CHECK_IF(inputLayoutQueryPtrStr != "BSND",
OP_LOGE("QuantLightningIndexerV2",
"candidate_mode=1 only supports layout_q=BSND, but got %s.",
inputLayoutQueryPtrStr.c_str()),
return GRAPH_FAILED);
const int64_t *candidate_topk_blocks = attrs->GetAttrPointer<int64_t>(ATTR_CANDIDATE_TOPK_BLOCKS_INDEX);
int64_t candBlocks = (candidate_topk_blocks != nullptr) ? *candidate_topk_blocks : CANDIDATE_TOPK_BLOCKS_FIX;
*candidateTopkIndexShape = *sparseIndicesShape;
candidateTopkIndexShape->SetDim(sparseIndicesShape->GetDimNum() - 1, candBlocks);
} else {
candidateTopkIndexShape->SetDimNum(1);
candidateTopkIndexShape->SetDim(0, 0);
}
}1ca0e0c to
0963d88
Compare
85e70ee to
90971c5
Compare
563b456 to
17a523e
Compare
Add Quant Lightning Indexer v2 enhancements, Sparse Flash MLA and metadata operators, Triton compressor helpers, and focused operator tests required by DeepSeek V4.1. Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
17a523e to
caa11e6
Compare
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
aa48a0c to
fd87bf4
Compare
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba * 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits) Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344) [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300) [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253) [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415) [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251) [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343) [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422) [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232) [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319) [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189) [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043) [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243) [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252) [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321) [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918) [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067) [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214) [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259) ...
…m-project#16422) ### What this PR does / why we need it? Adds the native and Triton operator layer required by DeepSeek V4.1 on Ascend: - Quant Lightning Indexer V2 and metadata operators - Sparse Flash MLA and metadata operators - compressor, index preparation, query quantization, Engram INT8, and required MoE operator updates - operator documentation, native examples, static checks, and operator-level unit/e2e tests This is the operator-only first part of the DeepSeek V4.1 enablement. Framework/model integration is intentionally split into dependent PR vllm-project#16423. ### Does this PR introduce _any_ user-facing change? Yes. It exposes the native operator capabilities required by DeepSeek V4.1 serving, but does not register or enable the model integration by itself. ### How was this patch tested? - operator static checks: 8 passed - focused related unit tests: 342 passed, 1 skipped - broad related unit tests: 855 passed, 12 skipped - clean native rebuild on Ascend A3: passed - final single-card NPU operator regression: 359 passed - pre-commit scoped checks: Ruff, Ruff format, codespell, typos, clang-format, markdownlint, and forbidden-import check passed End-to-end model validation is recorded in dependent PR vllm-project#16423. - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: tianming2009 <13246728590@163.com>
…m-project#16422) ### What this PR does / why we need it? Adds the native and Triton operator layer required by DeepSeek V4.1 on Ascend: - Quant Lightning Indexer V2 and metadata operators - Sparse Flash MLA and metadata operators - compressor, index preparation, query quantization, Engram INT8, and required MoE operator updates - operator documentation, native examples, static checks, and operator-level unit/e2e tests This is the operator-only first part of the DeepSeek V4.1 enablement. Framework/model integration is intentionally split into dependent PR vllm-project#16423. ### Does this PR introduce _any_ user-facing change? Yes. It exposes the native operator capabilities required by DeepSeek V4.1 serving, but does not register or enable the model integration by itself. ### How was this patch tested? - operator static checks: 8 passed - focused related unit tests: 342 passed, 1 skipped - broad related unit tests: 855 passed, 12 skipped - clean native rebuild on Ascend A3: passed - final single-card NPU operator regression: 359 passed - pre-commit scoped checks: Ruff, Ruff format, codespell, typos, clang-format, markdownlint, and forbidden-import check passed End-to-end model validation is recorded in dependent PR vllm-project#16423. - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: like-0517 <ithwlike@126.com>
vllm-project#16422 reintroduced SDK Shape ToString in QLIV2 and SparseFlashMLA tiling error logs. CANN 9.1.0 does not export that symbol, so A5 OPC fails to load liboptiling.so. Restore the vllm-project#16016 local formatter; tiling checks and kernel math are unchanged. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
### What this PR does / why we need it? Fix two A5 QLIV2 build regressions introduced by #16422: - Restore the local shape formatter from #16016. The [A5 Ubuntu amd64 image build](https://github.com/vllm-project/vllm-ascend/actions/runs/34829507531/job/103929405987) cannot load `liboptiling.so` because the 53 restored `Ops::Base::ToString(const gert::Shape&)` calls leave `_ZN3Ops4Base8ToStringERKN4gert5ShapeE` unresolved. Use the bracket-preserving `ShapeToStringForLog` wrapper around the existing `ToStringRaw`. - Match the A5 dependency with `"ascend950" IN_LIST ASCEND_COMPUTE_UNIT`. The standard build passes `ascend950;`; the string-equality check skips the sibling dependency and causes `lightning_indexer_v2_vector1_base.h` to be missing from the staged kernel sources. List membership also handles multi-target configurations while keeping the dependency A5-specific. ### Does this PR introduce _any_ user-facing change? Yes: affected A5 source builds can resolve these QLIV2 Host and kernel-source dependencies. Operator interfaces, numerical computation, and diagnostic formatting are unchanged. ### How was this patch tested? Rebased on main `fbb75a47436901616849da155828f755ec58b619` (including the CPU UT fix #16574). Current HEAD is `5793e3135057281e08f36ab66f807d2c229843f3`; `git range-diff` confirms both patches are unchanged. All 18 native lint hooks passed on this HEAD. A new [CI run](https://github.com/vllm-project/vllm-ascend/actions/runs/34960944860) was triggered. The full build evidence below belongs to the pre-rebase HEAD; the full build was not repeated after rebase. In the A5-107 aarch64 development container with CANN 9.1.0: - Compiled the complete QLIV2 Host translation unit before and after the formatter fix into separate objects. `nm -u` confirms the exact external Shape formatter reference in the baseline object and its absence after the fix. - Compared the actual local formatting helpers against the SDK formatter: all seven shape cases match. - Built the shared Host tiling library with the formatter fix and loaded it in a fresh process using `RTLD_NOW`; the library has neither the unresolved Shape formatter nor a direct `libops_base.so` dependency. - Reproduced the missing sibling header with an actual QLIV2 kernel compilation. After the CMake fix, standard reconfiguration records the dependency and stages the missing header; all five generated QLIV2 configurations (0–4) compile successfully in isolated output directories and produce their final object and JSON artifacts. - On pre-rebase HEAD `5837fe671ec6ef180e48a350e9ecc33d71547b5f`, `bash format.sh ci` passed all 18 hooks with a clean worktree; Gitleaks passed for the two-commit baseline-to-HEAD range. - On that pre-rebase HEAD, the standard `python3 setup.py build` completed with exit 0 (`SOC_VERSION=ascend950dt_9582`, `MAX_JOBS=8`, `OPS_CPU_NUMBER=8`, `VLLM_BATCH_INVARIANT=0`). All 212 A5 kernel targets completed, including all five QLIV2 targets; the OPP installer and Python C++ extension were generated successfully. - Verified the packaged output: 212 kernel objects, the sibling header byte-identical to its source, and `liboptiling.so` loading successfully in a fresh `RTLD_NOW` process without `LD_PRELOAD` or `ASCEND_CUSTOM_OPP_PATH`. The installed library has no unresolved Shape formatter reference or direct `libops_base.so` dependency. The final build reused completed targets from the initial formatter-only attempt and the compiler cache; it is not an all-cache-disabled rebuild. This is a native aarch64 container build, not a rerun of the amd64 Docker image CI. NPU numerical/performance tests were not run. No new Python unit test is added: the native compilation, symbol, library-loading, and package checks directly exercise these build failures. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: Foriv <2293567056@qq.com>
…#16544) ### What this PR does / why we need it? Add DeepSeek V4.1 framework support on top of the native operators introduced in #16422, porting the framework integration from #16423 and adapting it to main. - Register the V4.1 text and vision models and integrate model configuration, weight loading, ModelSlim quantization, and frontend encoding/reasoning/tool-call handling. - Add V4.1 DSA attention and context parallelism, compressed/shared KV cache allocation, ring-buffer state, and DSpark integration while preserving main's model-runner and padded-page cache behavior. - Support Engram HBM storage and optional host offload. Host offload keeps tables on CPU and stages selected BF16 rows through bounded pinned buffers, with a fused INT8 CPU lookup operator and fixed-address graph inputs. - Add regression coverage, the model tutorial, and A3 launch examples. Fix fresh-process model import ordering and the API-node/headless DP launch arguments. The branch contains two commits: framework integration (`4e07c7f`) and Engram host offload (`9038b37`). The integration baseline is main `b8db6e985`, paired with vLLM `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. ### Does this PR introduce _any_ user-facing change? Yes. Users can select the V4.1 model/frontend and configure Engram storage/offload. The tutorial and launch examples cover A3 deployment with DSpark and full decode graphs. Host offload adds a native CPU extension build dependency through the existing CMake build. ### How was this patch tested? Added tests cover V4.1 cache allocation and metadata, compressor/indexer and attention operator contracts, model configuration, frontend encoding, DSpark, fresh-process imports, Engram storage/offload, and integration with the worker/model runner. Earlier integration state `2ac02e6` with vLLM `a97dacb71` passed 400 framework unit tests, 3 fresh-process import cases, and 44 targeted regressions. An A3 single-node TP8/DP2/EP16 deployment with Engram disabled also exercised ordinary generation, thinking mode, and graph replay. These results predate the current main rebase and host-offload commit and are not claimed as full validation of this PR head. Current-head CI and complete end-to-end/GPQA accuracy results remain to be confirmed. No passing two-node Engram or host-offload accuracy result is claimed here. Reproduction entry points: - `examples/deepseek_v41/serve_a3.sh` - `examples/deepseek_v41/serve_a3_offload.sh` - `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md` - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: chenmenglong <chenmenglong1@huawei.com> Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com> Signed-off-by: drslark <slarksblood@qq.com> Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: Yizhou Liu <liu_yizhou@outlook.com> Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com> Signed-off-by: Angazenn <supperccell@163.com> Co-authored-by: Angazenn <supperccell@163.com> Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: jack <QwertyJack@users.noreply.github.com> Co-authored-by: nwpu-zxr <zhouxuerong2@huawei.com> Co-authored-by: pgzddxx <184603735+pgzddxx@users.noreply.github.com> Co-authored-by: Qi Mao <maomaoyu870@gmail.com> Co-authored-by: Qiu Chunshuo <qiuchunshuo@huawei.com> Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com> Co-authored-by: weijinqian0 <1184188277@qq.com> Co-authored-by: Zhu Yi Lin <GDzhu01@users.noreply.github.com> Co-authored-by: zhuyilin (A) <z00911894@china.huawei.com> Co-authored-by: ZT-AIA <1028681969@qq.com> Co-authored-by: drslark <slarksblood@qq.com> Co-authored-by: GPT-5.6 Codex <codex@openai.com> Co-authored-by: Yizhou Liu <liu_yizhou@outlook.com> Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com>
…#16925) ### What this PR does / why we need it? Re-submit the complete DeepSeek V4.1 framework integration from #16544 after its revert in #16905, updated for the current vllm-ascend main. This builds on the native operators from #16422. - Register the V4.1 text and vision models and integrate model configuration, weight loading, ModelSlim quantization, and frontend encoding, reasoning, and tool-call handling. - Add V4.1 DSA attention and context parallelism, compressed/shared KV cache allocation, ring-buffer state from #16890, and DSpark integration while preserving the current model-runner and padded-page cache behavior. - Support Engram HBM storage and optional CPU host offload with bounded pinned buffers, fused INT8 lookup, and fixed-address graph inputs. - Restore the regression tests and model tutorial, including its A3 launch commands. Resolve integration conflicts against current main, including its KV-cache dtype API and drafter typing contract. ### Does this PR introduce _any_ user-facing change? Yes. Users can select the V4.1 model/frontend and configure Engram storage/offload. The tutorial's launch commands cover A3 deployment with DSpark and full decode graphs. Host offload uses the native CPU extension provided by the operator prerequisite. ### How was this patch tested? The original #16544 description records framework-unit-test, fresh-process import, targeted-regression, and A3 single-node results from earlier integration states. Those results predate this resubmission and are not claimed as validation of this PR head. For this head, all changed Python files parse and `git diff --check` passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks pass. Selected A2/A3/310P hardware tests are pending. No end-to-end accuracy or two-node Engram/host-offload result is claimed. Reproduction entry point: `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`. Based on #16544. cc @dragondream-chen @GDzhu01 for review and attribution. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: pgzddxx <1697817735@qq.com> Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com> Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com> Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…vllm-project#16925) ### What this PR does / why we need it? Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544 after its revert in vllm-project#16905, updated for the current vllm-ascend main. This builds on the native operators from vllm-project#16422. - Register the V4.1 text and vision models and integrate model configuration, weight loading, ModelSlim quantization, and frontend encoding, reasoning, and tool-call handling. - Add V4.1 DSA attention and context parallelism, compressed/shared KV cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration while preserving the current model-runner and padded-page cache behavior. - Support Engram HBM storage and optional CPU host offload with bounded pinned buffers, fused INT8 lookup, and fixed-address graph inputs. - Restore the regression tests and model tutorial, including its A3 launch commands. Resolve integration conflicts against current main, including its KV-cache dtype API and drafter typing contract. ### Does this PR introduce _any_ user-facing change? Yes. Users can select the V4.1 model/frontend and configure Engram storage/offload. The tutorial's launch commands cover A3 deployment with DSpark and full decode graphs. Host offload uses the native CPU extension provided by the operator prerequisite. ### How was this patch tested? The original vllm-project#16544 description records framework-unit-test, fresh-process import, targeted-regression, and A3 single-node results from earlier integration states. Those results predate this resubmission and are not claimed as validation of this PR head. For this head, all changed Python files parse and `git diff --check` passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks pass. Selected A2/A3/310P hardware tests are pending. No end-to-end accuracy or two-node Engram/host-offload result is claimed. Reproduction entry point: `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`. Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and attribution. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: pgzddxx <1697817735@qq.com> Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com> Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com> Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…vllm-project#16925) Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544 after its revert in vllm-project#16905, updated for the current vllm-ascend main. This builds on the native operators from vllm-project#16422. - Register the V4.1 text and vision models and integrate model configuration, weight loading, ModelSlim quantization, and frontend encoding, reasoning, and tool-call handling. - Add V4.1 DSA attention and context parallelism, compressed/shared KV cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration while preserving the current model-runner and padded-page cache behavior. - Support Engram HBM storage and optional CPU host offload with bounded pinned buffers, fused INT8 lookup, and fixed-address graph inputs. - Restore the regression tests and model tutorial, including its A3 launch commands. Resolve integration conflicts against current main, including its KV-cache dtype API and drafter typing contract. Yes. Users can select the V4.1 model/frontend and configure Engram storage/offload. The tutorial's launch commands cover A3 deployment with DSpark and full decode graphs. Host offload uses the native CPU extension provided by the operator prerequisite. The original vllm-project#16544 description records framework-unit-test, fresh-process import, targeted-regression, and A3 single-node results from earlier integration states. Those results predate this resubmission and are not claimed as validation of this PR head. For this head, all changed Python files parse and `git diff --check` passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks pass. Selected A2/A3/310P hardware tests are pending. No end-to-end accuracy or two-node Engram/host-offload result is claimed. Reproduction entry point: `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`. Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and attribution. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: pgzddxx <1697817735@qq.com> Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com> Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com> Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…vllm-project#16925) ### What this PR does / why we need it? Re-submit the complete DeepSeek V4.1 framework integration from vllm-project#16544 after its revert in vllm-project#16905, updated for the current vllm-ascend main. This builds on the native operators from vllm-project#16422. - Register the V4.1 text and vision models and integrate model configuration, weight loading, ModelSlim quantization, and frontend encoding, reasoning, and tool-call handling. - Add V4.1 DSA attention and context parallelism, compressed/shared KV cache allocation, ring-buffer state from vllm-project#16890, and DSpark integration while preserving the current model-runner and padded-page cache behavior. - Support Engram HBM storage and optional CPU host offload with bounded pinned buffers, fused INT8 lookup, and fixed-address graph inputs. - Restore the regression tests and model tutorial, including its A3 launch commands. Resolve integration conflicts against current main, including its KV-cache dtype API and drafter typing contract. ### Does this PR introduce _any_ user-facing change? Yes. Users can select the V4.1 model/frontend and configure Engram storage/offload. The tutorial's launch commands cover A3 deployment with DSpark and full decode graphs. Host offload uses the native CPU extension provided by the operator prerequisite. ### How was this patch tested? The original vllm-project#16544 description records framework-unit-test, fresh-process import, targeted-regression, and A3 single-node results from earlier integration states. Those results predate this resubmission and are not claimed as validation of this PR head. For this head, all changed Python files parse and `git diff --check` passes. The pre-commit/mypy, CPU UT, docs, DCO, and CI rebase checks pass. Selected A2/A3/310P hardware tests are pending. No end-to-end accuracy or two-node Engram/host-offload result is claimed. Reproduction entry point: `docs/source/tutorials/models/DeepSeek-V4.1-Flash.md`. Based on vllm-project#16544. cc @dragondream-chen @GDzhu01 for review and attribution. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: pgzddxx <1697817735@qq.com> Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: Qiu Chunshuo <chunshuoq@gmail.com> Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: Qiu Chunshuo <chunshuoq@gmail.com> Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
What this PR does / why we need it?
Adds the native and Triton operator layer required by DeepSeek V4.1 on Ascend:
This is the operator-only first part of the DeepSeek V4.1 enablement. Framework/model integration is intentionally split into dependent PR #16423.
Does this PR introduce any user-facing change?
Yes. It exposes the native operator capabilities required by DeepSeek V4.1 serving, but does not register or enable the model integration by itself.
How was this patch tested?
End-to-end model validation is recorded in dependent PR #16423.