Repository navigation
Conversation
…ng (vllm-project#16146) ### What this PR does / why we need it? This PR aims to fix kimi k3 dspark kv grouping when not enabling prefix cache. The problem is that patch_mamba_config.py aligns ssm to k, which is from only target model without draft model. Since kimi k3 dspark may use gqa, kv cache layout thus can be problematic. That is, when aligning, mamba considers mla without gqa, so finally v in gqa can overlap ssm in mamba. Example is shown below: ``` Mamba: | Conv | SSM | Padding | MLA: | Padding | K(latent) | RoPE | GQA: | Padding | K | V | ``` When enabling prefix cache, kimi k3 has its own kv grouping logic where no tensor is shared by mamba and gqa. So we just fall back into it. ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? by ci - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Zetong Li <slippersss@126.com>
cursor
Bot
force-pushed
the
enable-mrv2-whitelist-deae
branch
2 times, most recently
from
September 10, 2026 02:43
675859b to
c27a1b5
Compare
…2.0.0 (vllm-project#15956) ### What this PR does / why we need it? ref to vllm-project#10072 Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds. Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation. Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes. ### Does this PR introduce _any_ user-facing change? A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration. ### How was this patch tested? - pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ``` ======================================================================= warnings summary ======================================================================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html =========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) =========================================================== [root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success. - Build the image with Dockerfile, and verification is OK. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com>
What this PR does / why we need it? Temporarily disable test_qwen3_6_mtp.py in CI by adding it to skip_tests. The test file is retained so it can be re-enabled later. Does this PR introduce any user-facing change? No. This only changes CI test scheduling. How was this patch tested? - Ran select_tests.py --explicit-e2e-tests tests/e2e/pull_request/two_card/model_runner_v2/test_qwen3_6_mtp.py and confirmed test_groups=[] and has_tests=false. - git diff --check passed. - Full formatting checks could not run because pre-commit was unavailable. No new tests were added because this change only updates the existing CI skip list. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: liaoqingqing <liaoqingqing3@h-partners.com> Co-authored-by: liaoqingqing <liaoqingqing3@h-partners.com>
### What this PR does / why we need it?
Currently DSACP with SP is implemented in a suboptimal way, Leading to
regressed performance during prefill.
We support further development for deepseek v4 with SP and DSACP.
SP chunk is implemented in model, while attention reduce scatter /
allgather is in decoder layer. This is a implementation similar to
upstream deepseek v4 SP.
### Does this PR introduce _any_ user-facing change?
As is required in the test, dsacp is only depend on itself. However,
current dsacp depend on flashcomm.
This PR enables dsacp to automatically enable sp.
### How was this patch tested?
**Current**
ruff check
DS v4 curl success
performance check
A/B test with main
Ablation with dsacp or SP
**On going**
### benchmark
model=DeepSeek-V4-Flash-w8a8
DP2, TP4, EP
```bash
vllm serve \
--model /mnt/weight/DeepSeek-V4-Flash-w8a8-mtp \
--served-model-name qwen \
--host 127.0.0.1 \
--port 8010 \
--max_model_len 25600 \
--data-parallel-size 2 \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.92 \
--max-num-seqs 8 \
--enable-expert-parallel \
--additional-config '{
"enable_flashcomm1":true,
"enable_shared_expert_dp":true,
"enable_dsa_cp":true,
"enable_cpu_binding":true,
"multistream_overlap_shared_expert":true
}'
```
```bash
python -m vllm.entrypoints.cli.main bench serve \
--backend openai \
--base-url http://127.0.0.1:8010 \
--endpoint /v1/completions \
--served-model-name qwen \
--dataset-name random \
--random-input-len 8192 \
--random-output-len 2048 \
--num-prompts 50 \
--num-warmups 5 \
--max-concurrency 8 \
--metric-percentiles 50,90,99 \
--seed 0
```
| branch | median TTFT (ms)| median TPOT (ms) |
|-------|-------|-------|
| this PR | 2796 | 43.75 |
| before PR vllm-project#13946 | 2777 | 42.93
| main | 3105 | 45.57 |
### Test Dspark and MTP
Tested with DP2, TP4, EP.
MTP model with w8a8 quantization.
DSpark model with w4a8 quantization.
| # | PR | spec | dsacp | accuracy | inference time |
|---|----|------|-------|----------|----------------|
| 1 | before | MTP=3 | off | 98.50 | 116s |
| 2 | after | MTP=3 | off | 98.50 | 123s |
| 3 | before | Dspark=7 | off | 98.50 | 200s |
| 4 | before | Dspark=7 | on | 98.50 | 188s |
| 5 | after | Dspark=7 | off | 98.00 | 200s |
| 6 | after | Dspark=7 | on | 98.00 | 188s |
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: MmMmaru <1843167357@qq.com>
### What this PR does / why we need it? This PR aligns upstream function interface and fixes the `apply_grammar_bitmask` Triton kernel to correctly handle adaptive verification and per-request logit offsets. It introduces `cu_num_logits` to resolve absolute logit rows on device and handles inactive mapping entries safely. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Tested with CI case. - vLLM main: vllm-project/vllm@e6bfe03 --------- Signed-off-by: zouzy <zouzongyu@huawei.com>
### What this PR does / why we need it? delete --headless for weekly cases ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? run the cases weekly - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: guxin108 <1252896542@qq.com>
### What this PR does / why we need it? Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV. ### Does this PR introduce _any_ user-facing change? Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1. ### How was this patch tested? GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector. | Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output cap hits + failed requests (limit) | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) | | Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) | | DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2 (8192) | | DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1 (8192) | Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: lxx <luxixiang@h-partners.com> Co-authored-by: lxx <luxixiang@h-partners.com>
…m-project#15737) ### What this PR does / why we need it? Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. ### Does this PR introduce _any_ user-facing change? Documentation only. No code behavior change. ### How was this patch tested? - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com>
…llm-project#16114) ### What this PR does / why we need it? Correct specific setup and first-inference ambiguities while preserving the existing documentation structure: - Pair the main development checkout with its verified vLLM commit, and pair release CPU test checkouts with matching release tags. - Identify commands that run inside the CPU test container and use the correct nightly image names for NPU examples. - Clarify that the manual 910B Ops installation example applies to A2 and refer other hardware to the official CANN guide. - Use `max_tokens` for completion requests, report HTTP failures with plain curl, and describe the expected nonempty completion. - Fix the shared hardware table indentation in Quick Start so each product series renders as a separate row. The existing chapters, hardware tabs, background-serving flow, startup/shutdown log examples, and four doctest stages are retained. No test infrastructure changes or new performance sections are included. ### Does this PR introduce _any_ user-facing change? Only the documented commands and explanations change. There are no runtime/API changes. The examples retain the existing background-serving and process-lookup commands. ### How was this patch tested? - `bash format.sh ci`: passed. - Strict English documentation build via `tools/rtd_build.sh`: passed. - 27 doctest helper tests passed; all eight online blocks extract and pass Bash syntax checks. - Verified that the generated Quick Start hardware table contains four separate product-series rows. The original PID lookup and stop commands are retained. - `git diff --check`: passed. AI assistance was used for editing and local validation. NPU inference, CANN installation, source builds, and hardware-dependent tests have not been run on this macOS host. No performance measurements are claimed. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: sugm <sugm0521@163.com>
### What this PR does / why we need it? 1.Qwen3.5-122B Use case:Container names are formed by concatenating test case names with other specific characters. Kubernetes restricts container names to a maximum of 64 characters. Overly long test case names cause Kubernetes scheduling failures. The current solution is to shorten the test case name length. 2.Qwen3.6-35B Use case: Modify the dataset configuration. ### Does this PR introduce _any_ user-facing change? NO ### How was this patch tested? by the running the test - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: weixin <murongfengerxch@163.com>
…st list for SWR (vllm-project#16130) ## What this PR does / why we need it? SWR rejects OCI image index (`application/vnd.oci.image.index.v1+json`) produced by `docker buildx imagetools create` when merging OCI-format single-arch images, returning `400 Bad Request: Invalid image, fail to parse 'manifest.json'`. **Root cause**: `docker buildx imagetools create` selects the output media type based on the source manifests. When all sources are Docker v2 schema2 (`application/vnd.docker.*`), it produces a **Docker manifest list** (`application/vnd.docker.distribution.manifest.list.v2+json`) which SWR accepts. With OCI sources it produces an OCI image index, which SWR rejects. **Fix**: add `oci-mediatypes=false` to the build output so single-arch images are built as Docker v2 schema2. This lets `docker buildx imagetools create` produce a Docker manifest list directly, and eliminates the skopeo-based format conversion entirely (which required `sudo`/root not available on the self-hosted runner). ### Changes - **Build output**: add `oci-mediatypes=false` - **build-push-digest "Tag digest to prevent GC"**: replace `skopeo copy` with `docker buildx imagetools create` - **merge-image**: replace `skopeo copy` + `docker manifest create/push` with direct `docker buildx imagetools create` - **merge-image-temp**: same as merge-image - **"Clean up temp tags"** (both jobs): remove `sudo apt-get install skopeo`; guard `skopeo delete` with a `command -v` check so the job no longer fails when skopeo is unavailable ## Does this PR introduce any user-facing change? No. Only the CI image build/push workflow is changed. ## How was this patch tested? - CI workflow changes verified locally (YAML valid, job structure intact) - No skopeo/sudo/apt-get dependencies remain in the merge path, so the previous `skopeo: command not found` failure on the self-hosted (non-root, no-sudo) runner is eliminated - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: JavaPythonAIForBAT <wuhejun@h-partners.com>
…ct#16121) ### What this PR does / why we need it? This PR fixes the MiniMax-M3 prefill index top-k score preparation path on Ascend A3. On A3 with Triton-Ascend 3.2.2, the previous `_prepare_prefill_topk_scores_kernel` could fail during sustained MiniMax-M3 serving workloads such as TerminalBench. Two device-side failure modes were observed during investigation: * The original query-tiled 2-D masked tail-store implementation could trigger an MTE/SMMU fault with `The DDR address of the MTE instruction is out of range`. * A follow-up implementation that converted the stores to query-direction 1-D stores could trigger a Vector Core timeout/trap during real serving workloads. The original implementation vectorized score updates across the query dimension, while the score tensor layout is: ```text [num_index_heads, total_query_tokens, score_blocks] ``` where `score_blocks` is the contiguous dimension. This PR changes the prefill score preparation to follow the tensor layout directly: * Initialize the score buffer to `-inf`, so invalid/padded score blocks no longer need a separate tail-fill path. * Keep the existing prefill QK score computation unchanged. * Treat one `(query, index_head)` row as the logical preparation unit. * Apply init/local forced priorities using block-contiguous stores along the last dimension. * Keep query handling as task scheduling rather than using query as the vectorized memory-access dimension. * Use an AIV-oriented program mapping so large/ragged prefills do not create one launched program per `(query, head)` and exceed the Ascend `coreDim <= 65535` launch limit. * Preserve the existing public operator APIs and priority semantics: * init blocks: `1e30` * local blocks: `1e29` * local priority overrides init priority when they overlap * invalid blocks remain `-inf` This removes the dynamic invalid-tail store path and avoids both the previous `[query, block]` masked GM store pattern and query-direction strided Vector stores. The issue was reproduced on MiniMax-M3 running on Ascend A3 with DP=2 / TP=8 across 16 devices under sustained TerminalBench workloads. The workaround is based on behavior observed with Triton-Ascend 3.2.2. The exact compiler/backend root cause has not been proven, so the implementation should be re-evaluated after Triton-Ascend upgrades. ### Does this PR introduce *any* user-facing change? No. The public APIs, model behavior, output layout, top-k semantics, and sparse-attention semantics are unchanged. This PR only changes the internal Triton implementation and task mapping used by MiniMax-M3 prefill index-score preparation on Ascend. ### How was this patch tested? The patch was validated with standalone correctness cases, TerminalBench-style workload sweeps, and real MiniMax-M3 serving workloads on Ascend A3. Correctness coverage includes: * sparse-block boundaries: query lengths `127`, `128`, and `129` * ragged query batches * long invalid score tails * init/local overlap * `2048` / `2049` prefill boundary cases * large ragged prefill workloads * block counts matching the production workload range, including `480` and `528` The regression tests verify the complete prefill index path: ```text index score → score preparation → torch.topk → invalid-index masking ``` Representative TerminalBench-style end-to-end results with 528 score blocks: ```text q=2048, prefix=0: baseline p50 = 2392.265 us patched p50 = 803.281 us q=2048, prefix=32768: baseline p50 = 2006.960 us patched p50 = 847.873 us q=2048, prefix=65512: baseline p50 = 1628.907 us patched p50 = 1537.166 us q=4096, prefix=32768: baseline p50 = 1362.001 us patched p50 = 1315.651 us q=8192, prefix=32768: baseline p50 = 2416.103 us patched p50 = 2336.767 us q=32768, prefix=0: baseline p50 = 4760.416 us patched p50 = 4365.349 us ``` A large ragged TerminalBench-style workload also passed: ```text query lengths: [2048, 2048, 4096, 8192, 16384] prefix lengths: [0, 8192, 16384, 32768, 0] total query tokens: 32768 score blocks: 528 baseline p50 = 4809.989 us patched p50 = 4567.960 us ``` This case previously exposed the `coreDim` limitation for the direct one-query-per-program A5-style launch (`81920 > 65535`). The patched task mapping completes successfully. Additional regression guards were added for the failure boundaries and large/ragged workloads. Large TerminalBench-style cases are kept as nightly/stability guards rather than regular lightweight CI cases. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
yjyang62
force-pushed
the
enable-mrv2-whitelist-deae
branch
from
September 10, 2026 09:38
c27a1b5 to
4261b44
Compare
…lm-project#16110) # [Bugfix] Fix MXFP4 MoE scale layout and dtype handling ### What this PR does / why we need it? This PR fixes several general compatibility issues in the Ascend W4A4 MXFP4 path. #### Explicit MXFP4 dtype interpretation for grouped matmul MXFP4 activations and weights are physically stored as `torch.uint8`. Without explicit logical dtype arguments, `npu_grouped_matmul` may interpret them as an unsupported `DT_UINT8`/`DT_UINT8` pair. This PR explicitly passes: ~~~python x_dtype=torch_npu.float4_e2m1fn_x2 weight_dtype=torch_npu.float4_e2m1fn_x2 ~~~ The tensors are not converted or copied. These arguments tell the NPU operator to interpret the packed `uint8` data as MXFP4. #### Support odd MXFP4 scale group counts The existing scale conversion assumes that the final scale dimension is even before reshaping it into pairs: ~~~text [G, N, K] -> [G, N, K // 2, 2] ~~~ Some valid tensor shapes produce an odd number of scale groups. Reshaping these scales directly causes a tensor shape mismatch. This PR pads the final scale dimension with one zero when necessary before applying the existing paired layout conversion. Existing checkpoints with an even number of scale groups continue to use the original path. #### Support uneven row-parallel MX scale shards Row-parallel MX scale parameters use a ceil-divided local group count. When the global number of scale groups is not divisible by the tensor-parallel size, the final shard can contain fewer valid groups than the allocated local parameter. This PR adds an Ascend MX scale loader wrapper that pads the global scale tensor before delegating to the existing row-parallel weight loader. Padding is applied only when: ~~~python 0 < padding < tp_size ~~~ This limits the behavior to the natural remainder introduced by tensor-parallel ceil division. Unexpected checkpoint shape mismatches are not hidden and continue to fail through the existing loader. The compatibility logic applies only to: - Ascend row-parallel linear layers - MX quantization methods - Per-group scale parameters The generic vLLM weight loader remains unchanged. ### Does this PR introduce _any_ user-facing change? Yes. W4A4 MXFP4 models can now load and execute correctly when they use: - Packed MXFP4 tensors requiring explicit logical dtype metadata - An odd number of scale groups - Scale dimensions that are not evenly divisible across tensor-parallel ranks There are no API, command-line interface, or configuration changes. ### How was this patch tested? Unit tests were added for: - Odd MXFP4 MoE scale group padding - Explicit MXFP4 logical dtype arguments passed to grouped matmul - Uneven tensor-parallel MX scale padding Static validation was performed with: ~~~bash git diff --check ~~~ The unit tests could not be executed in the local development environment because PyTorch and an Ascend NPU runtime are unavailable. The relevant tests can be run in a configured vLLM-Ascend development environment with: ~~~bash pytest -q tests/ut/quantization/methods/test_w4a4_mxfp4.py pytest -q tests/ut/quantization/test_method_adapters.py \ -k mx_scale_weight_loader ~~~ End-to-end validation should cover W4A4 MXFP4 model loading and inference with multiple tensor-parallel configurations, including configurations where scale groups cannot be divided evenly across ranks. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: “yzlbig” <2582195355@qq.com> Signed-off-by: yzlbig <2582195355@qq.com>
…16220) Use the verified vLLM main commit from pr_test.yaml, and allow the coverage workflow to run on pull requests. ### What this PR does / why we need it? The version of the scheduled task vllm is changed to be the same as that of the PR. ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? The version of the VLLM used by the coverage rate scheduled task is the same as that used by the PR. - vLLM main: vllm-project/vllm@b2f6858
### What this PR does / why we need it? Fix a query-length mismatch during PCP speculative graph replay by rebuilding draft prefill metadata only for DSA and SFA. ### Does this PR introduce _any_ user-facing change? Yes, fixes speculative decoding graph replay failures. ### How was this patch tested? - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: leolee <yihao.li@huawei.com>
…mmand not found' (vllm-project#16272) ## Summary - `Trigger quay.io sync` job failed with `gh: command not found` (exit code 127) because the `linux-amd64-cpu-4` runner lacks the gh CLI - Replace `run: gh workflow run ...` with `actions/github-script@v7` which uses the GitHub API directly and doesn't depend on gh being installed on the runner - Verified: target workflow `sync-vllm-ascend.yml` exists and is active in `ascend-gha-runners/sync-tools` ## Does this PR introduce any user-facing change? No. CI-only change. ## How was this patch tested? - YAML syntax validated - Target workflow verified active via API - createWorkflowDispatch API is the correct method for triggering workflow_dispatch - Full validation requires a real image build CI run - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: JavaPythonAIForBAT <wuhejun@h-partners.com>
…ct#16204) ## Summary - Prefiller/decoder use `--data-parallel-start-rank` with multi-node DP. On vLLM 0.28, a set start-rank infers hybrid LB; with `local=1` that autoswitches to external LB, which forbids `--headless` on remotes. - **Keep** `--data-parallel-start-rank`; **remove** `--headless` on both prefiller and decoder secondary nodes so remotes expose API servers (PD proxy can target them). ## Related failure - https://github.com/vllm-project/vllm-ascend/actions/runs/34372289942/job/102544223304 ## Test plan - [ ] DeepSeek-V3.2-W8A8-EP starts without `Remote engine must not use --headless` - vLLM main: vllm-project/vllm@b2f6858 --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…ction sampling kernels (vllm-project#16222) ## What this PR does / why we need it? Fixes sporadic device error **507035** (MTE illegal GM address access, fault symbol `load_gm_to_ubuf_1d_int32_t+0x6c`) in `rejection_greedy_sample_triton`, and eliminates three more instances of the same latent pattern. **Root cause:** `tl.where` is a select instruction, not control flow — *both arms are always evaluated*. The pattern ```python start_idx = tl.where(offset == 0, 0, tl.load(cu_num_draft_tokens_ptr + offset - 1, is_greedy_mask)) ``` therefore reads one int32 **before the buffer start** whenever lane `offset == 0` is active in the load mask (e.g. all-greedy batches where `is_greedy=None`). Whether this faults depends on the allocator layout: in production it crashed only when `cu_num_draft_tokens` happened to sit at an allocation boundary (`ptr - 4` falling into an unmapped page), which explains the sporadic, TP-dependent failures. The fault was reproduced in isolation: with `cu` placed at a boundary address (`0x400040000000`, faulting address `0x40003ffffffc`), the old kernel reports 507035 at synchronize while the fixed kernel passes 3/3 runs; with padded placement the old kernel also passes — confirming the crash is layout-dependent, not data-dependent. **Changes:** 1. **Mask the prev-entry load itself** in the three tensor-path kernels (lane 0's load is now predicated off at the hardware level, so no memory transaction is issued): - `rejection_greedy_sample_triton`: `mask=is_greedy_mask & (offset > 0), other=0` - `rejection_random_sample_kernel`: `mask=not_greedy_mask & (offsets > 0), other=0` - `expand_kernel`: `mask=len_mask & (offset > 0), other=0` (lane 0 was *unconditionally* active here — same exposure level as the crashed greedy kernel) 2. **`sample_recovered_tokens_kernel`** (scalar `req_idx` path): clamp the index with `tl.maximum(req_idx - 1, 0)` so the load address always stays inside the buffer; `tl.where` then merely discards the (in-bounds, unused) value for `req_idx == 0`. A 0-d masked load is not a verified construct on triton-ascend, so the clamp is the safer equivalent. 3. **`other=0` on the `end_idx` loads**: without it, masked lanes receive an *undefined* value (Triton only guarantees the memory is not accessed). The per-request token loops (`for i in range(num_tokens1)`) rely on `num_draft_tokens == 0` to skip inactive lanes, so an undefined `end_idx` could turn into an out-of-bounds loop with both OOB reads *and writes*. `other=0` makes this invariant deterministic instead of dependent on backend implementation details. Note: this matches the masking idiom already used by `rejection_random_sample_block_verify_kernel` in the same file (`prev_mask = not_greedy_mask & (offsets > 0)`). ## Does this PR introduce _any_ user-facing change? No API or behavior change. For all in-bounds inputs the kernels produce bit-identical results — the only eliminated behaviors are an out-of-bounds read and reliance on undefined masked-load values. Users benefit from the disappearance of sporadic 507035 device errors / hangs in speculative decoding workloads (observed on an 8-machine PD deployment under D-model graph mode; after the fix, 3 consecutive rounds of 64-concurrent requests completed 192/192 with no device errors). ## How was this patch tested? **Unit/e2e tests added** (`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py`, 9 new parameterized cases, device=npu, compared against the PyTorch reference implementations): - `test_rejection_greedy_sample_triton_boundary` — 4 patterns: `is_greedy=None` all-greedy (the always-OOB lane-0 path), all-greedy tensor, mixed batch with first request greedy vs sampling; batch includes all-match (bonus), first-token reject, mid reject, and zero-draft-token requests; also asserts non-greedy rows stay untouched (sentinel preserved) - `test_rejection_greedy_sample_triton_boundary_multilane` — batch=256 randomized mixed batch forcing `BLOCK_SIZE > 1`, so lane 0 shares a program block with other lanes; first request greedy/sampling both covered - `test_rejection_random_sample_boundary` — 3 patterns: all-random (lane-0 always active in this kernel), mixed with first request sampling vs greedy; includes forced all-accept (bonus), a `-1` draft placeholder (always rejected → recovered token), and a zero-draft-token request **Existing tests** for `expand_kernel` and `sample_recovered_tokens_kernel` in the same file already exercise the req0/lane0 path on every launch and serve as regression coverage for fixes `expand_kernel` and `sample_recovered_tokens_kernel`. **Fault-level reproduction** (isolated single-op experiment, documented in the linked investigation): boundary-address allocation of `cu` reproduces 507035 with the old kernel and passes with the fixed kernel, for both the all-accept and rejection branches. ```bash # Commands used pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py -k "boundary" pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py ruff check && ruff format --check ``` > Note: the boundary tests verify semantic equivalence and lane-0 masking behavior; they cannot deterministically force the allocator to place tensors at mapping boundaries, so the 507035 reproduction itself relies on the dedicated single-op repro rather than pytest. ``` tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_expand_kernel [14:59:22] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_sample_recovered_tokens_kernel [14:59:23] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-1-False] [14:59:23] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-1-True] [14:59:24] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-2-False] [14:59:24] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-2-True] [14:59:24] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-3-False] [14:59:25] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-3-True] [14:59:25] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-1-False] [14:59:26] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-1-True] [14:59:26] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-2-False] [14:59:26] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-2-True] [14:59:27] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-3-False] [14:59:27] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-3-True] [14:59:27] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-1-False] [14:59:28] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-1-True] [14:59:28] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-2-False] [14:59:29] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-2-True] [14:59:29] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-3-False] [14:59:29] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-3-True] [14:59:30] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-1-False] [14:59:30] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-1-True] [14:59:31] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-2-False] [14:59:31] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-2-True] [14:59:31] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-3-False] [14:59:32] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-3-True] [14:59:32] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_spec_len_1_triton_kernel[False] [14:59:32] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_spec_len_1_triton_kernel[True] [14:59:33] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_kernel[False] [14:59:33] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_kernel[True] [14:59:34] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[all_greedy_none] [14:59:34] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[all_greedy] [14:59:34] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[mixed_first_greedy] [14:59:35] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[mixed_first_random] [14:59:35] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary_multilane[True] [14:59:36] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary_multilane[False] [14:59:36] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[all_random] [14:59:36] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[mixed_first_random] [14:59:37] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[mixed_first_greedy] [14:59:37] PASSED tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_sampler_block_verify_triton_kernel[5-3-7-is_greedy0-uniform_probs0-recovered_token_ids0-bonus_token_ids0-target_probs0-None-draft_token_ids0-cu_num_draft_tokens0] [14:59:38] PASSED[INFO] Model cache cleared ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: liyishi <1252651434@qq.com> Co-authored-by: Zetong Li <slippersss@126.com>
### What this PR does / why we need it?
This PR adds block-major, range-based layerwise KV-cache transfer
support for the Mooncake backend in `AscendStoreConnector`.
Previously, Mooncake layerwise transfer stored every layer of each KV
block as an independent remote object:
```text
block × layer × rank
```
For large models, this creates a large number of remote keys and
objects, increasing:
- Remote metadata usage.
- Cache existence-check overhead.
- Mooncake put/get control-plane overhead.
- Object lifecycle and lease-management costs.
This PR changes the remote object layout so that all KV-cache layers
belonging to the same block and rank are stored in one Mooncake object:
```text
block × rank
```
The total KV-cache payload is not reduced. Instead, the layer dimension
is moved from the remote key into byte ranges inside a whole-block
object.
When a layer becomes ready, only the byte ranges belonging to that layer
are transferred. This preserves layer-by-layer transfer and
computation/communication overlap while reducing the number of remote
objects by approximately the number of KV-cache layers.
The main changes are:
- Add block-major Mooncake keys for scheduler-side cache lookup.
- Store all layers of one KV block and rank in a single remote object.
- Calculate the whole-object size and per-layer ranges from the actual
KV-cache layout.
- Add range-session APIs to the AscendStore backend abstraction.
- Map the new abstraction to the following Mooncake APIs:
- `batch_put_session_start`
- `batch_put_from_multi_buffer_ranges`
- `batch_put_session_end`
- `batch_put_session_revoke`
- `batch_get_session_start`
- `batch_get_into_multi_buffer_ranges`
- `batch_get_session_end`
- Add `MooncakeSessionTracker` to manage put/get sessions across
chunked-prefill steps.
- Track multiple requests sharing the same prefix-cache key.
- Commit a remote object only after all required layer ranges have been
written successfully.
- Revoke incomplete remote objects when a layer transfer fails.
- Keep get sessions alive until the last request using the shared key
releases it.
- Preserve continuous-prefix semantics by requiring every saving rank of
a block to be available.
- Split large range transfers according to:
- `layerwise_max_transfer_blocks`
- `layerwise_max_transfer_bytes`
- Add fail-fast validation for required Mooncake range-session APIs.
- Reject currently unsupported topology and KV-cache layouts with
explicit errors.
- Add optional range-transfer audit logging through
`VLLM_ASCEND_KVPOOL_RANGE_DEBUG`.
The Mooncake write lifecycle becomes:
```text
batch_put_session_start
→ layer 0 range put
→ layer 1 range put
→ ...
→ final layer range put
→ batch_put_session_end
```
If any required range transfer fails:
```text
batch_put_session_start
→ one or more layer range puts
→ transfer failure
→ batch_put_session_revoke
```
The Mooncake read lifecycle becomes:
```text
batch_get_session_start
→ layer 0 range get
→ layer 1 range get
→ ...
→ final owner releases the key
→ batch_get_session_end
```
The change is limited to the Mooncake layerwise KV Pool path. Existing
non-layerwise and Memcache data paths are not changed.
### Does this PR introduce *any* user-facing change?
Yes.
Users can enable the new Mooncake block-major layerwise path with:
```json
{
"kv_connector": "AscendStoreConnector",
"kv_connector_extra_config": {
"backend": "mooncake",
"use_layerwise": true
}
}
```
When this configuration is enabled:
- Mooncake stores one whole-block object per KV block and saving rank
instead of one object per block, layer, and rank.
- KV-cache data is still transferred layer by layer, but through ranges
inside the whole-block object.
- Remote object and metadata counts are reduced from approximately
`block × layer × rank` to `block × rank`.
- Startup fails early if the installed Mooncake client does not provide
the required range-session APIs.
- Unsupported configurations are rejected with explicit errors instead
of failing during transfer.
The currently unsupported configurations include:
- Prefill/Decode TP mismatch.
- Pipeline parallel size greater than one.
- Prefill context parallel size greater than one.
- Decode context parallel size greater than one.
- Hybrid or multiple KV-cache groups.
Detailed range-transfer audit logging can be enabled with:
```bash
export VLLM_ASCEND_KVPOOL_RANGE_DEBUG=1
```
There is no behavior change unless Mooncake layerwise transfer is
explicitly enabled.
### How was this patch tested?
#### Unit tests
The patch adds and updates CPU unit tests covering:
- Whole-block object sizing based on the actual KV-cache layout.
- Block-major key generation.
- Cache-hit validation across all saving ranks.
- Per-layer object offsets, local addresses, and transferred byte
counts.
- Key-major range construction.
- Splitting transfers by maximum block count.
- Splitting large individual ranges by maximum byte count.
- Put-session creation.
- Per-layer range writes.
- Final-layer commit ordering.
- Revocation of incomplete objects after transfer failures.
- Get-session creation.
- Shared-key ownership across multiple requests.
- Closing a get session only after its last owner releases it.
- Failed get attempts and retry cleanup.
- Chunked-prefill session persistence.
- Commit retry and terminal request cleanup.
- Full remote-cache hit handling.
- Partial-block replacement by a completed full-block key.
- Reloading committed prefixes during subsequent chunks.
- Required Mooncake range-session API validation.
- Validation of aligned per-key result counts.
- Rejection of unsupported TP-mismatch and topology configurations.
- Strict parsing and default behavior of
`VLLM_ASCEND_KVPOOL_RANGE_DEBUG`.
The main test files are:
- `tests/ut/distributed/ascend_store/test_mooncake_layerwise.py`
- `tests/ut/distributed/ascend_store/test_backend.py`
- `tests/ut/distributed/ascend_store/test_pool_scheduler.py`
- `tests/ut/distributed/ascend_store/test_pool_worker.py`
- `tests/ut/test_envs.py`
CI validation on the current PR head includes:
- `pre-commit`: passed
- CPU unit tests: passed
#### End-to-end performance test
The implementation was also tested with the following workload:
- Model: `Qwen3-30B`
- Data type: `W4A4`
- Input length: `16K`
- Output length: `500`
- External prefix-cache ratio: `75%`
- Concurrency: `1`
- Total requests per run: `32`
- KV Pool enabled: yes
Lower TTFT is better.
| Run | Model | Cards | Data type | Layerwise | KV Pool | Input | Output
| External prefix cache | Prefix cache | Prefill parallelism | Decode
parallelism | Concurrency | Total requests | TTFT (s) |
|---:|---|---:|---|:---:|:---:|---:|---:|---:|:---:|:---:|:---:|---:|---:|---:|
| 1 | Qwen3-30B | 2 | W4A4 | No | Yes | 16K | 500 | 75% | N/A | 1 | 1 |
1 | 32 | 0.44 |
| 2 | Qwen3-30B | 2 | W4A4 | Yes | Yes | 16K | 500 | 75% | N/A | 1 | 1 |
1 | 32 | 0.36 |
| 3 | Qwen3-30B | 4 | W4A4 | Yes | Yes | 16K | 500 | 75% | N/A | TP2 |
DP2 | 1 | 32 | 0.40 |
For the matching two-card configuration:
```text
Non-layerwise TTFT: 0.44 s
Layerwise TTFT, run : 0.36 s
```
The average TTFT improvement is:
```text
(0.44 - 0.36) / 0.44 × 100% = 18.2%
```
The two layerwise runs improve TTFT by approximately:
- `18.2%` for the `0.36 s` run.
The four-card `TP2/DP2` layerwise configuration achieved a TTFT of `0.40
s`. It is reported separately because there is no matching four-card
non-layerwise baseline in the current test data, so a direct speedup
comparison is not made.
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: fangrongcan <17343701736@163.com>
) Reverts vllm-project#12599 - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: hejianping-00178005 <44997374+winson-00178005@users.noreply.github.com>
…lm-project#16224) ### What this PR does / why we need it? This PR refactors the hardcoded `quant_mode` values in the DSA indexer to be dynamically retrieved via `DeviceOperator.get_dsa_indexer_quant_mode()`. This allows different device adaptors (e.g., Non-A5 using INT8 and A5 using FP8) to return their respective quantization modes, enabling proper support for different hardware configurations. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Tested with existing unit tests in `tests/ut/models/test_deepseek_v4_indexer.py` which were updated to use the new dynamic method. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: lenghuixing <lenghuixing@huawei.com> Co-authored-by: lenghuixing <lenghuixing@huawei.com>
…llm-project#14496) ### What this PR does / why we need it? Come from vllm-project#13773 In that PR, they refactor moe_mlp's methods into differnet quantization method, so this RP,we refactor moe-mlp and quantization methods for 310p. ### Does this PR introduce _any_ user-facing change? No,this is an internal refactor of the fused-moe and method for 310p and its test coverage. ### How was this patch tested? The following models are tested and can be curled successfully: 300I DUO: Qwen3.5-35B-A3B-mtp - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: z00599867 <zhulinang@huawei.com> Co-authored-by: z00599867 <zhulinang@huawei.com>
…6054) Switch the scheduled main2main workflow from the a2 pool to the a3-16 pool: - run on `linux-aarch64-a3-16-sh-001` with container `cann:9.1.0-a3-ubuntu22.04-py3.12` - `main2main_tests.json` now holds the reviewed 15-case selection <img width="2044" height="1811" alt="workflow" src="https://github.com/user-attachments/assets/03e8123c-bf10-44a3-8479-af876ed358b4" /> - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: wjunLu <wjunlu217@gmail.com>
…6294) ### What this PR does / why we need it? Adds an experimental deployment tutorial and supported-model matrix entry for `DeepSeek-V4.1-Flash`. The tutorial follows the official model release and the current vLLM Ascend adaptation work. It documents: - the released model name and public Hugging Face/ModelScope checkpoints; - the two-node Atlas 800 A3 and four-node Atlas 800 A2 W8A8 topologies (TP8/DP4/EP32); - INT8 Engram storage, DSpark speculative decoding, and `FULL_DECODE_ONLY` ACL Graph; - Docker, source-build, multi-node launch, health, text, and image verification commands; - explicit accuracy, performance, context-length, and deployment limitations. The pre-release `Aurora` name and internal checkpoint paths are intentionally omitted. The guide uses placeholders for the Ascend-quantized checkpoint path and keeps the documented claims within the current validation boundary. Related adaptation work: https://github.com/GDzhu01/vllm-ascend-v41-private ### Does this PR introduce _any_ user-facing change? Yes. It adds an experimental DeepSeek-V4.1-Flash deployment guide and advertises the documented support boundary in the model support matrix. ### How was this patch tested? - `markdownlint-cli@0.45.0` passes for both changed Markdown files. - `git diff --check` passes. - `docs/hooks/nav_titles.py` passes Python syntax compilation. - The new supported-model matrix row has the same 20 columns as its header. - The tutorial is registered exactly once in `mkdocs.yml` and has matching English and Chinese navigation titles. - The Hugging Face, ModelScope, ModelSlim, and vLLM benchmark links return HTTP 200. - The branch contains one documentation-only commit and changes only four documentation/navigation files. Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…ansfer (vllm-project#15854) ## What this PR does / why we need it? The layerwise KV pool transfer path stores **every** KV block of reachability-limited groups (sliding-window raw KV and compressor-state caches) to the external pool, while the non-layerwise path only stores blocks the KV cache managers consider reachable (segment-tail windows). For DeepSeek-V4-Flash (block_size=32, 43+1 layers, 6 KV cache groups) on 910C, a single 131072-token request allocates **91,168 pool keys (~45.6 GB DRAM)**, 98% of which sits in state/SWA groups whose reachable subset is only ~1.4 GB — a ~32x redundancy that floods the pool and the interconnect (~1.8 GB/s of writes during prefill). This PR aligns the layerwise path with the non-layerwise reachable-store semantics on all three sides, reusing the existing `AscendStoreCoordinator` masks: - **Save** (`_alloc_gvas_for_save` / `_process_save_for_layer_batch`): per-group store masks are computed once per scheduler step; keys are allocated and transfer ranges generated only for reachable blocks (the trailing partial block rides the last mask-allowed run). - **Hit check** (`_get_layerwise_hit_tokens`): only reachable blocks are queried per group, and the hit length is derived via `coordinator.find_longest_cache_hit`, so sparse storage still yields correct prefix hits (a contiguous walk would report 0 hits). - **Load** (`_prepare_load_gvas` / `_process_load_for_layer_batch`): key info is fetched only for the load-mask subset, so blocks that are deliberately not pooled no longer trigger the multi-group load failure path; GVA positions are filled per queried block index. Full-attention groups (mask=None), the trailing partial block, and the non-hybrid / single-group fallbacks keep their previous behavior. Measured effect: a 131072-token request drops from ~45.6 GB to ~1.4 GB of pool DRAM with identical hit semantics. **Note**: this PR also includes the two commits from vllm-project#15830 (`feat(kv_pool): pass lease TTL to memcache batch_alloc` and `fix(kv_pool): align MTP final-block trim to lcm block boundary`); it supersedes both vllm-project#15830 and the closed vllm-project#15672, so a single review covers all three changes. ## Does this PR introduce _any_ user-facing change? Yes: layerwise mode with a hybrid KV cache layout (e.g. DSV4 compressed + SWA groups) now persists only reachable blocks to the external KV pool, drastically reducing pool memory usage and write bandwidth. Cache-hit behavior is unchanged. ## How was this patch tested? - Unit tests added/updated (all pass; 3 pre-existing environment failures on a no-vllm/no-npu Windows box are unrelated — verified failing on a pristine tree): - `tests/ut/distributed/ascend_store/test_metadata.py`: `masked_block_runs` helper (sparse/None/empty/out-of-range masks) - `tests/ut/distributed/ascend_store/test_pool_worker.py`: mask-driven key allocation, multi-run transfer ranges, partial-block riding the last run, load-mask key queries, load range splitting - `tests/ut/distributed/ascend_store/test_pool_scheduler.py`: reachability-aware hit check (sparse SWA storage still yields full hits, hit stops where a stored tail is missing, no-pool → 0, non-hybrid fallback unchanged) - Command: `python -m pytest -sv tests/ut/distributed/ascend_store/` - NPU e2e validation is in progress on an Ascend cluster (draft until verified): plan is to re-run the DSV4-Flash layerwise deployment and compare the `alloc_gvas` key counts per group against the reachable subset (expected drop: group-4 state cache 65536 → ~128 keys per 131072-token request), plus long-context hit-rate and accuracy checks. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: tyy0829 <1455207791@qq.com> Signed-off-by: tyy0829 <87685049+tyy0829@users.noreply.github.com>
…vllm-project#16317) ### What this PR does / why we need it? Allow external cache lookup for prompts shorter than two retention intervals. With a retention interval and transfer granularity of 4096, a 5083-token prompt can reuse an existing 4096-token checkpoint. The current early return reports zero without querying the cache. The production change only removes this seven-line retention-length guard. Existing checks for requests shorter than one transfer unit, actual cache availability, and MTP tail recomputation remain unchanged. Requests previously excluded by the guard can now incur a lookup even when the cache is empty. The MTP implementation is identical to the upstream base, including its aligned `final_block_start` calculation. MTP boundary tests confirm existing behavior; they are not evidence of another newly fixed MTP bug. ### Does this PR introduce _any_ user-facing change? Yes. Requests shorter than two retention intervals can reuse valid external checkpoints instead of being forced to miss. No API or configuration changes. ### How was this patch tested? - **Unit tests:** 50 passed in `tests/ut/distributed/ascend_store/test_pool_scheduler.py`. The retention regression covers MTP enabled/disabled, stored and missing checkpoints, and the existing early return below one transfer unit. Additional boundary coverage confirms the unchanged MTP behavior. - **Baseline control:** On the upstream base without this fix, the retention lookup regression fails (0 instead of 4096), while the added MTP boundary test already passes. - **Static checks:** Ruff check, Ruff format check, and `git diff --check` pass for the changed files. The full `bash format.sh ci` hook suite could not run because the runtime lacks `pre-commit`. - **NPU experiments rerun with upstream MTP logic:** DeepSeek-V4-Flash-w8a8-mtp, TP4 × DP2, MTP=1, retention interval 4096. The runtime uses the upstream MTP block verbatim and removes the retention guard. Five frozen AISBench inputs were replayed with four warmup requests and four serial measured requests per length, one output token per request. Measured external hits/queries and outputs match both the previous implementation and the previously measured local-prefix-only mode. | AISBench input_len | Actual prompt tokens/request | External hits/queries | Previously measured local prefix hits/queries | Hit rate | | --- | --- | --- | --- | --- | | 4096 | 4179 | 16384/16716 | 16384/16716 | 98.01% | | 5000 | 5083 | 16384/20332 | 16384/20332 | 80.58% | | 5004 | 5087 | 16384/20348 | 16384/20348 | 80.52% | | 8192 | 8275 | 32768/33100 | 32768/33100 | 99.00% | | 9000 | 9083 | 32768/36332 | 32768/36332 | 90.19% | The chat template adds 83 tokens in this setup. Counts exclude warmup. Direct AISBench commands were also rerun for all five lengths, with four successful requests and zero failures in both warmup and measured stages: ```bash python3 aisbench_test.py --input_len 5000 --output_len 1 --data_num 4 \ --concurrency 1 --request_rate 0 --dataset_type prefix_cache \ --repeat_rate 1.0 --prefix_test --dp 4 ``` Token-ID input tests reconfirmed that actual lengths 4096/5004/8192/9000 reuse 0/4096/4096/8192 tokens per request. A separate 5083-token input produced identical 64-token output for an external miss and subsequent hit. **Validation limitations:** Clean-checkout unit tests use the directory's existing `_mock_deps` stubs and `--confcutdir=tests/ut/distributed/ascend_store`; the runtime's parent conftest has an unavailable `vllm.v1.attention.ops.pcp` dependency. The NPU runtime is the existing `bdab8bde86aad5eba3b362f923a3ff03b2c0e772` deployment with the upstream aligned MTP block and the retention guard removed, retaining its existing memcache backend changes; it is not a full deployment of the latest main. Cold concurrent warmup can differ between local and external caching due to publication timing and DP cache scope. These checks do not establish general model accuracy, lookahead KV safety for different continuations, graph-mode coverage, or statistical performance equivalence. This PR leaves upstream MTP handling unchanged. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
## Auto-Translation Summary Translated **13** file(s): - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/community/slash-commands.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/context_parallel.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/quantization.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/testing.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/getting_started/installation/install_vllm_ascend.inc.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_multi_node.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_single_node.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V4-Pro.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/configuration/additional_config.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/batch_invariance.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/context_parallel.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/kv_pool.po</code> - <code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/speculative_decoding.po</code> --- [Workflow run](https://github.com/vllm-project/vllm-ascend/actions/runs/34436986441) - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: wangxiyuan <wangxiyuan@users.noreply.github.com> Co-authored-by: wangxiyuan <wangxiyuan@users.noreply.github.com>
…ast (vllm-project#16094) ### What this PR does / why we need it? Add KV layer parallelism (KVPP). Each target layer's KV cache is persisted on one TP rank, and its full contiguous memory buffer is broadcast before computation. Two reusable scratch buffers enable one-layer-ahead prefetch without adding a custom operator. Supported features: - Model Runner V1 and V2. - TP, EP, and PP with layer ownership assigned within each PP stage. - Chunked prefill, prefix caching, and asynchronous scheduling. - LI-C8 and SFA-C8 cache layouts. - Fixed-step MTP, with MTP layers excluded from KVPP partitioning. This PR does not include graph mode, PCP, or DCP support. Prefill/decode disaggregation is not supported yet; its integration will be submitted after the Mooncake refactor is complete. ### Does this PR introduce _any_ user-facing change? KVPP is disabled by default and can be enabled through `enable_kvpp` in the additional configuration. ### How was this patch tested? Validated multiple combinations of the features above, including prefix cache hits and cache-block reuse, PP layer partitioning, and MTP. GPQA accuracy evaluation is in progress. Accuracy: | Dataset | Correct | Incorrect | Accuracy | |---|---:|---:|---:| | GPQA Diamond (194 scored questions) | 169 | 29 | 85.35% | Test configuration: GLM-5.2 on Ascend A3 (16 dies), TP16 + EP + DSA-CP, chunked prefill with `max_num_batched_tokens=32768`. | Input length | KVPP disabled | Full-layer Broadcast | Overhead | Increase | |---|---:|---:|---:|---:| | 64K | 8.41 s | 8.54 s | 0.13 s | 1.55% | | 128K | 17.20 s | 18.35 s | 1.15 s | 6.69% | Configuration: GLM-5.2-W4A8C8, TP16 + EP, chunk size 16K, max model length 32K, max sequences 32, MTP1, LI-C8/SFA-C8, prefix caching, FULL_DECODE_ONLY, GPU memory utilization 90%. Cache capacity was automatically calculated without a block-count override. | Configuration | KV cache capacity (tokens) | Capacity multiplier | |---|---:|---:| | KVPP disabled | 505,472 | 1.00× | | Full-layer Broadcast | 4,131,968 | 8.17× | Equivalent KV memory reduction at the same token capacity: **87.77%**. This is not a reduction in total device memory usage. ### Performance on Ascend A5 KV cache capacity with the same configuration as the pooling benchmark above, using `gpu_memory_utilization=0.90`. Capacity was automatically calculated without a block-count override. | Configuration | KV cache capacity (tokens) | Capacity multiplier | |---|---:|---:| | KVPP disabled | 235,008 | 1.00× | | Full-layer Broadcast | 1,467,904 | 6.25× | Equivalent KV memory reduction at the same token capacity, estimated from the capacity ratio: **83.99%**. This is not a reduction in total device memory usage. Test configuration: GLM-5.2-W4A8C8 on a single Ascend A5 node (8 dies), TP8 + EP + DSA-CP, chunked prefill, prefix caching, asynchronous scheduling, LI-C8, Model Runner V1, eager mode, and `max_num_seqs=12`. MTP, SFA-C8, and FlashComm1 were not enabled. #### TTFT Input lengths: 32K / 64K / 128K; output length: 1 token; prefix cache hit rate: 0%. Each result is the mean of 4 requests at client concurrency 1. | Token budget | Input length | KVPP disabled | Full-layer Broadcast | Overhead | Increase | |---|---|---:|---:|---:|---:| | 16K | 32K | 2.650 s | 3.046 s | +0.395 s | +14.92% | | 16K | 64K | 5.371 s | 6.518 s | +1.147 s | +21.35% | | 16K | 128K | 11.017 s | 13.582 s | +2.565 s | +23.28% | | 32K | 32K | 2.623 s | 2.614 s | -0.009 s | -0.33% | | 32K | 64K | 5.295 s | 5.375 s | +0.080 s | +1.52% | | 32K | 128K | 10.901 s | 11.268 s | +0.367 s | +3.36% | Overhead and percentage changes are calculated from unrounded measurements. The 32K token budget achieved lower TTFT for both configurations and was used for the throughput tests. #### Prefill throughput Configuration: `max_num_batched_tokens=32768`, `max_num_seqs=12`, client concurrency 12, 40 requests, 128K input tokens and 1 output token per request. KV pooling was disabled. | Prefix cache hit rate | KVPP disabled (input tokens/s) | Full-layer Broadcast (input tokens/s) | Change | |---|---:|---:|---:| | 0% | 12,492.4 | 12,205.7 | -2.30% | | 90% | 116,516.2 | 110,224.8 | -5.40% | Throughput is calculated as total input tokens divided by benchmark duration, including cached input tokens. The measured cache hit rate in the 90% scenario was 89.9414% for both configurations. All four throughput runs completed 40/40 requests successfully, with zero preemptions. Test configuration: GLM-5.2-W4A8C8 on a single Ascend A5 node (8 dies), TP8 + EP + DSA-CP, chunked prefill, prefix caching, asynchronous scheduling, LI-C8, Model Runner V1, and eager mode. `max_num_batched_tokens=32768`, `max_num_seqs=12`. KV pooling was enabled in both configurations using `AscendStoreConnector` with the Memcache backend, `kv_role=kv_both`, `use_layerwise=false`, and `load_async=true`. CPU pool capacity was configured as 32 GB per die. Test scenario: prefill throughput with multiple reusable prefixes. Sixteen independent prefix families were warmed before measurement, exceeding the HBM KV cache capacity in both configurations. The workload contained 40 requests: eight frequently reused families with four requests each, plus eight families with one request each. Each request had 128K input tokens, approximately 90% shared prefix, and 1 output token, at client concurrency 12. Both configurations used identical request data and ordering, retaining HBM cache after warmup to compare HBM residency and CPU pool loading. | Prefix cache hit rate | Configuration | Throughput (input tokens/s) | HBM cache hit rate | CPU pool cache hit rate | Throughput change | |---|---|---:|---:|---:|---:| | 90% (multiple prefixes) | KVPP disabled | 66,362.5 | 8.46% | 81.48% | — | | 90% (multiple prefixes) | Full-layer Broadcast | 97,318.3 | 45.46% | 44.48% | +46.65% | Throughput includes cached input tokens. Cache hit rates are relative to total input tokens; the combined hit rate was 89.94% in both configurations. Both runs completed 40/40 requests without preemption. This benchmark measures prefill throughput with 1 output token per request. ### Performance on Ascend A5 (dual-node PP) Test configuration: GLM-5.2-W8A8C8-mxfp8 on two Ascend A5 nodes (8 dies per node), TP8 + PP2 with a 38/40 layer split, EP, DSA-CP, chunked prefill, prefix caching, asynchronous scheduling, LI-C8, Model Runner V1, and eager mode. `max_num_batched_tokens=32768`, `max_num_seqs=12`, and `gpu_memory_utilization=0.90`. MTP, SFA-C8, and FlashComm1 were not enabled. Both configurations used vllm-ascend `493ec2b` with the pooling restriction removed, and vLLM `b2f6858`, matching the single-node comparison baseline. #### TTFT Input lengths: 32K / 64K / 128K; output length: 1 token; prefix cache hit rate: 0%. Each result is the mean of 4 requests at client concurrency 1. | Token budget | Input length | KVPP disabled | Full-layer Broadcast | Overhead | Increase | |---|---|---:|---:|---:|---:| | 32K | 32K | 2.900 s | 2.905 s | +0.006 s | +0.19% | | 32K | 64K | 4.503 s | 4.887 s | +0.384 s | +8.53% | | 32K | 128K | 7.805 s | 8.665 s | +0.859 s | +11.01% | Overhead and percentage changes are calculated from unrounded measurements. #### Prefill throughput Configuration: client concurrency 12, 40 requests, 128K input tokens and 1 output token per request. KV pooling was disabled. | Prefix cache hit rate | KVPP disabled (input tokens/s) | Full-layer Broadcast (input tokens/s) | Change | |---|---:|---:|---:| | 0% | 21,395.2 | 18,742.3 | -12.40% | | 90% | 177,375.4 | 154,717.3 | -12.77% | Throughput is calculated as total input tokens divided by benchmark duration, including cached input tokens. The measured cache hit rate in the 90% scenario was 89.9414% for both configurations. All four throughput runs completed 40/40 requests successfully, with zero preemptions. #### Pooled prefill throughput KV pooling was enabled in both configurations using `AscendStoreConnector` with the Memcache backend, `kv_role=kv_both`, `use_layerwise=false`, and `load_async=true`. CPU pool capacity was configured as 32 GB per device. Test scenario: prefill throughput with multiple reusable prefixes. Thirty-two independent prefix families were warmed before measurement, totaling 3,772,416 tokens and exceeding the HBM KV cache capacity in both configurations. The measured workload contained 40 requests: sixteen frequently reused families with two requests each, plus eight families with one request each. The remaining eight families were used only during warmup. Each request had 128K input tokens, approximately 90% shared prefix, and 1 output token, at client concurrency 12. Both configurations used identical request data and ordering, retaining HBM cache after warmup to compare HBM residency and CPU pool loading. | Prefix cache hit rate | Configuration | Throughput (input tokens/s) | HBM cache hit rate | CPU pool cache hit rate | Throughput change | |---|---|---:|---:|---:|---:| | 90% (multiple prefixes) | KVPP disabled | 86,438.8 | 0.59% | 89.35% | — | | 90% (multiple prefixes) | Full-layer Broadcast | 125,699.3 | 53.87% | 36.07% | +45.42% | Throughput includes cached input tokens. Cache hit rates are relative to total input tokens; the combined hit rate was 89.94% in both configurations. Both runs completed 40/40 requests without preemption. #### KV cache capacity Capacity was obtained from the non-pooling service startup logs, using `gpu_memory_utilization=0.90`, without a block-count override. | Configuration | KV cache capacity (tokens) | Capacity multiplier | |---|---:|---:| | KVPP disabled | 555,392 | 1.00× | | Full-layer Broadcast | 2,972,416 | 5.35× | Equivalent KV memory reduction at the same token capacity, estimated from the capacity ratio: **81.32%**. This is not a reduction in total device memory usage. Environment note: a storage-link issue required temporary NFS forwarding through the second node. Both configurations used the same storage route, and model loading and warmup completed before measurement. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: chengruiqi (C) <c00913489@china.huawei.com> Signed-off-by: recky-c <ruiqicheng510@gmail.com> Co-authored-by: chengruiqi (C) <c00913489@china.huawei.com>
…A3 machine and obtains baseline results (vllm-project#15083) ### What this PR does / why we need it? This PR migrates the nightly single-node a3 test cases to the 560T A3 machine and obtains baseline results ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? run the cases nightly - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: guxin108 <1252896542@qq.com>
### What this PR does / why we need it? vLLM #47692 changed the hybrid DP load-balancing inference to treat an explicit `--data-parallel-start-rank 0` as meaningful. As a result, the primary node in these multi-node `internal_dp` configurations is now inferred as hybrid LB, while the remote node is intended to run headless for internal LB. PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to satisfy the resulting handshake, but that changed the topology covered by the tests. The regular multi-node internal-DP runner sends benchmark traffic only to the primary node. In disaggregated-prefill cases, the PD proxy was explicitly designed to exclude headless nodes and target one API endpoint per DP group. Switching these deployments to hybrid/external LB therefore changes the intended test coverage. Restore the intended internal-LB topology in all 17 affected nightly and weekly configurations: - remove `--data-parallel-start-rank 0` from the primary node of each DP group, allowing rank 0 to be inferred without enabling hybrid LB; - restore `--headless` on the remote node while retaining its non-zero DP rank offset. This keeps one API endpoint on the primary node and lets it schedule requests across all local and remote DP ranks. ### Does this PR introduce _any_ user-facing change? No. This only updates nightly test deployment configurations. ### How was this patch tested? - `git diff --check` - Parsed all 17 modified YAML files with PyYAML and verified their internal-DP node roles. - Multi-node NPU nightly and weekly jobs should validate the complete deployments. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
…5488) ### What does this PR do / why we need it? In the W4A8 MXFP quantization path of `npu_grouped_matmul_swiglu_quant` (`vllm_ascend/device/device_op.py`), the call to `torch.ops._C_ascend.npu_swiglu_group_quant` now passes a `group_index` instead of `None`. `npu_swiglu_group_quant`'s `group_index` input currently only supports the `count` type, and internally the operator uses `group_index` only for summation. We therefore pass the last value of the cumsum tensor (`group_list[-1:]`, i.e. the total token count across groups) as the `group_index`, which provides the operator with the count it needs. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Tested with DSv4 Flash - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: lcfenglinwan <lcfenglin@qq.com>
…rv1 (vllm-project#15196)" (vllm-project#16331) ### What this PR does / why we need it? revert "[feature] reduce cases that required DP padding in mrv1 (vllm-project#15196)" ### Does this PR introduce _any_ user-facing change? None. ### How was this patch tested? - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: zzzzwwjj <1183291235@qq.com>
…project#16241) ### What this PR does / why we need it? Refs vllm-project#15665 Support sparse MLA attention without RoPE inputs on A2/A3, as used by GLM-5.3-Flash. Update host tiling and the cube/vector kernel paths to handle a zero RoPE dimension, keeping the existing RoPE path available. This PR contains only the sparse-flash-attention operator change: 5 files, +309/-64 lines, one signed-off commit on main. Model integration and other pending PRs are excluded. ### Does this PR introduce _any_ user-facing change? Yes. The existing sparse flash attention operator accepts the NoPE query/key layout on A2/A3. No serving configuration or dependency-version change is introduced. ### How was this patch tested? - Python AST/Ruff, C++ clang-format, and `git diff --check` passed for the relevant changes. - The operator changes were included in an integrated A3 native-extension build. This is not a standalone numerical validation of this branch. - Added code comments only: the NoPE path is documented in English, no functional change beyond the operator update above. Standalone NPU numerical validation is still pending. - The A3 cache build failed during CMake's Torch import because backend autoload initialized torch_npu/Triton in the build container. The build environment issue and an A3 rebuild remain pending. A2 and 310P cache builds passed. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
### What this PR does / why we need it? update a3-560t for nightly_config.yaml ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? run the case nightly - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: guxin108 <1252896542@qq.com>
…m-project#13045) ### What this PR does / why we need it? Adds **Gemma4 MTP speculative decoding** support on Ascend NPU (A5 / 950PR). Gemma4 (`gemma-4-31B-it` + `gemma-4-31B-it-assistant`) is a **multi-group KV cache** model: layers are split into a `sliding` group (head_dim=256) and a `full_attention` group (global head_dim=512). The MTP draft reuses the target's architecture and **shares the target's per-layer KV cache** (sliding for layer 58, global for layer 59), so each draft attention group must read K/V from its own physical block layout. Upstream vLLM supports this on CUDA via `Gemma4Proposer`. This PR adds the Ascend path as a **thin wrapper** of the upstream proposer — no rewrite of vLLM internals, no forking of base-class draft-loop methods. | file | change | role | |---|---|---| | `spec_decode/gemma4_proposer.py` | **+new (150)** | `AscendGemma4Proposer` — multiple inheritance of upstream `Gemma4Proposer` + `AscendSpecDecodeBaseProposer`. Overrides only: eager centroids sampling, `_sync_kv_sharing_target_to_impl`, per-group block table swap (`attn_update_stack_num_spec_norm` wrapper), `build_draft_attn_metadata` with FIA SpecDecoding state. | | `spec_decode/llm_base_proposer.py` | mod (+31/-11) | (1) `constant_draft_positions` guards in `_run_merged_draft` and `attn_update_stack_num_spec_norm` — restores upstream `SpecDecodeBaseProposer` semantics (vllm/v1/spec_decode/llm_base_proposer.py:635,693,705), no-op for existing proposers (flag defaults False upstream). (2) `_build_multi_group_graph_capture_metadata` hook (returns None for existing proposers) for per-group ACL graph capture. (3) `_propose` step-update loop iterates `attn_group.layer_names`, matching `build_draft_attn_metadata`'s existing iteration. | | `ops/rotary_embedding.py` | mod (+22/-1) | Q-only RoPE: when `key is None` (Gemma4 MTP sliding layers, K/V from target cache), use a throwaway key buffer with the regular rotary implementation (zero-kv-head dummy on the Triton path). | | `worker/model_runner_v1.py` | mod (+19/-2) | Wire `AscendGemma4Proposer` into the drafter type union, per-group block table capture, and kv-cache / cudagraph-key asserts. | | `spec_decode/__init__.py` | mod (+3) | Route `mtp` + `use_gemma4_mtp()` to `AscendGemma4Proposer`. | | `tests/ut/spec_decode/test_gemma4_proposer.py` | **+new (130)** | Unit tests: proposer routing, KV-sharing sync, per-group metadata. | | `tests/ut/ops/test_rotary_embedding.py` | mod (+18) | Q-only RoPE test. | ### Does this PR introduce _any_ user-facing change? Additive only. Users can now run Gemma4 MTP on Ascend with `--speculative-config '{"method":"mtp",...}'`. No change to existing models. ### How was this patch tested? #### Acceptance Gemma4-31B-it MTP, k=3, greedy (temp=0), `FULL_DECODE_ONLY` graph mode, coding prompt suite (5 prompts × 3 = 15 reqs). Metric: `vllm:spec_decode_num_accepted_tokens_per_pos` ÷ `vllm:spec_decode_num_drafts`, scraped from `/metrics`. | platform | backend | TP | pos0 | pos1 | pos2 | agg | accepted/step | |---|---|---|---|---|---|---|---| | **Ascend A5 (950PR) —TP4DP1** | vllm-ascend | 1 | **98.2%** | **95.4%** | **90.8%** | **94.8%** | **2.85/3** | | **Ascend A5 (950DT) —TP1DP4** | vllm-ascend | 1×4 | 100.0% | 98.8% | 91.9% | 96.9% | 2.91/3 | | NVIDIA L20 — CUDA reference | upstream vLLM | 2 | 98.2% | 95.5% | 90.0% | 94.6% | 2.84/3 | Ascend now matches the CUDA reference. Here is the prompts: ```python #{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef binary_search(arr, target):\n \"\"\"Return the index of target in sorted array arr, or -1 if not found.\"\"\"", "max_out_len": 512, "answer": ""} #{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef merge_sort(arr):\n \"\"\"Sort the array arr in ascending order and return it.\"\"\"", "max_out_len": 512, "answer": ""} #{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef lru_cache_get(cache, key):\n \"\"\"Return value for key from an OrderedDict-backed LRU cache, or None. Mark as recently used on hit.\"\"\"", "max_out_len": 512, "answer": ""} #{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef is_balanced(s):\n \"\"\"Return True if the string s has balanced parentheses/brackets/braces, else False.\"\"\"", "max_out_len": 512, "answer": ""} #{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef flatten(nested):\n \"\"\"Yield elements from an arbitrarily nested list, depth-first.\"\"\"", "max_out_len": 512, "answer": ""} ``` #### Accuracy - Matches the raw target performance - Tested on the **gpqa_diamond** dataset by using Evalscope 1. Accuracy Metrics:**Score: 0.7778**,Sample Size: 198 samples 2. Average Latency: 15.0612 seconds 3. Average Throughput: 65.59 tokens/second (Avg Thpt) #### Performance (950PR TP=4,DP=1) **Overall Conclusion:** Tested on 4K input/1K output by Evalscope. The model demonstrates excellent scalability and stability. As concurrency increases, the overall throughput grows significantly while maintaining a consistent speculative acceptance rate, proving that the MTP (Multi-Token Prediction) mechanism is functioning effectively. ##### **1. Key Latency & Efficiency Metrics** * **TPOT (Time Per Output Token):** * At low concurrency (**Conc=1**), the average TPOT is very low (**15.0 ms**), indicating extremely fast generation. * As concurrency increases to **64**, the average TPOT rises to **81.0 ms**, which is expected due to increased system load, but remains within an acceptable range for high-throughput serving. * **Speculative Acceptance Rate:** * The acceptance rate remains remarkably stable across all concurrency levels, hovering around **71% - 72%**. This indicates that the MTP model's predictions are consistently accurate regardless of the system load. ##### **2. Throughput and Scalability** * **Completion Throughput (Output Speed):** * There is a clear linear increase in output speed as concurrency rises. * **Conc 1:** 59.46 tok/s $\rightarrow$ **Conc 64:** 345.05 tok/s. * **Total System Throughput:** * The **Overall Total Prompt throughput** scales from **237.85 tok/s** (Conc 1) up to **1380.35 tok/s** (Conc 64). * The **"Last 30s"** peak throughput reaches nearly **3,924 tok/s** at maximum concurrency, showing the model's ability to handle heavy bursts of traffic. ##### **3. Summary Table** | Concurrency | Avg TPOT (ms) | Spec. Accept Rate | Completion Throughput | Total Prompt Throughput | | :--- | :--- | :--- | :--- | :--- | | **1** | 15.0 | 71.5% | 59.46 tok/s | 237.85 tok/s | | **8** | 30.9 | 71.6% | 236.27 tok/s | 945.17 tok/s | | **16** | 43.3 | 71.7% | 299.15 tok/s | 1196.73 tok/s | | **32** | 79.4 | 71.5% | 330.30 tok/s | 1321.32 tok/s | | **64** | 81.0 | 71.6% | 345.05 tok/s | 1380.35 tok/s | #### Additional: Real-world Agent Prompt Validation (950DT TP=1,DP=4) Tested with a ~3.4K-token real tel-sales agent system prompt (300 output tokens, `vllm bench serve`, greedy, ignore-eos, TP=1): | Conc | TPOT p99 (ms) | Output tok/s | Spec accept | |---|---|---|---| | 1 | 9.6 | 102 | 63.4% | | 8 | 10.9 | 732 | 63.6% | | 16 | 12.2 | 1290 | 63.6% | | 32 | 14.3 | 2267 | 63.5% | | 48 | 16.0 | 2981 | 63.6% | | 64 | 17.2 | 3469 | 63.9% | Acceptance stays flat across concurrency; throughput scales linearly to 3469 tok/s at conc=64. Zero errors across the full sweep. Repro (copy-paste-able): ```bash ASCEND_RT_VISIBLE_DEVICES=0 vllm serve <gemma-4-31B-it> \ --served-model-name gemma4-31b-it-mtp \ --tensor-parallel-size 1 \ --speculative-config '{"method":"mtp","model":<gemma-4-31B-it-assistant>,"num_speculative_tokens":3}' \ --max-model-len 32000 --gpu-memory-utilization 0.85 \ --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \ --trust-remote-code --host 0.0.0.0 --port 8831 # then drive the server and scrape /metrics for per-position acceptance ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Wu Meng <wumeng@ascend-debug.local> Co-authored-by: Wu Meng <wumeng@ascend-debug.local> Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: XUE TONGYAO <xuetongyao2001@gmail.com>
### What this PR does / why we need it? The op has had no callers since vllm-project#7557 reverted the small-batch GMM optimization introduced in vllm-project#7100 (qwen3-next gsm8k 98 -> 91). All grouped-matmul paths run on torch_npu.npu_grouped_matmul. - delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api) - drop schema/impl registration from csrc/torch_binding.cpp - drop meta registration from csrc/torch_binding_meta.cpp - drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93) ### Does this PR introduce _any_ user-facing change? Yes, the `moe_grouped_matmul` operator is no longer available in the PyTorch bindings. ### How was this patch tested? No new tests were added as this is a code removal PR. Existing CI tests should pass. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: TangPeng <85704592@qq.com>
### What this PR does / why we need it?
#### Summary
Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.
| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |
Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.
#### Scope and version handling
- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.
#### File-by-file changes and upstream evidence
Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.
##### Dependency pin (1 file)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Device test adaptations (4 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
##### Existing CPU test adaptations (24 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Runtime adaptations (32 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
#### Interface review and limitations
The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.
That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.
### Does this PR introduce _any_ user-facing change?
Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.
### How was this patch tested?
- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.
[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290
[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
…lm-project#16387) Revert of PR vllm-project#16107 (merged onto `main`). Original PR: vllm-project#16107 Original author: @xqchen7 Merge commit: `fed28d4ce692173954cb887284af45dfa2284a4d` --- ### What this PR does / why we need it? adapt resource got a5 multi nightly,add openlibing.secret parameter to enable a5 multi nightly run ### Does this PR introduce _any_ user-facing change? eable developer run a5 multi nightly ### How was this patch tested? - vLLM main: vllm-project/vllm@b2f6858 Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…rate_deepseekv4-flash-w8a8-mtp.yaml ### What this PR does / why we need it? we add nightly kimi25_w4a8_step_3_7.yaml kimi-k2.6_max_model_len.yaml acceptace_rate_deepseekv4-flash-w8a8-mtp.yaml ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? run case nightly - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: wanglei727 <naie.wanglei@h-partners.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
… weights (vllm-project#16259) ### What this PR does / why we need it? Serving a native DeepSeek-V4 checkpoint whose config sets `num_hash_layers > 0` fails during weight loading: ``` KeyError: 'model.layers.0.mlp.gate.e_score_correction_bias' File "vllm_ascend/models/deepseek_v4/model.py", line 1319, in load_weights param = params_dict[name] ``` `DeepseekV4MoE` gives hash-router layers a `tid2eid` lookup table and deliberately leaves `e_score_correction_bias` unset, because those layers route text tokens purely through the table rather than through router scores: ```python if self.hash: self.gate.tid2eid = nn.Parameter(...) self.gate.e_score_correction_bias = None ``` The checkpoints, however, still ship a `gate.bias` tensor for every MoE layer. `load_weights` renames it to `gate.e_score_correction_bias` and then indexes `params_dict` unconditionally, so the unused hash-layer bias aborts startup. This PR skips the tensor when the model does not own the target parameter. The existing `gate.bias_vl` guard directly above is unaffected. Reproduced on `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` (`num_hash_layers: 3`, `n_routed_experts: 256`) on Ascend 950PR. The checkpoint carries all three tensors for the hash layers: ``` layers.0.ffn.gate.tid2eid layers.0.ffn.gate.bias layers.0.ffn.gate.bias_vl layers.0.ffn.gate.weight ``` ### Does this PR introduce _any_ user-facing change? No new flags or interfaces. Native DeepSeek-V4 checkpoints with hash-router layers now finish loading instead of raising `KeyError` at startup. ### How was this patch tested? Unit test added: `tests/ut/models/test_deepseek_v4_moe.py::test_hash_layer_router_bias_is_skipped_when_unused`. It feeds a router bias for one hash layer and one dense layer and asserts only the dense one is consumed, so the previous `KeyError` is a regression failure. ```bash pytest -sv tests/ut/models/test_deepseek_v4_moe.py ``` Manual check: `vllm serve` of `DeepSeek-V4-Flash-Vision-Exp` on one Ascend 950PR node reaches `Application startup complete` and answers `/v1/completions` requests, where it previously failed in `WorkerProc` init on every rank. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com> Co-authored-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com>
…llm-project#16214) ### What this PR does / why we need it? Refs vllm-project#15665 This PR enables Multi-Token Prediction (MTP) speculative decoding for GLM-5.3-Flash (`glm5_next`) on Ascend. The `Glm5NextMTPModel` implementation already exists in the GLM-5.3-Flash model code. This PR completes the integration required to select and run that model through vLLM's existing `deepseek_mtp` speculative decoding flow. The main changes include: - Register `glm5_next_mtp` as a supported MTP model type. - Normalize GLM-5.3-Flash draft configurations: - Convert `glm5_next` and `glm5_next_text` to `glm5_next_mtp`. - Select the existing `Glm5NextMTPModel` architecture. - Propagate `num_nextn_predict_layers` as `n_predict`. - Correct main-model graph selection for speculative decoding: - Treat a speculative batch as uniform decode only when every request has finished its prompt. - Prevent partial-prefill requests from incorrectly replaying a uniform decode graph. - Keep the main model graph-enabled while the MTP draft model runs eagerly through `enforce_eager`. - Add ModelSlim W8A8 quantization adaptations: - Register the fused MLP, MoE, MLA, and KDA projection mappings for `glm5_next_mtp`. - Map runtime prefixes containing `.mtp_block.` back to checkpoint-style quantization prefixes. - Resolve MTP expert prefixes stored under `model.language_model.layers.*`. - Support multimodal checkpoints whose text-model quantization entries use different `model.language_model.*` or `language_model.model.*` namespaces. - Reuse the `glm5_next` packed-module mapping for `glm5_next_text`. - Preserve expert layers marked as `FLOAT` as unquantized. - Add multimodal GLM-5.3-Flash weight mappings: - Flatten ModelSlim KDA `forget_gate` parameters to the runtime module layout. - Map attention and FFN hyper-connection parameters. - Map visual, language-model, embedding, and LM-head prefixes. - Ignore the standalone `rot.*` tensor because the exported rotation has already been folded into `model.visual.merger.down_proj.weight`. - Add unit tests covering speculative-config registration, graph dispatch, multimodal weight mappings, and ModelSlim quantization-prefix resolution. This allows GLM-5.3-Flash to use one-token MTP speculative decoding with the existing `deepseek_mtp` method. ### Does this PR introduce *any* user-facing change? Yes. Users can enable MTP speculative decoding for GLM-5.3-Flash with: ```bash --speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}' ``` For ModelSlim W8A8 checkpoints, the Ascend quantization method must also be specified explicitly: ```bash --quantization ascend \ --speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}' ``` The MTP draft model runs eagerly, while the main model can continue using decode graphs such as `FULL_DECODE_ONLY`. ### How was this patch tested? Formatting and pre-commit checks: ```bash bash format.sh ``` All hooks passed, including ruff, codespell, typos, clang-format, actionlint, gitleaks, shellcheck, and the repository-specific checks. Unit tests: ```bash pytest -q \ tests/ut/models/test_glm5_next_mtp.py \ tests/ut/worker/test_model_runner_v1.py ``` Result: ```text 51 passed, 2 skipped ``` ModelSlim unit tests: ```bash pytest -q tests/ut/quantization/configs/test_modelslim_config.py ``` Result: ```text 60 passed ``` Runtime validation was performed with: - GLM-5.3-Flash ModelSlim W8A8 checkpoint - Tensor parallel size: 8 - Expert parallel enabled - One speculative token - MTP draft model in eager mode - Main-model graph mode: `FULL_DECODE_ONLY` - Explicit `--quantization ascend` Example runtime options: ```bash --tensor-parallel-size 8 \ --enable-expert-parallel \ --quantization ascend \ --speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}' \ --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,8,16,32,64,96,128]}' ``` A single chat-completions request completed successfully and returned the expected answer `4` for `2 + 2`. The speculative-decoding metrics confirmed that MTP was active: ```text draft tokens: 31 accepted tokens: 28 ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: lrf-vm <lirongfan168@gmail.com> Signed-off-by: root <root@localhost.localdomain> Co-authored-by: root <root@localhost.localdomain>
…asks (vllm-project#16067) ### What this PR does / why we need it? Reduce small-operator overhead around Kimi K3 recurrent KDA on `main` by optimizing preprocessing copies and removing redundant framework padding cleanup after KDA and the output norm gate. - **Before KDA:** consume supported non-contiguous Q/K/V views from the fused projection directly, eliminating forced contiguous copies. Independent token/head strides flow through ACLNN, tiling, and the A2/A3 and A5 kernels. Unsupported layouts still materialize. - **After KDA and norm:** delete `_zero_padded_recurrent_output` on spec/non-spec outputs and `_zero_padded_output` after `o_norm`, including the now-unused live-token count calculation. These removals eliminate the corresponding `arange / comparison / where` chains, profiled as `Range / Less / Fill / SelectV2`. Norm output is copied directly into the existing result buffer. - Preserve state updates, mixed-token index copies, idle/dummy handling, and buffer initialization. No new kernel interface, switch, or replacement mask is needed. The deleted masks sanitize padding after recurrent state computation; the norm gate operates independently per token/head row. Padding rows inside the graph-shaped region no longer have a zero-value guarantee. Effective token computation is unchanged. ### Does this PR introduce _any_ user-facing change? Performance optimization of Kimi K3 KDA preprocessing and postprocessing. The Torch operator schema and contiguous-input behavior remain unchanged. ### How was this patch tested? - Accuracy validation (reported by the project owner): **GPQA Avg@5: 93.03**. - Added spec/decode/mixed framework regression cases with NaN padding, checking effective output rows, state updates, and output-buffer tail initialization. - All three cases passed on each branch in an isolated CPU harness executing the actual `_forward` and `_prepare_beta` bodies with mocked kernel/context dependencies (six cases total). This is not a full package or NPU test run. - Ruff check/format, Python AST parsing, `git diff --check`, and repository forbidden-import, boolean-context-manager, and long-function checks passed. - The project owner previously bypassed the padding masks and reported successful runtime validation. Full-model/NPU validation and profiling were not rerun for these commits. - Full `format.sh ci` was attempted on main and stopped during actionlint environment installation; the full lint suite did not complete. - The direct recurrent-operator regression covers distinct Q/K/V strides, non-contiguous state pools, guard holes, and numerical reference checks. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
…oject#15918) ### What this PR does / why we need it? This PR enables the bundled AscendC `MsaIndexScore` path for MiniMax-M3 on Ascend 950 (A5) with native FP8 index query/key tensors and an FP8 index-key paged cache. - Sync the operator implementation from [cann/ops-transformer#10672](https://gitcode.com/cann/ops-transformer/pull/10672) at `2e685d24ed2e54ea1f1547dc6e5ad29862f152ef`. - Add the Ascend 950 `arch35` FP8 kernels, tiling, Catlass dependencies, dtype registration, examples, and documentation. - Add `msa_index_score` to the Ascend 950 custom-op build list. - Route A5 prefill index scoring through `torch.ops._C_ascend.npu_msa_index_score`, while retaining the existing A5 Triton TopK and decode paths. - Cast Index-Q to E4M3 when the index-key cache is E4M3 and the query dtype does not already match. - Include the A5 wide-`block_table` fix: score columns beyond the 256-column UB window are flushed in windows, with width-257 FP8/BF16 regression coverage. The MiniMax-M3 A5 FP8 contract used by this integration is: query/key use `torch.float8_e4m3fn`, `scale=None`, and the operator emits FP32 scores. ### Does this PR introduce _any_ user-facing change? Yes. MiniMax-M3 FP8 inference on Ascend 950 uses the bundled AscendC MSA index-score implementation for prefill, while decode continues to use the existing A5 Triton path. The public Python API is unchanged. ### How was this patch tested? Static validation completed: - Verified all 54 imported operator files match PR vllm-project#10672 head `2e685d24` byte-for-byte (excluding its standalone `torch_extension`; vLLM Ascend keeps its in-tree Torch adapter). - `ruff check` passed for the modified MiniMax-M3 Python files and unit tests. - `ruff format --check` passed. - Python AST/compile checks passed. - `bash -n csrc/build_aclnn.sh` passed. - `git diff --check` passed. - Added unit assertions for A5 FP8 registration/build wiring and the width-257 windowed-flush regression. #### A5 prefill IndexScore operator performance Compared the AscendC `MsaIndexScore` kernel with the Triton `_index_block_score_kernel` on MiniMax-M3-MXFP8, using Ascend 950, TP4 × DP2, and the same P+D mixed-load request orchestration. Only prefill IndexScore device time is reported. Each value is the mean device duration per IndexScore kernel call across 8 ranks. MiniMax-M3 invokes IndexScore in 57 sparse-attention layers, so 57 calls form one prefill chunk. For 128K, only chunks captured by both implementations are compared to keep the context lengths matched. | Input length | Matched prefill scope | AscendC | Triton | Duration reduction | Speedup | | --- | --- | ---: | ---: | ---: | ---: | | 16K | Chunk 0 | 165.646 µs/call | 214.812 µs/call | 22.89% | 1.297× | | 128K | Chunk 0 | 166.359 µs/call | 213.549 µs/call | 22.10% | 1.284× | | 128K | Chunk 1 | 363.359 µs/call | 529.388 µs/call | 31.36% | 1.457× | | 128K | Chunk 2 | 560.355 µs/call | 837.372 µs/call | 33.08% | 1.494× | | 128K | First 3 matched chunks combined | 363.358 µs/call | 526.770 µs/call | 31.02% | 1.450× | These are matched operator-level profiling results, not end-to-end TTFT, throughput, or serving-latency improvements. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: zonghaoxin <z00946994@china.huawei.com> Signed-off-by: Haoxin Zong <534687988@qq.com> Co-authored-by: zonghaoxin <z00946994@china.huawei.com>
…m-project#16321) ### What this PR does / why we need it? Refs vllm-project#15665 Replace decomposed PyTorch mHC reductions with the existing Ascend `npu_hc_pre_v2` and `npu_hc_post` operators for GLM-5.3-Flash. Preserve post-mixing scale, optional input RMSNorm, output shapes, and deferred residual mixing across layers. The model and native operators are already in main. This PR has no pending Flash source dependency; full-model validation combines the other Flash integration PRs and their prerequisites. ### Does this PR introduce _any_ user-facing change? Yes. GLM-5.3-Flash uses native fused mHC operators on Ascend, reducing decomposed operator launches without changing serving options. ### How was this patch tested? - Four cases in `tests/ut/models/test_glm5next_mhc.py` cover optional RMSNorm, post-mixing scales, deferred mixing, and preservation of the input residual. All 182 focused regression tests passed in the integrated tree. - A3 numerical comparisons against the original Torch path used actual GLM mHC weights and BF16 inputs at 1, 4, 17, and 3500 tokens. All 16 output comparisons passed; maximum NRMSE was 2.674e-5. - GLM-5.3-Flash-w8a8 passed eight smoke cases on A3 with TP8, expert parallelism, MTP=3, and FULL_DECODE_ONLY, including a 5264-token chunked prefill. A 3500-input/1500-output request also completed. - Profiling confirmed HcPre/HcPost execution during prefill and decode. - GitHub pre-commit, CPU unit tests, and the CI gate passed. Selected device jobs were skipped in that CI run. Hardware validation covers the GLM BF16 configuration with hc_mult=4 and hidden_size=4096. These results do not establish an end-to-end speedup or a full GSM8K result. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…ject#16252) ### What this PR does / why we need it? Refs vllm-project#15665 Extend the shared SFA backend to execute sparse MLA without RoPE. Add NoPE preprocessing and forward execution, physical-page cache addressing, graph-stable metadata buffers, and indexer-provided visible top-k lengths. Keep cache ownership with the indexer and mask unwritten graph-padding rows before output projections. Reuse the existing hardware-specific operators: - A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now merged. - A5: the DeepSeek-V4 sparse MLA interface in `vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN `cann_ops_transformer` operators. This PR contains shared attention integration. GLM KeyPool model routing is handled separately in vllm-project#16253. ### Does this PR introduce _any_ user-facing change? Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when the corresponding operators are available. Existing RoPE behavior and cache allocation/grouping are preserved. ### How was this patch tested? - The following files passed in a CPU integration run with 319 total passing tests, covering metadata/addressing, operator dispatch, NoPE forward behavior, and existing SFA regressions: ```bash pytest -q \ tests/ut/attention/test_sfa_nope_metadata.py \ tests/ut/attention/test_sfa_nope_forward.py \ tests/ut/attention/test_sfa_v1.py ``` - A focused rerun of the same three files on A3 passed 60 tests. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate on the current head. The focused unit tests use mocked operator calls and do not establish native numerical accuracy. A5 native operator validation remains outstanding. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…lm-project#16243) ### What this PR does / why we need it? Refs vllm-project#15665 Add Triton KeyPool state compression, paged cache writes, and pooled-key selection for GLM-5.3-Flash. Allow `IndexerWrapper` to select a model-provided backend while preserving checkpoint parameter names. Both operators reuse the shared `next_power_of_2` helper. Include single-operator accuracy tests and documentation covering formulas, parameters, layout constraints, graph replay, and test commands. Model-specific backend integration is handled separately. ### Does this PR introduce _any_ user-facing change? Yes. Provides the Triton pooled-indexing operators and a backend-selection hook. This PR does not enable the model backend by itself or change cache allocation and grouping. ### How was this patch tested? Validated on Atlas A3 with PyTorch 2.10.0, torch-npu 2.10.0.post4, and Triton-Ascend 3.2.0: - 10 tests passed in `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_glm5next_kpool_triton.py`: CPU-reference accuracy, historical windows, rollback, paging, noncontiguous storage, invalid slots, graph padding, empty inputs, and eager/graph replay. - 9 tests passed in `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_glm5next_pool_key_indexer_triton.py`: multiple requests, pool capacities, paged caches, token chunking, causal tails, and eager/graph replay with changing inputs. - 16 tests passed across `tests/ut/models/test_glm5next_indexer_backend.py` and `tests/ut/ops/test_mla.py`. The CI follow-up adds a type annotation to the test reference data without changing test behavior. Hardware coverage is limited to Atlas A3; these tests establish operator accuracy, not model-level throughput. - vLLM main: vllm-project/vllm@a97dacb Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) Revert of PR vllm-project#15908 (merged onto `main`). Original PR: vllm-project#15908 Original author: @zhiyu-wa Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920` --- **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@a97dacb Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…-project#16043) ### What this PR does / why we need it? Enable **MTP speculative decoding on Ascend 310P under Model Runner v2** for Qwen3.5 (dense / MoE), aligned with MRv1 contracts where possible. - Triton-free CPU paths: rejection sampling, draft input prep, host draft-step under ACLGraph capture - Target + draft **FULL_DECODE_ONLY** (K=1 SpecDecoding capture; K>1 per-step draft-decode FULL) - Hybrid GDN / prefix-cache CoW fixes for concurrent SpecDecoding FULL - Lean UT + e2e smoke (MRv1 eager + MRv2 FULL K=1) RFC: vllm-project#15577 ### Does this PR introduce _any_ user-facing change? No. On 310P with `VLLM_USE_V2_MODEL_RUNNER=1`, MTP can be enabled via: `--speculative-config '{"method":"mtp","num_speculative_tokens":K}'` ### How was this patch tested? - UT: `tests/ut/_310p/spec_decode/test_mtp_mrv2_310.py`, MTP gate / CoW in `test_model_runner_v2_310p.py` - E2E: `tests/e2e/pull_request/one_card/_310p/test_spec_decode_mtp_310p.py` - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: Thiagor2002 <13476117628@163.com>
…ct#16189) ### What this PR does / why we need it? Nightly `qwen3-vl-235b-a22b-instruct-w8a8` (A3, `FULL_DECODE_ONLY` ACLGraph) fails because internal-router MoE recomputes logits with: `gate.weight.to(torch.float32)` That emits a per-forward `aclop Cast`. DeepSeek-V4 and 310P already avoid this by setting `gate.precast_fp32_weight = True` so `AscendUnquantizedLinearMethod.process_weights_after_loading` materializes `weight_fp32` at load time. This PR applies the same precast in `AscendMoERunner.__init__` for all Ascend platforms, and routes the hot path through `_gate_weight_fp32()` so ACLGraph capture no longer depends on a live Cast. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - Code-path review against 310P `AscendMoERunner310` and DeepSeek-V4 `precast_fp32_weight`. - Nightly should re-run `Qwen3-VL-235B-A22B-Instruct-W8A8`. Made with [Cursor](https://cursor.com) - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: LQD <1107297340@qq.com> Signed-off-by: liaoqidan <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
### What this PR does / why we need it? enable mrv2 dspark e2e test ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? CI passed with new added/existing test. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: wxsIcey <1790571317@qq.com>
…oject#16232) ### What this PR does / why we need it? Documents the scheduling limitations of batch invariance in the user guide: chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported. These features are not disabled automatically, so the guide now instructs users to explicitly disable chunked prefill and prefix caching, and pairs `--block-size 128` with the chunked prefill disabling. The online/offline inference examples and the batch invariance e2e test are updated with the same config. ### Does this PR introduce _any_ user-facing change? Documentation-only update plus an e2e test configuration adjustment. ### How was this patch tested? Doc-only change; the e2e test config (`test_batch_invariant_tp4.py`) follows the documented settings. CI verification is sufficient. - VLLM_WORKER_MULTIPROC_METHOD=spawn pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ```bash [DEBUG] Final model_config: {'model_name': ..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} [DEBUG] vllm_runner fixture - model_config: {'model_name': '..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} ... PASSEDINFO 09-10 20:42:39 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1350] [shutdown] EngineCore: trigger received signal=SIGTERM (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1501] [shutdown] EngineCore: start mode=abort timeout=0s (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1532] [shutdown] EngineCore: request processing complete; starting resource teardown (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1363] [shutdown] EngineCore: exiting busy loop (EngineCore pid=64849) INFO 09-10 20:42:39 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=4 (Worker_TP0 pid=65187) INFO 09-10 20:42:39 [multiproc_executor.py:836] Parent process exited, terminating worker queues WARNING 09-10 20:42:44 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1 WARNING 09-10 20:42:44 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 64849 (EngineCore pid=64849) WARNING 09-10 20:42:44 [multiproc_executor.py:484] [shutdown] Executor: workers still running after grace period; sending SIGTERM count=3 [INFO] Model cache cleared =============================== warnings summary =============================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html ================== 1 passed, 15 warnings in 118.88s (0:01:58) ================== /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com>
…m-project#16422) ### What this PR does / why we need it? Adds the native and Triton operator layer required by DeepSeek V4.1 on Ascend: - Quant Lightning Indexer V2 and metadata operators - Sparse Flash MLA and metadata operators - compressor, index preparation, query quantization, Engram INT8, and required MoE operator updates - operator documentation, native examples, static checks, and operator-level unit/e2e tests This is the operator-only first part of the DeepSeek V4.1 enablement. Framework/model integration is intentionally split into dependent PR vllm-project#16423. ### Does this PR introduce _any_ user-facing change? Yes. It exposes the native operator capabilities required by DeepSeek V4.1 serving, but does not register or enable the model integration by itself. ### How was this patch tested? - operator static checks: 8 passed - focused related unit tests: 342 passed, 1 skipped - broad related unit tests: 855 passed, 12 skipped - clean native rebuild on Ascend A3: passed - final single-card NPU operator regression: 359 passed - pre-commit scoped checks: Ruff, Ruff format, codespell, typos, clang-format, markdownlint, and forbidden-import check passed End-to-end model validation is recorded in dependent PR vllm-project#16423. - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Resolve the patch_use_v2_model_runner conflict by keeping the Ascend whitelist default (apply_v2_model_runner_config_patch) and folding in main's v0.28.0 PCP unsupported-feature exception. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main.
Replace the env-only
use_v2_model_runneroverride with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). ExplicitVLLM_USE_V2_MODEL_RUNNERstill wins when set.Conflict resolution vs current
main: keep the whitelist default (apply_v2_model_runner_config_patch) and fold in the v0.28.0 PCP unsupported-feature exception from main. Do not restore the env-only property or the non-310P upstream_validate_v2_model_runnerpath; Ascend already owns that decision inmrv2_utils.Upstream PR: vllm-project#16203
Does this PR introduce any user-facing change?
Yes. Model Runner V2 is enabled by default for whitelisted architectures/features when Triton is available and the platform is not 310P.
VLLM_USE_V2_MODEL_RUNNERstill overrides that default.How was this patch tested?
tests/ut/test_mrv2_utils.py,tests/ut/patch/platform/test_patch_use_v2_model_runner.pyupstream/maininto this head; targeted CPU UTs for the whitelist and patch files