Skip to content

[Feat][Platform] Enable model runner V2 by default via whitelists - #144

Closed
yjyang62 wants to merge 76 commits into
cursor/vllm-ascend-main-deaefrom
enable-mrv2-whitelist-deae
Closed

yjyang62 wants to merge 76 commits into
cursor/vllm-ascend-main-deaefrom
enable-mrv2-whitelist-deae

Conversation

@yjyang62

@yjyang62 yjyang62 commented Sep 10, 2026 •

Copy link
Copy Markdown
Owner

What this PR does / why we need it?

Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set.

Conflict resolution vs current main: keep the whitelist default (apply_v2_model_runner_config_patch) and fold in the v0.28.0 PCP unsupported-feature exception from main. Do not restore the env-only property or the non-310P upstream _validate_v2_model_runner path; Ascend already owns that decision in mrv2_utils.

Upstream PR: vllm-project#16203

Does this PR introduce any user-facing change?

Yes. Model Runner V2 is enabled by default for whitelisted architectures/features when Triton is available and the platform is not 310P. VLLM_USE_V2_MODEL_RUNNER still overrides that default.

How was this patch tested?

  • Unit tests added/updated: tests/ut/test_mrv2_utils.py, tests/ut/patch/platform/test_patch_use_v2_model_runner.py
  • Merge-conflict resolution: local merge of upstream/main into this head; targeted CPU UTs for the whitelist and patch files
Open in Web Open in Cursor 

…ng (vllm-project#16146)

### What this PR does / why we need it?
This PR aims to fix kimi k3 dspark kv grouping when not enabling prefix
cache. The problem is that patch_mamba_config.py aligns ssm to k, which
is from only target model without draft model. Since kimi k3 dspark may
use gqa, kv cache layout thus can be problematic. That is, when
aligning, mamba considers mla without gqa, so finally v in gqa can
overlap ssm in mamba. Example is shown below:

```
Mamba: |   Conv   |               SSM               | Padding |
MLA:   | Padding  |            K(latent)            |  RoPE   |
GQA:   |    Padding     |        K         |        V         |
```

When enabling prefix cache, kimi k3 has its own kv grouping logic where
no tensor is shared by mamba and gqa. So we just fall back into it.

### Does this PR introduce _any_ user-facing change?
N/A

### How was this patch tested?
by ci

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Zetong Li <slippersss@126.com>
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch 2 times, most recently from 675859b to c27a1b5 Compare September 10, 2026 02:43
wangx700 and others added 13 commits September 10, 2026 10:59
…2.0.0 (vllm-project#15956)

### What this PR does / why we need it?

ref to vllm-project#10072

Upgrade the batch invariance operator run package and PyTorch extension
archive to 2.0.0. Map ascend950 variants to the 950 package and enable
operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced
PIECEWISE configuration from the online/offline examples and the TP4
batch invariance test, preserving its capture sizes and updating the
graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other
dtypes fall back to the saved native torch.sum implementation, avoiding
recursive calls through the patched Tensor.sum. Add parameterized
regression tests for supported and unsupported NPU dtypes.

### Does this PR introduce _any_ user-facing change?
A5 source-built images install the batch invariance operators. Users
still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4
consistency test use the default graph configuration.

### How was this patch tested?
- pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
```
- VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation
--no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple
--trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after
installation run test_batch_invariant_tp4.py to verify success.
- Build the image with Dockerfile, and verification is OK.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
What this PR does / why we need it?
Temporarily disable test_qwen3_6_mtp.py in CI by adding it to
skip_tests. The test file is retained so it can be re-enabled later.
Does this PR introduce any user-facing change?
No. This only changes CI test scheduling.
How was this patch tested?
- Ran select_tests.py --explicit-e2e-tests
tests/e2e/pull_request/two_card/model_runner_v2/test_qwen3_6_mtp.py and
confirmed test_groups=[] and has_tests=false.
- git diff --check passed.
- Full formatting checks could not run because pre-commit was
unavailable.
No new tests were added because this change only updates the existing CI
skip list.
- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: liaoqingqing <liaoqingqing3@h-partners.com>
Co-authored-by: liaoqingqing <liaoqingqing3@h-partners.com>
### What this PR does / why we need it?
Currently DSACP with SP is implemented in a suboptimal way, Leading to
regressed performance during prefill.
We support further development for deepseek v4 with SP and DSACP.

SP chunk is implemented in model, while attention reduce scatter /
allgather is in decoder layer. This is a implementation similar to
upstream deepseek v4 SP.
### Does this PR introduce _any_ user-facing change?
As is required in the test, dsacp is only depend on itself. However,
current dsacp depend on flashcomm.
This PR enables dsacp to automatically enable sp.

### How was this patch tested?
**Current**
ruff check
DS v4 curl success
performance check
A/B test with main
Ablation with dsacp or SP
**On going**
### benchmark
model=DeepSeek-V4-Flash-w8a8
DP2, TP4, EP
```bash
vllm serve \
    --model /mnt/weight/DeepSeek-V4-Flash-w8a8-mtp \
    --served-model-name qwen \
    --host 127.0.0.1 \
    --port 8010 \
    --max_model_len 25600 \
    --data-parallel-size 2 \
    --tensor-parallel-size 4 \
    --gpu-memory-utilization 0.92 \
    --max-num-seqs 8 \
    --enable-expert-parallel \
    --additional-config '{
        "enable_flashcomm1":true,
        "enable_shared_expert_dp":true,
        "enable_dsa_cp":true,
        "enable_cpu_binding":true,
        "multistream_overlap_shared_expert":true
    }'
```
```bash
python -m vllm.entrypoints.cli.main bench serve \
  --backend openai \
  --base-url http://127.0.0.1:8010 \
  --endpoint /v1/completions \
  --served-model-name qwen \
  --dataset-name random \
  --random-input-len 8192 \
  --random-output-len 2048 \
  --num-prompts 50 \
  --num-warmups 5 \
  --max-concurrency 8 \
  --metric-percentiles 50,90,99 \
  --seed 0
```
| branch | median TTFT (ms)| median TPOT (ms) |
|-------|-------|-------|
| this PR | 2796 | 43.75 |
| before PR vllm-project#13946 | 2777 | 42.93
| main | 3105 | 45.57 |
### Test Dspark and MTP
Tested with DP2, TP4, EP.
MTP model with w8a8 quantization.
DSpark model with w4a8 quantization.
| # | PR | spec | dsacp | accuracy | inference time |
|---|----|------|-------|----------|----------------|
| 1 | before | MTP=3 | off | 98.50 | 116s |
| 2 | after | MTP=3 | off | 98.50 | 123s |
| 3 | before | Dspark=7 | off | 98.50 | 200s |
| 4 | before | Dspark=7 | on | 98.50 | 188s |
| 5 | after | Dspark=7 | off | 98.00 | 200s |
| 6 | after | Dspark=7 | on | 98.00 | 188s |


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: MmMmaru <1843167357@qq.com>
### What this PR does / why we need it?
This PR aligns upstream function interface and fixes the
`apply_grammar_bitmask` Triton kernel to correctly handle adaptive
verification and per-request logit offsets. It introduces
`cu_num_logits` to resolve absolute logit rows on device and handles
inactive mapping entries safely.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with CI case.


- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: zouzy <zouzongyu@huawei.com>
### What this PR does / why we need it?
 delete --headless for weekly cases

### Does this PR introduce _any_ user-facing change?
no

### How was this patch tested?
run the cases weekly

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: guxin108 <1252896542@qq.com>
### What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill
replica. Reuse TP/group routing and existing completion tracking to
release KV.

### Does this PR introduce _any_ user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D
disaggregation using MooncakeConnectorV1.

### How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses
TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides
with MooncakeHybridConnector.

| Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output
cap hits + failed requests (limit) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) |
| Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2
(8192) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1
(8192) |

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and
cleanup were verified, with no duplicate DONEs or timeout reclamations.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it?

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: lxx <luxixiang@h-partners.com>
Co-authored-by: lxx <luxixiang@h-partners.com>
…m-project#15737)

### What this PR does / why we need it?

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

### Does this PR introduce _any_ user-facing change?

Documentation only. No code behavior change.

### How was this patch tested?

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
…llm-project#16114)

### What this PR does / why we need it?

Correct specific setup and first-inference ambiguities while preserving
the existing documentation structure:

- Pair the main development checkout with its verified vLLM commit, and
pair release CPU test checkouts with matching release tags.
- Identify commands that run inside the CPU test container and use the
correct nightly image names for NPU examples.
- Clarify that the manual 910B Ops installation example applies to A2
and refer other hardware to the official CANN guide.
- Use `max_tokens` for completion requests, report HTTP failures with
plain curl, and describe the expected nonempty completion.
- Fix the shared hardware table indentation in Quick Start so each
product series renders as a separate row.

The existing chapters, hardware tabs, background-serving flow,
startup/shutdown log examples, and four doctest stages are retained. No
test infrastructure changes or new performance sections are included.

### Does this PR introduce _any_ user-facing change?

Only the documented commands and explanations change. There are no
runtime/API changes. The examples retain the existing background-serving
and process-lookup commands.

### How was this patch tested?

- `bash format.sh ci`: passed.
- Strict English documentation build via `tools/rtd_build.sh`: passed.
- 27 doctest helper tests passed; all eight online blocks extract and
pass Bash syntax checks.
- Verified that the generated Quick Start hardware table contains four
separate product-series rows. The original PID lookup and stop commands
are retained.
- `git diff --check`: passed.

AI assistance was used for editing and local validation. NPU inference,
CANN installation, source builds, and hardware-dependent tests have not
been run on this macOS host. No performance measurements are claimed.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: sugm <sugm0521@163.com>
### What this PR does / why we need it?
1.Qwen3.5-122B Use case:Container names are formed by concatenating test
case names with other specific characters. Kubernetes restricts
container names to a maximum of 64 characters. Overly long test case
names cause Kubernetes scheduling failures. The current solution is to
shorten the test case name length.
2.Qwen3.6-35B Use case: Modify the dataset configuration.

### Does this PR introduce _any_ user-facing change?
NO

### How was this patch tested?
by the running the test

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: weixin <murongfengerxch@163.com>
…st list for SWR (vllm-project#16130)

## What this PR does / why we need it?

SWR rejects OCI image index (`application/vnd.oci.image.index.v1+json`)
produced by `docker buildx imagetools create` when merging OCI-format
single-arch images, returning `400 Bad Request: Invalid image, fail to
parse 'manifest.json'`.

**Root cause**: `docker buildx imagetools create` selects the output
media type based on the source manifests. When all sources are Docker v2
schema2 (`application/vnd.docker.*`), it produces a **Docker manifest
list** (`application/vnd.docker.distribution.manifest.list.v2+json`)
which SWR accepts. With OCI sources it produces an OCI image index,
which SWR rejects.

**Fix**: add `oci-mediatypes=false` to the build output so single-arch
images are built as Docker v2 schema2. This lets `docker buildx
imagetools create` produce a Docker manifest list directly, and
eliminates the skopeo-based format conversion entirely (which required
`sudo`/root not available on the self-hosted runner).

### Changes
- **Build output**: add `oci-mediatypes=false`
- **build-push-digest "Tag digest to prevent GC"**: replace `skopeo
copy` with `docker buildx imagetools create`
- **merge-image**: replace `skopeo copy` + `docker manifest create/push`
with direct `docker buildx imagetools create`
- **merge-image-temp**: same as merge-image
- **"Clean up temp tags"** (both jobs): remove `sudo apt-get install
skopeo`; guard `skopeo delete` with a `command -v` check so the job no
longer fails when skopeo is unavailable

## Does this PR introduce any user-facing change?

No. Only the CI image build/push workflow is changed.

## How was this patch tested?

- CI workflow changes verified locally (YAML valid, job structure
intact)
- No skopeo/sudo/apt-get dependencies remain in the merge path, so the
previous `skopeo: command not found` failure on the self-hosted
(non-root, no-sudo) runner is eliminated

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: JavaPythonAIForBAT <wuhejun@h-partners.com>
…ct#16121)

### What this PR does / why we need it?

This PR fixes the MiniMax-M3 prefill index top-k score preparation path
on Ascend A3.

On A3 with Triton-Ascend 3.2.2, the previous
`_prepare_prefill_topk_scores_kernel` could fail during sustained
MiniMax-M3 serving workloads such as TerminalBench.

Two device-side failure modes were observed during investigation:

* The original query-tiled 2-D masked tail-store implementation could
trigger an MTE/SMMU fault with `The DDR address of the MTE instruction
is out of range`.
* A follow-up implementation that converted the stores to
query-direction 1-D stores could trigger a Vector Core timeout/trap
during real serving workloads.

The original implementation vectorized score updates across the query
dimension, while the score tensor layout is:

```text
[num_index_heads, total_query_tokens, score_blocks]
```

where `score_blocks` is the contiguous dimension.

This PR changes the prefill score preparation to follow the tensor
layout directly:

* Initialize the score buffer to `-inf`, so invalid/padded score blocks
no longer need a separate tail-fill path.
* Keep the existing prefill QK score computation unchanged.
* Treat one `(query, index_head)` row as the logical preparation unit.
* Apply init/local forced priorities using block-contiguous stores along
the last dimension.
* Keep query handling as task scheduling rather than using query as the
vectorized memory-access dimension.
* Use an AIV-oriented program mapping so large/ragged prefills do not
create one launched program per `(query, head)` and exceed the Ascend
`coreDim <= 65535` launch limit.
* Preserve the existing public operator APIs and priority semantics:

  * init blocks: `1e30`
  * local blocks: `1e29`
  * local priority overrides init priority when they overlap
  * invalid blocks remain `-inf`

This removes the dynamic invalid-tail store path and avoids both the
previous `[query, block]` masked GM store pattern and query-direction
strided Vector stores.

The issue was reproduced on MiniMax-M3 running on Ascend A3 with DP=2 /
TP=8 across 16 devices under sustained TerminalBench workloads.

The workaround is based on behavior observed with Triton-Ascend 3.2.2.
The exact compiler/backend root cause has not been proven, so the
implementation should be re-evaluated after Triton-Ascend upgrades.

### Does this PR introduce *any* user-facing change?

No.

The public APIs, model behavior, output layout, top-k semantics, and
sparse-attention semantics are unchanged.

This PR only changes the internal Triton implementation and task mapping
used by MiniMax-M3 prefill index-score preparation on Ascend.

### How was this patch tested?

The patch was validated with standalone correctness cases,
TerminalBench-style workload sweeps, and real MiniMax-M3 serving
workloads on Ascend A3.

Correctness coverage includes:

* sparse-block boundaries: query lengths `127`, `128`, and `129`
* ragged query batches
* long invalid score tails
* init/local overlap
* `2048` / `2049` prefill boundary cases
* large ragged prefill workloads
* block counts matching the production workload range, including `480`
and `528`

The regression tests verify the complete prefill index path:

```text
index score
→ score preparation
→ torch.topk
→ invalid-index masking
```

Representative TerminalBench-style end-to-end results with 528 score
blocks:

```text
q=2048, prefix=0:
baseline  p50 = 2392.265 us
patched   p50 =  803.281 us

q=2048, prefix=32768:
baseline  p50 = 2006.960 us
patched   p50 =  847.873 us

q=2048, prefix=65512:
baseline  p50 = 1628.907 us
patched   p50 = 1537.166 us

q=4096, prefix=32768:
baseline  p50 = 1362.001 us
patched   p50 = 1315.651 us

q=8192, prefix=32768:
baseline  p50 = 2416.103 us
patched   p50 = 2336.767 us

q=32768, prefix=0:
baseline  p50 = 4760.416 us
patched   p50 = 4365.349 us
```

A large ragged TerminalBench-style workload also passed:

```text
query lengths:
[2048, 2048, 4096, 8192, 16384]

prefix lengths:
[0, 8192, 16384, 32768, 0]

total query tokens:
32768

score blocks:
528

baseline  p50 = 4809.989 us
patched   p50 = 4567.960 us
```

This case previously exposed the `coreDim` limitation for the direct
one-query-per-program A5-style launch (`81920 > 65535`). The patched
task mapping completes successfully.

Additional regression guards were added for the failure boundaries and
large/ragged workloads. Large TerminalBench-style cases are kept as
nightly/stability guards rather than regular lightweight CI cases.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
@yjyang62
yjyang62 force-pushed the enable-mrv2-whitelist-deae branch from c27a1b5 to 4261b44 Compare September 10, 2026 09:38
yzlbig and others added 13 commits September 10, 2026 19:10
…lm-project#16110)

# [Bugfix] Fix MXFP4 MoE scale layout and dtype handling

### What this PR does / why we need it?

This PR fixes several general compatibility issues in the Ascend W4A4
MXFP4 path.

#### Explicit MXFP4 dtype interpretation for grouped matmul

MXFP4 activations and weights are physically stored as `torch.uint8`.
Without explicit logical dtype arguments, `npu_grouped_matmul` may
interpret them as an unsupported `DT_UINT8`/`DT_UINT8` pair.

This PR explicitly passes:

~~~python
x_dtype=torch_npu.float4_e2m1fn_x2
weight_dtype=torch_npu.float4_e2m1fn_x2
~~~

The tensors are not converted or copied. These arguments tell the NPU
operator to interpret the packed `uint8` data as MXFP4.

#### Support odd MXFP4 scale group counts

The existing scale conversion assumes that the final scale dimension is
even before reshaping it into pairs:

~~~text
[G, N, K] -> [G, N, K // 2, 2]
~~~

Some valid tensor shapes produce an odd number of scale groups.
Reshaping these scales directly causes a tensor shape mismatch.

This PR pads the final scale dimension with one zero when necessary
before applying the existing paired layout conversion.

Existing checkpoints with an even number of scale groups continue to use
the original path.

#### Support uneven row-parallel MX scale shards

Row-parallel MX scale parameters use a ceil-divided local group count.
When the global number of scale groups is not divisible by the
tensor-parallel size, the final shard can contain fewer valid groups
than the allocated local parameter.

This PR adds an Ascend MX scale loader wrapper that pads the global
scale tensor before delegating to the existing row-parallel weight
loader.

Padding is applied only when:

~~~python
0 < padding < tp_size
~~~

This limits the behavior to the natural remainder introduced by
tensor-parallel ceil division. Unexpected checkpoint shape mismatches
are not hidden and continue to fail through the existing loader.

The compatibility logic applies only to:

- Ascend row-parallel linear layers
- MX quantization methods
- Per-group scale parameters

The generic vLLM weight loader remains unchanged.

### Does this PR introduce _any_ user-facing change?

Yes.

W4A4 MXFP4 models can now load and execute correctly when they use:

- Packed MXFP4 tensors requiring explicit logical dtype metadata
- An odd number of scale groups
- Scale dimensions that are not evenly divisible across tensor-parallel
ranks

There are no API, command-line interface, or configuration changes.

### How was this patch tested?

Unit tests were added for:

- Odd MXFP4 MoE scale group padding
- Explicit MXFP4 logical dtype arguments passed to grouped matmul
- Uneven tensor-parallel MX scale padding

Static validation was performed with:

~~~bash
git diff --check
~~~

The unit tests could not be executed in the local development
environment because PyTorch and an Ascend NPU runtime are unavailable.

The relevant tests can be run in a configured vLLM-Ascend development
environment with:

~~~bash
pytest -q tests/ut/quantization/methods/test_w4a4_mxfp4.py
pytest -q tests/ut/quantization/test_method_adapters.py \
  -k mx_scale_weight_loader
~~~

End-to-end validation should cover W4A4 MXFP4 model loading and
inference with multiple tensor-parallel configurations, including
configurations where scale groups cannot be divided evenly across ranks.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: “yzlbig” <2582195355@qq.com>
Signed-off-by: yzlbig <2582195355@qq.com>
…16220)

Use the verified vLLM main commit from pr_test.yaml, and allow the
coverage workflow to run on pull requests.

### What this PR does / why we need it?
The version of the scheduled task vllm is changed to be the same as that
of the PR.
### Does this PR introduce _any_ user-facing change?
no
### How was this patch tested?
The version of the VLLM used by the coverage rate scheduled task is the
same as that used by the PR.
- vLLM main:
vllm-project/vllm@b2f6858
### What this PR does / why we need it?

Fix a query-length mismatch during PCP speculative graph replay by
rebuilding draft prefill metadata only for DSA and SFA.

### Does this PR introduce _any_ user-facing change?

Yes, fixes speculative decoding graph replay failures.

### How was this patch tested?


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: leolee <yihao.li@huawei.com>
…mmand not found' (vllm-project#16272)

## Summary

- `Trigger quay.io sync` job failed with `gh: command not found` (exit
code 127) because the `linux-amd64-cpu-4` runner lacks the gh CLI
- Replace `run: gh workflow run ...` with `actions/github-script@v7`
which uses the GitHub API directly and doesn't depend on gh being
installed on the runner
- Verified: target workflow `sync-vllm-ascend.yml` exists and is active
in `ascend-gha-runners/sync-tools`

## Does this PR introduce any user-facing change?

No. CI-only change.

## How was this patch tested?

- YAML syntax validated
- Target workflow verified active via API
- createWorkflowDispatch API is the correct method for triggering
workflow_dispatch
- Full validation requires a real image build CI run

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: JavaPythonAIForBAT <wuhejun@h-partners.com>
…ct#16204)

## Summary
- Prefiller/decoder use `--data-parallel-start-rank` with multi-node DP.
On vLLM 0.28, a set start-rank infers hybrid LB; with `local=1` that
autoswitches to external LB, which forbids `--headless` on remotes.
- **Keep** `--data-parallel-start-rank`; **remove** `--headless` on both
prefiller and decoder secondary nodes so remotes expose API servers (PD
proxy can target them).

## Related failure
-
https://github.com/vllm-project/vllm-ascend/actions/runs/34372289942/job/102544223304

## Test plan
- [ ] DeepSeek-V3.2-W8A8-EP starts without `Remote engine must not use
--headless`

- vLLM main:
vllm-project/vllm@b2f6858

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…ction sampling kernels (vllm-project#16222)

## What this PR does / why we need it?

Fixes sporadic device error **507035** (MTE illegal GM address access,
fault symbol `load_gm_to_ubuf_1d_int32_t+0x6c`) in
`rejection_greedy_sample_triton`, and eliminates three more instances of
the same latent pattern.

**Root cause:** `tl.where` is a select instruction, not control flow —
*both arms are always evaluated*. The pattern

```python
start_idx = tl.where(offset == 0, 0, tl.load(cu_num_draft_tokens_ptr + offset - 1, is_greedy_mask))
```

therefore reads one int32 **before the buffer start** whenever lane
`offset == 0` is active in the load mask (e.g. all-greedy batches where
`is_greedy=None`). Whether this faults depends on the allocator layout:
in production it crashed only when `cu_num_draft_tokens` happened to sit
at an allocation boundary (`ptr - 4` falling into an unmapped page),
which explains the sporadic, TP-dependent failures. The fault was
reproduced in isolation: with `cu` placed at a boundary address
(`0x400040000000`, faulting address `0x40003ffffffc`), the old kernel
reports 507035 at synchronize while the fixed kernel passes 3/3 runs;
with padded placement the old kernel also passes — confirming the crash
is layout-dependent, not data-dependent.

**Changes:**

1. **Mask the prev-entry load itself** in the three tensor-path kernels
(lane 0's load is now predicated off at the hardware level, so no memory
transaction is issued):
- `rejection_greedy_sample_triton`: `mask=is_greedy_mask & (offset > 0),
other=0`
- `rejection_random_sample_kernel`: `mask=not_greedy_mask & (offsets >
0), other=0`
- `expand_kernel`: `mask=len_mask & (offset > 0), other=0` (lane 0 was
*unconditionally* active here — same exposure level as the crashed
greedy kernel)
2. **`sample_recovered_tokens_kernel`** (scalar `req_idx` path): clamp
the index with `tl.maximum(req_idx - 1, 0)` so the load address always
stays inside the buffer; `tl.where` then merely discards the (in-bounds,
unused) value for `req_idx == 0`. A 0-d masked load is not a verified
construct on triton-ascend, so the clamp is the safer equivalent.
3. **`other=0` on the `end_idx` loads**: without it, masked lanes
receive an *undefined* value (Triton only guarantees the memory is not
accessed). The per-request token loops (`for i in range(num_tokens1)`)
rely on `num_draft_tokens == 0` to skip inactive lanes, so an undefined
`end_idx` could turn into an out-of-bounds loop with both OOB reads *and
writes*. `other=0` makes this invariant deterministic instead of
dependent on backend implementation details.

Note: this matches the masking idiom already used by
`rejection_random_sample_block_verify_kernel` in the same file
(`prev_mask = not_greedy_mask & (offsets > 0)`).

## Does this PR introduce _any_ user-facing change?

No API or behavior change. For all in-bounds inputs the kernels produce
bit-identical results — the only eliminated behaviors are an
out-of-bounds read and reliance on undefined masked-load values. Users
benefit from the disappearance of sporadic 507035 device errors / hangs
in speculative decoding workloads (observed on an 8-machine PD
deployment under D-model graph mode; after the fix, 3 consecutive rounds
of 64-concurrent requests completed 192/192 with no device errors).

## How was this patch tested?

**Unit/e2e tests added**
(`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py`,
9 new parameterized cases, device=npu, compared against the PyTorch
reference implementations):

- `test_rejection_greedy_sample_triton_boundary` — 4 patterns:
`is_greedy=None` all-greedy (the always-OOB lane-0 path), all-greedy
tensor, mixed batch with first request greedy vs sampling; batch
includes all-match (bonus), first-token reject, mid reject, and
zero-draft-token requests; also asserts non-greedy rows stay untouched
(sentinel preserved)
- `test_rejection_greedy_sample_triton_boundary_multilane` — batch=256
randomized mixed batch forcing `BLOCK_SIZE > 1`, so lane 0 shares a
program block with other lanes; first request greedy/sampling both
covered
- `test_rejection_random_sample_boundary` — 3 patterns: all-random
(lane-0 always active in this kernel), mixed with first request sampling
vs greedy; includes forced all-accept (bonus), a `-1` draft placeholder
(always rejected → recovered token), and a zero-draft-token request

**Existing tests** for `expand_kernel` and
`sample_recovered_tokens_kernel` in the same file already exercise the
req0/lane0 path on every launch and serve as regression coverage for
fixes `expand_kernel` and `sample_recovered_tokens_kernel`.

**Fault-level reproduction** (isolated single-op experiment, documented
in the linked investigation): boundary-address allocation of `cu`
reproduces 507035 with the old kernel and passes with the fixed kernel,
for both the all-accept and rejection branches.

```bash
# Commands used
pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py -k "boundary"
pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py
ruff check && ruff format --check
```

> Note: the boundary tests verify semantic equivalence and lane-0
masking behavior; they cannot deterministically force the allocator to
place tensors at mapping boundaries, so the 507035 reproduction itself
relies on the dedicated single-op repro rather than pytest.

```
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_expand_kernel [14:59:22]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_sample_recovered_tokens_kernel [14:59:23]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-1-False] [14:59:23]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-1-True] [14:59:24]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-2-False] [14:59:24]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-2-True] [14:59:24]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-3-False] [14:59:25]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1-1024-3-True] [14:59:25]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-1-False] [14:59:26]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-1-True] [14:59:26]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-2-False] [14:59:26]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-2-True] [14:59:27]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-3-False] [14:59:27]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[256-1024-3-True] [14:59:27]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-1-False] [14:59:28]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-1-True] [14:59:28]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-2-False] [14:59:29]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-2-True] [14:59:29]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-3-False] [14:59:29]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[512-1024-3-True] [14:59:30]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-1-False] [14:59:30]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-1-True] [14:59:31]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-2-False] [14:59:31]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-2-True] [14:59:31]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-3-False] [14:59:32]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample[1024-1024-3-True] [14:59:32]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_spec_len_1_triton_kernel[False] [14:59:32]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_spec_len_1_triton_kernel[True] [14:59:33]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_kernel[False] [14:59:33]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_kernel[True] [14:59:34]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[all_greedy_none] [14:59:34]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[all_greedy] [14:59:34]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[mixed_first_greedy] [14:59:35]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary[mixed_first_random] [14:59:35]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary_multilane[True] [14:59:36]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_greedy_sample_triton_boundary_multilane[False] [14:59:36]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[all_random] [14:59:36]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[mixed_first_random] [14:59:37]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_random_sample_boundary[mixed_first_greedy] [14:59:37]
PASSED
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_rejection_sample.py::test_rejection_sampler_block_verify_triton_kernel[5-3-7-is_greedy0-uniform_probs0-recovered_token_ids0-bonus_token_ids0-target_probs0-None-draft_token_ids0-cu_num_draft_tokens0] [14:59:38]
PASSED[INFO] Model cache cleared
```

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: liyishi <1252651434@qq.com>
Co-authored-by: Zetong Li <slippersss@126.com>
### What this PR does / why we need it?

This PR adds block-major, range-based layerwise KV-cache transfer
support for the Mooncake backend in `AscendStoreConnector`.

Previously, Mooncake layerwise transfer stored every layer of each KV
block as an independent remote object:

```text
block × layer × rank
```

For large models, this creates a large number of remote keys and
objects, increasing:

- Remote metadata usage.
- Cache existence-check overhead.
- Mooncake put/get control-plane overhead.
- Object lifecycle and lease-management costs.

This PR changes the remote object layout so that all KV-cache layers
belonging to the same block and rank are stored in one Mooncake object:

```text
block × rank
```

The total KV-cache payload is not reduced. Instead, the layer dimension
is moved from the remote key into byte ranges inside a whole-block
object.

When a layer becomes ready, only the byte ranges belonging to that layer
are transferred. This preserves layer-by-layer transfer and
computation/communication overlap while reducing the number of remote
objects by approximately the number of KV-cache layers.

The main changes are:

- Add block-major Mooncake keys for scheduler-side cache lookup.
- Store all layers of one KV block and rank in a single remote object.
- Calculate the whole-object size and per-layer ranges from the actual
KV-cache layout.
- Add range-session APIs to the AscendStore backend abstraction.
- Map the new abstraction to the following Mooncake APIs:
  - `batch_put_session_start`
  - `batch_put_from_multi_buffer_ranges`
  - `batch_put_session_end`
  - `batch_put_session_revoke`
  - `batch_get_session_start`
  - `batch_get_into_multi_buffer_ranges`
  - `batch_get_session_end`
- Add `MooncakeSessionTracker` to manage put/get sessions across
chunked-prefill steps.
- Track multiple requests sharing the same prefix-cache key.
- Commit a remote object only after all required layer ranges have been
written successfully.
- Revoke incomplete remote objects when a layer transfer fails.
- Keep get sessions alive until the last request using the shared key
releases it.
- Preserve continuous-prefix semantics by requiring every saving rank of
a block to be available.
- Split large range transfers according to:
  - `layerwise_max_transfer_blocks`
  - `layerwise_max_transfer_bytes`
- Add fail-fast validation for required Mooncake range-session APIs.
- Reject currently unsupported topology and KV-cache layouts with
explicit errors.
- Add optional range-transfer audit logging through
`VLLM_ASCEND_KVPOOL_RANGE_DEBUG`.

The Mooncake write lifecycle becomes:

```text
batch_put_session_start
    → layer 0 range put
    → layer 1 range put
    → ...
    → final layer range put
    → batch_put_session_end
```

If any required range transfer fails:

```text
batch_put_session_start
    → one or more layer range puts
    → transfer failure
    → batch_put_session_revoke
```

The Mooncake read lifecycle becomes:

```text
batch_get_session_start
    → layer 0 range get
    → layer 1 range get
    → ...
    → final owner releases the key
    → batch_get_session_end
```

The change is limited to the Mooncake layerwise KV Pool path. Existing
non-layerwise and Memcache data paths are not changed.

### Does this PR introduce *any* user-facing change?

Yes.

Users can enable the new Mooncake block-major layerwise path with:

```json
{
  "kv_connector": "AscendStoreConnector",
  "kv_connector_extra_config": {
    "backend": "mooncake",
    "use_layerwise": true
  }
}
```

When this configuration is enabled:

- Mooncake stores one whole-block object per KV block and saving rank
instead of one object per block, layer, and rank.
- KV-cache data is still transferred layer by layer, but through ranges
inside the whole-block object.
- Remote object and metadata counts are reduced from approximately
`block × layer × rank` to `block × rank`.
- Startup fails early if the installed Mooncake client does not provide
the required range-session APIs.
- Unsupported configurations are rejected with explicit errors instead
of failing during transfer.

The currently unsupported configurations include:

- Prefill/Decode TP mismatch.
- Pipeline parallel size greater than one.
- Prefill context parallel size greater than one.
- Decode context parallel size greater than one.
- Hybrid or multiple KV-cache groups.

Detailed range-transfer audit logging can be enabled with:

```bash
export VLLM_ASCEND_KVPOOL_RANGE_DEBUG=1
```

There is no behavior change unless Mooncake layerwise transfer is
explicitly enabled.

### How was this patch tested?

#### Unit tests

The patch adds and updates CPU unit tests covering:

- Whole-block object sizing based on the actual KV-cache layout.
- Block-major key generation.
- Cache-hit validation across all saving ranks.
- Per-layer object offsets, local addresses, and transferred byte
counts.
- Key-major range construction.
- Splitting transfers by maximum block count.
- Splitting large individual ranges by maximum byte count.
- Put-session creation.
- Per-layer range writes.
- Final-layer commit ordering.
- Revocation of incomplete objects after transfer failures.
- Get-session creation.
- Shared-key ownership across multiple requests.
- Closing a get session only after its last owner releases it.
- Failed get attempts and retry cleanup.
- Chunked-prefill session persistence.
- Commit retry and terminal request cleanup.
- Full remote-cache hit handling.
- Partial-block replacement by a completed full-block key.
- Reloading committed prefixes during subsequent chunks.
- Required Mooncake range-session API validation.
- Validation of aligned per-key result counts.
- Rejection of unsupported TP-mismatch and topology configurations.
- Strict parsing and default behavior of
`VLLM_ASCEND_KVPOOL_RANGE_DEBUG`.

The main test files are:

- `tests/ut/distributed/ascend_store/test_mooncake_layerwise.py`
- `tests/ut/distributed/ascend_store/test_backend.py`
- `tests/ut/distributed/ascend_store/test_pool_scheduler.py`
- `tests/ut/distributed/ascend_store/test_pool_worker.py`
- `tests/ut/test_envs.py`

CI validation on the current PR head includes:

- `pre-commit`: passed
- CPU unit tests: passed

#### End-to-end performance test

The implementation was also tested with the following workload:

- Model: `Qwen3-30B`
- Data type: `W4A4`
- Input length: `16K`
- Output length: `500`
- External prefix-cache ratio: `75%`
- Concurrency: `1`
- Total requests per run: `32`
- KV Pool enabled: yes

Lower TTFT is better.

| Run | Model | Cards | Data type | Layerwise | KV Pool | Input | Output
| External prefix cache | Prefix cache | Prefill parallelism | Decode
parallelism | Concurrency | Total requests | TTFT (s) |

|---:|---|---:|---|:---:|:---:|---:|---:|---:|:---:|:---:|:---:|---:|---:|---:|
| 1 | Qwen3-30B | 2 | W4A4 | No | Yes | 16K | 500 | 75% | N/A | 1 | 1 |
1 | 32 | 0.44 |
| 2 | Qwen3-30B | 2 | W4A4 | Yes | Yes | 16K | 500 | 75% | N/A | 1 | 1 |
1 | 32 | 0.36 |
| 3 | Qwen3-30B | 4 | W4A4 | Yes | Yes | 16K | 500 | 75% | N/A | TP2 |
DP2 | 1 | 32 | 0.40 |

For the matching two-card configuration:

```text
Non-layerwise TTFT:       0.44 s
Layerwise TTFT, run :    0.36 s
```

The average TTFT improvement is:

```text
(0.44 - 0.36) / 0.44 × 100% = 18.2%
```

The two layerwise runs improve TTFT by approximately:

- `18.2%` for the `0.36 s` run.

The four-card `TP2/DP2` layerwise configuration achieved a TTFT of `0.40
s`. It is reported separately because there is no matching four-card
non-layerwise baseline in the current test data, so a direct speedup
comparison is not made.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: fangrongcan <17343701736@163.com>
)

Reverts vllm-project#12599

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: hejianping-00178005 <44997374+winson-00178005@users.noreply.github.com>
…lm-project#16224)

### What this PR does / why we need it?
This PR refactors the hardcoded `quant_mode` values in the DSA indexer
to be dynamically retrieved via
`DeviceOperator.get_dsa_indexer_quant_mode()`. This allows different
device adaptors (e.g., Non-A5 using INT8 and A5 using FP8) to return
their respective quantization modes, enabling proper support for
different hardware configurations.

### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with existing unit tests in
`tests/ut/models/test_deepseek_v4_indexer.py` which were updated to use
the new dynamic method.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: lenghuixing <lenghuixing@huawei.com>
Co-authored-by: lenghuixing <lenghuixing@huawei.com>
…llm-project#14496)

### What this PR does / why we need it?
Come from vllm-project#13773
In that PR, they refactor moe_mlp's methods into differnet quantization
method, so this RP,we refactor moe-mlp and quantization methods for
310p.
### Does this PR introduce _any_ user-facing change?
No,this is an internal refactor of the fused-moe and method for 310p and
its test coverage.
### How was this patch tested?
The following models are tested and can be curled successfully:
300I DUO:
Qwen3.5-35B-A3B-mtp

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: z00599867 <zhulinang@huawei.com>
Co-authored-by: z00599867 <zhulinang@huawei.com>
…6054)

Switch the scheduled main2main workflow from the a2 pool to the a3-16
pool:
- run on `linux-aarch64-a3-16-sh-001` with container
`cann:9.1.0-a3-ubuntu22.04-py3.12`
- `main2main_tests.json` now holds the reviewed 15-case selection 

<img width="2044" height="1811" alt="workflow"
src="https://github.com/user-attachments/assets/03e8123c-bf10-44a3-8479-af876ed358b4"
/>


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: wjunLu <wjunlu217@gmail.com>
…6294)

### What this PR does / why we need it?

Adds an experimental deployment tutorial and supported-model matrix
entry for `DeepSeek-V4.1-Flash`.

The tutorial follows the official model release and the current vLLM
Ascend adaptation work. It documents:

- the released model name and public Hugging Face/ModelScope
checkpoints;
- the two-node Atlas 800 A3 and four-node Atlas 800 A2 W8A8 topologies
(TP8/DP4/EP32);
- INT8 Engram storage, DSpark speculative decoding, and
`FULL_DECODE_ONLY` ACL Graph;
- Docker, source-build, multi-node launch, health, text, and image
verification commands;
- explicit accuracy, performance, context-length, and deployment
limitations.

The pre-release `Aurora` name and internal checkpoint paths are
intentionally omitted. The guide uses placeholders for the
Ascend-quantized checkpoint path and keeps the documented claims within
the current validation boundary.

Related adaptation work:
https://github.com/GDzhu01/vllm-ascend-v41-private

### Does this PR introduce _any_ user-facing change?

Yes. It adds an experimental DeepSeek-V4.1-Flash deployment guide and
advertises the documented support boundary in the model support matrix.

### How was this patch tested?

- `markdownlint-cli@0.45.0` passes for both changed Markdown files.
- `git diff --check` passes.
- `docs/hooks/nav_titles.py` passes Python syntax compilation.
- The new supported-model matrix row has the same 20 columns as its
header.
- The tutorial is registered exactly once in `mkdocs.yml` and has
matching English and Chinese navigation titles.
- The Hugging Face, ModelScope, ModelSlim, and vLLM benchmark links
return HTTP 200.
- The branch contains one documentation-only commit and changes only
four documentation/navigation files.

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…ansfer (vllm-project#15854)

## What this PR does / why we need it?

The layerwise KV pool transfer path stores **every** KV block of
reachability-limited groups (sliding-window raw KV and compressor-state
caches) to the external pool, while the non-layerwise path only stores
blocks the KV cache managers consider reachable (segment-tail windows).
For DeepSeek-V4-Flash (block_size=32, 43+1 layers, 6 KV cache groups) on
910C, a single 131072-token request allocates **91,168 pool keys (~45.6
GB DRAM)**, 98% of which sits in state/SWA groups whose reachable subset
is only ~1.4 GB — a ~32x redundancy that floods the pool and the
interconnect (~1.8 GB/s of writes during prefill).

This PR aligns the layerwise path with the non-layerwise reachable-store
semantics on all three sides, reusing the existing
`AscendStoreCoordinator` masks:

- **Save** (`_alloc_gvas_for_save` / `_process_save_for_layer_batch`):
per-group store masks are computed once per scheduler step; keys are
allocated and transfer ranges generated only for reachable blocks (the
trailing partial block rides the last mask-allowed run).
- **Hit check** (`_get_layerwise_hit_tokens`): only reachable blocks are
queried per group, and the hit length is derived via
`coordinator.find_longest_cache_hit`, so sparse storage still yields
correct prefix hits (a contiguous walk would report 0 hits).
- **Load** (`_prepare_load_gvas` / `_process_load_for_layer_batch`): key
info is fetched only for the load-mask subset, so blocks that are
deliberately not pooled no longer trigger the multi-group load failure
path; GVA positions are filled per queried block index.

Full-attention groups (mask=None), the trailing partial block, and the
non-hybrid / single-group fallbacks keep their previous behavior.
Measured effect: a 131072-token request drops from ~45.6 GB to ~1.4 GB
of pool DRAM with identical hit semantics.

**Note**: this PR also includes the two commits from vllm-project#15830
(`feat(kv_pool): pass lease TTL to memcache batch_alloc` and
`fix(kv_pool): align MTP final-block trim to lcm block boundary`); it
supersedes both vllm-project#15830 and the closed vllm-project#15672, so a single review covers
all three changes.

## Does this PR introduce _any_ user-facing change?

Yes: layerwise mode with a hybrid KV cache layout (e.g. DSV4 compressed
+ SWA groups) now persists only reachable blocks to the external KV
pool, drastically reducing pool memory usage and write bandwidth.
Cache-hit behavior is unchanged.

## How was this patch tested?

- Unit tests added/updated (all pass; 3 pre-existing environment
failures on a no-vllm/no-npu Windows box are unrelated — verified
failing on a pristine tree):
- `tests/ut/distributed/ascend_store/test_metadata.py`:
`masked_block_runs` helper (sparse/None/empty/out-of-range masks)
- `tests/ut/distributed/ascend_store/test_pool_worker.py`: mask-driven
key allocation, multi-run transfer ranges, partial-block riding the last
run, load-mask key queries, load range splitting
- `tests/ut/distributed/ascend_store/test_pool_scheduler.py`:
reachability-aware hit check (sparse SWA storage still yields full hits,
hit stops where a stored tail is missing, no-pool → 0, non-hybrid
fallback unchanged)
- Command: `python -m pytest -sv tests/ut/distributed/ascend_store/`
- NPU e2e validation is in progress on an Ascend cluster (draft until
verified): plan is to re-run the DSV4-Flash layerwise deployment and
compare the `alloc_gvas` key counts per group against the reachable
subset (expected drop: group-4 state cache 65536 → ~128 keys per
131072-token request), plus long-context hit-rate and accuracy checks.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: tyy0829 <1455207791@qq.com>
Signed-off-by: tyy0829 <87685049+tyy0829@users.noreply.github.com>
zouyida2052 and others added 29 commits September 11, 2026 18:17
…vllm-project#16317)

### What this PR does / why we need it?

Allow external cache lookup for prompts shorter than two retention
intervals. With a retention interval and transfer granularity of 4096, a
5083-token prompt can reuse an existing 4096-token checkpoint. The
current early return reports zero without querying the cache.

The production change only removes this seven-line retention-length
guard. Existing checks for requests shorter than one transfer unit,
actual cache availability, and MTP tail recomputation remain unchanged.
Requests previously excluded by the guard can now incur a lookup even
when the cache is empty.

The MTP implementation is identical to the upstream base, including its
aligned `final_block_start` calculation. MTP boundary tests confirm
existing behavior; they are not evidence of another newly fixed MTP bug.

### Does this PR introduce _any_ user-facing change?

Yes. Requests shorter than two retention intervals can reuse valid
external checkpoints instead of being forced to miss. No API or
configuration changes.

### How was this patch tested?

- **Unit tests:** 50 passed in
`tests/ut/distributed/ascend_store/test_pool_scheduler.py`. The
retention regression covers MTP enabled/disabled, stored and missing
checkpoints, and the existing early return below one transfer unit.
Additional boundary coverage confirms the unchanged MTP behavior.
- **Baseline control:** On the upstream base without this fix, the
retention lookup regression fails (0 instead of 4096), while the added
MTP boundary test already passes.
- **Static checks:** Ruff check, Ruff format check, and `git diff
--check` pass for the changed files. The full `bash format.sh ci` hook
suite could not run because the runtime lacks `pre-commit`.
- **NPU experiments rerun with upstream MTP logic:**
DeepSeek-V4-Flash-w8a8-mtp, TP4 × DP2, MTP=1, retention interval 4096.
The runtime uses the upstream MTP block verbatim and removes the
retention guard. Five frozen AISBench inputs were replayed with four
warmup requests and four serial measured requests per length, one output
token per request. Measured external hits/queries and outputs match both
the previous implementation and the previously measured
local-prefix-only mode.

| AISBench input_len | Actual prompt tokens/request | External
hits/queries | Previously measured local prefix hits/queries | Hit rate
|
| --- | --- | --- | --- | --- |
| 4096 | 4179 | 16384/16716 | 16384/16716 | 98.01% |
| 5000 | 5083 | 16384/20332 | 16384/20332 | 80.58% |
| 5004 | 5087 | 16384/20348 | 16384/20348 | 80.52% |
| 8192 | 8275 | 32768/33100 | 32768/33100 | 99.00% |
| 9000 | 9083 | 32768/36332 | 32768/36332 | 90.19% |

The chat template adds 83 tokens in this setup. Counts exclude warmup.
Direct AISBench commands were also rerun for all five lengths, with four
successful requests and zero failures in both warmup and measured
stages:

```bash
python3 aisbench_test.py --input_len 5000 --output_len 1 --data_num 4 \
  --concurrency 1 --request_rate 0 --dataset_type prefix_cache \
  --repeat_rate 1.0 --prefix_test --dp 4
```

Token-ID input tests reconfirmed that actual lengths 4096/5004/8192/9000
reuse 0/4096/4096/8192 tokens per request. A separate 5083-token input
produced identical 64-token output for an external miss and subsequent
hit.

**Validation limitations:** Clean-checkout unit tests use the
directory's existing `_mock_deps` stubs and
`--confcutdir=tests/ut/distributed/ascend_store`; the runtime's parent
conftest has an unavailable `vllm.v1.attention.ops.pcp` dependency. The
NPU runtime is the existing `bdab8bde86aad5eba3b362f923a3ff03b2c0e772`
deployment with the upstream aligned MTP block and the retention guard
removed, retaining its existing memcache backend changes; it is not a
full deployment of the latest main. Cold concurrent warmup can differ
between local and external caching due to publication timing and DP
cache scope. These checks do not establish general model accuracy,
lookahead KV safety for different continuations, graph-mode coverage, or
statistical performance equivalence. This PR leaves upstream MTP
handling unchanged.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
## Auto-Translation Summary

Translated **13** file(s):

-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/community/slash-commands.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/context_parallel.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/quantization.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/testing.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/getting_started/installation/install_vllm_ascend.inc.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_multi_node.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_single_node.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V4-Pro.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/configuration/additional_config.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/batch_invariance.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/context_parallel.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/kv_pool.po</code>
-
<code>/home/runner/_work/vllm-ascend/vllm-ascend/docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/speculative_decoding.po</code>

---

[Workflow
run](https://github.com/vllm-project/vllm-ascend/actions/runs/34436986441)
- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: wangxiyuan <wangxiyuan@users.noreply.github.com>
Co-authored-by: wangxiyuan <wangxiyuan@users.noreply.github.com>
…ast (vllm-project#16094)

### What this PR does / why we need it?

Add KV layer parallelism (KVPP). Each target layer's KV cache is
persisted on one TP rank, and its full contiguous memory buffer is
broadcast before computation. Two reusable scratch buffers enable
one-layer-ahead prefetch without adding a custom operator.

Supported features:
- Model Runner V1 and V2.
- TP, EP, and PP with layer ownership assigned within each PP stage.
- Chunked prefill, prefix caching, and asynchronous scheduling.
- LI-C8 and SFA-C8 cache layouts.
- Fixed-step MTP, with MTP layers excluded from KVPP partitioning.

This PR does not include graph mode, PCP, or DCP support. Prefill/decode
disaggregation is not supported yet; its integration will be submitted
after the Mooncake refactor is complete.

### Does this PR introduce _any_ user-facing change?

KVPP is disabled by default and can be enabled through `enable_kvpp` in
the additional configuration.

### How was this patch tested?

Validated multiple combinations of the features above, including prefix
cache hits and cache-block reuse, PP layer partitioning, and MTP.

GPQA accuracy evaluation is in progress. Accuracy:
| Dataset | Correct | Incorrect | Accuracy |
|---|---:|---:|---:|
| GPQA Diamond (194 scored questions) | 169 | 29 | 85.35% |

Test configuration: GLM-5.2 on Ascend A3 (16 dies), TP16 + EP + DSA-CP,
chunked prefill with `max_num_batched_tokens=32768`.

| Input length | KVPP disabled | Full-layer Broadcast | Overhead |
Increase |
|---|---:|---:|---:|---:|
| 64K | 8.41 s | 8.54 s | 0.13 s | 1.55% |
| 128K | 17.20 s | 18.35 s | 1.15 s | 6.69% |

Configuration: GLM-5.2-W4A8C8, TP16 + EP, chunk size 16K, max model
length 32K, max sequences 32, MTP1, LI-C8/SFA-C8, prefix caching,
FULL_DECODE_ONLY, GPU memory utilization 90%. Cache capacity was
automatically calculated without a block-count override.

| Configuration | KV cache capacity (tokens) | Capacity multiplier |
|---|---:|---:|
| KVPP disabled | 505,472 | 1.00× |
| Full-layer Broadcast | 4,131,968 | 8.17× |

Equivalent KV memory reduction at the same token capacity: **87.77%**.
This is not a reduction in total device memory usage.

### Performance on Ascend A5

KV cache capacity with the same configuration as the pooling benchmark
above, using `gpu_memory_utilization=0.90`. Capacity was automatically
calculated without a block-count override.

| Configuration | KV cache capacity (tokens) | Capacity multiplier |
|---|---:|---:|
| KVPP disabled | 235,008 | 1.00× |
| Full-layer Broadcast | 1,467,904 | 6.25× |

Equivalent KV memory reduction at the same token capacity, estimated
from the capacity ratio: **83.99%**. This is not a reduction in total
device memory usage.

Test configuration: GLM-5.2-W4A8C8 on a single Ascend A5 node (8 dies),
TP8 + EP + DSA-CP, chunked prefill, prefix caching, asynchronous
scheduling, LI-C8, Model Runner V1, eager mode, and `max_num_seqs=12`.
MTP, SFA-C8, and FlashComm1 were not enabled.

#### TTFT

Input lengths: 32K / 64K / 128K; output length: 1 token; prefix cache
hit rate: 0%. Each result is the mean of 4 requests at client
concurrency 1.

| Token budget | Input length | KVPP disabled | Full-layer Broadcast |
Overhead | Increase |
|---|---|---:|---:|---:|---:|
| 16K | 32K | 2.650 s | 3.046 s | +0.395 s | +14.92% |
| 16K | 64K | 5.371 s | 6.518 s | +1.147 s | +21.35% |
| 16K | 128K | 11.017 s | 13.582 s | +2.565 s | +23.28% |
| 32K | 32K | 2.623 s | 2.614 s | -0.009 s | -0.33% |
| 32K | 64K | 5.295 s | 5.375 s | +0.080 s | +1.52% |
| 32K | 128K | 10.901 s | 11.268 s | +0.367 s | +3.36% |

Overhead and percentage changes are calculated from unrounded
measurements. The 32K token budget achieved lower TTFT for both
configurations and was used for the throughput tests.

#### Prefill throughput

Configuration: `max_num_batched_tokens=32768`, `max_num_seqs=12`, client
concurrency 12, 40 requests, 128K input tokens and 1 output token per
request. KV pooling was disabled.

| Prefix cache hit rate | KVPP disabled (input tokens/s) | Full-layer
Broadcast (input tokens/s) | Change |
|---|---:|---:|---:|
| 0% | 12,492.4 | 12,205.7 | -2.30% |
| 90% | 116,516.2 | 110,224.8 | -5.40% |

Throughput is calculated as total input tokens divided by benchmark
duration, including cached input tokens. The measured cache hit rate in
the 90% scenario was 89.9414% for both configurations. All four
throughput runs completed 40/40 requests successfully, with zero
preemptions.

Test configuration: GLM-5.2-W4A8C8 on a single Ascend A5 node (8 dies),
TP8 + EP + DSA-CP, chunked prefill, prefix caching, asynchronous
scheduling, LI-C8, Model Runner V1, and eager mode.
`max_num_batched_tokens=32768`, `max_num_seqs=12`.

KV pooling was enabled in both configurations using
`AscendStoreConnector` with the Memcache backend, `kv_role=kv_both`,
`use_layerwise=false`, and `load_async=true`. CPU pool capacity was
configured as 32 GB per die.

Test scenario: prefill throughput with multiple reusable prefixes.
Sixteen independent prefix families were warmed before measurement,
exceeding the HBM KV cache capacity in both configurations. The workload
contained 40 requests: eight frequently reused families with four
requests each, plus eight families with one request each. Each request
had 128K input tokens, approximately 90% shared prefix, and 1 output
token, at client concurrency 12. Both configurations used identical
request data and ordering, retaining HBM cache after warmup to compare
HBM residency and CPU pool loading.

| Prefix cache hit rate | Configuration | Throughput (input tokens/s) |
HBM cache hit rate | CPU pool cache hit rate | Throughput change |
|---|---|---:|---:|---:|---:|
| 90% (multiple prefixes) | KVPP disabled | 66,362.5 | 8.46% | 81.48% |
— |
| 90% (multiple prefixes) | Full-layer Broadcast | 97,318.3 | 45.46% |
44.48% | +46.65% |

Throughput includes cached input tokens. Cache hit rates are relative to
total input tokens; the combined hit rate was 89.94% in both
configurations. Both runs completed 40/40 requests without preemption.
This benchmark measures prefill throughput with 1 output token per
request.

### Performance on Ascend A5 (dual-node PP)

Test configuration: GLM-5.2-W8A8C8-mxfp8 on two Ascend A5 nodes (8 dies
per node), TP8 + PP2 with a 38/40 layer split, EP, DSA-CP, chunked
prefill, prefix caching, asynchronous scheduling, LI-C8, Model Runner
V1, and eager mode. `max_num_batched_tokens=32768`, `max_num_seqs=12`,
and `gpu_memory_utilization=0.90`. MTP, SFA-C8, and FlashComm1 were not
enabled.

Both configurations used vllm-ascend `493ec2b` with the pooling
restriction removed, and vLLM `b2f6858`, matching the single-node
comparison baseline.

#### TTFT

Input lengths: 32K / 64K / 128K; output length: 1 token; prefix cache
hit rate: 0%. Each result is the mean of 4 requests at client
concurrency 1.

| Token budget | Input length | KVPP disabled | Full-layer Broadcast |
Overhead | Increase |
|---|---|---:|---:|---:|---:|
| 32K | 32K | 2.900 s | 2.905 s | +0.006 s | +0.19% |
| 32K | 64K | 4.503 s | 4.887 s | +0.384 s | +8.53% |
| 32K | 128K | 7.805 s | 8.665 s | +0.859 s | +11.01% |

Overhead and percentage changes are calculated from unrounded
measurements.

#### Prefill throughput

Configuration: client concurrency 12, 40 requests, 128K input tokens and
1 output token per request. KV pooling was disabled.

| Prefix cache hit rate | KVPP disabled (input tokens/s) | Full-layer
Broadcast (input tokens/s) | Change |
|---|---:|---:|---:|
| 0% | 21,395.2 | 18,742.3 | -12.40% |
| 90% | 177,375.4 | 154,717.3 | -12.77% |

Throughput is calculated as total input tokens divided by benchmark
duration, including cached input tokens. The measured cache hit rate in
the 90% scenario was 89.9414% for both configurations. All four
throughput runs completed 40/40 requests successfully, with zero
preemptions.

#### Pooled prefill throughput

KV pooling was enabled in both configurations using
`AscendStoreConnector` with the Memcache backend, `kv_role=kv_both`,
`use_layerwise=false`, and `load_async=true`. CPU pool capacity was
configured as 32 GB per device.

Test scenario: prefill throughput with multiple reusable prefixes.
Thirty-two independent prefix families were warmed before measurement,
totaling 3,772,416 tokens and exceeding the HBM KV cache capacity in
both configurations. The measured workload contained 40 requests:
sixteen frequently reused families with two requests each, plus eight
families with one request each. The remaining eight families were used
only during warmup.

Each request had 128K input tokens, approximately 90% shared prefix, and
1 output token, at client concurrency 12. Both configurations used
identical request data and ordering, retaining HBM cache after warmup to
compare HBM residency and CPU pool loading.

| Prefix cache hit rate | Configuration | Throughput (input tokens/s) |
HBM cache hit rate | CPU pool cache hit rate | Throughput change |
|---|---|---:|---:|---:|---:|
| 90% (multiple prefixes) | KVPP disabled | 86,438.8 | 0.59% | 89.35% |
— |
| 90% (multiple prefixes) | Full-layer Broadcast | 125,699.3 | 53.87% |
36.07% | +45.42% |

Throughput includes cached input tokens. Cache hit rates are relative to
total input tokens; the combined hit rate was 89.94% in both
configurations. Both runs completed 40/40 requests without preemption.

#### KV cache capacity

Capacity was obtained from the non-pooling service startup logs, using
`gpu_memory_utilization=0.90`, without a block-count override.

| Configuration | KV cache capacity (tokens) | Capacity multiplier |
|---|---:|---:|
| KVPP disabled | 555,392 | 1.00× |
| Full-layer Broadcast | 2,972,416 | 5.35× |

Equivalent KV memory reduction at the same token capacity, estimated
from the capacity ratio: **81.32%**. This is not a reduction in total
device memory usage.

Environment note: a storage-link issue required temporary NFS forwarding
through the second node. Both configurations used the same storage
route, and model loading and warmup completed before measurement.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: chengruiqi (C) <c00913489@china.huawei.com>
Signed-off-by: recky-c <ruiqicheng510@gmail.com>
Co-authored-by: chengruiqi (C) <c00913489@china.huawei.com>
…A3 machine and obtains baseline results (vllm-project#15083)

### What this PR does / why we need it?
This PR migrates the nightly single-node a3 test cases to the 560T A3
machine and obtains baseline results

### Does this PR introduce _any_ user-facing change?
no

### How was this patch tested?
run the cases nightly


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: guxin108 <1252896542@qq.com>
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
…5488)

### What does this PR do / why we need it?

In the W4A8 MXFP quantization path of `npu_grouped_matmul_swiglu_quant`
  (`vllm_ascend/device/device_op.py`), the call to
`torch.ops._C_ascend.npu_swiglu_group_quant` now passes a `group_index`
  instead of `None`.

`npu_swiglu_group_quant`'s `group_index` input currently only supports
the
  `count` type, and internally the operator uses `group_index` only for
  summation. We therefore pass the last value of the cumsum tensor
  (`group_list[-1:]`, i.e. the total token count across groups) as the
  `group_index`, which provides the operator with the count it needs.

  ### Does this PR introduce any user-facing change?

  No. 

  ### How was this patch tested?

  Tested with DSv4 Flash


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: lcfenglinwan <lcfenglin@qq.com>
…rv1 (vllm-project#15196)" (vllm-project#16331)

### What this PR does / why we need it?

revert "[feature] reduce cases that required DP padding in mrv1
(vllm-project#15196)"

### Does this PR introduce _any_ user-facing change?

None.

### How was this patch tested?

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: zzzzwwjj <1183291235@qq.com>
…project#16241)

### What this PR does / why we need it?

Refs vllm-project#15665

Support sparse MLA attention without RoPE inputs on A2/A3, as used by
GLM-5.3-Flash. Update host tiling and the cube/vector kernel paths to
handle a zero RoPE dimension, keeping the existing RoPE path available.

This PR contains only the sparse-flash-attention operator change: 5
files, +309/-64 lines, one signed-off commit on main. Model integration
and other pending PRs are excluded.

### Does this PR introduce _any_ user-facing change?

Yes. The existing sparse flash attention operator accepts the NoPE
query/key layout on A2/A3. No serving configuration or
dependency-version change is introduced.

### How was this patch tested?

- Python AST/Ruff, C++ clang-format, and `git diff --check` passed for
the relevant changes.
- The operator changes were included in an integrated A3
native-extension build. This is not a standalone numerical validation of
this branch.
- Added code comments only: the NoPE path is documented in English, no
functional change beyond the operator update above.

Standalone NPU numerical validation is still pending.

- The A3 cache build failed during CMake's Torch import because backend
autoload initialized torch_npu/Triton in the build container. The build
environment issue and an A3 rebuild remain pending. A2 and 310P cache
builds passed.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
### What this PR does / why we need it?
update a3-560t for nightly_config.yaml

### Does this PR introduce _any_ user-facing change?
no

### How was this patch tested?
run the case nightly

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: guxin108 <1252896542@qq.com>
…m-project#13045)

### What this PR does / why we need it?

Adds **Gemma4 MTP speculative decoding** support on Ascend NPU (A5 /
950PR).

Gemma4 (`gemma-4-31B-it` + `gemma-4-31B-it-assistant`) is a
**multi-group KV
cache** model: layers are split into a `sliding` group (head_dim=256)
and a
`full_attention` group (global head_dim=512). The MTP draft reuses the
target's architecture and **shares the target's per-layer KV cache**
(sliding
for layer 58, global for layer 59), so each draft attention group must
read
K/V from its own physical block layout.

Upstream vLLM supports this on CUDA via `Gemma4Proposer`. This PR adds
the
Ascend path as a **thin wrapper** of the upstream proposer — no rewrite
of
vLLM internals, no forking of base-class draft-loop methods.

| file | change | role |
|---|---|---|
| `spec_decode/gemma4_proposer.py` | **+new (150)** |
`AscendGemma4Proposer` — multiple inheritance of upstream
`Gemma4Proposer` + `AscendSpecDecodeBaseProposer`. Overrides only: eager
centroids sampling, `_sync_kv_sharing_target_to_impl`, per-group block
table swap (`attn_update_stack_num_spec_norm` wrapper),
`build_draft_attn_metadata` with FIA SpecDecoding state. |
| `spec_decode/llm_base_proposer.py` | mod (+31/-11) | (1)
`constant_draft_positions` guards in `_run_merged_draft` and
`attn_update_stack_num_spec_norm` — restores upstream
`SpecDecodeBaseProposer` semantics
(vllm/v1/spec_decode/llm_base_proposer.py:635,693,705), no-op for
existing proposers (flag defaults False upstream). (2)
`_build_multi_group_graph_capture_metadata` hook (returns None for
existing proposers) for per-group ACL graph capture. (3) `_propose`
step-update loop iterates `attn_group.layer_names`, matching
`build_draft_attn_metadata`'s existing iteration. |
| `ops/rotary_embedding.py` | mod (+22/-1) | Q-only RoPE: when `key is
None` (Gemma4 MTP sliding layers, K/V from target cache), use a
throwaway key buffer with the regular rotary implementation
(zero-kv-head dummy on the Triton path). |
| `worker/model_runner_v1.py` | mod (+19/-2) | Wire
`AscendGemma4Proposer` into the drafter type union, per-group block
table capture, and kv-cache / cudagraph-key asserts. |
| `spec_decode/__init__.py` | mod (+3) | Route `mtp` +
`use_gemma4_mtp()` to `AscendGemma4Proposer`. |
| `tests/ut/spec_decode/test_gemma4_proposer.py` | **+new (130)** | Unit
tests: proposer routing, KV-sharing sync, per-group metadata. |
| `tests/ut/ops/test_rotary_embedding.py` | mod (+18) | Q-only RoPE
test. |

### Does this PR introduce _any_ user-facing change?

Additive only. Users can now run Gemma4 MTP on Ascend with
`--speculative-config '{"method":"mtp",...}'`. No change to existing
models.

### How was this patch tested?

#### Acceptance
Gemma4-31B-it MTP, k=3, greedy (temp=0), `FULL_DECODE_ONLY` graph mode,
coding
prompt suite (5 prompts × 3 = 15 reqs). Metric:
`vllm:spec_decode_num_accepted_tokens_per_pos` ÷
`vllm:spec_decode_num_drafts`,
scraped from `/metrics`.

| platform | backend | TP | pos0 | pos1 | pos2 | agg | accepted/step |
|---|---|---|---|---|---|---|---|
| **Ascend A5 (950PR) —TP4DP1** | vllm-ascend | 1 | **98.2%** |
**95.4%** | **90.8%** | **94.8%** | **2.85/3** |
| **Ascend A5 (950DT) —TP1DP4** | vllm-ascend | 1×4 | 100.0% | 98.8% |
91.9% | 96.9% | 2.91/3 |
| NVIDIA L20 — CUDA reference | upstream vLLM | 2 | 98.2% | 95.5% |
90.0% | 94.6% | 2.84/3 |

Ascend now matches the CUDA reference.

Here is the prompts:
```python
#{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef binary_search(arr, target):\n    \"\"\"Return the index of target in sorted array arr, or -1 if not found.\"\"\"", "max_out_len": 512, "answer": ""}
#{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef merge_sort(arr):\n    \"\"\"Sort the array arr in ascending order and return it.\"\"\"", "max_out_len": 512, "answer": ""}
#{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef lru_cache_get(cache, key):\n    \"\"\"Return value for key from an OrderedDict-backed LRU cache, or None. Mark as recently used on hit.\"\"\"", "max_out_len": 512, "answer": ""}
#{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef is_balanced(s):\n    \"\"\"Return True if the string s has balanced parentheses/brackets/braces, else False.\"\"\"", "max_out_len": 512, "answer": ""}
#{"question": "Complete the following Python function. Only output the function body, no explanation.\n\ndef flatten(nested):\n    \"\"\"Yield elements from an arbitrarily nested list, depth-first.\"\"\"", "max_out_len": 512, "answer": ""}
```

#### Accuracy
 - Matches the raw target performance
 - Tested on the **gpqa_diamond** dataset by using Evalscope
1. Accuracy Metrics:**Score: 0.7778**,Sample Size: 198 samples
2. Average Latency: 15.0612 seconds
3. Average Throughput: 65.59 tokens/second (Avg Thpt)

#### Performance (950PR TP=4,DP=1)
**Overall Conclusion:** 
Tested on 4K input/1K output by Evalscope.

The model demonstrates excellent scalability and stability. As
concurrency increases, the overall throughput grows significantly while
maintaining a consistent speculative acceptance rate, proving that the
MTP (Multi-Token Prediction) mechanism is functioning effectively.

##### **1. Key Latency & Efficiency Metrics**
*   **TPOT (Time Per Output Token):** 
* At low concurrency (**Conc=1**), the average TPOT is very low (**15.0
ms**), indicating extremely fast generation.
* As concurrency increases to **64**, the average TPOT rises to **81.0
ms**, which is expected due to increased system load, but remains within
an acceptable range for high-throughput serving.
*   **Speculative Acceptance Rate:** 
* The acceptance rate remains remarkably stable across all concurrency
levels, hovering around **71% - 72%**. This indicates that the MTP
model's predictions are consistently accurate regardless of the system
load.

##### **2. Throughput and Scalability**
*   **Completion Throughput (Output Speed):**
* There is a clear linear increase in output speed as concurrency rises.
    *   **Conc 1:** 59.46 tok/s $\rightarrow$ **Conc 64:** 345.05 tok/s.
*   **Total System Throughput:**
* The **Overall Total Prompt throughput** scales from **237.85 tok/s**
(Conc 1) up to **1380.35 tok/s** (Conc 64).
* The **"Last 30s"** peak throughput reaches nearly **3,924 tok/s** at
maximum concurrency, showing the model's ability to handle heavy bursts
of traffic.

##### **3. Summary Table**

| Concurrency | Avg TPOT (ms) | Spec. Accept Rate | Completion
Throughput | Total Prompt Throughput |
| :--- | :--- | :--- | :--- | :--- |
| **1** | 15.0 | 71.5% | 59.46 tok/s | 237.85 tok/s |
| **8** | 30.9 | 71.6% | 236.27 tok/s | 945.17 tok/s |
| **16** | 43.3 | 71.7% | 299.15 tok/s | 1196.73 tok/s |
| **32** | 79.4 | 71.5% | 330.30 tok/s | 1321.32 tok/s |
| **64** | 81.0 | 71.6% | 345.05 tok/s | 1380.35 tok/s |

#### Additional: Real-world Agent Prompt Validation (950DT TP=1,DP=4)
Tested with a ~3.4K-token real tel-sales agent system prompt (300 output
tokens,
`vllm bench serve`, greedy, ignore-eos, TP=1):

| Conc | TPOT p99 (ms) | Output tok/s | Spec accept |
|---|---|---|---|
| 1 | 9.6 | 102 | 63.4% |
| 8 | 10.9 | 732 | 63.6% |
| 16 | 12.2 | 1290 | 63.6% |
| 32 | 14.3 | 2267 | 63.5% |
| 48 | 16.0 | 2981 | 63.6% |
| 64 | 17.2 | 3469 | 63.9% |

Acceptance stays flat across concurrency; throughput scales linearly to
3469 tok/s at conc=64. Zero errors across the full sweep.

Repro (copy-paste-able):
```bash

ASCEND_RT_VISIBLE_DEVICES=0 vllm serve <gemma-4-31B-it> \
   --served-model-name gemma4-31b-it-mtp \
   --tensor-parallel-size 1 \
   --speculative-config '{"method":"mtp","model":<gemma-4-31B-it-assistant>,"num_speculative_tokens":3}' \
   --max-model-len 32000 --gpu-memory-utilization 0.85 \
   --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
   --trust-remote-code --host 0.0.0.0 --port 8831
# then drive the server and scrape /metrics for per-position acceptance
```


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Wu Meng <wumeng@ascend-debug.local>
Co-authored-by: Wu Meng <wumeng@ascend-debug.local>
Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: XUE TONGYAO <xuetongyao2001@gmail.com>
### What this PR does / why we need it?

The op has had no callers since vllm-project#7557 reverted the small-batch GMM
optimization introduced in vllm-project#7100 (qwen3-next gsm8k 98 -> 91). All
grouped-matmul paths run on torch_npu.npu_grouped_matmul.

- delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api)
- drop schema/impl registration from csrc/torch_binding.cpp
- drop meta registration from csrc/torch_binding_meta.cpp
- drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93)

### Does this PR introduce _any_ user-facing change?
Yes, the `moe_grouped_matmul` operator is no longer available in the
PyTorch bindings.

### How was this patch tested?
No new tests were added as this is a code removal PR. Existing CI tests
should pass.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: TangPeng <85704592@qq.com>
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
…lm-project#16387)

Revert of PR vllm-project#16107 (merged onto `main`).

Original PR: vllm-project#16107
Original author: @xqchen7
Merge commit: `fed28d4ce692173954cb887284af45dfa2284a4d`

---
### What this PR does / why we need it?
adapt resource got a5 multi nightly,add openlibing.secret parameter to
enable a5 multi nightly run

### Does this PR introduce _any_ user-facing change?
eable developer run a5 multi nightly

### How was this patch tested?


- vLLM main:
vllm-project/vllm@b2f6858

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…rate_deepseekv4-flash-w8a8-mtp.yaml

### What this PR does / why we need it?
we add nightly kimi25_w4a8_step_3_7.yaml kimi-k2.6_max_model_len.yaml
acceptace_rate_deepseekv4-flash-w8a8-mtp.yaml

### Does this PR introduce _any_ user-facing change?
no

### How was this patch tested?
run  case nightly


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: wanglei727 <naie.wanglei@h-partners.com>
)

### What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer
introduced by vllm-project#14560, and fix import-presence checks for deleted or
annotation-only names. Changes stay under `tests/e2e/vllm_interface/`.

The pytest entry, PR base/head resolution, complete source indexing,
three analysis branches, deterministic summary, and failure policy
remain in place. No persistent cache, pickle files, new environment
variables, additional CI jobs, or main2main/monkey-patch analysis are
introduced.

#### Execution flow and implementation details

1. **Resolve and verify inputs.** The existing pytest entry resolves the
vLLM PR base/head and records the vllm-ascend revision. Commit and
checkout verification remains enabled.
2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the
complete source trees and finalizes the indexes. Single-path namespace
merges copy an already-normalized state instead of repeatedly rebuilding
sorted keys and alternative sets. State dictionaries remain independent;
immutable binding tuples can be shared.
3. **Resolve aliases and namespace states.**
`RepositoryIndex.canonical_name()` checks successively shorter
dot-delimited prefixes in the live alias table instead of
sorting/scanning every alias per lookup. Longest-prefix behavior and
cycle protection remain. Always-bound module names are derived from the
final namespace already computed. Decorator and direct-call resolution
reuse identical statement-prefix states through a per-module, 256-entry
LRU keyed by actual AST statements and guard names. The memo returns
independent dictionaries; expression resolution and caller-specific
fallback still run for each query. It is not shared across revisions or
CI runs.
4. **Discover dependencies.** Override and exact-call discovery retains
affected call sites and inherited implementation locations. Import
discovery reuses the already-parsed vllm-ascend ASTs, with a
source-reading fallback for files absent from the index. Triton launch
invocation kinds remain distinct.
5. **Read base/head source.** Each analysis branch retains separate
base/head `GitSnapshot` instances. Each snapshot lazily uses one `git
cat-file --batch` process rather than launching `git show` per file.
Binary response framing preserves empty files, CRLF and non-ASCII
source. Invalid or failed reads raise analysis errors, not
missing-symbol findings. An `ExitStack` closes processes and temporary
stderr streams on success or exceptions. This is not a persistent source
cache.
6. **Resolve contracts and compare usage.** Snapshots reuse module/class
namespace states, named bindings, owner lookup and resolved API
contracts. Call-contract keys include target expression, access kind,
receiver type, member and invocation kind. Argument binding and
return-use checks still execute separately at every call site; one
site's compatibility result is never reused for another.
7. **Check import presence.** Import checks resolve presence before
constructing unused signature/return details. Presence now follows the
final namespace: a later `del` removes a binding; an annotation without
a value does not create a name or erase an existing value. Constants and
re-exports that remain bound are retained. A P1 requires proven presence
at the base and proven absence at the head. Ambiguous bindings, star
imports, dynamic module exports and unparseable source are not treated
as proven removals. Package-submodule fallback is preserved. Full
endpoint details are still built for findings.
8. **Merge and print.** The three branches merge and sort findings
deterministically. The CI entry prints the summary directly; it does not
create report artifacts. Historical findings remain subject to the
existing PR-only filtering policy.

### Does this PR introduce _any_ user-facing change?

Reduced analyzer runtime and more accurate detection of imports whose
names were deleted or became annotation-only. This is therefore not
exclusively a performance-only change. Analyzer version advances to
`2.1.1`. CLI options and CI log wording are unchanged; existing
dynamic-dispatch and return-dataflow limitations remain.

### How was this patch tested?

#### Local regressions

- 123 local tests passed: Git batch-reader lifecycle/error handling,
alias equivalence, final import bindings, callable/return contracts,
serial/parallel fixture equivalence, pytest-entry integration, and
scope-prefix isolation/eviction.
- Included 12,000 synthetic alias comparisons and 196 generated
scope-flow combinations. Deleted and annotation-only imports were
exercised through actual CLI runs against temporary Git repositories.
- The pytest integration fixtures launch the analyzer subprocess; the
analyzer is not mocked. Historical replay scans described below call the
analyzer core directly, not the hardware E2E suite.
- Ruff, formatting, compileall and mypy checks passed in local
validation. Full repository formatting was previously attempted; some
hooks require Linux shell tools unavailable on this Windows host. Full
Linux CI and NPU workloads are not claimed as locally verified.
- Regression/replay harnesses remain outside the E2E collection path, as
requested; no analyzer unit-test directory is reintroduced into vLLM PR
jobs.

#### Latest paired performance measurement

Windows / Python 3.11, four indexing processes and three analysis
threads, fixed source SHAs, fresh processes, sequential scans and no
persistent analyzer cache. Times below are analyzer time, excluding
network range discovery and hardware tests.

| vLLM PR | Before latest scope reuse | After | Reduction | Result on
both sides |
| --- | ---: | ---: | ---: | --- |
| #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 |
| #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause |
| #50620 | 100.4 s | 81.8 s | 18.5% | PASS |

This baseline already includes the earlier Git batching, alias
optimization and import fix; it is **not** the original pre-optimization
implementation. Full reports matched after removing only timing
metadata, stdout matched byte-for-byte, and expected exit codes were
1/1/0. These are single paired measurements, not statistical benchmarks
or Linux CI speed guarantees.

#### Expanded accuracy replay

Eight additional cases were each run against the frozen pre-scope-reuse
baseline and current code: 16 actual scans. All eight full reports
matched except timing, and all eight stdout logs were byte-identical.

| Fixed historical input | P1 impacts | P1 roots | Review findings |
| --- | ---: | ---: | ---: |
| vllm-ascend vllm-project#11709 range | 5 | 5 | 6 |
| vllm-ascend vllm-project#12020 range | 2 | 1 | 0 |
| vllm-ascend vllm-project#12420 range | 0 | 0 | 0 |
| vllm-ascend vllm-project#12502 range | 33 | 20 | 0 |
| vllm-ascend vllm-project#12648 range | 0 | 0 | 5 |
| vllm-ascend vllm-project#13358 range | 5 | 4 | 11 |
| vLLM #47808 | 2 | 2 | 5 |
| vLLM #50504 | 0 | 0 | 0 |

The upgrade ranges are fixed CI-analyzer inputs, not a full main2main
mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases
use their pre-adaptation baselines. The earlier three cases above were
reused only after verifying current source hashes, giving 11 distinct
cases overall. There were not 22 new scans in the expanded round.

29 semantic/history assertions passed. Previously confirmed P1 findings
remained. #47808 matches the newer August 28 historical log (2 P1), not
the older August 17 report (1 P1). vllm-project#11709's six review items already
existed in the frozen baseline; they are not introduced by scope reuse.

Known limitations are not hidden by these results: vllm-project#12648's
`compute_slot_mappings(out=...)` remains review-only because its
dispatch requires monkey-patch/field propagation; broader return-value
consumption is still outside the exact dependency model. Regression
equivalence does not establish universal precision/recall.

#### Fresh verification of the published PR head

After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched
`refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it
out in a separate, clean detached worktree. The fetched commit and all
nine analyzer source hashes matched the validated version.

Three additional fresh scans used that fetched code:

| vLLM PR | Result | Analyzer time | Process wall time | Exit code |
| --- | --- | ---: | ---: | ---: |
| #39568 | 1 P1 | 59.8 s | 61.0 s | 1 |
| #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 |
| #50620 | PASS | 79.8 s | 81.3 s | 0 |

All three complete reports were identical to the pre-push reports after
excluding only timing metadata; raw stdout was byte-identical. Exit 1 is
the expected detected-break outcome, not an analyzer crash. These runs
are separate from the earlier paired measurements and expanded
regression round.

The 123 local tests passed again before pushing. Ruff, formatting,
spelling and Markdown checks passed; several shell-based repository
hooks could not execute on Windows. The fetched code also passed mypy
with the Python 3.10/Linux target. This is static type checking, not
execution on Linux.

These historical scans execute `analyze_range()` and the production
summary renderer directly. They do not run the full upstream pytest job
or NPU workloads. The actual pytest entry is covered separately by the
local integration fixtures.

Published-head verification details:
vllm-project#16364 (comment)

#### Reproduce a pinned case

With this PR's analyzer checked out, prepare separate clean source
worktrees at the following head/revision:

| vLLM PR | vLLM base | vLLM head | vllm-ascend revision |
| --- | --- | --- | --- |
| #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` |
`c7560af42487b1570c4e6f4cea5df1605a4d59fc` |
`60f0238b0eec4c91fe466497ae8862daf521aecc` |
| #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` |
`c05d75aaa95cf89f547503044c1921625905085d` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |
| #50620 | `c05d75aaa95cf89f547503044c1921625905085d` |
`653cc6faca6885e36760bf35a25bb63442519b14` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |

The analyzer checkout supplies the tool; `--vllm-root` and
`--ascend-root` supply the pinned source trees being analyzed. They do
not need to be the analyzer checkout itself. Existing repositories can
be reused when the head/revision and required base objects match the
table; no model installation or NPU is required for this static CLI
scan.

For example, run #39568 from this PR's analyzer checkout after preparing
the two source paths at the revisions above:

```bash
VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \
  --vllm-root /path/to/vllm-source \
  --ascend-root /path/to/vllm-ascend-source \
  --old ce29c26 \
  --new c7560af \
  --expect-ascend-sha 60f0238 \
  --index-workers 4 --analysis-workers 3 --fail-on introduced
```

PowerShell equivalent (replace the two source directory paths):

```powershell
$env:VLLM_INTERFACE_TIMINGS = "1"
python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range `
  --vllm-root "C:\path\to\vllm-source" `
  --ascend-root "C:\path\to\vllm-ascend-source" `
  --old ce29c26 `
  --new c7560af `
  --expect-ascend-sha 60f0238 `
  --index-workers 4 --analysis-workers 3 --fail-on introduced
Write-Host "Analyzer exit code: $LASTEXITCODE"
```

Timing diagnostics are printed as phases finish, followed by the
compatibility summary; this is not per-file progress. #39568 reports
`BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`,
still called at `vllm_ascend/core/recompute_scheduler.py:907` in the
pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for
#50620.

These commands use the PR's CLI, not pytest: they bypass PR-range
network discovery and parent `conftest.py` dependency loading while
exercising the same analysis engine. Base objects must already exist
locally to avoid Git fetching missing objects. The SHAs are replay
inputs, not analyzer constants. Raw logs and full comparison evidence
are retained locally; the CI entry itself still prints logs without
creating report artifacts.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
… weights (vllm-project#16259)

### What this PR does / why we need it?

Serving a native DeepSeek-V4 checkpoint whose config sets
`num_hash_layers > 0` fails during weight loading:

```
KeyError: 'model.layers.0.mlp.gate.e_score_correction_bias'
  File "vllm_ascend/models/deepseek_v4/model.py", line 1319, in load_weights
    param = params_dict[name]
```

`DeepseekV4MoE` gives hash-router layers a `tid2eid` lookup table and
deliberately leaves `e_score_correction_bias` unset, because those
layers route text tokens purely through the table rather than through
router scores:

```python
if self.hash:
    self.gate.tid2eid = nn.Parameter(...)
    self.gate.e_score_correction_bias = None
```

The checkpoints, however, still ship a `gate.bias` tensor for every MoE
layer. `load_weights` renames it to `gate.e_score_correction_bias` and
then indexes `params_dict` unconditionally, so the unused hash-layer
bias aborts startup.

This PR skips the tensor when the model does not own the target
parameter. The existing `gate.bias_vl` guard directly above is
unaffected.

Reproduced on `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp`
(`num_hash_layers: 3`, `n_routed_experts: 256`) on Ascend 950PR. The
checkpoint carries all three tensors for the hash layers:

```
layers.0.ffn.gate.tid2eid
layers.0.ffn.gate.bias
layers.0.ffn.gate.bias_vl
layers.0.ffn.gate.weight
```

### Does this PR introduce _any_ user-facing change?

No new flags or interfaces. Native DeepSeek-V4 checkpoints with
hash-router layers now finish loading instead of raising `KeyError` at
startup.

### How was this patch tested?

Unit test added:
`tests/ut/models/test_deepseek_v4_moe.py::test_hash_layer_router_bias_is_skipped_when_unused`.
It feeds a router bias for one hash layer and one dense layer and
asserts only the dense one is consumed, so the previous `KeyError` is a
regression failure.

```bash
pytest -sv tests/ut/models/test_deepseek_v4_moe.py
```

Manual check: `vllm serve` of `DeepSeek-V4-Flash-Vision-Exp` on one
Ascend 950PR node reaches `Application startup complete` and answers
`/v1/completions` requests, where it previously failed in `WorkerProc`
init on every rank.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com>
Co-authored-by: yiminghub2024 <202503791+yiminghub2024@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com>
…llm-project#16214)

### What this PR does / why we need it?

Refs vllm-project#15665

This PR enables Multi-Token Prediction (MTP) speculative decoding for
GLM-5.3-Flash (`glm5_next`) on Ascend.

The `Glm5NextMTPModel` implementation already exists in the
GLM-5.3-Flash model code. This PR completes the integration required to
select and run that model through vLLM's existing `deepseek_mtp`
speculative decoding flow.

The main changes include:

- Register `glm5_next_mtp` as a supported MTP model type.
- Normalize GLM-5.3-Flash draft configurations:
  - Convert `glm5_next` and `glm5_next_text` to `glm5_next_mtp`.
  - Select the existing `Glm5NextMTPModel` architecture.
  - Propagate `num_nextn_predict_layers` as `n_predict`.
- Correct main-model graph selection for speculative decoding:
- Treat a speculative batch as uniform decode only when every request
has finished its prompt.
- Prevent partial-prefill requests from incorrectly replaying a uniform
decode graph.
- Keep the main model graph-enabled while the MTP draft model runs
eagerly through `enforce_eager`.
- Add ModelSlim W8A8 quantization adaptations:
- Register the fused MLP, MoE, MLA, and KDA projection mappings for
`glm5_next_mtp`.
- Map runtime prefixes containing `.mtp_block.` back to checkpoint-style
quantization prefixes.
- Resolve MTP expert prefixes stored under
`model.language_model.layers.*`.
- Support multimodal checkpoints whose text-model quantization entries
use different `model.language_model.*` or `language_model.model.*`
namespaces.
  - Reuse the `glm5_next` packed-module mapping for `glm5_next_text`.
  - Preserve expert layers marked as `FLOAT` as unquantized.
- Add multimodal GLM-5.3-Flash weight mappings:
- Flatten ModelSlim KDA `forget_gate` parameters to the runtime module
layout.
  - Map attention and FFN hyper-connection parameters.
  - Map visual, language-model, embedding, and LM-head prefixes.
- Ignore the standalone `rot.*` tensor because the exported rotation has
already been folded into `model.visual.merger.down_proj.weight`.
- Add unit tests covering speculative-config registration, graph
dispatch, multimodal weight mappings, and ModelSlim quantization-prefix
resolution.

This allows GLM-5.3-Flash to use one-token MTP speculative decoding with
the existing `deepseek_mtp` method.

### Does this PR introduce *any* user-facing change?

Yes.

Users can enable MTP speculative decoding for GLM-5.3-Flash with:

```bash
--speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}'
```

For ModelSlim W8A8 checkpoints, the Ascend quantization method must also
be specified explicitly:

```bash
--quantization ascend \
--speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}'
```

The MTP draft model runs eagerly, while the main model can continue
using decode graphs such as `FULL_DECODE_ONLY`.

### How was this patch tested?

Formatting and pre-commit checks:

```bash
bash format.sh
```

All hooks passed, including ruff, codespell, typos, clang-format,
actionlint, gitleaks, shellcheck, and the repository-specific checks.

Unit tests:

```bash
pytest -q \
  tests/ut/models/test_glm5_next_mtp.py \
  tests/ut/worker/test_model_runner_v1.py
```

Result:

```text
51 passed, 2 skipped
```

ModelSlim unit tests:

```bash
pytest -q tests/ut/quantization/configs/test_modelslim_config.py
```

Result:

```text
60 passed
```

Runtime validation was performed with:

- GLM-5.3-Flash ModelSlim W8A8 checkpoint
- Tensor parallel size: 8
- Expert parallel enabled
- One speculative token
- MTP draft model in eager mode
- Main-model graph mode: `FULL_DECODE_ONLY`
- Explicit `--quantization ascend`

Example runtime options:

```bash
--tensor-parallel-size 8 \
--enable-expert-parallel \
--quantization ascend \
--speculative-config '{"num_speculative_tokens":1,"method":"deepseek_mtp","enforce_eager":true}' \
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,8,16,32,64,96,128]}'
```

A single chat-completions request completed successfully and returned
the expected answer `4` for `2 + 2`.

The speculative-decoding metrics confirmed that MTP was active:

```text
draft tokens:    31
accepted tokens: 28
```

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: lrf-vm <lirongfan168@gmail.com>
Signed-off-by: root <root@localhost.localdomain>
Co-authored-by: root <root@localhost.localdomain>
…asks (vllm-project#16067)

### What this PR does / why we need it?

Reduce small-operator overhead around Kimi K3 recurrent KDA on `main` by
optimizing preprocessing copies and removing redundant framework padding
cleanup after KDA and the output norm gate.

- **Before KDA:** consume supported non-contiguous Q/K/V views from the
fused projection directly, eliminating forced contiguous copies.
Independent token/head strides flow through ACLNN, tiling, and the A2/A3
and A5 kernels. Unsupported layouts still materialize.
- **After KDA and norm:** delete `_zero_padded_recurrent_output` on
spec/non-spec outputs and `_zero_padded_output` after `o_norm`,
including the now-unused live-token count calculation. These removals
eliminate the corresponding `arange / comparison / where` chains,
profiled as `Range / Less / Fill / SelectV2`. Norm output is copied
directly into the existing result buffer.
- Preserve state updates, mixed-token index copies, idle/dummy handling,
and buffer initialization. No new kernel interface, switch, or
replacement mask is needed.

The deleted masks sanitize padding after recurrent state computation;
the norm gate operates independently per token/head row. Padding rows
inside the graph-shaped region no longer have a zero-value guarantee.
Effective token computation is unchanged.

### Does this PR introduce _any_ user-facing change?

Performance optimization of Kimi K3 KDA preprocessing and
postprocessing. The Torch operator schema and contiguous-input behavior
remain unchanged.

### How was this patch tested?

- Accuracy validation (reported by the project owner): **GPQA Avg@5:
93.03**.

- Added spec/decode/mixed framework regression cases with NaN padding,
checking effective output rows, state updates, and output-buffer tail
initialization.
- All three cases passed on each branch in an isolated CPU harness
executing the actual `_forward` and `_prepare_beta` bodies with mocked
kernel/context dependencies (six cases total). This is not a full
package or NPU test run.
- Ruff check/format, Python AST parsing, `git diff --check`, and
repository forbidden-import, boolean-context-manager, and long-function
checks passed.
- The project owner previously bypassed the padding masks and reported
successful runtime validation. Full-model/NPU validation and profiling
were not rerun for these commits.
- Full `format.sh ci` was attempted on main and stopped during
actionlint environment installation; the full lint suite did not
complete.
- The direct recurrent-operator regression covers distinct Q/K/V
strides, non-contiguous state pools, guard holes, and numerical
reference checks.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
…oject#15918)

### What this PR does / why we need it?

This PR enables the bundled AscendC `MsaIndexScore` path for MiniMax-M3
on Ascend 950 (A5) with native FP8 index query/key tensors and an FP8
index-key paged cache.

- Sync the operator implementation from
[cann/ops-transformer#10672](https://gitcode.com/cann/ops-transformer/pull/10672)
at `2e685d24ed2e54ea1f1547dc6e5ad29862f152ef`.
- Add the Ascend 950 `arch35` FP8 kernels, tiling, Catlass dependencies,
dtype registration, examples, and documentation.
- Add `msa_index_score` to the Ascend 950 custom-op build list.
- Route A5 prefill index scoring through
`torch.ops._C_ascend.npu_msa_index_score`, while retaining the existing
A5 Triton TopK and decode paths.
- Cast Index-Q to E4M3 when the index-key cache is E4M3 and the query
dtype does not already match.
- Include the A5 wide-`block_table` fix: score columns beyond the
256-column UB window are flushed in windows, with width-257 FP8/BF16
regression coverage.

The MiniMax-M3 A5 FP8 contract used by this integration is: query/key
use `torch.float8_e4m3fn`, `scale=None`, and the operator emits FP32
scores.

### Does this PR introduce _any_ user-facing change?

Yes. MiniMax-M3 FP8 inference on Ascend 950 uses the bundled AscendC MSA
index-score implementation for prefill, while decode continues to use
the existing A5 Triton path. The public Python API is unchanged.

### How was this patch tested?

Static validation completed:

- Verified all 54 imported operator files match PR vllm-project#10672 head
`2e685d24` byte-for-byte (excluding its standalone `torch_extension`;
vLLM Ascend keeps its in-tree Torch adapter).
- `ruff check` passed for the modified MiniMax-M3 Python files and unit
tests.
- `ruff format --check` passed.
- Python AST/compile checks passed.
- `bash -n csrc/build_aclnn.sh` passed.
- `git diff --check` passed.
- Added unit assertions for A5 FP8 registration/build wiring and the
width-257 windowed-flush regression.

#### A5 prefill IndexScore operator performance

Compared the AscendC `MsaIndexScore` kernel with the Triton
`_index_block_score_kernel` on MiniMax-M3-MXFP8, using Ascend 950, TP4 ×
DP2, and the same P+D mixed-load request orchestration. Only prefill
IndexScore device time is reported.

Each value is the mean device duration per IndexScore kernel call across
8 ranks. MiniMax-M3 invokes IndexScore in 57 sparse-attention layers, so
57 calls form one prefill chunk. For 128K, only chunks captured by both
implementations are compared to keep the context lengths matched.

| Input length | Matched prefill scope | AscendC | Triton | Duration
reduction | Speedup |
| --- | --- | ---: | ---: | ---: | ---: |
| 16K | Chunk 0 | 165.646 µs/call | 214.812 µs/call | 22.89% | 1.297× |
| 128K | Chunk 0 | 166.359 µs/call | 213.549 µs/call | 22.10% | 1.284× |
| 128K | Chunk 1 | 363.359 µs/call | 529.388 µs/call | 31.36% | 1.457× |
| 128K | Chunk 2 | 560.355 µs/call | 837.372 µs/call | 33.08% | 1.494× |
| 128K | First 3 matched chunks combined | 363.358 µs/call | 526.770
µs/call | 31.02% | 1.450× |

These are matched operator-level profiling results, not end-to-end TTFT,
throughput, or serving-latency improvements.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: zonghaoxin <z00946994@china.huawei.com>
Signed-off-by: Haoxin Zong <534687988@qq.com>
Co-authored-by: zonghaoxin <z00946994@china.huawei.com>
…m-project#16321)

### What this PR does / why we need it?

Refs vllm-project#15665

Replace decomposed PyTorch mHC reductions with the existing Ascend
`npu_hc_pre_v2` and `npu_hc_post` operators for GLM-5.3-Flash. Preserve
post-mixing scale, optional input RMSNorm, output shapes, and deferred
residual mixing across layers.

The model and native operators are already in main. This PR has no
pending Flash source dependency; full-model validation combines the
other Flash integration PRs and their prerequisites.

### Does this PR introduce _any_ user-facing change?

Yes. GLM-5.3-Flash uses native fused mHC operators on Ascend, reducing
decomposed operator launches without changing serving options.

### How was this patch tested?

- Four cases in `tests/ut/models/test_glm5next_mhc.py` cover optional
RMSNorm, post-mixing scales, deferred mixing, and preservation of the
input residual. All 182 focused regression tests passed in the
integrated tree.
- A3 numerical comparisons against the original Torch path used actual
GLM mHC weights and BF16 inputs at 1, 4, 17, and 3500 tokens. All 16
output comparisons passed; maximum NRMSE was 2.674e-5.
- GLM-5.3-Flash-w8a8 passed eight smoke cases on A3 with TP8, expert
parallelism, MTP=3, and FULL_DECODE_ONLY, including a 5264-token chunked
prefill. A 3500-input/1500-output request also completed.
- Profiling confirmed HcPre/HcPost execution during prefill and decode.
- GitHub pre-commit, CPU unit tests, and the CI gate passed. Selected
device jobs were skipped in that CI run.

Hardware validation covers the GLM BF16 configuration with hc_mult=4 and
hidden_size=4096. These results do not establish an end-to-end speedup
or a full GSM8K result.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…ject#16252)

### What this PR does / why we need it?

Refs vllm-project#15665

Extend the shared SFA backend to execute sparse MLA without RoPE. Add
NoPE preprocessing and forward execution, physical-page cache
addressing, graph-stable metadata buffers, and indexer-provided visible
top-k lengths. Keep cache ownership with the indexer and mask unwritten
graph-padding rows before output projections.

Reuse the existing hardware-specific operators:

- A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now
merged.
- A5: the DeepSeek-V4 sparse MLA interface in
`vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN
`cann_ops_transformer` operators.

This PR contains shared attention integration. GLM KeyPool model routing
is handled separately in vllm-project#16253.

### Does this PR introduce _any_ user-facing change?

Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when
the corresponding operators are available. Existing RoPE behavior and
cache allocation/grouping are preserved.

### How was this patch tested?

- The following files passed in a CPU integration run with 319 total
passing tests, covering metadata/addressing, operator dispatch, NoPE
forward behavior, and existing SFA regressions:

  ```bash
  pytest -q \
    tests/ut/attention/test_sfa_nope_metadata.py \
    tests/ut/attention/test_sfa_nope_forward.py \
    tests/ut/attention/test_sfa_v1.py
  ```

- A focused rerun of the same three files on A3 passed 60 tests.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate on the current head.

The focused unit tests use mocked operator calls and do not establish
native numerical accuracy. A5 native operator validation remains
outstanding.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…lm-project#16243)

### What this PR does / why we need it?

Refs vllm-project#15665

Add Triton KeyPool state compression, paged cache writes, and pooled-key
selection for GLM-5.3-Flash. Allow `IndexerWrapper` to select a
model-provided backend while preserving checkpoint parameter names. Both
operators reuse the shared `next_power_of_2` helper.

Include single-operator accuracy tests and documentation covering
formulas, parameters, layout constraints, graph replay, and test
commands. Model-specific backend integration is handled separately.

### Does this PR introduce _any_ user-facing change?

Yes. Provides the Triton pooled-indexing operators and a
backend-selection hook. This PR does not enable the model backend by
itself or change cache allocation and grouping.

### How was this patch tested?

Validated on Atlas A3 with PyTorch 2.10.0, torch-npu 2.10.0.post4, and
Triton-Ascend 3.2.0:

- 10 tests passed in
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_glm5next_kpool_triton.py`:
CPU-reference accuracy, historical windows, rollback, paging,
noncontiguous storage, invalid slots, graph padding, empty inputs, and
eager/graph replay.
- 9 tests passed in
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_glm5next_pool_key_indexer_triton.py`:
multiple requests, pool capacities, paged caches, token chunking, causal
tails, and eager/graph replay with changing inputs.
- 16 tests passed across
`tests/ut/models/test_glm5next_indexer_backend.py` and
`tests/ut/ops/test_mla.py`.

The CI follow-up adds a type annotation to the test reference data
without changing test behavior. Hardware coverage is limited to Atlas
A3; these tests establish operator accuracy, not model-level throughput.

- vLLM main:
vllm-project/vllm@a97dacb

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)

Revert of PR vllm-project#15908 (merged onto `main`).

Original PR: vllm-project#15908
Original author: @zhiyu-wa
Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920`

---
**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2
 
**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.


- vLLM main:
vllm-project/vllm@a97dacb

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…-project#16043)

### What this PR does / why we need it?

Enable **MTP speculative decoding on Ascend 310P under Model Runner v2**
for Qwen3.5 (dense / MoE), aligned with MRv1 contracts where possible.
- Triton-free CPU paths: rejection sampling, draft input prep, host
draft-step under ACLGraph capture
- Target + draft **FULL_DECODE_ONLY** (K=1 SpecDecoding capture; K>1
per-step draft-decode FULL)
- Hybrid GDN / prefix-cache CoW fixes for concurrent SpecDecoding FULL
- Lean UT + e2e smoke (MRv1 eager + MRv2 FULL K=1)

RFC: vllm-project#15577 

### Does this PR introduce _any_ user-facing change?

No. On 310P with `VLLM_USE_V2_MODEL_RUNNER=1`, MTP can be enabled via:
`--speculative-config '{"method":"mtp","num_speculative_tokens":K}'`

### How was this patch tested?

- UT: `tests/ut/_310p/spec_decode/test_mtp_mrv2_310.py`, MTP gate / CoW
in `test_model_runner_v2_310p.py`
- E2E:
`tests/e2e/pull_request/one_card/_310p/test_spec_decode_mtp_310p.py`

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: Thiagor2002 <13476117628@163.com>
…ct#16189)

### What this PR does / why we need it?

Nightly `qwen3-vl-235b-a22b-instruct-w8a8` (A3, `FULL_DECODE_ONLY`
ACLGraph) fails because internal-router MoE recomputes logits with:

`gate.weight.to(torch.float32)`

That emits a per-forward `aclop Cast`. DeepSeek-V4 and 310P already
avoid this by setting `gate.precast_fp32_weight = True` so
`AscendUnquantizedLinearMethod.process_weights_after_loading`
materializes `weight_fp32` at load time.

This PR applies the same precast in `AscendMoERunner.__init__` for all
Ascend platforms, and routes the hot path through `_gate_weight_fp32()`
so ACLGraph capture no longer depends on a live Cast.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Code-path review against 310P `AscendMoERunner310` and DeepSeek-V4
`precast_fp32_weight`.
- Nightly should re-run `Qwen3-VL-235B-A22B-Instruct-W8A8`.

Made with [Cursor](https://cursor.com)

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: LQD <1107297340@qq.com>
Signed-off-by: liaoqidan <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
### What this PR does / why we need it?
enable mrv2 dspark e2e test

### Does this PR introduce _any_ user-facing change?
N/A
### How was this patch tested?
CI passed with new added/existing test.


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: wxsIcey <1790571317@qq.com>
…oject#16232)

### What this PR does / why we need it?
Documents the scheduling limitations of batch invariance in the user
guide: chunked prefill, prefix caching, and request preemption (eviction
and recomputation) are not supported. These features are not disabled
automatically, so the guide now instructs users to explicitly disable
chunked prefill and prefix caching, and pairs `--block-size 128` with
the chunked prefill disabling. The online/offline inference examples and
the batch invariance e2e test are updated with the same config.

### Does this PR introduce _any_ user-facing change?
Documentation-only update plus an e2e test configuration adjustment.

### How was this patch tested?
Doc-only change; the e2e test config (`test_batch_invariant_tp4.py`)
follows the documented settings. CI verification is sufficient.
- VLLM_WORKER_MULTIPROC_METHOD=spawn pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```bash
[DEBUG] Final model_config: {'model_name': ..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}}
[DEBUG] vllm_runner fixture - model_config: {'model_name': '..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}}
...
PASSEDINFO 09-10 20:42:39 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore
(EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1350] [shutdown] EngineCore: trigger received signal=SIGTERM
(EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1501] [shutdown] EngineCore: start mode=abort timeout=0s
(EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1532] [shutdown] EngineCore: request processing complete; starting resource teardown
(EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1363] [shutdown] EngineCore: exiting busy loop
(EngineCore pid=64849) INFO 09-10 20:42:39 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=4
(Worker_TP0 pid=65187) INFO 09-10 20:42:39 [multiproc_executor.py:836] Parent process exited, terminating worker queues
WARNING 09-10 20:42:44 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1
WARNING 09-10 20:42:44 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 64849
(EngineCore pid=64849) WARNING 09-10 20:42:44 [multiproc_executor.py:484] [shutdown] Executor: workers still running after grace period; sending SIGTERM count=3
[INFO] Model cache cleared

=============================== warnings summary ===============================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================== 1 passed, 15 warnings in 118.88s (0:01:58) ==================
/usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '

```


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
…m-project#16422)

### What this PR does / why we need it?

Adds the native and Triton operator layer required by DeepSeek V4.1 on
Ascend:

- Quant Lightning Indexer V2 and metadata operators
- Sparse Flash MLA and metadata operators
- compressor, index preparation, query quantization, Engram INT8, and
required MoE operator updates
- operator documentation, native examples, static checks, and
operator-level unit/e2e tests

This is the operator-only first part of the DeepSeek V4.1 enablement.
Framework/model integration is intentionally split into dependent PR
vllm-project#16423.

### Does this PR introduce _any_ user-facing change?

Yes. It exposes the native operator capabilities required by DeepSeek
V4.1 serving, but does not register or enable the model integration by
itself.

### How was this patch tested?

- operator static checks: 8 passed
- focused related unit tests: 342 passed, 1 skipped
- broad related unit tests: 855 passed, 12 skipped
- clean native rebuild on Ascend A3: passed
- final single-card NPU operator regression: 359 passed
- pre-commit scoped checks: Ruff, Ruff format, codespell, typos,
clang-format, markdownlint, and forbidden-import check passed

End-to-end model validation is recorded in dependent PR vllm-project#16423.

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Resolve the patch_use_v2_model_runner conflict by keeping the Ascend
whitelist default (apply_v2_model_runner_config_patch) and folding in
main's v0.28.0 PCP unsupported-feature exception.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@yjyang62 yjyang62 closed this Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.