Conversation
Rebase the PR onto latest main as a linear history so CI's `git rebase $BASE_SHA` no longer replays old commits onto the Keep default-V2 selection, the Ascend feature blacklist, and the SFA C8 DCP hardware capability check from vllm-project#16656. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
V2 init_speculator does not implement ngram/ngram_gpu, so those methods now default to V1. The MLA precision e2e exercises the V1 kernel and writes num_tokens onto ForwardContext, so pin it to V1 like the SFA V1 precision tests. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Restore the original DYNAMIC_EPLB configs and baselines, and pin those nightlies with VLLM_USE_V2_MODEL_RUNNER=0. Do not rewrite them to --enable-eplb in this default-V2 change. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Apply the 706eaf1 rewrite to this PD nightly only: drop V1 DYNAMIC_EPLB/EXPERT_MAP_RECORD and use --enable-eplb plus load_collection_phase on Model Runner V2. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Default these models and the 310P platform to Model Runner V1. VLLM_USE_V2_MODEL_RUNNER still overrides the blacklist. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Whisper encoder-decoder fails V2 dummy compile when input_ids is None. KVPP combined-features mismatches greedy results on V2. Default both to V1 unless VLLM_USE_V2_MODEL_RUNNER is set. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Keep the same default-V2 blacklist and env override, but collapse one-line wrappers, use architecture prefixes, and group the selection flow as constants, probes, decision, then the config patch. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
V2 encoder_runner.capture() uses upstream CUDA graph_capture, which asserts on the Ascend TP communicator. Default cudagraph_mm_encoder configs to V1; explicit VLLM_USE_V2_MODEL_RUNNER still wins. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port vllm-project#16920: salt the RNG position when resampling a random residual so residual draws do not reuse the rejection-conditioned noise. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port 2ce7480: insert perf_warmup before perf for GLM-5.1/5.2 W8A8 A3 dual-node nightlies. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port 3c33450: set reasoning_effort: low on the W8A8 A3 GPQA case. Keep thinking: true, indexer_kv_dtype int8, and the performance case. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Match 706eaf1 on the seven nightlies still pinned to V1 DYNAMIC_EPLB. QWEN3_235B_PD.yaml already matched. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
KVPP has a V2 runtime (KVPPRuntime). Stop forcing V1 via the blacklist; unset VLLM_USE_V2_MODEL_RUNNER now selects V2 when additional_config.enable_kvpp is true. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
KV pool (AscendStoreConnector / memcache) is no longer a V2 blacklist reason. Unset VLLM_USE_V2_MODEL_RUNNER now selects V2 for KV pool. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
VLLM_USE_V1 is ignored by current vLLM and does not pin Model Runner V1. These five DeepSeek-V3.2 configs should use the default Model Runner V2 path instead of carrying a dead engine-era env var. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and vllm-project#15514 --kv-cache-dtype/--indexer_kv_dtype fp8 without the retired enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor
Bot
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
from
September 21, 2026 13:31
0aa2400 to
c8c5487
Compare
yjyang62
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
from
September 21, 2026 15:26
c8c5487 to
ddffb7a
Compare
cursor
Bot
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
2 times, most recently
from
September 23, 2026 01:23
afa3fd7 to
1133a5f
Compare
yjyang62
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
2 times, most recently
from
September 23, 2026 04:30
c7c66ec to
9422be7
Compare
cursor
Bot
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
from
September 23, 2026 11:43
c47a3f5 to
550cce2
Compare
yjyang62
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
2 times, most recently
from
September 24, 2026 04:28
3c7c95b to
a25d0f3
Compare
cursor
Bot
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
from
September 24, 2026 10:00
5d8d901 to
f1f7356
Compare
yjyang62
force-pushed
the
expand-mrv2-whitelist-copy-df29
branch
2 times, most recently
from
September 26, 2026 05:55
43eaa34 to
bfee6c6
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
vllm-project#16848 re-applied the vllm-project#16626 EPLB V2 nightly yaml snapshot and overwrote three later main commits. Keep V2
--enable-eplb, restore those overlays:Kimi-K2.6-w4a8-A3.yaml: baseline1433.4454→1296([Test] update Kimi-K2.6-w4a8-A3 baseline vllm-project/vllm-ascend#16900)GLM5_1_W4A4_A5.yaml:HCCL_INTRA_ROCE_ENABLE0→1([Test] update Qwen3-32B-QuaRot-eagle3 baseline vllm-project/vllm-ascend#16815)--kv-cache-dtype fp8/--attention_config.indexer_kv_dtype fp8and drop retiredenable_sparse_sfa_c8/enable_sparse_li_c8additional-config keys ([Refactor][Quantization]KVCache quantization dtype specified by the --kv-quant-dtype parameter vllm-project/vllm-ascend#15514; C8 is now derived from those CLI dtypes)Does this PR introduce any user-facing change?
No. Nightly test config only.
How was this patch tested?
Inspected the two yaml files against vllm-project#16900 / vllm-project#16815 / vllm-project#15514 and the vllm-project#16626 EPLB V2 conversion. No NPU run in this change.