Skip to content

[BugFix][CI] Restore post-16626 Kimi baseline and GLM C8 nightly flags - #158

Draft
yjyang62 wants to merge 16 commits into
expand-mrv2-whitelist-copy-df29from
cursor/restore-post-16626-nightly-yaml-1985
Draft

yjyang62 wants to merge 16 commits into
expand-mrv2-whitelist-copy-df29from
cursor/restore-post-16626-nightly-yaml-1985

Conversation

@yjyang62

Copy link
Copy Markdown
Owner

What this PR does / why we need it?

vllm-project#16848 re-applied the vllm-project#16626 EPLB V2 nightly yaml snapshot and overwrote three later main commits. Keep V2 --enable-eplb, restore those overlays:

Does this PR introduce any user-facing change?

No. Nightly test config only.

How was this patch tested?

Inspected the two yaml files against vllm-project#16900 / vllm-project#16815 / vllm-project#15514 and the vllm-project#16626 EPLB V2 conversion. No NPU run in this change.

Open in Web Open in Cursor 

cursoragent and others added 16 commits September 20, 2026 16:24
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
V2 init_speculator does not implement ngram/ngram_gpu, so those
methods now default to V1. The MLA precision e2e exercises the V1
kernel and writes num_tokens onto ForwardContext, so pin it to V1
like the SFA V1 precision tests.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Restore the original DYNAMIC_EPLB configs and baselines, and pin
those nightlies with VLLM_USE_V2_MODEL_RUNNER=0. Do not rewrite
them to --enable-eplb in this default-V2 change.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Apply the 706eaf1 rewrite to this PD nightly only: drop V1
DYNAMIC_EPLB/EXPERT_MAP_RECORD and use --enable-eplb plus
load_collection_phase on Model Runner V2.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Default these models and the 310P platform to Model Runner V1.
VLLM_USE_V2_MODEL_RUNNER still overrides the blacklist.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Whisper encoder-decoder fails V2 dummy compile when input_ids is
None. KVPP combined-features mismatches greedy results on V2.
Default both to V1 unless VLLM_USE_V2_MODEL_RUNNER is set.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Keep the same default-V2 blacklist and env override, but collapse
one-line wrappers, use architecture prefixes, and group the
selection flow as constants, probes, decision, then the config patch.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
V2 encoder_runner.capture() uses upstream CUDA graph_capture, which
asserts on the Ascend TP communicator. Default cudagraph_mm_encoder
configs to V1; explicit VLLM_USE_V2_MODEL_RUNNER still wins.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port 2ce7480: insert perf_warmup before perf for GLM-5.1/5.2
W8A8 A3 dual-node nightlies.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Port 3c33450: set reasoning_effort: low on the W8A8 A3 GPQA case.
Keep thinking: true, indexer_kv_dtype int8, and the performance case.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Match 706eaf1 on the seven nightlies still pinned to V1 DYNAMIC_EPLB.
QWEN3_235B_PD.yaml already matched.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
KVPP has a V2 runtime (KVPPRuntime). Stop forcing V1 via the
blacklist; unset VLLM_USE_V2_MODEL_RUNNER now selects V2 when
additional_config.enable_kvpp is true.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
KV pool (AscendStoreConnector / memcache) is no longer a V2 blacklist
reason. Unset VLLM_USE_V2_MODEL_RUNNER now selects V2 for KV pool.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
VLLM_USE_V1 is ignored by current vLLM and does not pin Model Runner V1.
These five DeepSeek-V3.2 configs should use the default Model Runner V2
path instead of carrying a dead engine-era env var.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
vllm-project#15514 --kv-cache-dtype/--indexer_kv_dtype fp8 without the retired
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
@cursor
cursor Bot force-pushed the expand-mrv2-whitelist-copy-df29 branch from 0aa2400 to c8c5487 Compare September 21, 2026 13:31
@yjyang62
yjyang62 force-pushed the expand-mrv2-whitelist-copy-df29 branch from c8c5487 to ddffb7a Compare September 21, 2026 15:26
@cursor
cursor Bot force-pushed the expand-mrv2-whitelist-copy-df29 branch 2 times, most recently from afa3fd7 to 1133a5f Compare September 23, 2026 01:23
@yjyang62
yjyang62 force-pushed the expand-mrv2-whitelist-copy-df29 branch 2 times, most recently from c7c66ec to 9422be7 Compare September 23, 2026 04:30
@cursor
cursor Bot force-pushed the expand-mrv2-whitelist-copy-df29 branch from c47a3f5 to 550cce2 Compare September 23, 2026 11:43
@yjyang62
yjyang62 force-pushed the expand-mrv2-whitelist-copy-df29 branch 2 times, most recently from 3c7c95b to a25d0f3 Compare September 24, 2026 04:28
@cursor
cursor Bot force-pushed the expand-mrv2-whitelist-copy-df29 branch from 5d8d901 to f1f7356 Compare September 24, 2026 10:00
@yjyang62
yjyang62 force-pushed the expand-mrv2-whitelist-copy-df29 branch 2 times, most recently from 43eaa34 to bfee6c6 Compare September 26, 2026 05:55

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants