Skip to content

[Feature][MRV2] Adapt extract_hidden_states for Model Runner V2 - #141

Draft
yjyang62 wants to merge 10 commits into
mainfrom
mrv2-extract-hidden-states-5b6a
Draft

yjyang62 wants to merge 10 commits into
mainfrom
mrv2-extract-hidden-states-5b6a

Conversation

@yjyang62

@yjyang62 yjyang62 commented Sep 9, 2026

Copy link
Copy Markdown
Owner

What this PR does / why we need it?

Rebase mrv2-extract-hidden-states-5b6a onto latest main so upstream PR vllm-project#14699 can pass CI.

cpu-ut failed because CI rebases the original feature commits onto main, and vllm_ascend/worker/v2/attn_utils.py conflicted: main imports vllm_version_is while the first extract_hidden_states commit adds is_hidden_state_cache_spec. The branch is now linear on main with both imports.

Does this PR introduce any user-facing change?

No. This is the same extract_hidden_states MRV2 work, rewritten onto current main.

How was this patch tested?

  • Reproduced the CI rebase conflict locally against vllm-project/vllm-ascend main

  • git rebase upstream-main is clean after the linearization

  • ruff check passed on the touched Python files

  • CPU unit tests were not run here (torch is not installed in this environment)

  • vLLM main: vllm-project/vllm@b2f6858

Open in Web Open in Cursor 

yjyang62 and others added 10 commits September 9, 2026 07:59
Port extract_hidden_states to Model Runner V2 on the 0828 pin
(e6bfe03ad / vLLM #51718). Dispatch upstream ExtractHiddenStatesSpeculator,
keep HiddenStateCacheSpec on private [B, H, N, C] buffers so dumps cannot
overlay hybrid Attention/Mamba backing, and cover allocate/reshape plus
speculator dispatch in unit tests.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The speculator crashed with "Expected 3 auxiliary hidden states, got 2"
on TP workers. NPU torch.compile graph-breaks at TP collectives can drop
early Python-list aux appends, and example ids such as [2, 18, 34] are
silently skipped on models with fewer than 35 layers.

Disable compile on the target backbone after load, reject out-of-range
layer ids with the EAGLE3 default, enable MiniMax aux collection for
this method, and tolerate residual=None in the DeepSeek V2 aux path.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the Ascend speculator subclass and call upstream
ExtractHiddenStatesSpeculator from init_speculator. Keep only NPU glue:
disable target compile so aux list collection survives TP graph-breaks,
and reject out-of-range eagle_aux_hidden_state_layer_ids after load.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Yesterday's 0828 port already reused upstream ExtractHiddenStatesSpeculator
and only needed private HiddenStateCacheSpec allocate/reshape after
vLLM #51718. The later helper, load_model hook, MiniMax method list, and
DeepSeek residual=None patch were added for a compile/OOB path that the
proven Qwen3.5-35B-A3B TP=8 enforce-eager dump does not need.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
CacheOnlyAttentionBackend dropped get_kv_cache_shape in #51718 and has
not restored it. Reshape the dump cache from HiddenStateCacheSpec
[B, H, N, C] properties at the call site instead of a helper that
could pick up a pre-51718 backend layout.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The cache_only_layers name check was a V1 fallback for shadowed spec
classes. V2 get_kv_cache_spec does not rewrite CacheOnly layers, and
allocate already keys off HiddenStateCacheSpec, so the layer-name OR
is redundant.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Cover skip_tokenizer_init + TokensPrompt on the generic generate path
and on extract_hidden_states dumps, including Model Runner V2.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep token-in/token-out coverage inside extract_hidden_states e2e, but
remove the extra one-card case that was not part of the MRV2 dump work.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3.5-0.8B is multimodal, so skip_tokenizer_init leaves tokenizer=None
and Qwen3VLProcessor crashes during LLM init. Keep the token-in/token-out
path, but cover it with dummy Qwen3-8B instead of the hybrid VL model.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants