Repository navigation
Conversation
Port extract_hidden_states to Model Runner V2 on the 0828 pin (e6bfe03ad / vLLM #51718). Dispatch upstream ExtractHiddenStatesSpeculator, keep HiddenStateCacheSpec on private [B, H, N, C] buffers so dumps cannot overlay hybrid Attention/Mamba backing, and cover allocate/reshape plus speculator dispatch in unit tests. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The speculator crashed with "Expected 3 auxiliary hidden states, got 2" on TP workers. NPU torch.compile graph-breaks at TP collectives can drop early Python-list aux appends, and example ids such as [2, 18, 34] are silently skipped on models with fewer than 35 layers. Disable compile on the target backbone after load, reject out-of-range layer ids with the EAGLE3 default, enable MiniMax aux collection for this method, and tolerate residual=None in the DeepSeek V2 aux path. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the Ascend speculator subclass and call upstream ExtractHiddenStatesSpeculator from init_speculator. Keep only NPU glue: disable target compile so aux list collection survives TP graph-breaks, and reject out-of-range eagle_aux_hidden_state_layer_ids after load. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Yesterday's 0828 port already reused upstream ExtractHiddenStatesSpeculator and only needed private HiddenStateCacheSpec allocate/reshape after vLLM #51718. The later helper, load_model hook, MiniMax method list, and DeepSeek residual=None patch were added for a compile/OOB path that the proven Qwen3.5-35B-A3B TP=8 enforce-eager dump does not need. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
CacheOnlyAttentionBackend dropped get_kv_cache_shape in #51718 and has not restored it. Reshape the dump cache from HiddenStateCacheSpec [B, H, N, C] properties at the call site instead of a helper that could pick up a pre-51718 backend layout. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The cache_only_layers name check was a V1 fallback for shadowed spec classes. V2 get_kv_cache_spec does not rewrite CacheOnly layers, and allocate already keys off HiddenStateCacheSpec, so the layer-name OR is redundant. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Cover skip_tokenizer_init + TokensPrompt on the generic generate path and on extract_hidden_states dumps, including Model Runner V2. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep token-in/token-out coverage inside extract_hidden_states e2e, but remove the extra one-card case that was not part of the MRV2 dump work. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3.5-0.8B is multimodal, so skip_tokenizer_init leaves tokenizer=None and Qwen3VLProcessor crashes during LLM init. Keep the token-in/token-out path, but cover it with dummy Qwen3-8B instead of the hybrid VL model. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
Rebase
mrv2-extract-hidden-states-5b6aonto latestmainso upstream PR vllm-project#14699 can pass CI.cpu-utfailed because CI rebases the original feature commits ontomain, andvllm_ascend/worker/v2/attn_utils.pyconflicted:mainimportsvllm_version_iswhile the first extract_hidden_states commit addsis_hidden_state_cache_spec. The branch is now linear onmainwith both imports.Does this PR introduce any user-facing change?
No. This is the same extract_hidden_states MRV2 work, rewritten onto current
main.How was this patch tested?
Reproduced the CI rebase conflict locally against
vllm-project/vllm-ascendmaingit rebase upstream-mainis clean after the linearizationruff checkpassed on the touched Python filesCPU unit tests were not run here (
torchis not installed in this environment)vLLM main: vllm-project/vllm@b2f6858