[Test] Add focused hybrid MTP prefix-cache regressions - #53189
Merged
vllm-bot merged 1 commit intoAug 22, 2026
Merged
Conversation
mgoin
force-pushed
the
mgoin/hybrid-mtp-prefix-cache-tests
branch
from
August 21, 2026 02:38
681be24 to
0aaf018
Compare
Member
Author
|
/ci run |
|
✅ Triggered Buildkite CI #84942 for commit |
Signed-off-by: mgoin <mgoin64@gmail.com>
mgoin
force-pushed
the
mgoin/hybrid-mtp-prefix-cache-tests
branch
from
August 21, 2026 02:48
0aaf018 to
30d2072
Compare
mgoin
marked this pull request as ready for review
August 21, 2026 14:17
Member
Author
|
/ci run |
|
✅ Triggered Buildkite CI #85079 for commit |
am-cohere
pushed a commit
to am-cohere/vllm
that referenced
this pull request
Sep 1, 2026
…53189) Signed-off-by: mgoin <mgoin64@gmail.com>
mikeshawcode
pushed a commit
to mikeshawcode/vllm
that referenced
this pull request
Sep 1, 2026
…53189) Signed-off-by: mgoin <mgoin64@gmail.com> Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Related to #43559.
Why this is not a duplicate
#48970 adds a separate, broadly calibrated correctness harness with model-specific control and liveness machinery. This patch is a compact alternative: it reuses
test_hybrid_chunked_prefill.py, constructs the exact one-block versus two-block boundary, directly asserts the cached-token invariant, and checks warm and uncached output correctness.The focused open-PR searches found #48970 but no other PR modifying this existing suite for the same cases.
Tests
The YAML parsed successfully. The full Nemotron real-weight e2e was not run locally; exact-checkpoint TP=1 dummy initialization consumed 75.36 GiB and fit under a conservative B200 memory budget. The optional B200 job provides the real-weight run.
Model evaluation is not applicable because this changes only tests and CI configuration; it does not change model output or serving behavior.
AI assistance
AI assistance was used to develop and test this change. The human submitter reviewed every changed line, understands the implementation and validation, and is responsible for defending the change end to end.