Repository navigation
Conversation
Upstream vLLM traced intermittent silent output corruption on 2x DGX Spark (MTP + prefix caching + CUDA graphs) to code this fork carries as perf backports: - #51318 reverted #50004: the C128A metadata builder writes packed rows at the live batch's stride while FULL-graph consumers keep the capture-time stride, so rows >= 1 read stale slot ids and attention lands on the wrong context slices. The 0.1.1 image's stock code is the exact pre-#50004 state, so removing the backport (compose chain, launcher status echo, worker sync list, test inventory) restores upstream's post-revert state. - #52492 kept #49486 but barred the short-context indexer shortcut during stream capture: a graph captured with the shortcut baked in replays it against longer cached prefixes and returns candidates 0..topk-1 unscored. The port now carries the guard verbatim and --status reports it; eager-step behavior is unchanged. The third fix from the same investigation (#52836, reverting #49236's eager scratch pool) does not apply: this fork never backported #49236. CPU gate: bash scripts/ci-validate.sh passes.
Collaborator
|
Maintainer integration is #121. It preserves this PR author attribution and verified #50004/#52492 changes while replaying them onto current main so #112/#113 are present in the exact tested head. Independent two-rank recreate validation and CPU evidence are posted there. This cross-repository PR will remain unmerged and can close after #121 lands; no defect is attributed to the contributor implementation. |
Collaborator
|
Superseded by maintainer integration #121, which preserves contributor credit and merged after exact-head hosted CI, independent review, and two-rank live validation. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #115.
Upstream vLLM's investigation of intermittent silent output corruption on 2× DGX Spark (MTP + prefix caching + CUDA graphs — NVIDIA forum thread) identified two of this fork's live perf backports. This PR mirrors upstream's post-fix state:
1. Remove the #50004 backport (vllm#51318 reverted it)
The adaptive C128A packing makes the metadata builder write packed rows at the live batch's stride while FULL-graph consumers keep the capture-time stride; rows ≥ 1 read stale slot ids and attention lands on the wrong context slices. Upstream chose full revert — adaptive stride is structurally incompatible with capture-stable layout. The 0.1.1 image's stock code is the exact pre-#50004 state, so dropping the patch restores upstream's post-revert state.
Removed from: compose
for _hf inchain (which fail-closes on missing listed files since #103), launcher status echo, launcher worker-sync list,test-hotfix-atomic-transaction.pyinventory/cases (SEVENrenamed toCHAIN),bench-patches.shexpectations,docs/vllm-027-new-patches.md(row kept, marked removed with rationale). The historical CHANGELOG entry describing the backport is retained as history; a new entry documents the removal.bench-baseline-issue22-only.sh'sactive_topk_widthgrep still expects 0 and is unchanged.2. Add the #52492 capture guard to the #49486 port (vllm#52492)
The port's header asserted "No CUDA-graph guard (faithful to upstream): max_seq_len is a per-captured-batch constant, so the branch is stable inside a capture." Upstream refuted exactly that: a graph captured with the shortcut baked in (short dummy metadata) replays against longer cached prefixes and returns candidates 0..topk-1 unscored. The guard is upstream verbatim —
and not torch.cuda.is_current_stream_capturing()— eager-step behavior unchanged,--statusnow reports the guard, frozen hunk digest updated.Not touched
hotfix-dsv4-flashmla-workspace-50298.sh: vllm#52836 reverts #49236's model-wide eager scratch pool (cross-stream reuse without allocator tracking). This fork never backported #49236; the #50298 mechanism (per-forwardget_simultaneousslices with capture-time reservation) is different and not implicated.Validation
bash scripts/ci-validate.shpasses (includes the reshapedtest-hotfix-atomic-transaction.py: chain apply/idempotence/fault-rollback, INVENTORY digest for the modified #49486 hunk).9e165c30,nvfp4_ds_mla, MTP=5, APC,FULL_AND_PIECEWISE): boot applies the six-patch chain fail-closed;--statusreports the #52492 guard;active_topk_widthabsent fromsparse_mla.py.Perf cost, stated exactly