Skip to content

fix(patches): drop the #50004 backport; add the #52492 capture guard to #49486 (silent corruption derisk) - #116

Closed
de1tydev wants to merge 1 commit into
MiaAI-Lab:mainfrom
de1tydev:fix/graph-capture-corruption-derisk
Closed

de1tydev wants to merge 1 commit into
MiaAI-Lab:mainfrom
de1tydev:fix/graph-capture-corruption-derisk

Conversation

@de1tydev

Copy link
Copy Markdown
Contributor

Closes #115.

Upstream vLLM's investigation of intermittent silent output corruption on 2× DGX Spark (MTP + prefix caching + CUDA graphs — NVIDIA forum thread) identified two of this fork's live perf backports. This PR mirrors upstream's post-fix state:

1. Remove the #50004 backport (vllm#51318 reverted it)

The adaptive C128A packing makes the metadata builder write packed rows at the live batch's stride while FULL-graph consumers keep the capture-time stride; rows ≥ 1 read stale slot ids and attention lands on the wrong context slices. Upstream chose full revert — adaptive stride is structurally incompatible with capture-stable layout. The 0.1.1 image's stock code is the exact pre-#50004 state, so dropping the patch restores upstream's post-revert state.

Removed from: compose for _hf in chain (which fail-closes on missing listed files since #103), launcher status echo, launcher worker-sync list, test-hotfix-atomic-transaction.py inventory/cases (SEVEN renamed to CHAIN), bench-patches.sh expectations, docs/vllm-027-new-patches.md (row kept, marked removed with rationale). The historical CHANGELOG entry describing the backport is retained as history; a new entry documents the removal. bench-baseline-issue22-only.sh's active_topk_width grep still expects 0 and is unchanged.

2. Add the #52492 capture guard to the #49486 port (vllm#52492)

The port's header asserted "No CUDA-graph guard (faithful to upstream): max_seq_len is a per-captured-batch constant, so the branch is stable inside a capture." Upstream refuted exactly that: a graph captured with the shortcut baked in (short dummy metadata) replays against longer cached prefixes and returns candidates 0..topk-1 unscored. The guard is upstream verbatim — and not torch.cuda.is_current_stream_capturing() — eager-step behavior unchanged, --status now reports the guard, frozen hunk digest updated.

Not touched

  • hotfix-dsv4-flashmla-workspace-50298.sh: vllm#52836 reverts #49236's model-wide eager scratch pool (cross-stream reuse without allocator tracking). This fork never backported #49236; the #50298 mechanism (per-forward get_simultaneous slices with capture-time reservation) is different and not implicated.
  • All other chain members (#50312, #48407, #48957, grammar-advance).

Validation

  • bash scripts/ci-validate.sh passes (includes the reshaped test-hotfix-atomic-transaction.py: chain apply/idempotence/fault-rollback, INVENTORY digest for the modified #49486 hunk).
  • Live on a 2× DGX Spark TP=2 pair (Anemll 0.1.1, official 0731 @ 9e165c30, nvfp4_ds_mla, MTP=5, APC, FULL_AND_PIECEWISE): boot applies the six-patch chain fail-closed; --status reports the #52492 guard; active_topk_width absent from sparse_mla.py.
  • A/B (fresh boot vs same-day fresh-boot baseline, streaming TTFT/decode + metrics-differential acceptance): Same-day fresh-boot A/B on our experimental 2× Spark TP=2 pair (streaming client, 384-token code-gen task, cache-busting random prefixes; acceptance from metrics deltas). Before = that morning's fresh boot with both backports live; after = fresh boot with this change. Greedy decode 8K 67.35 → 68.90 tok/s, 32K 64.41 → 66.95 tok/s; temp0.6 8K 3-run mean 67.44 → 67.21 tok/s; TTFT 8K 3.41 → 3.67 s, 32K 13.33 → 13.51 s; acceptance 0.674/0.695 → 0.702/0.703 (greedy/temp0.6), ~4.5 tokens/step in both arms. Every delta sits inside same-arm run-to-run spread (a prior same-config 3-run range was ~10 tok/s): no regression, and the #50004 removal's claimed 1.0% is invisible at this noise level. RULER-lite spot check (sniah + vartrack at 8K/32K) 4/4 PASS on the patched boot.

Perf cost, stated exactly

  • #50004 claimed ~1.0% E2E upstream; removing it gives that back.
  • The #52492 guard removes the shortcut only from captured graphs; eager steps keep it. Short-context (≤2048 tokens) requests running through captured graphs lose the skip — the same trade upstream accepted for correctness.

Upstream vLLM traced intermittent silent output corruption on 2x DGX
Spark (MTP + prefix caching + CUDA graphs) to code this fork carries as
perf backports:

- #51318 reverted #50004: the C128A metadata builder writes packed rows
  at the live batch's stride while FULL-graph consumers keep the
  capture-time stride, so rows >= 1 read stale slot ids and attention
  lands on the wrong context slices. The 0.1.1 image's stock code is the
  exact pre-#50004 state, so removing the backport (compose chain,
  launcher status echo, worker sync list, test inventory) restores
  upstream's post-revert state.
- #52492 kept #49486 but barred the short-context indexer shortcut
  during stream capture: a graph captured with the shortcut baked in
  replays it against longer cached prefixes and returns candidates
  0..topk-1 unscored. The port now carries the guard verbatim and
  --status reports it; eager-step behavior is unchanged.

The third fix from the same investigation (#52836, reverting #49236's
eager scratch pool) does not apply: this fork never backported #49236.

CPU gate: bash scripts/ci-validate.sh passes.
@plotarmordev

Copy link
Copy Markdown
Collaborator

Maintainer integration is #121. It preserves this PR author attribution and verified #50004/#52492 changes while replaying them onto current main so #112/#113 are present in the exact tested head. Independent two-rank recreate validation and CPU evidence are posted there. This cross-repository PR will remain unmerged and can close after #121 lands; no defect is attributed to the contributor implementation.

@plotarmordev

Copy link
Copy Markdown
Collaborator

Superseded by maintainer integration #121, which preserves contributor credit and merged after exact-head hosted CI, independent review, and two-rank live validation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants