Skip to content

[Bugfix] Backport DSpark graph replay safety fix - #12

Open
Code4me2 wants to merge 1 commit into
Anemll:mainfrom
Code4me2:fix/dspark-graph-replay-safety
Open

Code4me2 wants to merge 1 commit into
Anemll:mainfrom
Code4me2:fix/dspark-graph-replay-safety

Conversation

@Code4me2

@Code4me2 Code4me2 commented Aug 26, 2026 •

Copy link
Copy Markdown

Summary

  • backport vLLM commit dfecbb52ce1801d12e99c1d8bfb6e37a5530cd20 from vLLM PR #51538
  • apply downstream vLLM patches after the exact locked checkout and before the repository overlay
  • verify patch applicability, whitespace, and Python syntax against the pinned vLLM commit in CI

Rationale

The vLLM revision pinned by this repository predates the upstream DSpark graph-replay safety fix. Its capture path still zero-initializes sample_idx_mapping, which can make padded capture rows scatter into request slot 0. Its draft-input kernel can also route rejected or non-resident rows into physical KV block 0 rather than PAD_SLOT_ID.

Those defects are relevant to the intermittent first-request DSpark hang and TP-rank divergence reported in #11. Upstream #51538 specifically describes unsafe DSpark graph replay and draft-KV writes, and was subsequently validated by another contributor on a two-node GB10 TP=2 DSpark deployment.

This is intentionally not a backport of vLLM #48167. That change targets generic non-causal FlashInfer attention on SM100, while this repository uses its custom SM121 sparse-MLA path.

Scope

The patch is an unmodified, path-limited export of the upstream commit and applies cleanly to VLLM_COMMIT=752a3a504485790a2e8491cacbb35c137339ad34. The vLLM, FlashInfer, and image version pins are unchanged.

Validation

Static validation:

Live validation used a runtime-equivalent image derived from 0.1.1 with the exact PR-patched Python module installed on both ranks. The module SHA-256 matched across ranks.

On two GB10 nodes with TP=2, DSpark k=5 probabilistic, native 1,048,576 context, concurrency 1, chunked prefill, prefix caching, and FULL_AND_PIECEWISE CUDA graphs:

  • the two upstream GPU regression cases passed on the pinned API
  • mandatory first-request inference passed
  • 120/120 serial decode requests passed, producing 7,680 completion tokens
  • a 28,820-token cold prompt and cached repeat both passed
  • a fresh cold start with an exactly 660,000-token Pi-session prompt as the first inference completed in 718.495 seconds with HTTP 200, zero prefix-cache hits, and returned exactly cold-660k-ok
  • DSpark drafted and accepted tokens on the 660K request
  • paired normalized NCCL sequences matched exactly across 16,591 collective records per rank
  • both GPUs returned to 0% after inference and after teardown
  • no traceback, EngineDeadError, RPC timeout, NCCL error, illegal-memory error, OOM, or persistent GPU wedge occurred

The released 0.1.1 image with DSpark disabled was restored after testing and passed real inference.

AI assistance

AI assistance was used for investigation, patch integration, and validation. The submitted runtime patch itself is the signed upstream vLLM change.

Refs #11

Apply the upstream vLLM DSpark fix at build time without changing the pinned vLLM revision. Validate patch applicability against the exact lock in CI.

Refs Anemll#11

Upstream: vllm-project/vllm#51538
Signed-off-by: code4me2 <velvetmoon222999@gmail.com>
@Code4me2

Copy link
Copy Markdown
Author

Live dual-GB10 validation is complete.

Configuration

  • TP=2 across two GB10 nodes
  • DSpark, k=5, probabilistic drafting
  • native 1,048,576-token context
  • max_num_seqs=1, max_num_batched_tokens=4096
  • chunked prefill + prefix caching
  • FULL_AND_PIECEWISE CUDA graphs (the affected graph path remained enabled)

The runtime candidate was derived from 0.1.1 with the exact PR-patched module installed. Both ranks reported the same module SHA-256.

Results

  • Two pinned-API adaptations of the upstream GPU regression tests: passed
  • Mandatory first-request smoke inference: passed
  • Serial decode soak: 120/120 passed, 7,680 completion tokens
  • 28,820-token cold/cached pair: passed (23.456 s / 1.459 s)
  • Fresh cold start, exactly 660,000 prompt tokens as the first inference, zero prefix-cache hits: passed
    • HTTP 200
    • 718.495 seconds
    • 7 completion tokens
    • exact response: cold-660k-ok
    • DSpark drafted 5 and accepted 5 tokens
  • Paired NCCL traces after the cold 660K request: 16,591 normalized collective records per rank, sequences exactly equal, no mismatch
  • Both GPUs returned to 0% after inference and teardown
  • No traceback, EngineDeadError, RPC timeout, NCCL error, illegal-memory error, OOM, or persistent GPU wedge

For transparency, an initial 660K calibration used a client timeout of 1,200 seconds and canceled before prefill completed. Both ranks remained synchronized and cancellation returned both GPUs cleanly to 0%. A cached follow-up completed in 52.925 seconds, after which the decisive test above was repeated from another fresh cold start with a 1,800-second timeout.

The released DSpark-off service was restored after testing and passed real inference. Based on this result, I am marking the PR ready for review.

@Code4me2
Code4me2 marked this pull request as ready for review August 26, 2026 06:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant