Repository navigation
[Bugfix][Attention] Keep FlashInfer's plan() staging buffer pageable - #59493
Sahil170595 wants to merge 1 commit into
Conversation
FlashInfer's plan() writes its scheduling metadata into a host buffer that each wrapper reuses, then copies it to the GPU asynchronously. The buffer is pinned, so the copy reads it only when the GPU reaches the copy. The drafter plans the same decode wrapper once per draft step, so unless something in between waits for the GPU, a launch still queued behind earlier work runs with the next step's metadata under its own plan_info offsets, and the split-KV merge can index out of range (vllm-project#40756). Replace the staging buffer with a pageable one whenever vLLM creates a wrapper, as vllm-project#32799 and vllm-project#54660 already do for vLLM's own reused host buffers. A copy from pageable memory reads its source when it is issued, so no sync is needed. Signed-off-by: Sahil Kadadekar <sahilkadadekar@gmail.com>
|
@benchislett on #42603 you said forcing a synchronization was an unacceptable fix until a root cause was identified. This PR has one for the staging-buffer race reported in #40756: FlashInfer's |
Purpose
Fixes a race where FlashInfer can overwrite a CPU planning buffer before the GPU has finished reading the previous plan. With MTP speculative decoding, the GPU can then use the wrong scheduling data and crash or hang. This addresses the staging-buffer race reported in #40756.
The change replaces FlashInfer's pinned CPU staging buffer with ordinary, pageable CPU memory when vLLM creates an attention wrapper. This lets CUDA stage the source data before the next plan reuses the buffer. It covers decode, prefill, DCP child wrappers and cascade child wrappers, and adds a regression test for both eager and CUDA graph decode.
Why it is needed
Repeated draft steps can plan the same wrapper before its previous GPU copy runs. A
seq_lens.cpu()wait masks that overlap in the recorded main baseline, but the performance change in #57352 removes the wait for drafting. This fix protects the planning buffer independently of that wait. The fused drafting path in #58371 does not cover the native FlashInfer path that re-plans each step.Why this is not a duplicate
Rechecked #40756 and open PR searches for the issue,
_pin_memory_int_workspace_bufferand FlashInfer planning. No other matching vLLM staging-buffer fix was found. #42603 used an explicit stream wait; #57352 removes a wait. The other PRs referencing #40756 (#45005, #51508, #53450) address different failure mechanisms.Validation
Previously recorded local results on an RTX 4080 Laptop under WSL2:
These are the existing author-recorded results; GPU tests were not rerun during this source review. Hosted pre-commit is currently skipped because the repository's contributor eligibility check fails.
Limits
The patch adds no explicit stream synchronization. Pageable copies may still block or synchronize inside the CUDA driver, so the measurements below are not a guarantee for every device or driver (CUDA synchronization behavior). FlashInfer MLA wrappers are outside this change and have not been validated on Hopper.
Test commands, reproductions and model measurements
Test Plan
tests/v1/attention/test_flashinfer_plan_staging.py: builds aFlashInferMetadataBuilder, holds the GPU withtorch.cuda._sleep, plans its decode wrapper twice throughfast_plan_decode, and compares the workspace the first launch reads against a clean first plan. It also asserts the second plan was issued while the first copy was still queued, so it cannot pass vacuously. Covers the eager and CUDA graph decode wrappers.plan()host time with a pinned vs pageable staging buffer while the GPU is held busy.pre-commit.All runs on an RTX 4080 Laptop (sm_89) under WSL2.
Test Result
(1) On this branch (FlashInfer 0.7.0.post1): fails without the fix for both wrappers, at
assert torch.equal(seen_by_a, plan_a)(the first launch reads the second plan), and passes with it.(2) Identical on FlashInfer 0.6.16.post3 and 0.7.0:
(3) Median host time per
plan(), 200 back-to-back plans behind a ~1 s GPU stall (the GPU was still stalled after the loop, so no call waited on the stream):(4)
tests/v1/attention/test_attention_backends.py -k flashinfer,tests/v1/attention/test_flashinfer_dcp_spec_reorder.pyand the new test: 31 passed, 5 skipped (run locally with the ungatedNousResearch/Meta-Llama-3-8Bconfig in place of the gatedmeta-llama/Meta-Llama-3-8B).pre-commitpasses on the changed files.(5) Run on baseline 4bb804c with FlashInfer 0.7.0 and Model Runner V2, with #57352 applied in the local comparison build at 5cf5be4c2. #57352 is still open upstream; "main" in the table means that recorded baseline. "Queued" counts plans issued while the same wrapper's previous copy had not run yet, "changed" counts those that changed the staged bytes, and "layout changed" counts those that also changed
plan_info(only measured in the last workload).seq_lens.cpu()wait from [Deprecation] Deprecate items scheduled for 0.29 #55353.plan_infowhile the previous copy was still queued.Credit to @brasrox for identifying pinned planning-buffer reuse in #40756, and to @sempi for corroborating runs.
AI assistance (Claude) was used for the investigation, patch and tests. Codex assisted with the subsequent source review and description cleanup.
Essential Elements of an Effective PR Description Checklist