Repository navigation
Conversation
Signed-off-by: z1ying <tzzying@outlook.com>
There was a problem hiding this comment.
Code Review
This pull request introduces a stream synchronization call in llm_base_proposer.py to ensure that input data writes to CUDA graph buffers are completed before downstream kernels are launched. This change aims to maintain data consistency with minimal CPU overhead compared to global synchronization. There were no review comments provided, so I have no additional feedback to offer.
…uential-decode-loop
There was a problem hiding this comment.
These operations are all on the same stream, no? So they should operate sequentially, i.e. no race condition. If there were actually a race condition somewhere, I suspect that this change just artificially slows down the execution enough for the copy to finish.
You mention #41162 is a related fix and this is a workaround "while the upstream fix propagates". 41162 has landed -- does it fix this issue? Are you still seeing this failure on the latest ToT? Were you using model runner V2? 41162 only touches model runner V2.
Finally, in the future, please disclose when you are using AI for PR contributions.
|
Hi @MatthewBonanni, Thanks for reviewing! You are right—there’s another factor I need to consider. I will dive deeper to gather more persuasive evidence and make the PR clearer. Just sharing some findings I discovered earlier:
Additional observations:
Added probes in self.input_ids[:batch_size] = input_ids
self.hidden_states[:batch_size] = hidden_states
I’m still investigating the FLASHINFER source to confirm its actual execution order; this will take some time. |
|
While I don't doubt that there may be a bug, forcing a synchronization is an unacceptable fix. Please migrate to a github issue, creating one if needed, until a root-cause is identified. |
|
@benchislett Do you have a suggestion for how to tackle this issue? Seems to be similar to another speculative decoding issue I'm seeing with Gemma MTP #42572 |
Just to clarify, I'm referring to an issue in vLLM v1 |
Summary
Fix #40756
A race condition in
vllm/v1/spec_decode/llm_base_proposer.pythat can triggercudaErrorIllegalAddressunder high concurrency when using MTP speculative decoding.Root Cause
The draft token generation loop writes shared GPU buffers
(
input_idsandhidden_states) and immediately launches the nextmodel forward pass using those buffers:
Because CUDA operations are asynchronous, the buffer writes are not guaranteed to be visible to the next kernel before it begins execution. Under high concurrency this can cause FlashInfer attention kernels to observe stale or partially-written buffer contents, causing illegal memory access.
The issue disappears with
CUDA_LAUNCH_BLOCKING=1, suggesting amissing synchronization between the buffer writes and the subsequent
draft model forward pass.
Fix
Add a stream-level synchronization after the buffer writes, before the model forward:
Compared to
torch.cuda.synchronize(), this approach is preferred because:Stream-scoped synchronization
Only synchronizes the current stream, avoiding full-device synchronization and reducing overhead.
Backend portability
Uses the
torch.acceleratorabstraction, maintaining compatibility across CUDA, ROCm, and XPU.Reproduction
Server startup command
Load test script (64 concurrent requests)
Crashes within minutes without the fix. Add
CUDA_LAUNCH_BLOCKING=1to confirm race condition — crash disappears.Evidence
Before patch — crash log
After patch — load test passes
Community Validation
Independently verified by two users across different hardware and vLLM versions:
Validator 1 — RTX PRO 6000 96GB (Blackwell, sm_120), vLLM 0.20.1, Qwen3.6-27B-FP8, PyTorch 2.11.0+cu130, CUDA 13.2, 40 concurrent agents × 2048 max_tokens:
Validator 2 — 4× RTX 3090 (TP=4), vLLM 0.20.2, Qwen3.6-27B-FP16:
Bug confirmed in v0.20.1 and v0.20.2. Fix confirmed across Blackwell and Ampere architectures, single-GPU and TP=4, FP8 and FP16.
Test Results
Regression Tests (48 passed, 7 skipped, 0 failed)
Tests run against all
SpecDecodeBaseProposersubclasses that inherit the modifiedpropose()method:test_mtp.pytest_eagle.pytest_speculators_dflash.pyAll tests passed