Repository navigation
Conversation
for nixl enable staging buffer. Signed-off-by: Fei Teng <fteng@nvidia.com>
fteng-NV
requested review from
ByronHsu,
Duyi-Wang,
HaiShaw,
ShangmingCai,
hnyls2002 and
sogalin
as code owners
September 3, 2026 02:59
iyastreb
approved these changes
Sep 3, 2026
This was referenced Sep 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
for nixl enable staging buffer.
Motivation
While investigating SGLang PD performance with nixl backend, a reproducible decode-side hung issue appeared, introduced by staging buffer.
Issue description
SGLang asymmetric PD (prefill tp=4, decode tp=8) with nixl backend, staging buffer enabled, ril=16384, chunked prefill size=8192 (i.e. 2 chunks per request).
When num_prompts >=64, there is a high probability that decode instance hung until router time out.
Root cause analysis
For decode instance, it checks if all chunks completed in staging buffer by _maybe_submit_last_scatter() with condition is_last_chunk is True. This means it assumes the last chunk's notification arrives last.
However when staging buffer enabled, for multiple chunks, there is a high chance that chunks arriving out of order, which means the last chunk cannot guarantee to be the last arrival.
If an earlier chunk arrived later than the last chunk, there is no chance to call _maybe_submit_last_scatter() so decode instance stuck by chunks never completed until router time out.
Modifications
The solution is simple, just move the logic calling _maybe_submit_last_scatter() out of "is_last_chunk is True” condition.
Accuracy Tests
I have verified after the change, hung issue disappeared and decode instance could work normally.
CI States
Latest PR Test (Base): ❌ Run #33709685390
Latest PR Test (Extra): ❌ Run #33709685197
Latest PR Test (AMD ROCm 7.2): ❌ Run #33709685368