[Scheduler] Make request-timeout aborts rank-consistent to fix TP collective hangs - #37143
Merged
hnyls2002 merged 5 commits intoSep 8, 2026
Merged
Conversation
nogumanov
requested review from
Ying1123,
hnyls2002,
merrymercy and
xiezhq-hermann
as code owners
August 30, 2026 15:29
This was referenced Sep 1, 2026
hnyls2002
reviewed
Sep 5, 2026
hnyls2002
left a comment
Collaborator
There was a problem hiding this comment.
also please resolve the conflicts.
| @@ -66,6 +66,12 @@ class SchedulerRequestReceiver: | |||
| stream_output: Callable[..., None] | |||
| get_last_batch: Callable[[], Any] | |||
| scripted_scheduler_hook: Optional[ScriptedSchedulerHook] = None | |||
| # Called on the request-pulling rank only; returns AbortReq objects for | |||
Collaborator
There was a problem hiding this comment.
Too many comments, normally we only use 1 or 2 lines for this regular field.
nogumanov
force-pushed
the
fix-rank-consistent-timeout-aborts
branch
from
September 5, 2026 13:19
2a34a94 to
4d148cc
Compare
nogumanov
force-pushed
the
fix-rank-consistent-timeout-aborts
branch
from
September 7, 2026 10:02
4d148cc to
9c773d1
Compare
Collaborator
|
/tag-and-rerun-ci |
Collaborator
|
/rerun-test test_scheduler_timeouts.py test_scheduler_chunked_req_gate.py test_scheduler_decision_batch_params.py test_pp_cp_rank_offsets.py test_mm_shm_error_consensus.py test_scheduler_control.py |
Contributor
|
Results for 🚀 🚀 |
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Sep 24, 2026
…r part, sgl-project#36638) into the B1 staging line Agent B: dispatched requests are aborted on handler failure/disconnect (except BaseException as upstream), the waiter holds the ReqState from construction (no KeyError for batch requests). Not taken with reason: abort_sent dedup, rid filter in the disconnect task (would reopen weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143 not applicable (timeouts -1). Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
SGLANG_REQ_WAITING_TIMEOUT/SGLANG_REQ_RUNNING_TIMEOUTare enforced by awall-clock scan that runs independently on every TP rank. A request sitting at
the timeout boundary gets aborted on whichever rank's clock crosses the
deadline first (sub-millisecond skew), so the waiting-queue composition
diverges across ranks, the extend-vs-decode admission decision splits, ranks
enter mismatched collectives, and the server hangs until the detokenizer
heartbeat health check kills it.
Repro conditions: TP > 1, timeout enabled, queue waits longer than the timeout
(long-context traffic under saturation). We hit it reliably on GLM NVFP4,
B200 TP4, 60k-300k token prompts,
SGLANG_REQ_WAITING_TIMEOUT=45.Evidence from per-rank scheduler tracing (identical recv/token/kv state on all
ranks up to the divergence): one run split 2+2 (two ranks admitted a chunked
prefill, two stayed in decode), another split 1+3 — in both, one rank had
dropped a timed-out request from its queue one iteration earlier than the
others. Setting the timeout to 0 on the same image eliminated the hang; with
this fix and the timeout re-enabled we saw 0 hangs over 21.6h of saturated
long-context benchmarks (vs. reproducible hangs within 16-90 min before).
Fix
(
Scheduler._poll_timeout_aborts, invoked fromSchedulerRequestReceiver.recv_requestsbefore_broadcast_reqs_across_ranks). It only emitsAbortReqs — no localdeletion.
AbortReqs join the recv stream and ride the existingbroadcast, so every rank removes the same requests in the same scheduler
iteration through the deterministic
abort_request()path (waiting queue:pop; running batch:
to_finish). No new collectives.abort_request()now propagates the initiator'sfinished_reason/abort_message, so clients still receive503 "Request waiting/running timeout reached." (previously the reason was
dropped on the waiting-queue echo path).
Notes
mechanism moved.
abort_request()also picks up its fuller cleanup(disaggregation KV / metadata / beam retire) that the old per-rank scan
skipped.
AbortReqs; tested on pp_size=1 only.test_scheduler_timeouts.pynow drives_poll_timeout_aborts(plus a dedup case for a request present in bothrunning and last batch).
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ⏳ Run #34174272146
Latest PR Test (Extra): ❌ Run #34174271992
Latest PR Test (AMD ROCm 7.2): ⏳ Run #34174272125