Skip to content

[Scheduler] Make request-timeout aborts rank-consistent to fix TP collective hangs - #37143

Merged
hnyls2002 merged 5 commits into
sgl-project:mainfrom
nogumanov:fix-rank-consistent-timeout-aborts
Sep 8, 2026
Merged

hnyls2002 merged 5 commits into
sgl-project:mainfrom
nogumanov:fix-rank-consistent-timeout-aborts

Conversation

@nogumanov

@nogumanov nogumanov commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Problem

SGLANG_REQ_WAITING_TIMEOUT / SGLANG_REQ_RUNNING_TIMEOUT are enforced by a
wall-clock scan that runs independently on every TP rank. A request sitting at
the timeout boundary gets aborted on whichever rank's clock crosses the
deadline first (sub-millisecond skew), so the waiting-queue composition
diverges across ranks, the extend-vs-decode admission decision splits, ranks
enter mismatched collectives, and the server hangs until the detokenizer
heartbeat health check kills it.

Repro conditions: TP > 1, timeout enabled, queue waits longer than the timeout
(long-context traffic under saturation). We hit it reliably on GLM NVFP4,
B200 TP4, 60k-300k token prompts, SGLANG_REQ_WAITING_TIMEOUT=45.

Evidence from per-rank scheduler tracing (identical recv/token/kv state on all
ranks up to the divergence): one run split 2+2 (two ranks admitted a chunked
prefill, two stayed in decode), another split 1+3 — in both, one rank had
dropped a timed-out request from its queue one iteration earlier than the
others. Setting the timeout to 0 on the same image eliminated the hang; with
this fix and the timeout re-enabled we saw 0 hangs over 21.6h of saturated
long-context benchmarks (vs. reproducible hangs within 16-90 min before).

Fix

  • The timeout scan runs only on the request-pulling rank
    (Scheduler._poll_timeout_aborts, invoked from
    SchedulerRequestReceiver.recv_requests before
    _broadcast_reqs_across_ranks). It only emits AbortReqs — no local
    deletion.
  • The emitted AbortReqs join the recv stream and ride the existing
    broadcast, so every rank removes the same requests in the same scheduler
    iteration through the deterministic abort_request() path (waiting queue:
    pop; running batch: to_finish). No new collectives.
  • abort_request() now propagates the initiator's finished_reason /
    abort_message, so clients still receive
    503 "Request waiting/running timeout reached." (previously the reason was
    dropped on the waiting-queue echo path).

Notes

  • Timeout semantics and env defaults are unchanged; only the enforcement
    mechanism moved.
  • Routing through abort_request() also picks up its fuller cleanup
    (disaggregation KV / metadata / beam retire) that the old per-rank scan
    skipped.
  • With PP > 1 the injected aborts follow the same path as tokenizer-initiated
    AbortReqs; tested on pp_size=1 only.
  • Unit tests adapted: test_scheduler_timeouts.py now drives
    _poll_timeout_aborts (plus a dedup case for a request present in both
    running and last batch).

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ⏳ Run #34174272146
Latest PR Test (Extra): ❌ Run #34174271992
Latest PR Test (AMD ROCm 7.2): ⏳ Run #34174272125

@hnyls2002 hnyls2002 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

also please resolve the conflicts.

@@ -66,6 +66,12 @@ class SchedulerRequestReceiver:
stream_output: Callable[..., None]
get_last_batch: Callable[[], Any]
scripted_scheduler_hook: Optional[ScriptedSchedulerHook] = None
# Called on the request-pulling rank only; returns AbortReq objects for

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Too many comments, normally we only use 1 or 2 lines for this regular field.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done

@nogumanov
nogumanov force-pushed the fix-rank-consistent-timeout-aborts branch from 2a34a94 to 4d148cc Compare September 5, 2026 13:19
@nogumanov
nogumanov force-pushed the fix-rank-consistent-timeout-aborts branch from 4d148cc to 9c773d1 Compare September 7, 2026 10:02
@hnyls2002

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@hnyls2002

Copy link
Copy Markdown
Collaborator

/rerun-test test_scheduler_timeouts.py test_scheduler_chunked_req_gate.py test_scheduler_decision_batch_params.py test_pp_cp_rank_offsets.py test_mm_shm_error_consensus.py test_scheduler_control.py

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_scheduler_timeouts.py test_scheduler_chunked_req_gate.py test_scheduler_decision_batch_params.py test_pp_cp_rank_offsets.py test_mm_shm_error_consensus.py test_scheduler_control.py:

🚀 ubuntu-latest (5 tests): ✅ View workflow run

cd test/ && python3 registered/unit/managers/test_scheduler_timeouts.py
cd test/ && python3 registered/unit/managers/test_scheduler_chunked_req_gate.py
cd test/ && python3 registered/unit/managers/test_scheduler_decision_batch_params.py
cd test/ && python3 registered/unit/managers/test_pp_cp_rank_offsets.py
cd test/ && python3 registered/unit/managers/test_mm_shm_error_consensus.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/scheduler/test_scheduler_control.py

@hnyls2002
hnyls2002 merged commit 792543f into sgl-project:main Sep 8, 2026
167 of 222 checks passed
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
…r part, sgl-project#36638) into the B1 staging line

Agent B: dispatched requests are aborted on handler failure/disconnect
(except BaseException as upstream), the waiter holds the ReqState from
construction (no KeyError for batch requests). Not taken with reason:
abort_sent dedup, rid filter in the disconnect task (would reopen
weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143
not applicable (timeouts -1).
Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants