Skip to content

[PD] Add the missing Prefill bootstrap timeout for NIXL - #34692

Merged
ShangmingCai merged 1 commit into
sgl-project:mainfrom
jambow0320:nixl-bootstrap-timeout
Aug 13, 2026
Merged

ShangmingCai merged 1 commit into
sgl-project:mainfrom
jambow0320:nixl-bootstrap-timeout

Conversation

@jambow0320

@jambow0320 jambow0320 commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Background

RFC #33861 proposes gradually consolidating the duplicated PD request/room protocol logic in Mooncake, NIXL, and Mori into a single common protocol layer, while keeping third-party engine-specific behavior in each backend Transport.

Before extracting the common protocol layer, Step 1 of the implementation plan in #34510 aligns clear, non-controversial semantic gaps through small, independent, backend-local PRs. This PR addresses the first gap: the missing bootstrap timeout in the NIXL Prefill Sender.

The bootstrap timeout covers the following case:

Prefill has created the Sender/room for a request, but Decode destination metadata never arrives. The Sender should not remain in KVPoll.Bootstrapping indefinitely; it should transition to KVPoll.Failed after the existing SGLANG_DISAGGREGATION_BOOTSTRAP_TIMEOUT deadline.

Current Problem

CommonKVSender already provides _check_bootstrap_timeout():

# python/sglang/srt/disaggregation/common/conn.py
def _check_bootstrap_timeout(self) -> Optional[KVPoll]:
    if self.init_time is None:
        return None
    elapsed = time.time() - self.init_time
    if elapsed < self.kv_mgr.bootstrap_timeout:
        return None

    self.kv_mgr.record_failure(
        self.bootstrap_room,
        f"Request {self.bootstrap_room} timed out after {elapsed:.1f}s "
        f"in KVPoll.Bootstrapping",
    )
    self.kv_mgr.update_status(self.bootstrap_room, KVPoll.Failed)
    return KVPoll.Failed

This helper:

  1. Computes the bootstrap wait time from the Sender's init_time;
  2. Returns None while the request remains within the deadline;
  3. Records a failure reason after the deadline;
  4. Updates the room to KVPoll.Failed;
  5. Returns KVPoll.Failed.

However, the current NIXL Sender has two missing pieces:

  1. NixlKVSender.__init__() does not record the start of the Prefill bootstrap deadline;
  2. NixlKVSender.poll() does not call the existing helper while the room is in KVPoll.Bootstrapping.

NIXL currently records _transfer_start_time only for actual KV/state transfer latency:

if self._transfer_start_time is None and (
    len(kv_indices) > 0 or state_indices is not None
):
    self._transfer_start_time = time.perf_counter()

That timer starts when the first meaningful KV/state chunk is submitted. It does not include the bootstrap phase spent waiting for Decode metadata, so it cannot replace init_time.

Similarly, the init_time set by NixlKVReceiver.send_metadata() belongs to the Decode Receiver waiting timeout. It is not the Prefill Sender bootstrap deadline.

As a result, if Decode destination metadata never arrives, a NIXL Prefill room can remain in KVPoll.Bootstrapping indefinitely.

Existing Behavior in the Other Backends

Mooncake

Mooncake records the bootstrap start time when creating the Sender:

# python/sglang/srt/disaggregation/mooncake/conn.py
super().__init__(
    mgr,
    bootstrap_addr,
    bootstrap_room,
    dest_tp_ranks,
    pp_rank,
    req_has_disagg_prefill_dp_rank,
)
self.conclude_state = None
self.init_time = time.time()
self._init_trace_ctx()

Its poll() calls the common helper while the room remains in KVPoll.Bootstrapping:

# python/sglang/srt/disaggregation/mooncake/conn.py
elif status == KVPoll.Bootstrapping:
    timeout_result = self._check_bootstrap_timeout()
    if timeout_result is not None:
        return timeout_result

Mooncake therefore cannot wait indefinitely for missing Decode metadata.

Mori

Mori also records the bootstrap start time when creating the Sender:

# python/sglang/srt/disaggregation/mori/conn.py
super().__init__(
    mgr,
    bootstrap_addr,
    bootstrap_room,
    dest_tp_ranks,
    pp_rank,
    req_has_disagg_prefill_dp_rank,
)
self.transfer_statuses = []
self.pending_infos = None
self.conclude_state = None
self.status_notified = False
self.init_time = time.time()

Mori does not call _check_bootstrap_timeout() directly. Instead, it performs the equivalent check inline in its own poll():

# python/sglang/srt/disaggregation/mori/conn.py
if status == KVPoll.Bootstrapping:
    elapsed = time.time() - self.init_time
    if elapsed >= self.kv_mgr.bootstrap_timeout:
        reason = (
            f"Request {self.bootstrap_room} timed out after {elapsed:.1f}s "
            "in KVPoll.Bootstrapping"
        )
        sent_status, _ = self._finalize_failure(reason)
        return sent_status
    return status

Mori uses an inline implementation because its Sender currently owns backend-specific terminalization. In addition to updating the local room state, _finalize_failure():

  • Records the Mori failure reason;
  • Sets conclude_state;
  • Uses _notify_lock/status_notified to emit the terminal status at most once;
  • Notifies Decode through the Mori control channel when destination information is already available.

The common _check_bootstrap_timeout() helper only records a local failure and updates the Manager status. It does not understand Mori's remote notification or terminal-once state. Mori therefore implements the same deadline semantics while retaining its backend-local failure finalization.

This PR only aligns NIXL with the bootstrap deadline already implemented by Mooncake and Mori. It does not change Mori's terminalization behavior.

Changes

This PR only changes NixlKVSender.

1. Record the bootstrap start time when creating the Sender

# python/sglang/srt/disaggregation/nixl/conn.py
super().__init__(
    mgr,
    bootstrap_addr,
    bootstrap_room,
    dest_tp_ranks,
    pp_rank,
    req_has_disagg_prefill_dp_rank,
)
self.init_time = time.time()

2. Call the existing timeout helper while Bootstrapping

# python/sglang/srt/disaggregation/nixl/conn.py
status = self.kv_mgr.check_status(self.bootstrap_room)
if status == KVPoll.Bootstrapping:
    timeout_result = self._check_bootstrap_timeout()
    if timeout_result is not None:
        return timeout_result

The timeout check runs only when status == KVPoll.Bootstrapping. Once enough Decode metadata has arrived and the room transitions to WaitingForInput, this deadline no longer applies.

Behavior After This Change

Before:

Create NixlKVSender
→ request_status[room] = Bootstrapping
→ Decode metadata never arrives
→ poll() returns Bootstrapping indefinitely

After:

Create NixlKVSender
→ init_time = current time
→ request_status[room] = Bootstrapping
→ Decode metadata does not arrive before the deadline
→ _check_bootstrap_timeout()
→ record_failure(...)
→ request_status[room] = Failed
→ poll() returns Failed

The deadline continues to use the existing environment variable:

SGLANG_DISAGGREGATION_BOOTSTRAP_TIMEOUT=300

Users can continue to relax the deadline through the existing environment variable. This PR adds no new configuration.

Testing

To keep the implementation PR diff minimal, the CPU regression test is currently stored on a dedicated branch in the fork:

branch: https://github.com/jambow0320/sglang/tree/rfc-pd-test
path: test/registered/unit/disaggregation/rfc-test/test_nixl_sender_bootstrap_timeout.py

Test scenario:

Sender creation time: 10s
Current poll time: 20s
bootstrap_timeout: 5s
Decode metadata: missing

Assertions:

  • sender.init_time == 10.0;
  • sender.poll() == KVPoll.Failed;
  • request_status[room] == KVPoll.Failed;
  • The failure reason contains timed out.

Test results:

Test from the dedicated test branch + source from this PR:
1 passed

The same test + source before this fix:
1 failed
Failure: sender.init_time is None

CI States

Latest PR Test (Base): ❌ Run #31676012627
Latest PR Test (Extra): ❌ Run #31676012346

Co-authored-by: Cursor <cursoragent@cursor.com>

@ShangmingCai ShangmingCai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good

@ShangmingCai ShangmingCai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If it appears on all backends now, should we just move it into common? Also, I remember NIXL has his own timeout notification mechanism in the C++ player, should we respect that?

@jambow0320

Copy link
Copy Markdown
Contributor Author

If it appears on all backends now, should we just move it into common? Also, I remember NIXL has his own timeout notification mechanism in the C++ player, should we respect that?

Thanks! I agree that this behavior should ultimately be unified in the common protocol layer.

However, Mori still differs in this path because its bootstrap timeout must go through _finalize_failure() to preserve its backend-specific terminalization, terminal-once guard, and Decode notification behavior. Given that difference, I would prefer not to introduce a larger Common extraction in this PR.

At the current stage, I am focusing on low-risk defensive parity fixes: aligning clearly missing behavior without changing the existing backend success paths, wire format, or terminal semantics. I plan to submit a dedicated follow-up PR for the remaining safe defensive alignments. Once those straightforward gaps are closed, I will move to the next stage and start extracting the shared protocol logic with the backend-specific terminal behavior explicitly accounted for.

Regarding the NIXL timeout mechanism, this timeout occurs before a NIXL transfer is created. While the Sender is in Bootstrapping, Prefill is still waiting for the per-room Decode metadata through SGLang's bootstrap channel; there is no NIXL transfer handle or completion notification yet.

I also did a preliminary scan of the timeout mechanisms in the NIXL v1.3.0 and latest C++ sources. The mechanisms I found apply to progress-thread wakeups, metadata transport, connection/handshake handling, active transfers, or NIXL-EP operations. They do not appear to cover SGLang's pre-transfer Bootstrapping phase, so I do not think they need to be considered for this patch.

@ShangmingCai

Copy link
Copy Markdown
Collaborator

/rerun-group disaggregation

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group disaggregation:

🚀 4-gpu-gb300 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_aarch64.py

🚀 2-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_basic.py
cd test/ && python3 registered/disaggregation/test_disaggregation_decode_offload.py
cd test/ && python3 registered/disaggregation/test_disaggregation_optimistic_prefill.py
cd test/ && python3 registered/disaggregation/test_disaggregation_rust_server.py

🚀 8-gpu-h20 (5 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_decode_radix_cache.py
cd test/ && python3 registered/disaggregation/test_disaggregation_different_tp.py
cd test/ && python3 registered/disaggregation/test_disaggregation_dp_attention.py
cd test/ && python3 registered/disaggregation/test_disaggregation_nixl.py
cd test/ && python3 registered/disaggregation/test_disaggregation_pp.py

🚀 8-gpu-h200 (3 tests): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_dsv4.py
cd test/ && python3 registered/disaggregation/test_disaggregation_hisparse.py
cd test/ && python3 registered/disaggregation/test_disaggregation_hybrid_attention.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_dwdp_gpt_oss.py

🚀 4-gpu-h100 (3 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_kimi_linear.py
cd test/ && python3 registered/disaggregation/test_disaggregation_unified_memory.py
cd test/ && python3 registered/disaggregation/test_epd_disaggregation.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_xpu.py

🚀 8-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_kimi_linear_pd_dcp4.py

@ShangmingCai

Copy link
Copy Markdown
Collaborator

Failed tests are irrelevant.

@ShangmingCai
ShangmingCai merged commit a82f8e1 into sgl-project:main Aug 13, 2026
105 of 121 checks passed
saturn-acc pushed a commit to saturn-acc/sglang that referenced this pull request Aug 16, 2026
Atituiset pushed a commit to Atituiset/sglang that referenced this pull request Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants