Skip to content

fix(nixl): prevent early staging buffer reuse and duplicate sends - #38100

Open
fly-go-run wants to merge 1 commit into
sgl-project:mainfrom
fly-go-run:codex/nixl-staging-send-safety
Open

fly-go-run wants to merge 1 commit into
sgl-project:mainfrom
fly-go-run:codex/nixl-staging-send-safety

Conversation

@fly-go-run

@fly-go-run fly-go-run commented Sep 5, 2026 •

Copy link
Copy Markdown

Closes #38099.

Motivation

Summary

While investigating incorrect output in a production deployment with separate prefill and decode workers, we found two defects in NIXL staging sends.

A worker uses one staging buffer to prepare the key/value (KV) cache data for each destination. It can prepare B's data while A's asynchronous transfer is still reading that buffer. If B instead has to wait for staging space, the worker puts the entire data chunk back in its queue and can resend A, including its attached state and metadata.

This PR fixes those two send-path defects. They are plausible contributors to the production symptoms, but the particular production failure has not been reproduced with this patch alone.

Modifications

  • Wait for a staged KV transfer to finish before the next destination can overwrite its source buffer.
  • Wait for all transfers already submitted before continuing after a staging retry, including additional state, metadata and transfers spanning different memory types.
  • Record completed destinations on the existing TransferKVChunk, so a retry processes only unfinished destinations.
  • Add three regression tests for source overwrite, duplicate sends and unfinished transfers at the next queue iteration.

The change uses the existing NIXL completion states. Other paths that prepare data in separate buffer regions still wait for their transfers before reusing those regions. Handling unfinished transfers after a transport error remains separate work in #36612 and #36707.

Accuracy Tests

Validation

Tested against db89f639ef475821e6669958cfe22caf914df022:

  • The three added regressions fail on the original code and pass with the fix.
  • test_nixl_backend_basic.py: 47 passed, plus 2 subtests.
  • test_nixl_deferred_kv_release.py, test_nixl_sender_failure_cleanup.py and test_disaggregation_wire.py: 42 passed, plus 10 subtests.
  • All applicable pre-commit hooks passed for the three changed files, including formatting, lint and registered-test checks. git diff --check also passed.

An additional local review ran four checks through the actual staged-send method, including another request running between retries and repeated waits for space. All four passed with the fix. Three multi-destination cases failed on the original code; the single-destination case passed on both. These extra checks are local review artifacts, not part of the PR's test file.

These tests check the real worker logic on macOS/Python 3.13. Triton imports are stubbed for the CPU tests, torch.compile is disabled, and GPU operations and transfers are simulated. Separately, a local implementation of these fixes has already been deployed to production and load-tested on real GPUs, with normal operation reported. That deployment includes other fixes; the exact minimal patch in this PR has not had a standalone GPU comparison.

Speed Tests and Profiling

The deployed local implementation has undergone production GPU load testing. An isolated before/after throughput comparison for this exact upstream patch has not been run.

Waiting for completion reduces overlap between destinations sharing the source buffer. The current placement also waits for staged KV before sending the final state and metadata, so a single-destination request may lose some overlap too.

A targeted GPU test with a destination temporarily lacking staging space would provide additional coverage for the retry behavior. Separate source buffers could restore overlap, but would require a larger change.

Checklist

  • Format the changed code and run the repository's lint checks.
  • Add regression tests and run the relevant CPU tests.
  • Deploy the local implementation and run production GPU load tests.
  • Run the targeted multi-destination regressions on real GPUs for this exact upstream patch.
  • Provide an isolated before/after GPU throughput comparison with the same model and deployment settings.
  • Link the companion bug report: [Bug] NIXL staging reuses the send buffer too early and repeats completed sends #38099.

Related: #36893 (similar Mooncake retry issue), #36612 and #36707 (error handling), and #37697 (out-of-order decode completion).


CI States

Latest PR Test (Base): ❌ Run #33943413770
Latest PR Test (Extra): ❌ Run #33943413643
Latest PR Test (AMD ROCm 7.2): ❌ Run #33943413697

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] NIXL staging reuses the send buffer too early and repeats completed sends

1 participant