Skip to content

[Fix] Release consumed residual contributions - #41749

Merged
ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/release-consumed-contribution
Sep 29, 2026
Merged

ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/release-consumed-contribution

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is part of a stack (oldest at bottom):

Motivation

Under attention DP, prefill CUDA graph capture uses about 62 MiB more memory than before the layer boundary refactor (#41547–#41557). The FFN all-reduce of a dense layer is now deferred to the next layer's input, so the unreduced output — gathered over the DP ranks, 2N rows for N local tokens — is stored in the residual stream's pending contribution. The next stage's prepare() consumes it, but the model loop still holds the returned handle until the next layer returns, and the handle keeps the tensor alive. Inside the graph memory pool that block cannot be reused.

Modifications

  • Contribution.release() drops a contribution's tensors.
  • ResidualStream.write() releases the contribution it replaces. It is called only where a pending contribution is consumed: a successful prepare(), fold() and branch merge. complete_output(), snapshot() and export (to_pp(), the final norm) do not release. Stale handles were already rejected by the stream, so nothing can read a released contribution.
  • Test: a consumed partial is freed while the caller still holds its handle. It fails on the parent commit.

Accuracy Tests

B200, --tp-size 2 --dp-size 2 --enable-dp-attention, python -m sglang.benchmark.one_batch --correctness-test: the printed prefill logits and all three generations are identical to the parent commit for Qwen3-0.6B and for Qwen1.5-MoE-A2.7B.

Speed Tests and Profiling

Prefill CUDA graph capture memory, same B200 configuration (58 capture shapes, 4–8192 tokens):

Tree Capture memory
Before the refactor (6582425854) 1124.0 MiB
Parent of this PR (f731e82f09) 1186.0 MiB
This PR 1124.0 MiB

Qwen1.5-MoE-A2.7B, same flags: 2.51 GB on the parent (the same on a repeat run), 2.26 GB with this PR, 2.26 GB before the refactor.

For Qwen3-0.6B the graph pool size after every shape equals the pre-refactor tree. Kernels and collectives are unchanged, so no latency run was made.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #36623301447
Latest PR Test (Extra): ❌ Run #36623301195
Latest PR Test (AMD ROCm 10): ❌ Run #36623301173

After the next stage's prepare consumed a pending contribution, the
caller's handle still referenced its tensors until the next layer
returned. Under attention DP a deferred FFN sum therefore kept its
gathered rows alive through the whole next layer, which raises prefill
CUDA graph memory. Writing the residual now drops the consumed
contribution's tensors; stale handles were already rejected.
@ch-wan

ch-wan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_residual_stream.py

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_residual_stream.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/layer_boundary/test_residual_stream.py

@ch-wan
ch-wan merged commit 8d2d874 into main Sep 29, 2026
99 of 115 checks passed
@ch-wan
ch-wan deleted the cheng/hot-fix/release-consumed-contribution branch September 29, 2026 23:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant