Skip to content

[Refactor] Remove TCPiecewise Cuda Graph - #41634

Open
Oasis-Git wants to merge 12 commits into
sgl-project:mainfrom
Oasis-Git:refactor/remove-tcpcg
Open

Oasis-Git wants to merge 12 commits into
sgl-project:mainfrom
Oasis-Git:refactor/remove-tcpcg

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Modifications

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #36651021053
Latest PR Test (Extra): ❌ Run #36651020873
Latest PR Test (AMD ROCm 10): ❌ Run #36651021020

@mingfeima

Copy link
Copy Markdown
Collaborator

@CaoE take a look at this one. sglang is changing from PCG to BCG

arathi-hlab added a commit to arathi-hlab/sglang that referenced this pull request Oct 5, 2026
XPU CI (stage-a-test-1-gpu-xpu) fails on every PR in
test/registered/xpu/test_xpu_graph.py: the tc_piecewise prefill graph
compiles the Qwen2 decoder with fullgraph=True, and since sgl-project#42301 the
forward builds msgspec.Struct layer-boundary values that Dynamo (torch
2.13) cannot construct inside a compiled region.

tc_piecewise is being removed in sgl-project#41634, so rather than reworking the
layer-boundary classes for a graph mode that is going away, skip the test
case for now and refactor the XPU graph test once sgl-project#41634 lands.
arathi-hlab added a commit to arathi-hlab/sglang that referenced this pull request Oct 5, 2026
XPU CI (stage-a-test-1-gpu-xpu) fails on every PR in
test/registered/xpu/test_xpu_graph.py: the tc_piecewise prefill graph
compiles the Qwen2 decoder with fullgraph=True, and since sgl-project#42301 the
forward builds msgspec.Struct layer-boundary values that Dynamo (torch
2.13) cannot construct inside a compiled region.

tc_piecewise is being removed in sgl-project#41634, so rather than reworking the
layer-boundary classes for a graph mode that is going away, mark the
test disabled in its CI registration (run_suite.py then skips the file)
and refactor the XPU graph test once sgl-project#41634 lands.
nvpohanh added a commit to nvpohanh/sglang that referenced this pull request Oct 5, 2026
The file is also registered in the per-commit AMD suite
stage-b-test-1-gpu-large-amd, where it fails with the same
object.__new__(Contribution) error since the Qwen2 decoder was built
from stage boundaries, and its failure cancels the sibling partitions.
Both of its tests exist only to exercise tc_piecewise, which is being
removed (sgl-project#41634), so there is nothing to keep or move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
thanhhao98 pushed a commit to thanhhao98/sglang that referenced this pull request Oct 5, 2026
tc_piecewise is being removed (sgl-project#41634) and M3 now defaults to breakable, so
restore these files to main.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 deepseek diffusion SGLang Diffusion documentation Improvements or additions to documentation jit-kernel Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants