Skip to content

[CI] Drop the redundant and broken tc_piecewise tests - #42571

Merged
ch-wan merged 3 commits into
sgl-project:mainfrom
nvpohanh:claude/drop-tc-piecewise-spec-tests
Oct 6, 2026
Merged

ch-wan merged 3 commits into
sgl-project:mainfrom
nvpohanh:claude/drop-tc-piecewise-spec-tests

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

[by Claude Code]

Motivation

The tc_piecewise tests for Llama and Qwen2 models fail on main while capturing the prefill CUDA graph:

torch._dynamo.exc.Unsupported: Unsupported object.__new__ user-defined class construction
  class=<class 'sglang.srt.layers.layer_boundary.residual.stream.Contribution'>,
  error=object.__new__(Contribution) is not safe, use Contribution.__new__()

tc_piecewise compiles the decoder with fullgraph=True. Since the Qwen2 and Llama decoders were built from stage boundaries (#42307, #42308), their forward constructs msgspec.Struct layer-boundary values that Dynamo cannot construct. Failing runs on main:

Per discussion with @ch-wan, tc_piecewise will be removed soon (#41634), so these tests should move to breakable or be removed rather than be fixed.

Modifications

Test Change Why
piecewise/test_pcg_with_speculative_decoding_dflash.py (weekly) Deleted Under breakable it duplicates the per-commit spec/dflash/test_dflash.py overlap variant: same Llama-3.1-8B + DFlash draft, flashinfer, page size 1, 64 running requests, --mem-fraction-static 0.7, decode batch sizes 1-64, and thresholds 0.75 / 2.8.
TestPCGWithSTANDALONE (weekly) Deleted Duplicates the per-commit spec/test_spec_standalone.py: same Llama-3.1-8B + 3.2-1B pair, with stricter thresholds (0.69 / 3.6).
TestPCGWithNGRAM (weekly) Deleted Duplicates the per-commit spec/test_spec_ngram.py (+ _extra): same Qwen2.5-Coder-7B, with stricter thresholds (0.79 / 1.8).
TestPCGWithMTP (weekly) Moved to breakable/test_bcg_with_speculative_decoding_extra.py as TestBCGWithMTP, config unchanged, est_time 450 -> 240 It never selected tc_piecewise and already ran on breakable; it passed in the weekly job above (score 0.97, accept length 3.46). It is the only Qwen3.5-35B-A3B FP8 TP2 MTP check.
piecewise/test_piecewise_cuda_graph_support_1_gpu.py (CUDA nightly + AMD per-commit) Deleted Both of its tests exist only to exercise tc_piecewise: GSM8K with tc_piecewise forced, and an embedding comparison between tc_piecewise and no prefill graph. It cannot move to breakable, because Qwen2.5-VL is not on the multimodal breakable allowlist. It fails on both CUDA and AMD.

piecewise/test_pcg_with_speculative_decoding.py (EAGLE3 on Qwen3-30B-A3B) is left for #41634. It still passes because Qwen3-MoE is not built from stage boundaries yet, and its breakable twin breakable/test_bcg_with_speculative_decoding.py already runs per-commit.

Accuracy Tests

This PR changes tests only. Validation:

  • pre-commit run --all-files passes.
  • ci_register.collect_tests registers MTP as weekly / 4-gpu-h100 / 240 s. No registered test references the removed files or classes.
  • On 1x H100 80GB HBM3 (lmsysorg/sglang:dev-cu13, SGLang f70e8c6, torch 2.14.1+cu130), the deleted tc_piecewise tests all fail with the error above. The same configs on breakable (the prefill backend their per-commit equivalents use) pass:
Config tc_piecewise breakable: GSM8K score / avg accept length
DFlash (Llama-3.1-8B) ❌ Contribution ✅ 0.845 / 3.92
STANDALONE (Llama-3.1-8B + 3.2-1B)* ❌ Contribution ✅ 0.870 / 3.33
NGRAM (Qwen2.5-Coder-7B) ❌ Contribution ✅ 0.780 / 1.93

* The draft was the ungated mirror unsloth/Llama-3.2-1B-Instruct, because the local HF token cannot access meta-llama/Llama-3.2-1B-Instruct.

Speed Tests and Profiling

N/A (test-only change).

Checklist

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #37301553658
Latest PR Test (Extra): ❌ Run #37301553150
Latest PR Test (AMD ROCm 10): ❌ Run #37301553443

…e nightly VLM one

tc_piecewise compiles the decoder with fullgraph=True, and since the Qwen2
and Llama decoders were built from stage boundaries (sgl-project#42307, sgl-project#42308)
Dynamo cannot construct their per-forward msgspec.Struct layer-boundary
values, so these launches die during prefill graph capture. tc_piecewise
is being removed (sgl-project#41634).

- Delete the weekly PCG DFLASH / STANDALONE / NGRAM tests. Under the
  default breakable backend they duplicate the per-commit
  spec/dflash/test_dflash.py, spec/test_spec_standalone.py and
  spec/test_spec_ngram.py, which use the same model pairs.
- Keep the MTP variant, which already ran on breakable, as
  breakable/test_bcg_with_speculative_decoding_extra.py.
- Disable the CUDA registration of the nightly Qwen2.5-VL tc_piecewise
  test; Qwen2.5-VL is not on the multimodal breakable allowlist, so it
  cannot move. The AMD registration is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The file is also registered in the per-commit AMD suite
stage-b-test-1-gpu-large-amd, where it fails with the same
object.__new__(Contribution) error since the Qwen2 decoder was built
from stage boundaries, and its failure cancels the sibling partitions.
Both of its tests exist only to exercise tc_piecewise, which is being
removed (sgl-project#41634), so there is nothing to keep or move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nvpohanh nvpohanh changed the title [CI] Drop the redundant tc_piecewise speculative tests and disable the nightly VLM one [CI] Drop the redundant and broken tc_piecewise tests Oct 5, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nvpohanh

nvpohanh commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/cuda_graph/breakable/test_bcg_with_speculative_decoding_extra.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_standalone.py test/registered/spec/test_spec_ngram.py test/registered/spec/test_spec_ngram_extra.py test/registered/spec/test_spec_standalone_extra.py test/registered/e2e/speculative/test_dflash_domino.py

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/cuda_graph/breakable/test_bcg_with_speculative_decoding_extra.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/spec/dflash/test_dflash.py

🚀 1-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/spec/test_spec_standalone.py
cd test/ && python3 registered/spec/test_spec_ngram.py
cd test/ && python3 registered/spec/test_spec_ngram_extra.py
cd test/ && python3 registered/spec/test_spec_standalone_extra.py

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/speculative/test_dflash_domino.py

@ch-wan
ch-wan merged commit 62af5f5 into sgl-project:main Oct 6, 2026
104 of 116 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants