Repository navigation
[CI] Drop the redundant and broken tc_piecewise tests - #42571
Merged
ch-wan merged 3 commits intoOct 6, 2026
Merged
Conversation
…e nightly VLM one tc_piecewise compiles the decoder with fullgraph=True, and since the Qwen2 and Llama decoders were built from stage boundaries (sgl-project#42307, sgl-project#42308) Dynamo cannot construct their per-forward msgspec.Struct layer-boundary values, so these launches die during prefill graph capture. tc_piecewise is being removed (sgl-project#41634). - Delete the weekly PCG DFLASH / STANDALONE / NGRAM tests. Under the default breakable backend they duplicate the per-commit spec/dflash/test_dflash.py, spec/test_spec_standalone.py and spec/test_spec_ngram.py, which use the same model pairs. - Keep the MTP variant, which already ran on breakable, as breakable/test_bcg_with_speculative_decoding_extra.py. - Disable the CUDA registration of the nightly Qwen2.5-VL tc_piecewise test; Qwen2.5-VL is not on the multimodal breakable allowlist, so it cannot move. The AMD registration is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The file is also registered in the per-commit AMD suite stage-b-test-1-gpu-large-amd, where it fails with the same object.__new__(Contribution) error since the Qwen2 decoder was built from stage boundaries, and its failure cancels the sibling partitions. Both of its tests exist only to exercise tc_piecewise, which is being removed (sgl-project#41634), so there is nothing to keep or move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
Author
|
/rerun-test test/registered/cuda_graph/breakable/test_bcg_with_speculative_decoding_extra.py test/registered/spec/dflash/test_dflash.py test/registered/spec/test_spec_standalone.py test/registered/spec/test_spec_ngram.py test/registered/spec/test_spec_ngram_extra.py test/registered/spec/test_spec_standalone_extra.py test/registered/e2e/speculative/test_dflash_domino.py |
Contributor
|
🚀 🚀 🚀 🚀 |
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[by Claude Code]
Motivation
The
tc_piecewisetests for Llama and Qwen2 models fail onmainwhile capturing the prefill CUDA graph:tc_piecewisecompiles the decoder withfullgraph=True. Since the Qwen2 and Llama decoders were built from stage boundaries (#42307, #42308), their forward constructsmsgspec.Structlayer-boundary values that Dynamo cannot construct. Failing runs onmain:TestPCGWithDFlash(Llama-3.1-8B): https://github.com/sgl-project/sglang/actions/runs/37166616108/job/111330769047TestPCGWithNGRAM(Qwen2.5-Coder-7B): https://github.com/sgl-project/sglang/actions/runs/37166616108/job/111330769003.TestPCGWithSTANDALONE(Llama) did not run in that job, because the run stopped at the NGRAM failure; it fails the same way locally (see below).TestPiecewiseCudaGraphQwen25VL: https://github.com/sgl-project/sglang/actions/runs/37131301136/job/111226858313stage-b-test-1-gpu-large-amd, which fails with the same error and cancels the sibling partitions: https://github.com/sgl-project/sglang/actions/runs/37253886036/job/111590592756Per discussion with @ch-wan,
tc_piecewisewill be removed soon (#41634), so these tests should move tobreakableor be removed rather than be fixed.Modifications
piecewise/test_pcg_with_speculative_decoding_dflash.py(weekly)breakableit duplicates the per-commitspec/dflash/test_dflash.pyoverlap variant: same Llama-3.1-8B + DFlash draft, flashinfer, page size 1, 64 running requests,--mem-fraction-static 0.7, decode batch sizes 1-64, and thresholds 0.75 / 2.8.TestPCGWithSTANDALONE(weekly)spec/test_spec_standalone.py: same Llama-3.1-8B + 3.2-1B pair, with stricter thresholds (0.69 / 3.6).TestPCGWithNGRAM(weekly)spec/test_spec_ngram.py(+_extra): same Qwen2.5-Coder-7B, with stricter thresholds (0.79 / 1.8).TestPCGWithMTP(weekly)breakable/test_bcg_with_speculative_decoding_extra.pyasTestBCGWithMTP, config unchanged,est_time450 -> 240tc_piecewiseand already ran onbreakable; it passed in the weekly job above (score 0.97, accept length 3.46). It is the only Qwen3.5-35B-A3B FP8 TP2 MTP check.piecewise/test_piecewise_cuda_graph_support_1_gpu.py(CUDA nightly + AMD per-commit)tc_piecewise: GSM8K withtc_piecewiseforced, and an embedding comparison betweentc_piecewiseand no prefill graph. It cannot move tobreakable, because Qwen2.5-VL is not on the multimodalbreakableallowlist. It fails on both CUDA and AMD.piecewise/test_pcg_with_speculative_decoding.py(EAGLE3 on Qwen3-30B-A3B) is left for #41634. It still passes because Qwen3-MoE is not built from stage boundaries yet, and itsbreakabletwinbreakable/test_bcg_with_speculative_decoding.pyalready runs per-commit.Accuracy Tests
This PR changes tests only. Validation:
pre-commit run --all-filespasses.ci_register.collect_testsregisters MTP as weekly /4-gpu-h100/ 240 s. No registered test references the removed files or classes.lmsysorg/sglang:dev-cu13, SGLangf70e8c6, torch 2.14.1+cu130), the deletedtc_piecewisetests all fail with the error above. The same configs onbreakable(the prefill backend their per-commit equivalents use) pass:tc_piecewisebreakable: GSM8K score / avg accept lengthContributionContributionContribution* The draft was the ungated mirror
unsloth/Llama-3.2-1B-Instruct, because the local HF token cannot accessmeta-llama/Llama-3.2-1B-Instruct.Speed Tests and Profiling
N/A (test-only change).
Checklist
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ❌ Run #37301553658
Latest PR Test (Extra): ❌ Run #37301553150
Latest PR Test (AMD ROCm 10): ❌ Run #37301553443