Skip to content

[Refactor][TCPCG] Remove obsolete runtime registries and split-op plumbing (8/9) - #42292

Open
Oasis-Git wants to merge 9 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-08-retire-machinery
Open

Oasis-Git wants to merge 9 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-08-retire-machinery

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 8/9, split from the original TCPCG removal PR #41634.

Depends on #42290 (6/9) and #42291 (7/9); merge those dependencies first. This branch targets upstream main, so its diff includes unmerged prerequisites. Review this step alone using the isolated step-8 diff.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — [Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285.
  2. Add explicit batch support to BCG eager regions — [Refactor][TCPCG] Bind live forward batches in BCG eager calls (2/9) #42286.
  3. Retire TCPCG backend selection, configuration, documentation, and tests — [Refactor][TCPCG] Retire backend selection and configuration (3/9) #42287.
  4. Consolidate Radix attention eager regions and isolate Inkling-specific behavior — [Refactor][TCPCG] Consolidate attention eager regions and isolate Inkling policy (4/9) #42288.
  5. Migrate DSA and DeepSeek eager regions — [Refactor][TCPCG] Migrate DSA and DeepSeek graph breaks to eager methods (5/9) #42289.
  6. Migrate MoE and Mamba eager regions — [Refactor][TCPCG] Migrate MoE and Mamba graph execution to eager methods (6/9) #42290.
  7. Migrate diffusion eager wrappers — [Refactor][TCPCG] Standardize diffusion eager graph wrappers (7/9) #42291. Depends only on step 2.
  8. Remove obsolete runtime contexts, registries, and split-op plumbing — this PR.
  9. Simplify quantization paths separately because of their torch.compile implications.

The disputed relocation of inactive compiler/reference files is excluded pending a separate decision.

Changes

After caller migration, remove the legacy runtime wiring that BCG no longer needs.

  • Stop publishing the TCPCG forward context and remove runtime module registries and split-op registration use.
  • Replace attention/MoE layer collections used for graph eligibility with an attention-layer count and companion-layer flag.
  • Remove obsolete graph branches and transitional eager-decorator compatibility; update runners, platform hooks, collectives, and test utilities together.
  • Keep the torch.compile framework and torch_compile_decoration.

The isolated comparison uses refactor/tcpcg-migrations-base, an integration base containing both the SRT migration through step 6 and the independent diffusion migration in step 7. It therefore shows only this cleanup, without repeating those migrations.

Inactive compiler/backend/context files remain at their original paths pending the separate reference-code placement decision. Quantization-specific context indirection is handled in step 9.

Validation

Pre-commit passed. The combined CPU suite at this stage passed 68 tests plus 16 subtests using the macOS import shim. Accelerator graph capture/replay, end-to-end model accuracy, and full main CI at this intermediate stage remain to be validated.


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 deepseek diffusion SGLang Diffusion documentation Improvements or additions to documentation Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant