Skip to content

[Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) - #42285

Merged
Oasis-Git merged 3 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-01-shared-helpers
Oct 6, 2026
Merged

Oasis-Git merged 3 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-01-shared-helpers

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 1/9, split from the original TCPCG removal PR #41634. It relocates shared helpers so BCG and torch.compile support can remain independent of the later TCPCG removal. Existing graph execution and TCPCG support remain in place at this step.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — this PR ([Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285).
  2. Add explicit batch support to BCG eager regions, with temporary compatibility for existing decorator callers.
  3. Retire TCPCG backend selection and configuration; remove TCPCG-only documentation and tests while retaining internal machinery needed by unmigrated callers.
  4. Consolidate Radix attention eager regions and isolate Inkling-specific behavior, including the dependent MLA wrapper migration.
  5. Migrate DSA and DeepSeek eager regions to explicit methods and batch arguments.
  6. Migrate MoE and Mamba eager regions.
  7. Migrate diffusion eager wrappers and standardize their names. This step depends only on step 2.
  8. Remove obsolete runtime contexts, layer registries, split-op plumbing, and temporary compatibility after caller migration.
  9. Simplify quantization paths separately because of their torch.compile implications.

Only step 1 is submitted upstream so far. Later steps will be submitted after review. This PR has no dependency on those later steps. The disputed relocation of inactive compiler/reference files is excluded from this series pending a separate decision.

Changes

  • Move weak_ref_tensor.py from compilation/ to model_executor/runner_backend_utils/ and update both the BCG and TCPCG imports.
  • Move the four DSA head-gate implementations into attention/dsa/head_gate.py, retaining their CUDA/ROCm guards, custom-op registration, and fake implementations.
  • Update the DSA consumer import and platform-guard test to use the new module.

The weak tensor helper is byte-identical to the original. All four relocated DSA helper implementations are AST-identical; only their surrounding module and comments change.

Validation

  • Pre-commit passed on the changed files.
  • Verified helper equivalence as described above.
  • GPU execution was not validated locally; the development machine is macOS. Upstream CI results are tracked below.

CI States

Latest PR Test (Base): ❌ Run #37249466699
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ❌ Run #37249466729

@Oasis-Git

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@Oasis-Git Oasis-Git added the bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) label Oct 5, 2026
@Oasis-Git

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@Oasis-Git

Copy link
Copy Markdown
Collaborator Author

@Oasis-Git
Oasis-Git merged commit 18ec2ec into sgl-project:main Oct 6, 2026
411 of 463 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) piecewise-cuda-graph run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant