Skip to content

[Refactor][TCPCG] Migrate DSA and DeepSeek graph breaks to eager methods (5/9) - #42289

Open
Oasis-Git wants to merge 5 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-05-dsa-eager
Open

Oasis-Git wants to merge 5 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-05-dsa-eager

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 5/9, split from the original TCPCG removal PR #41634.

Depends on #42288 (4/9), following #42285, #42286, and #42287. Merge those steps first. This branch targets upstream main, so its diff includes the preceding steps until they land. Review this step alone using the isolated step-5 diff.

Move DSA and DeepSeek graph execution from TCPCG split-op adapters and hidden context lookups into eager methods with explicit module and batch arguments. Retain the graph-owned output buffers needed by subsequent captured segments.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — [Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285.
  2. Add explicit batch support to BCG eager regions — [Refactor][TCPCG] Bind live forward batches in BCG eager calls (2/9) #42286.
  3. Retire TCPCG backend selection, configuration, documentation, and tests — [Refactor][TCPCG] Retire backend selection and configuration (3/9) #42287.
  4. Consolidate Radix attention eager regions and isolate Inkling-specific behavior — [Refactor][TCPCG] Consolidate attention eager regions and isolate Inkling policy (4/9) #42288.
  5. Migrate DSA and DeepSeek eager regions to explicit methods and batch arguments — this PR.
  6. Migrate MoE and Mamba eager regions.
  7. Migrate diffusion eager wrappers and standardize their names. This step depends only on step 2.
  8. Remove obsolete runtime contexts, layer registries, split-op plumbing, and temporary compatibility after caller migration.
  9. Simplify quantization paths separately because of their torch.compile implications.

Only steps 1–5 are submitted upstream so far. Later steps will be submitted after review. The disputed relocation of inactive compiler/reference files is excluded pending a separate decision.

Changes

  • Move DSA indexer and kpool prefill graph execution into class-owned eager methods, using explicit forward_batch arguments and the caller's padded output buffers.
  • Delete dsa_prefill_cuda_graph.py and kpool_prefill_cuda_graph.py after migrating their consumers. The head-gate custom ops and fake implementations relocated in step 1 remain available for torch.compile.
  • Migrate DeepSeek V2/V4 graph wrappers, including attention, source projections, and Engram hash-ID handling, away from TCPCG context lookups.
  • Replace legacy graph predicates in DSA and the related FlashInfer/MLA attention paths with the appropriate BCG or full-prefill checks. Keep the non-speculative prefill distinction where required.
  • Update the MLA wrapper migrated in step 4 to use the renamed DSA BCG-prefill predicate, and remove obsolete TCPCG mocks from the DSV4 indexer tests.

This step relies on replay-time batch binding from step 2 and the attention/ForwardContext migration in step 4. The remaining legacy runtime context is retained for callers migrated in later steps.

Validation

  • Pre-commit and Python syntax checks passed on the changed files.
  • CUDA/ROCm indexer execution, graph capture/replay, end-to-end model accuracy, and full main CI at this intermediate step remain to be validated. The local macOS checks do not establish accelerator correctness.

CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 deepseek documentation Improvements or additions to documentation Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant