Skip to content

[Refactor][TCPCG] Migrate MoE and Mamba graph execution to eager methods (6/9) - #42290

Open
Oasis-Git wants to merge 6 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-06-moe-mamba-eager
Open

Oasis-Git wants to merge 6 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-06-moe-mamba-eager

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 6/9, split from the original TCPCG removal PR #41634.

Depends on #42289 (5/9); merge those dependencies first. This branch targets upstream main, so its diff includes unmerged prerequisites. Review this step alone using the isolated step-6 diff.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — [Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285.
  2. Add explicit batch support to BCG eager regions — [Refactor][TCPCG] Bind live forward batches in BCG eager calls (2/9) #42286.
  3. Retire TCPCG backend selection, configuration, documentation, and tests — [Refactor][TCPCG] Retire backend selection and configuration (3/9) #42287.
  4. Consolidate Radix attention eager regions and isolate Inkling-specific behavior — [Refactor][TCPCG] Consolidate attention eager regions and isolate Inkling policy (4/9) #42288.
  5. Migrate DSA and DeepSeek eager regions — [Refactor][TCPCG] Migrate DSA and DeepSeek graph breaks to eager methods (5/9) #42289.
  6. Migrate MoE and Mamba eager regions — this PR.
  7. Migrate diffusion eager wrappers. Depends only on step 2.
  8. Remove obsolete runtime contexts, registries, and split-op plumbing.
  9. Simplify quantization paths separately because of their torch.compile implications.

The disputed relocation of inactive compiler/reference files is excluded pending a separate decision.

Changes

Move the remaining MoE and Mamba graph execution away from TCPCG split-op adapters and hidden batch/module lookups.

  • Use a class-owned eager method for expert all-to-all, preserving its capture stub and stable output buffer.
  • Run Nemotron Mamba through an eager method with an explicit batch, real-token slicing, padded-tail zeroing, and replay-time collective policy restoration.
  • Remove TCPCG-only fused-MoE dispatch and the Qwen3-Next TC-specific normalization choice.

Runtime registries remain until the cleanup step. This step uses the batch binding introduced in step 2.

Validation

Pre-commit and Python syntax checks passed on the changed files. Distributed MoE collectives, Mamba CUDA replay, model accuracy, and full main CI at this intermediate step remain to be validated; local macOS checks do not establish accelerator correctness.


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 deepseek documentation Improvements or additions to documentation Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant