Skip to content

[Refactor][TCPCG] Consolidate attention eager regions and isolate Inkling policy (4/9) - #42288

Open
Oasis-Git wants to merge 1 commit into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-04-attention-eager
Open

Oasis-Git wants to merge 1 commit into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-04-attention-eager

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 4/9, split from the original TCPCG removal PR #41634.

Steps 1–3 are merged. Rebased onto upstream main (58cb1a2be6); the PR now contains only step 4 as one commit. Review the step-4 files.

Replace attention's TCPCG split-op adapters and layer-registry lookups with class-owned eager methods. On BCG replay, these methods read the live prepared batch through get_forward_batch() from step 2's runner-owned scope. The eager decorator and replay callbacks do not capture or rebind a batch argument.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — [Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285 (merged).
  2. Simplify BCG eager replay and publish the live batch in runner utilities — [Refactor][TCPCG] Bind live forward batches in BCG eager calls (2/9) #42286 (merged).
  3. Retire TCPCG backend selection, configuration, documentation, and tests — [Refactor][TCPCG] Retire backend selection and configuration (3/9) #42287 (merged).
  4. Consolidate attention eager regions and isolate Inkling policy — this PR.
  5. Migrate DSA and DeepSeek eager regions — [Refactor][TCPCG] Migrate DSA and DeepSeek graph breaks to eager methods (5/9) #42289.
  6. Migrate MoE and Mamba eager regions — [Refactor][TCPCG] Migrate MoE and Mamba graph execution to eager methods (6/9) #42290.
  7. Migrate diffusion eager wrappers and standardize names — [Refactor][TCPCG] Standardize diffusion eager graph wrappers (7/9) #42291; depends only on step 2.
  8. Remove obsolete runtime contexts, layer registries, split-op plumbing, and temporary compatibility after caller migration — [Refactor][TCPCG] Remove obsolete runtime registries and split-op plumbing (8/9) #42292.
  9. Simplify quantization paths separately because of their torch.compile implications — [Refactor][TCPCG] Simplify quantization paths after TCPCG retirement (9/9) #42293.

Changes

  • Consolidate dense attention, optional LSE, sparse handling, and extra keyword arguments in RadixAttention._eager_attention. Preserve main’s sparse-BCG fix: sparse attention also runs in an eager region and reads the live batch during replay.
  • Share output allocation, token slicing, temporary batch-state restoration, and padded-tail zeroing through attention/graph_utils.py. Preserve the newer eager padded-extend handling, including independent query/KV extents and tuple outputs.
  • Migrate linear attention, HPC RoPE/store-KV, and MLA BMM attention to class-owned eager methods. Preserve the newer linear-attention path that stays inside BCG when the backend supports it.
  • Keep Inkling's attention bypass and token handling in the model, without making the eager decorator nesting-aware.
  • Keep Cheng's ForwardContext unchanged. Publish the full-prefill flag and raw token count in a separate runner-owned prefill_graph_scope, with restoration on exit.

The legacy TCPCG context and layer registries remain temporarily for unmigrated callers until step 8. This changes attention dispatch and eager replay, while preserving backend computation and graph-owned output buffers.

Validation

  • Pre-commit passed on all 14 files in the step-4 diff.
  • Attention, padded-extend, linear-attention, and Inkling fusion-gating CPU tests: 54 passed plus 24 subtests. These exercise real implementations with a macOS import shim for accelerator-dependent imports.
  • Radix attention's CI-style script entry (python file.py -f): 24 passed under the same shim.
  • CI collection sanity check: 2,243 test files / 2,741 registrations, including executable test-entry validation.
  • Added dense/LSE/sparse live-batch replay, nested scope restoration, and linear-attention in-graph dispatch coverage. Updated the GPU attention regression test for the new class method; it has not been run locally.
  • CUDA capture/replay, model accuracy, performance, and full CI remain unvalidated. This update is prepared for code review; it does not claim GPU validation.

CI States

Latest PR Test (Base): ❌ Run #38111399643
Latest PR Test (Extra): ❌ Run #38111399479
Latest PR Test (AMD ROCm 10): ❌ Run #38111399624

@Oasis-Git
Oasis-Git force-pushed the refactor/tcpcg-04-attention-eager branch from c5cabe3 to fed12ee Compare October 11, 2026 04:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek documentation Improvements or additions to documentation Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant