Skip to content

[DSv4.1] Enable DeepGEMM sparse prefill under CP - #40613

Open
Silas-Zeng wants to merge 3 commits into
sgl-project:mainfrom
Silas-Zeng:feat/dsv41-cp-sparse-prefill
Open

Silas-Zeng wants to merge 3 commits into
sgl-project:mainfrom
Silas-Zeng:feat/dsv41-cp-sparse-prefill

Conversation

@Silas-Zeng

@Silas-Zeng Silas-Zeng commented Sep 21, 2026 •

Copy link
Copy Markdown

Problem

This addresses item 3 of #40574 and follows #40352.

CP prefill currently scores every consumer index layer densely. The source publishes 2,048 blocks of 8 compressed positions per query row, but CP local rows do not line up with the shared page table, so consumers fall back to dense scoring. This change scores only the candidate blocks published by the indexer.

Change

  • Reuse the page table built from req_to_token and reordered by apply_cp_reindex, including compact tails.
  • Skip tail-table construction when a rank has no remaining tail queries.
  • Add coverage for row ordering, compression ratios 1 and 2, shuffled pages, request boundaries, tail publication, and signed weights.
  • Register the GPU tests with the B200 CI stage.

Validation

The GPU validation used 2x NVIDIA B200 (SM100), driver 580.126.09, Ubuntu 24.04, Python 3.12.3, CUDA 13.0, PyTorch 2.13.0+cu130, and sgl-deep-gemm 0.2.0.

  • CPU indexer and candidate-block tests: 28 tests passed, 68 subtests.
  • B200 sparse-prefill kernel suite: 9 passed, 0 skipped before the rebase.
  • The empty-tail regression fails with the base _publish_prefill implementation and passes with this change.
  • Boundary and ragged-layout stress: 40/40 positive-weight and 40/40 negative-weight cases passed.
  • Mixed signed-weight stress: 36/40 passed the existing dense-score cutoff.
  • Forced score chunking: 6 simulated CP ranks passed across CP2 and CP4 layouts.
  • Two-rank CP component checks covered compression ratios 1 and 2, seeds 17 and 101, and 96/97-row layouts with real CP padding: 8/8 structural, BF16 top-k, slot, and cache checks.
  • CUDA memcheck: 9 GPU tests passed with 0 errors before the rebase.

Existing score-floor thresholds were unchanged. The seed-17 failures match the dependency baseline at compression ratios 1 and 2. BF16 top-k, physical-slot, and cache checks pass in all eight scenarios. The rebase changed the in-tree test harness; CI should rerun the GPU suite on the final head.

Benchmark

Single-GPU component benchmark with random inputs and a simulated CP2 rank-0 layout. Five warmups and 30 synchronized timed runs per arm, using the same inputs and resident page tables/output buffers while alternating the dense and sparse arms. Peak memory is measured above the shared resident baseline:

Local query rows Context tokens Dense consumer (ms) Sparse consumer (ms) Dense peak delta (MiB) Sparse peak delta (MiB)
128 4,096 0.343 0.129 3.50 4.31
128 131,072 0.573 0.129 96.25 4.45
256 4,096 0.344 0.131 7.00 8.63
256 131,072 0.685 0.137 192.50 8.63

At 4K context, sparse scoring uses about 23% more temporary allocated memory.


CI States

Latest PR Test (Base): ❌ Run #35767171561
Latest PR Test (Extra): ❌ Run #35767170888
Latest PR Test (AMD ROCm 10): ❌ Run #35767171044

@Silas-Zeng
Silas-Zeng force-pushed the feat/dsv41-cp-sparse-prefill branch from d543204 to d7258b3 Compare September 22, 2026 18:11

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant