Skip to content

[DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios - #39921

Merged
hnyls2002 merged 18 commits into
mainfrom
lsyin/dsv4-ratio-generalization
Sep 17, 2026
Merged

hnyls2002 merged 18 commits into
mainfrom
lsyin/dsv4-ratio-generalization

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Make the set of compressed-KV ratios data owned by the KV pool (present_ratios) and give the shared code one accessor per concept keyed by ratio, instead of hard-coding c4 / c128 in field names, branches and PD payload code
  • No behavior change for V4: every existing path produces the same tensors and indices (see Verification); the new surface is consumed by dsv4.1: remaining model and runtime integration #38798

Changes

Attention metadata (DSV4AttnMetadata)

  • Rename c4_sparse_topk -> index_topk (the indexer's top-k, not a c4 property); add required present_ratios from the pool
  • Build compression metadata, FlashMLA schedules, compressor and indexer metadata only for present ratios; CP reindex treats per-ratio fields as optional; copy_ / breakable-CUDA-graph refresh check present_ratios
  • Add ratio-keyed sparse_page_indices(r), sparse_topk_lengths(r), sparse_raw_indices(r) and the writer set_sparse_topk(r, ...); call sites and the attention test kit use them, forward dispatches on compress_ratio != 0 (storage stays flat)

Sparse prefill (SparsePrefillChunkCache, combine_topk_swa_indices)

  • Kernel takes each row's absolute query_pos instead of assuming rows are the trailing extend tokens, and keeps -1 entries inside the top-k span as -1 instead of shifting them by compressed_base (both identities on V4 inputs)
  • build() takes query_lens / query_pos alongside extend_seq_lens (the SWA gather still spans the whole extend)
  • Keep per-ratio gather state in CompressedGather, keyed by ratio in cache.compressed (c128 included); both sparse-prefill forwards share cache.layer_inputs(...)

KV pool and PD state

  • Add DeepSeekV4TokenToKVPool.present_ratios; CompressStatePool.request_scoped (set by the pool factory) replaces ratio == 128 checks in the state-buffer bookkeeping
  • PD prefill/decode call request_state_transfer_indices(req_pool_idx, seq_len) / clear_request_scoped_state(req_pool_idx) on the pool instead of combining an env lookup, get_ring_size(128) and a c128 helper; the index arithmetic moves to deepseek_v4_compress_state.py (c4_state_transfer_indices, request_scoped_state_transfer_indices, CompressStatePool.transfer_indices)
  • Set mla_compression_ratios in setup_state_kv_args for both roles
  • Add get_swa_key_layout / get_extra_key_layout / get_*_bytes_per_token so the attention kernel views each cache with its own row width (the extra cache previously reused the SWA width)

Removed

  • is_dsv4_c128_online_enabled (the pool's online flag is the single source), CompressedGather.page_size (unused), get_dsv4_c4_state_indices / get_dsv4_c128_state_indices in disaggregation/utils.py (moved as above)

Verification

  • AST comparison against main per function: only the functions listed above differ
  • Differential dump of the V4 paths (metadata init for prefill / decode / trtllm, copy_, refresh, CP reindex; chunk-cache build, c128 / c4 gathers and combines; pool state-buffer bookkeeping and request-state clear; PD request-state indices) between main and this branch: 274 shared entries, 0 mismatches (main-only entries are the renamed c4_sparse_topk)
  • Registered tests pass: test_deepseek_v4.py (incl. real sparse-prefill c4 / c128 vs. torch reference), test_q8kv8_sparse_prefill_backend.py, test_disaggregation_wire.py, test_dsv4_c4_state_lifecycle.py, test_dsv4_compressed_pools.py, test_dsv4_unified_fp8_pool.py, test_deepseek_v4_compress_state_runtime_shapes.py
  • New: test/registered/kernel/attention/test_combine_topk_swa_indices.py (torch oracle: trailing extend, -1 holes, non-trailing query_pos with cross-chunk offset, SWA-only); metadata contract tests for present_ratios gating, accessor / writer routing and CP reindex; pool request-state transfer tests

CI States

Latest PR Test (Base): ✅ Run #35276391750
Latest PR Test (Extra): 🚫 Run #35276390533
Latest PR Test (AMD ROCm 10): ⏳ Run #35276390785

@hnyls2002 hnyls2002 added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
@hnyls2002 hnyls2002 added the run-ci-extra CI: also run the extra suite (requires run-ci) label Sep 17, 2026
@hnyls2002
hnyls2002 merged commit 1f0c73e into main Sep 17, 2026
193 of 230 checks passed
@hnyls2002
hnyls2002 deleted the lsyin/dsv4-ratio-generalization branch September 17, 2026 22:55
hnyls2002 added a commit that referenced this pull request Sep 17, 2026
zozyo pushed a commit to Phala-Network/sglang that referenced this pull request Sep 22, 2026
…r compress ratios (sgl-project#39921)

(cherry picked from commit 1f0c73e)
(cherry picked from commit 5648185)
pengwu22 added a commit to pengwu22/sglang that referenced this pull request Sep 24, 2026
…KV pool

sgl-project#39921 sized the compressed KV sub-pools with _num_dsv4_physical_kv_pages,
which reserves one FULL logical page (ratio physical pages) for the
allocator's dummy page. The c4 lightning-indexer pool kept the old
(size + page_size + 1) // page_size count, so the PD entries registered by
get_contiguous_buf_infos no longer share one page count: on the flash
layout (ratios 4/128, page 256) rows per entry went from 33 everywhere to
36 (c4 KV) / 33 (c4 indexer) / 160 (c128 KV).

The NIXL connector prepares one destination descriptor list per registered
entry from the slot count the decode advertises for entry 0 (the c4 KV
page count). The indexer entries are three pages shorter than that, the
descriptors walk past their registered region, and the prefill dies with
NIXL_ERR_NOT_FOUND in _prep_equal_tp_dlist. Mooncake does not prebuild
full-range lists and is unaffected.

Give DeepSeekV4IndexerPool the same global_page_size as
DeepSeekV4SingleKVPool and count its pages with the same helper, so the c4
KV and indexer entries are 36/36 again and every entry has at least the
c4 KV page count. Without a global_page_size the helper reduces to the old
count for page-aligned sizes, so the low-ratio index pools and the NPU
packed buffer keep their geometry. The NPU factory accepts the keyword.
pengwu22 added a commit to pengwu22/sglang that referenced this pull request Sep 25, 2026
…KV pool

sgl-project#39921 sized the compressed KV sub-pools with _num_dsv4_physical_kv_pages,
which reserves one FULL logical page (ratio physical pages) for the
allocator's dummy page. The c4 lightning-indexer pool kept the old
(size + page_size + 1) // page_size count, so the PD entries registered by
get_contiguous_buf_infos no longer share one page count: on the flash
layout (ratios 4/128, page 256) rows per entry went from 33 everywhere to
36 (c4 KV) / 33 (c4 indexer) / 160 (c128 KV).

The NIXL connector prepares one destination descriptor list per registered
entry from the slot count the decode advertises for entry 0 (the c4 KV
page count). The indexer entries are three pages shorter than that, the
descriptors walk past their registered region, and the prefill dies with
NIXL_ERR_NOT_FOUND in _prep_equal_tp_dlist. Mooncake does not prebuild
full-range lists and is unaffected.

Give DeepSeekV4IndexerPool the same global_page_size as
DeepSeekV4SingleKVPool and count its pages with the same helper, so the c4
KV and indexer entries are 36/36 again and every entry has at least the
c4 KV page count. Without a global_page_size the helper reduces to the old
count for page-aligned sizes, so the low-ratio index pools and the NPU
packed buffer keep their geometry. The NPU factory accepts the keyword.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail deepseek high priority jit-kernel memory-pool run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant