Repository navigation
[DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios - #39921
Merged
Merged
Conversation
hnyls2002
requested review from
BBuf,
ByronHsu,
DarkSharpness,
Duyi-Wang,
Fridge003,
HaiShaw,
HydraQYH,
Qiaolin-Yu,
ShangmingCai,
Ying1123,
alphabetc1,
celve,
hanming-lu,
hebiao064,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
sogalin,
xiezhq-hermann,
yizhang2077 and
yuan-luo
as code owners
September 17, 2026 07:06
…te helper as pool method; clear_request_scoped_state
…_INDEX_TOPK; comments
hnyls2002
requested review from
iforgetmyname,
ping1jing2 and
whybeyoung
as code owners
September 17, 2026 07:58
hnyls2002
added a commit
that referenced
this pull request
Sep 17, 2026
…; set_sparse_topk; swa layout from pool
…est-state transfer
hnyls2002
added a commit
that referenced
this pull request
Sep 17, 2026
This was referenced Sep 19, 2026
zozyo
pushed a commit
to Phala-Network/sglang
that referenced
this pull request
Sep 22, 2026
…r compress ratios (sgl-project#39921) (cherry picked from commit 1f0c73e) (cherry picked from commit 5648185)
pengwu22
added a commit
to pengwu22/sglang
that referenced
this pull request
Sep 24, 2026
…KV pool sgl-project#39921 sized the compressed KV sub-pools with _num_dsv4_physical_kv_pages, which reserves one FULL logical page (ratio physical pages) for the allocator's dummy page. The c4 lightning-indexer pool kept the old (size + page_size + 1) // page_size count, so the PD entries registered by get_contiguous_buf_infos no longer share one page count: on the flash layout (ratios 4/128, page 256) rows per entry went from 33 everywhere to 36 (c4 KV) / 33 (c4 indexer) / 160 (c128 KV). The NIXL connector prepares one destination descriptor list per registered entry from the slot count the decode advertises for entry 0 (the c4 KV page count). The indexer entries are three pages shorter than that, the descriptors walk past their registered region, and the prefill dies with NIXL_ERR_NOT_FOUND in _prep_equal_tp_dlist. Mooncake does not prebuild full-range lists and is unaffected. Give DeepSeekV4IndexerPool the same global_page_size as DeepSeekV4SingleKVPool and count its pages with the same helper, so the c4 KV and indexer entries are 36/36 again and every entry has at least the c4 KV page count. Without a global_page_size the helper reduces to the old count for page-aligned sizes, so the low-ratio index pools and the NPU packed buffer keep their geometry. The NPU factory accepts the keyword.
pengwu22
added a commit
to pengwu22/sglang
that referenced
this pull request
Sep 25, 2026
…KV pool sgl-project#39921 sized the compressed KV sub-pools with _num_dsv4_physical_kv_pages, which reserves one FULL logical page (ratio physical pages) for the allocator's dummy page. The c4 lightning-indexer pool kept the old (size + page_size + 1) // page_size count, so the PD entries registered by get_contiguous_buf_infos no longer share one page count: on the flash layout (ratios 4/128, page 256) rows per entry went from 33 everywhere to 36 (c4 KV) / 33 (c4 indexer) / 160 (c128 KV). The NIXL connector prepares one destination descriptor list per registered entry from the slot count the decode advertises for entry 0 (the c4 KV page count). The indexer entries are three pages shorter than that, the descriptors walk past their registered region, and the prefill dies with NIXL_ERR_NOT_FOUND in _prep_equal_tp_dlist. Mooncake does not prebuild full-range lists and is unaffected. Give DeepSeekV4IndexerPool the same global_page_size as DeepSeekV4SingleKVPool and count its pages with the same helper, so the c4 KV and indexer entries are 36/36 again and every entry has at least the c4 KV page count. Without a global_page_size the helper reduces to the old count for page-aligned sizes, so the low-ratio index pools and the NPU packed buffer keep their geometry. The NPU factory accepts the keyword.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
present_ratios) and give the shared code one accessor per concept keyed by ratio, instead of hard-coding c4 / c128 in field names, branches and PD payload codeChanges
Attention metadata (
DSV4AttnMetadata)c4_sparse_topk->index_topk(the indexer's top-k, not a c4 property); add requiredpresent_ratiosfrom the poolcopy_/ breakable-CUDA-graph refresh checkpresent_ratiossparse_page_indices(r),sparse_topk_lengths(r),sparse_raw_indices(r)and the writerset_sparse_topk(r, ...); call sites and the attention test kit use them,forwarddispatches oncompress_ratio != 0(storage stays flat)Sparse prefill (
SparsePrefillChunkCache,combine_topk_swa_indices)query_posinstead of assuming rows are the trailing extend tokens, and keeps-1entries inside the top-k span as-1instead of shifting them bycompressed_base(both identities on V4 inputs)build()takesquery_lens/query_posalongsideextend_seq_lens(the SWA gather still spans the whole extend)CompressedGather, keyed by ratio incache.compressed(c128 included); both sparse-prefill forwards sharecache.layer_inputs(...)KV pool and PD state
DeepSeekV4TokenToKVPool.present_ratios;CompressStatePool.request_scoped(set by the pool factory) replacesratio == 128checks in the state-buffer bookkeepingrequest_state_transfer_indices(req_pool_idx, seq_len)/clear_request_scoped_state(req_pool_idx)on the pool instead of combining an env lookup,get_ring_size(128)and a c128 helper; the index arithmetic moves todeepseek_v4_compress_state.py(c4_state_transfer_indices,request_scoped_state_transfer_indices,CompressStatePool.transfer_indices)mla_compression_ratiosinsetup_state_kv_argsfor both rolesget_swa_key_layout/get_extra_key_layout/get_*_bytes_per_tokenso the attention kernel views each cache with its own row width (the extra cache previously reused the SWA width)Removed
is_dsv4_c128_online_enabled(the pool'sonlineflag is the single source),CompressedGather.page_size(unused),get_dsv4_c4_state_indices/get_dsv4_c128_state_indicesindisaggregation/utils.py(moved as above)Verification
copy_, refresh, CP reindex; chunk-cache build, c128 / c4 gathers and combines; pool state-buffer bookkeeping and request-state clear; PD request-state indices) between main and this branch: 274 shared entries, 0 mismatches (main-only entries are the renamedc4_sparse_topk)test_deepseek_v4.py(incl. real sparse-prefill c4 / c128 vs. torch reference),test_q8kv8_sparse_prefill_backend.py,test_disaggregation_wire.py,test_dsv4_c4_state_lifecycle.py,test_dsv4_compressed_pools.py,test_dsv4_unified_fp8_pool.py,test_deepseek_v4_compress_state_runtime_shapes.pytest/registered/kernel/attention/test_combine_topk_swa_indices.py(torch oracle: trailing extend,-1holes, non-trailingquery_poswith cross-chunk offset, SWA-only); metadata contract tests forpresent_ratiosgating, accessor / writer routing and CP reindex; pool request-state transfer testsCI States
Latest PR Test (Base): ✅ Run #35276391750
Latest PR Test (Extra): 🚫 Run #35276390533
Latest PR Test (AMD ROCm 10): ⏳ Run #35276390785