Repository navigation
[Suggestion for #57169] Map kernel blocks per attention group instead of re-paging in builders - #7
Merged
andakai merged 1 commit intoOct 3, 2026
Conversation
Suggested alternative to the kernel-page machinery in this PR (kernel_page_size, get_strided_block_page_rows, builder page expansion, get_cache_view_spec). Block tables and slot mappings stay in KV cache manager blocks everywhere; an attention group whose backend needs smaller blocks gets its block table mapped to kernel blocks when its metadata is built, and a packed compressed cache (the kpool indexer) gets its bound view re-strided. BlockTable/BlockTables no longer know about kernel blocks. Carries over this PR's independent fixes: NIXL compat-hash backend dedupe, packed-MLA push region length, and SM90 top-k padding. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Suggestion for vllm-project#57169: map kernel blocks per attention group
This is a suggested alternative to the kernel-page machinery in vllm-project#57169:
kernel_page_sizeon the spec,get_strided_block_page_rows,select_common_block_size_for_layout, builder-side page expansion andget_cache_view_spec. It keeps vllm-project#57169's GLM-5.3-Flash move to the generic packed layout with theCircularBufferSpectail. The same change as a standalone PR againstmainis vllm-project#59297, for comparison.Idea. Block tables and slot mappings stay in KV cache manager blocks everywhere. When an attention group's backend needs smaller blocks, the group maps its own block table to kernel blocks at metadata build: kernel block
jof blockbisb * kernel_block_stride + j. The result goes into a persistent per-ubatch buffer.AttentionGroup.map_kv_cachere-strides the bound view over the packed blocks. The stride comes from the real tensor.block_stride_alignment = MultipleOf(max size). The engine resolves it against the final block size, and it's a no-op when the block runs whole.BlockTable/BlockTableslosekernel_block_sizes, the model no longer needs its own paging code, and the indexer backend doesn't re-page.Diff against this PR's head:
vllm/is 36 files, +461/−808; tests are 29 files, +263/−1139.test_kpool_page_geometry.pyis removed because it tests the replaced machinery; the replacement is covered intest_attn_utils.pyandtest_deepgemm_attention.py.Carried over from this PR's latest commits (independent fixes):
block_stride);_pack_topk_indicespadding with the first valid slot.Please double-check that nothing else from 9e1bf3c..8e214ad is needed. I believe the rest reworks the machinery this replaces, or is already present.
Test results (B300)
Real weights,
zai-org/GLM-5.3-FlashFP8 at TP=2, default config: 2176-token blocks in the packed BLHNC layout, the indexer mapped to seventeen 32-state pages per block, TRT-LLM sparse MLA on whole blocks. GSM8K (1319 questions, 5-shot,tests/evals/gsm8k/gsm8k_eval.py):Dummy weights, shrunk GLM-5.3-Flash: bitwise identical greedy tokens and top-5 logprobs against the reference for V1 and V2 with CUDA graphs at block 1024, V2 with MTP 3, and V2 eager at block 256.
Other models: GLM-4.7-Flash with MTP (FlashInfer MLA, block 256 split into 64s) and Qwen3-8B with EAGLE3 (FlashInfer in the HND layout, block 256 split into 64s) are bitwise identical to the reference on V1.
Unit tests, run one file at a time: worker, attention, KV-cache-utils, GLM and NIXL-geometry suites pass. Details are in [KV Cache] Map kernel blocks per attention group; GLM-5.3-Flash on the generic packed layout vllm-project/vllm#59297.
This suggestion was developed with AI assistance (Claude Code) and reviewed by the submitter.
🤖 Generated with Claude Code