Skip to content

[Suggestion for #57169] Map kernel blocks per attention group instead of re-paging in builders - #7

Merged
andakai merged 1 commit into
andakai:refactor/glm-releasefrom
LucasWilkinson:suggest/57169-kernel-blocks-per-attn-group
Oct 3, 2026
Merged

andakai merged 1 commit into
andakai:refactor/glm-releasefrom
LucasWilkinson:suggest/57169-kernel-blocks-per-attn-group

Conversation

@LucasWilkinson

Copy link
Copy Markdown

Suggestion for vllm-project#57169: map kernel blocks per attention group

This is a suggested alternative to the kernel-page machinery in vllm-project#57169: kernel_page_size on the spec, get_strided_block_page_rows, select_common_block_size_for_layout, builder-side page expansion and get_cache_view_spec. It keeps vllm-project#57169's GLM-5.3-Flash move to the generic packed layout with the CircularBufferSpec tail. The same change as a standalone PR against main is vllm-project#59297, for comparison.

Idea. Block tables and slot mappings stay in KV cache manager blocks everywhere. When an attention group's backend needs smaller blocks, the group maps its own block table to kernel blocks at metadata build: kernel block j of block b is b * kernel_block_stride + j. The result goes into a persistent per-ubatch buffer.

  • Packed indexer. For a packed compressed cache such as the kpool indexer, AttentionGroup.map_kv_cache re-strides the bound view over the packed blocks. The stride comes from the real tensor.
  • Alignment. A backend taking only fixed kernel sizes declares block_stride_alignment = MultipleOf(max size). The engine resolves it against the final block size, and it's a no-op when the block runs whole.
  • Removed as a result: BlockTable / BlockTables lose kernel_block_sizes, the model no longer needs its own paging code, and the indexer backend doesn't re-page.

Diff against this PR's head: vllm/ is 36 files, +461/−808; tests are 29 files, +263/−1139. test_kpool_page_geometry.py is removed because it tests the replaced machinery; the replacement is covered in test_attn_utils.py and test_deepgemm_attention.py.

Carried over from this PR's latest commits (independent fixes):

  • NIXL compat-hash backend dedupe;
  • NIXL packed-MLA push region length (block_stride);
  • SM90 _pack_topk_indices padding with the first valid slot.

Please double-check that nothing else from 9e1bf3c..8e214ad is needed. I believe the rest reworks the machinery this replaces, or is already present.

Test results (B300)

  • Real weights, zai-org/GLM-5.3-Flash FP8 at TP=2, default config: 2176-token blocks in the packed BLHNC layout, the indexer mapped to seventeen 32-state pages per block, TRT-LLM sparse MLA on whole blocks. GSM8K (1319 questions, 5-shot, tests/evals/gsm8k/gsm8k_eval.py):

    Branch Accuracy Invalid
    This suggestion 0.917 0.000
    Reference (worker mapping for the indexer only, block tables still split) 0.912 0.000
  • Dummy weights, shrunk GLM-5.3-Flash: bitwise identical greedy tokens and top-5 logprobs against the reference for V1 and V2 with CUDA graphs at block 1024, V2 with MTP 3, and V2 eager at block 256.

  • Other models: GLM-4.7-Flash with MTP (FlashInfer MLA, block 256 split into 64s) and Qwen3-8B with EAGLE3 (FlashInfer in the HND layout, block 256 split into 64s) are bitwise identical to the reference on V1.

  • Unit tests, run one file at a time: worker, attention, KV-cache-utils, GLM and NIXL-geometry suites pass. Details are in [KV Cache] Map kernel blocks per attention group; GLM-5.3-Flash on the generic packed layout vllm-project/vllm#59297.

This suggestion was developed with AI assistance (Claude Code) and reviewed by the submitter.

🤖 Generated with Claude Code

Suggested alternative to the kernel-page machinery in this PR
(kernel_page_size, get_strided_block_page_rows, builder page expansion,
get_cache_view_spec). Block tables and slot mappings stay in KV cache manager
blocks everywhere; an attention group whose backend needs smaller blocks gets
its block table mapped to kernel blocks when its metadata is built, and a
packed compressed cache (the kpool indexer) gets its bound view re-strided.
BlockTable/BlockTables no longer know about kernel blocks.

Carries over this PR's independent fixes: NIXL compat-hash backend dedupe,
packed-MLA push region length, and SM90 top-k padding.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@andakai
andakai merged commit f2c8793 into andakai:refactor/glm-release Oct 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants