Skip to content

[KV Cache] Map kernel blocks per attention group; GLM-5.3-Flash on the generic packed layout - #59297

Draft
LucasWilkinson wants to merge 1 commit into
vllm-project:mainfrom
LucasWilkinson:kv-kernel-blocks-per-attn-group
Draft

LucasWilkinson wants to merge 1 commit into
vllm-project:mainfrom
LucasWilkinson:kv-kernel-blocks-per-attn-group

Conversation

@LucasWilkinson

Copy link
Copy Markdown
Contributor

Purpose

Alternative to #57169 (and its earlier version #55219). It keeps #57169's GLM-5.3-Flash move to the generic packed KV layout with a CircularBufferSpec tail. It replaces the builder-side re-paging of the kpool indexer with a generic, worker-side kernel-block mapping per attention group. As a result, BlockTable / BlockTables stop knowing about kernel blocks.

Design

  • Block tables and slot mappings are always in KV cache manager blocks, in V1 and MRv2. kernel_block_sizes is removed from BlockTable, MultiGroupBlockTable, BlockTables, both input batches and both slot-mapping kernels. Slot values are unchanged: b*bs + off == (b*bpk + off//kbs)*kbs + off%kbs.
  • Mapping per attention group. When a group's backend needs blocks smaller than the manager's (AttentionGroup.kernel_block_size), AttentionGroup.build_metadata / build_metadata_for_cudagraph_capture / build_metadata_for_drafting map its block table to kernel blocks. Kernel block j of block b is b * kernel_block_stride + j. The mapped table goes into a persistent per-ubatch buffer (kernel_block_table), so CUDA graphs see a stable address.
  • Ordinary splits. In a dense (unpacked) layout the stride is blocks-per-kernel-block. KV views are allocated in kernel blocks as before.
  • Packed compressed caches. Example: GLM-5.3-Flash's kpool indexer, which DeepGEMM paged MQA reads in 32/64-state pages. AttentionGroup.map_kv_cache re-strides the bound view over the packed blocks and takes the stride from the real tensor. It raises for token caches, which are written through manager-unit slot mappings.
  • Alignment. A backend with only fixed kernel block sizes gets block_stride_alignment = MultipleOf(max size) (customize_attention_spec). KVCacheSpec.get_block_stride_alignment() resolves it in the engine against the final (possibly rescaled) block size, and returns 1 when the block runs whole.
  • Connectors keep manager-block views. NIXL and Mooncake inline the removed BlockTable.map_to_kernel_blocks.

Related:

This PR instead makes kernel blocks a per-attention-group concern in the worker.

Test Plan / Results

Run on B300 with .venv/bin/python -m pytest, one file at a time. Running several of these files in one process fails the same way on the base branch, from cross-file interference.

  • Unit tests, passing:
    • tests/v1/worker/test_attn_utils.py, including new tests for packed and split mapping, the token-cache error, and MultipleOf alignment resolution;
    • test_gpu_block_table.py, test_gpu_input_batch.py, test_gpu_model_runner.py, test_kv_block_zeroer.py, test_gpu_kpool_tail_slot_mapping.py, test_gpu_pcp_manager.py;
    • tests/v1/spec_decode/test_llm_base_proposer.py, test_dflash_prepare_inputs.py;
    • tests/v1/core/test_kv_cache_utils.py (129);
    • tests/v1/attention/test_sparse_mla_kv_cache_layout.py (42), test_kpool_tail_slot_mapping.py;
    • tests/v1/kv_connector/unit/test_nixl_desc_geometry.py (152 on its own);
    • tests/kernels/attention/test_deepgemm_attention.py;
    • tests/models/glm5next.
  • Unit tests, known failure: tests/v1/spec_decode/test_acceptance_length.py::test_eagle3_acceptance_length[FLASH_ATTN-...-gpt-oss-20b-eagle3] fails on this branch (0.483) and on the reference branch (0.486) against a 0.4864 minimum, so it's pre-existing on this machine.
  • End-to-end greedy equivalence. The reference is the same tree without the generic block-table change, which matches main for non-GLM models. FlashInfer autotuning was off and MoE was Triton, for determinism.
    • GLM-5.3-Flash, dummy weights, packed indexer mapping: bitwise identical tokens and top-5 logprobs for V2 CUDA graphs at block 1024, V1 CUDA graphs at block 1024, V2 with MTP 3 at block 1024, and V2 eager at block 256.
    • GLM-4.7-Flash with MTP, FLASHINFER_MLA, block 256 split into 64 (V1): bitwise identical, and reruns are stable.
    • Qwen3-8B with EAGLE3, FLASHINFER in the HND LBHNC layout, block 256 split into 64 (V1): bitwise identical.
    • V2 runner with MTP on GLM-4.7-Flash: not deterministic run to run, even on the reference branch, so it's not comparable.
  • Not yet run: accuracy evals on real GLM-5.3-Flash weights.

AI assistance

This PR was developed with AI assistance (Claude Code). The design was iterated with and reviewed by the submitter, who is responsible for every changed line. It is a draft pending full review and evals.

🤖 Generated with Claude Code

…e generic packed layout

Block tables and slot mappings stay in KV cache manager blocks everywhere.
An attention group whose backend needs smaller blocks gets its block table
mapped to kernel blocks when its metadata is built (kernel block j of block b
is b * stride + j), so BlockTable/BlockTables no longer know about kernel
blocks. The same mechanism lets a compressed cache in a block-outermost packed
layout (GLM-5.3-Flash's kpool indexer) run on its kernel's 32/64-state pages:
its bound KV view is re-strided over the packed blocks, and packed blocks are
kept whole kernel blocks apart via a MultipleOf block_stride_alignment that
the engine resolves against the final block size.

GLM-5.3-Flash moves to the generic packed KV layout with a CircularBufferSpec
tail (from vllm-project#57169).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@mergify

mergify Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @LucasWilkinson.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

Status: Todo
Status: No status
Status: Backlog

Development

Successfully merging this pull request may close these issues.

1 participant