Repository navigation
[KV Cache] Map kernel blocks per attention group; GLM-5.3-Flash on the generic packed layout - #59297
Draft
LucasWilkinson wants to merge 1 commit into
Draft
LucasWilkinson wants to merge 1 commit into
LucasWilkinson wants to merge 1 commit into
Conversation
…e generic packed layout Block tables and slot mappings stay in KV cache manager blocks everywhere. An attention group whose backend needs smaller blocks gets its block table mapped to kernel blocks when its metadata is built (kernel block j of block b is b * stride + j), so BlockTable/BlockTables no longer know about kernel blocks. The same mechanism lets a compressed cache in a block-outermost packed layout (GLM-5.3-Flash's kpool indexer) run on its kernel's 32/64-state pages: its bound KV view is re-strided over the packed blocks, and packed blocks are kept whole kernel blocks apart via a MultipleOf block_stride_alignment that the engine resolves against the final block size. GLM-5.3-Flash moves to the generic packed KV layout with a CircularBufferSpec tail (from vllm-project#57169). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
Closed
5 of 6 tasks
6 tasks done
6 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Alternative to #57169 (and its earlier version #55219). It keeps #57169's GLM-5.3-Flash move to the generic packed KV layout with a
CircularBufferSpectail. It replaces the builder-side re-paging of the kpool indexer with a generic, worker-side kernel-block mapping per attention group. As a result,BlockTable/BlockTablesstop knowing about kernel blocks.Design
kernel_block_sizesis removed fromBlockTable,MultiGroupBlockTable,BlockTables, both input batches and both slot-mapping kernels. Slot values are unchanged:b*bs + off == (b*bpk + off//kbs)*kbs + off%kbs.AttentionGroup.kernel_block_size),AttentionGroup.build_metadata/build_metadata_for_cudagraph_capture/build_metadata_for_draftingmap its block table to kernel blocks. Kernel blockjof blockbisb * kernel_block_stride + j. The mapped table goes into a persistent per-ubatch buffer (kernel_block_table), so CUDA graphs see a stable address.AttentionGroup.map_kv_cachere-strides the bound view over the packed blocks and takes the stride from the real tensor. It raises for token caches, which are written through manager-unit slot mappings.block_stride_alignment = MultipleOf(max size)(customize_attention_spec).KVCacheSpec.get_block_stride_alignment()resolves it in the engine against the final (possibly rescaled) block size, and returns 1 when the block runs whole.BlockTable.map_to_kernel_blocks.Related:
This PR instead makes kernel blocks a per-attention-group concern in the worker.
Test Plan / Results
Run on B300 with
.venv/bin/python -m pytest, one file at a time. Running several of these files in one process fails the same way on the base branch, from cross-file interference.tests/v1/worker/test_attn_utils.py, including new tests for packed and split mapping, the token-cache error, andMultipleOfalignment resolution;test_gpu_block_table.py,test_gpu_input_batch.py,test_gpu_model_runner.py,test_kv_block_zeroer.py,test_gpu_kpool_tail_slot_mapping.py,test_gpu_pcp_manager.py;tests/v1/spec_decode/test_llm_base_proposer.py,test_dflash_prepare_inputs.py;tests/v1/core/test_kv_cache_utils.py(129);tests/v1/attention/test_sparse_mla_kv_cache_layout.py(42),test_kpool_tail_slot_mapping.py;tests/v1/kv_connector/unit/test_nixl_desc_geometry.py(152 on its own);tests/kernels/attention/test_deepgemm_attention.py;tests/models/glm5next.tests/v1/spec_decode/test_acceptance_length.py::test_eagle3_acceptance_length[FLASH_ATTN-...-gpt-oss-20b-eagle3]fails on this branch (0.483) and on the reference branch (0.486) against a 0.4864 minimum, so it's pre-existing on this machine.mainfor non-GLM models. FlashInfer autotuning was off and MoE was Triton, for determinism.FLASHINFER_MLA, block 256 split into 64 (V1): bitwise identical, and reruns are stable.FLASHINFERin the HNDLBHNClayout, block 256 split into 64 (V1): bitwise identical.AI assistance
This PR was developed with AI assistance (Claude Code). The design was iterated with and reviewed by the submitter, who is responsible for every changed line. It is a draft pending full review and evals.
🤖 Generated with Claude Code