Repository navigation
[Bugfix] Keep the GLM-5.3 kpool tail and Qwen4 QSA ring out of the null block - #59528
Merged
njhill merged 5 commits intoOct 8, 2026
Merged
Conversation
ivanium
force-pushed
the
fix/glm53-kpool-tail-skips-null-block
branch
from
October 1, 2026 02:07
3c3bac9 to
4ac9ca0
Compare
ivanium
requested review from
AndreasKaratzas,
DarkLight1337,
WoosukKwon,
gau-nernst,
mgoin,
tlrmchlsmth,
yewentao256,
ywang96 and
zyongye
as code owners
October 1, 2026 02:17
ivanium
force-pushed
the
fix/glm53-kpool-tail-skips-null-block
branch
2 times, most recently
from
October 1, 2026 03:27
8e811e4 to
752ff37
Compare
ivanium
force-pushed
the
fix/glm53-kpool-tail-skips-null-block
branch
2 times, most recently
from
October 1, 2026 03:33
62f0e10 to
7b29aa4
Compare
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
JaredforReal
approved these changes
Oct 7, 2026
JaredforReal
left a comment
Contributor
There was a problem hiding this comment.
LGTM @ivanium
nit: tests for diffs in vllm/v1/attention/backends/mla/indexer.py also needed
The tail slot kernel maps every token to its request's tail block from column 0 of the block table. Dummy runs, CUDA graph capture and padding requests get the null block there, so their tail rows were written into block 0, which all KV cache groups share. Map tokens of a request without a tail block to PAD, as the generic slot mapping does for them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
The QSA circular ring maps a request's tokens to block_table[req, 0] and kept every block >= 0, but the runners mark rows that own no ring with the null block (0), not PAD. Dummy runs and CUDA graph capture therefore wrote ring rows into block 0, which all KV cache groups share. The compressed QSA path already keeps such dummy slots inert; do the same for the ring by requiring a block > 0. Tests that used block 0 as a real ring block now start real blocks at 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
The CPU padding test now ends with a padding request on the null block, and the Triton-vs-CPU test runs a mixed batch and an all-null dummy batch. Both check that tokens of a request on block 0 map to PAD. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
ivanium
force-pushed
the
fix/glm53-kpool-tail-skips-null-block
branch
from
October 7, 2026 23:41
7b29aa4 to
b96ce88
Compare
Contributor
Author
|
/ci run |
ywang96
approved these changes
Oct 8, 2026
|
❌ This PR is 1 commit behind upstream |
Member
|
/ci run |
njhill
enabled auto-merge (squash)
October 8, 2026 19:00
|
✅ Triggered Buildkite CI #93683 for commit |
WoosukKwon
approved these changes
Oct 8, 2026
WoosukKwon
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Block id 0 is the null block. It is shared by every KV cache group and is assumed to stay zero. Since #35431, the runners mark rows that own no state with 0 in the block table, not
PAD_SLOT_ID: dummy runs, CUDA graph capture and padding rows all carry 0. Two ring caches that address their pages through column 0 of the block table did not treat 0 as "no ring", so dummy runs wrote their rows into block 0:_kpool_tail_slot_mapping_kernelmaps every token toblock_table[req][0] * kpool + pos % kpool. Its torch fallback does the same.circular_qsa_slot_mappingand_build_qsa_metadata_kernelmap tokens toblock_table[req, 0]and kept every block>= 0. The compressed QSA path already keeps dummy slots inert (test_qsa_compressed_metadata_keeps_dummy_slots_inert); the ring path did not.Every scheduled request owns its ring block, so a request on the null block has none. Its tokens now map to PAD, as the generic slot mapping does. The ring writers already skip PAD slots.
This is the same pattern as #58560, which fixes the DeepSeek-V4.1 compressor ring.
Changes
vllm/v1/attention/backends/mla/indexer.py: the kpool tail slot mapping (Triton and torch) emits PAD for requests on the null block.vllm/models/qwen4_exp/common/qsa_cache.py: the QSA circular ring slot mapping (torch and Triton) requires block> 0._make_block_tablein the AMD pre-indexer test, which could randomly hand out block 0. For the kpool tail, the existing CPU padding test now ends with a padding request on the null block, and the existing Triton-vs-CPU test adds a mixed batch and an all-null dummy batch. Both check that the null-block request maps to PAD.Not a duplicate
CircularBufferSpec/ the packed layout but keep its slot kernel unchanged.Test Plan
Test Result
d6fe5dca68. 102 passed, 16 skipped on GB200. The 16 skips aretest_qsa_pre_indexer_amd.py, which requires ROCm.pre-commitpasses.BlockPoolpops it as the null block at init), so their mappings and outputs are unchanged.AI assistance (Claude Code) was used for this PR. The submitter reviewed the changes and ran the tests above.
🤖 Generated with Claude Code