Skip to content

Restore the per-group KV pool in the v0.31.0 seam - #31

Merged
arbi-dev merged 2084 commits into
mainfrom
upgrade/v0.31.0-pool
Oct 7, 2026
Merged

arbi-dev merged 2084 commits into
mainfrom
upgrade/v0.31.0-pool

Conversation

@arbi-dev

@arbi-dev arbi-dev commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Restores the per-group KV pool that #30 dropped. That drop was a mistake: #30 called the pools "conflicting with upstream's single-pool design". But the shared pool is exactly what wasted the KV budget on hybrid models, and why the seam had its own pools. turbo-attn ATTRIBUTION §7.3 measured Qwen3.5-0.8B at a 10 GB budget: 5,728 tokens with a shared pool, 355,696 with per-group pools.

Shape

The branch is v0.31.0 + one carry commit (e54368700: #30's carry with the pool restored), then git merge -s ours origin/main so it merges without force-pushing main. The merge leaves the tree untouched (verified equal). Diff vs main: 16 files, +949/−141. The fork now changes 20 vllm/ files vs v0.31.0. #30 changed 12.

How the pool is re-expressed on v0.31

v0.31 allocates one backing buffer and places a [B, H, N, C] view per layer (vllm#51718). On v0.28 the pools were separate tensors; that code can't be replayed. The capability is rebuilt on upstream's structures instead.

  • Eligibility (_mamba_pool_eligible): Mamba cache mode none/align; no HiSparse, KV connector, DCP/PCP or hidden-state layers. Those features address blocks across groups by one id space, so they keep upstream's shared pool.
  • Planning:
    • The Mamba/GDN state pages are no longer unified with attention pages. Attention layers are still unified among themselves.
    • The state pool holds max_num_seqs × Σ per-request blocks per group + 1 blocks.
    • That pool plus the sampler reserve (VLLM_SAMPLER_RESERVE_MIB) comes off the budget before attention is planned. So auto-fit, num_gpu_blocks_override, the admission check and the per-rank re-plan all plan the attention pool only.
    • Both pools are regions of the one buffer, each laid out by upstream's own placement (_layout_kv_cache_tensors, factored out unchanged). The regions never alias.
    • New fields: KVCacheConfig.mamba_pool_num_blocks / mamba_pool_offset / group_num_blocks(), and KVCacheTensor.num_blocks.
  • Scheduler:
    • The coordinator gets a second BlockPool for the state groups.
    • Admission checks each pool against its own free blocks (get_num_blocks_to_allocate_per_pool, KVCacheManager._fits). Watermark and reservations count main-pool blocks.
    • Usage reports the fuller pool, and events come from both pools.
    • Deferred frees already route by KVCacheBlock.pool upstream; a test confirms it.
  • Worker:
    • The two pools' block ids overlap. CoW copies of state blocks carry ids shifted past num_blocks, and copy_kv_cache_blocks_inplace applies them only to views inside the state region.
    • MRV2 profiling dummy contexts (set_dummy_context) and both warmups number each pool's blocks separately.

Per-hunk rows (changes vs #30)

file what verdict
v1/core/kv_cache_utils.py eligibility, split planner, region helper, grouping without state/attention unification, capacity excludes the state pool, get_kv_cache_configs fixed-cost accounting still needed, re-expressed on v0.31 placement
v1/kv_cache_interface.py pool fields on KVCacheConfig / KVCacheTensor still needed
v1/core/kv_cache_coordinator.py second BlockPool, per-pool block needs still needed
v1/core/kv_cache_manager.py per-pool admission, usage, events, CoW id shift still needed
v1/worker/utils.py, gpu_model_runner.py, gpu/model_runner.py CoW copies per pool still needed (CoW is new upstream since v0.28)
v1/worker/gpu/input_batch.py, gpu/warmup.py per-pool dummy/warmup block ids still needed (MRV2 is the default runner in v0.31)
envs.py VLLM_SAMPLER_RESERVE_MIB still needed (restored)
PROTOCOL.md new section doc

Not restored, with the reason:

  • Padded-page block fill (_fill_padded_page_block_size): it only worked around state/attention page unification, which the split removes.
  • Drain hook / on_kv_manager_created and the zeroer clamp: no turbo-attn consumer.

Composite (smart-mix) layout

aggregated_layer_count is unchanged from #30: every fused layer's view is the same dense page region of the attention pool. test_mamba_pool.py::test_composite_attention_shares_one_dense_page_beside_the_state_pool checks it.

The plugin side adds one fix (turbo-attn PR, linked below). It puts each composite layer's region view into runner.kv_caches in place of the composite view. Otherwise a prefix-cache CoW copy, which works by page index, would move other layers' bytes in the layer-major composite.

Verification

CPU only, in vllm/vllm-openai:v0.31.0 with the 20 files overlaid, on 10.1.0.202.

  • tests/v1/core/test_mamba_pool.py (new): 12 passed. It covers:
    • pool sizing, and 16.9× the attention blocks vs the shared pool on a bf16 Qwen3.5-like layout;
    • no aliasing, and the allocation bound;
    • the eligibility exclusions;
    • each pool allocating from itself, and admission refusing when the state pool is full even with attention blocks free;
    • deferred frees returning to their own pools, and usage;
    • CoW copy routing;
    • composite on the split layout.
  • Seam tests: all pass. test_kv_seam_invariants, test_aggregated_layer_count, test_mla_wrapper_selection, and the hook tests in test_gpu_worker.
  • Upstream suites: tests/v1/core, tests/v1/worker and tests/engine/test_arg_utils.py, stock v0.31.0 vs overlay. The overlay shows no new failures and 18 more passes (the seam's own tests). The remaining failures are CPU-container device-inference errors, identical on stock. The one per-run difference (test_prefix_cache_default) fails identically on both when run alone.
  • turbo-attn vLLM tests (with the plugin PR): 2099 passed. The 2 failures are already stale on turbo-attn main, and the 7 errors are GPU fixtures.

Not yet run

GPU validation of the engine image. Hybrid capacity needs measuring on a real boot (Qwen3.5-0.8B tkv: blocks reported vs v0.28 image).

🤖 Generated with Claude Code

Yejing-Lai and others added 30 commits September 24, 2026 02:05
…c value (vllm-project#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…-project#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
…vllm-project#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
…ting Sentence Transformers from applying chat template (vllm-project#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
…ect#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
…sor (vllm-project#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
… wo_a as a grouped FP8 GEMM (vllm-project#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…al tests (vllm-project#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
…ast` (vllm-project#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
…eneration-config (vllm-project#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
…on (vllm-project#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
…online shorthands (vllm-project#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…mori-io (vllm-project#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
…oken_ids (vllm-project#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
…roject#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
khluu and others added 29 commits September 28, 2026 20:18
…t#59118)

Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit aedaba8)
…t#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db)
…vllm-project#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c3)
…-project#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd)
…etention (vllm-project#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bdd)
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e)
…llm-project#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e0)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf5)
…cy kernel (vllm-project#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f66)
…m-project#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc)
vllm-project#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a)
…lm-project#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa13)
…pletion (vllm-project#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87c)
…vllm-project#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec)
…cy metrics (vllm-project#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae)
…mpt-end eviction in align mode (vllm-project#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e6)
…ct#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e)
…ine (vllm-project#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360)
…lm-project#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d)
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba)
…oject#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a69636)
…es (vllm-project#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2)
… a saturated GPU pool (vllm-project#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8a)
…ject#59309 backport

Resolving the vllm-project#59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from vllm-project#57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream vllm-project#59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>
…formers v5.18 (vllm-project#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f16)
…ject#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298)
Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2 #28, 76a3e6d, 6d4c756 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 27.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner
  - per-group KV pool for hybrid models, re-expressed on v0.31's single
    backing allocation: the O(1) Mamba/GDN state groups get their own
    BlockPool sized for max_num_seqs, a region of the allocation after the
    attention pool; their pages are not unified with attention pages;
    per-pool admission, events and usage; CoW copies and warmup/profiling
    block ids per pool; VLLM_SAMPLER_RESERVE_MIB headroom

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - padded-page block fill (the split pool removes the page unification
    it worked around), drain hook / on_kv_manager_created, KVBlockZeroer
    clamp: no turbo-attn consumer
  - rotary fast-path registry: upstream guards the import (1f60771,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arbi-dev
arbi-dev merged commit 1e0e0ab into main Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.