Repository navigation
Restore the per-group KV pool in the v0.31.0 seam - #31
Merged
Merged
Conversation
…c value (vllm-project#54514) Signed-off-by: Lai, Yejing <yejing.lai@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…-project#54508) Signed-off-by: priyansh jain <priyansh.jain2@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
…vllm-project#58378) Co-authored-by: Bugen Zhao <i@bugenzhao.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: reidliu41 <reid201711@gmail.com> Signed-off-by: Bugen Zhao <i@bugenzhao.com>
…ting Sentence Transformers from applying chat template (vllm-project#58117) Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Li, Jiang <jiang1.li@intel.com>
…ect#57528) Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
…sor (vllm-project#58460) Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
… wo_a as a grouped FP8 GEMM (vllm-project#58456) Signed-off-by: fai <fangzhouai@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…al tests (vllm-project#57347) Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
…ast` (vllm-project#58550) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
…lm-project#57918) Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
…eneration-config (vllm-project#50769) Signed-off-by: Chenglun Hu <chenglunhu@gmail.com> Signed-off-by: hclsys <chenglunhu@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Wauplin <lucainp@gmail.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
…on (vllm-project#58593) Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Kimi <noreply@moonshot.cn>
Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Kimi Code <noreply@moonshot.cn>
…llm-project#58427) Signed-off-by: mgoin <mgoin64@gmail.com>
…ect#58572) Signed-off-by: mgoin <mgoin64@gmail.com>
…online shorthands (vllm-project#53585) Signed-off-by: Felix Marty <Felix.Marty@amd.com> Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: mgoin <mgoin64@gmail.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…mori-io (vllm-project#51681) Signed-off-by: Vincent Cave <vincent.cave@amd.com> Signed-off-by: Shiksha Patel <shikpate@amd.com> Co-authored-by: Shiksha Patel <shikpate@amd.com> Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com> Co-authored-by: Andreas Karatzas <akaratza@amd.com>
…oken_ids (vllm-project#58216) Signed-off-by: Matt Mastracci <matthew@mastracci.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Misha Goin <mgoin64@gmail.com>
…roject#48521) Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com> Co-authored-by: Misha Goin <mgoin64@gmail.com>
…vllm-project#59068) Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> (cherry picked from commit ec5e0c3)
…-project#59137) Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com> (cherry picked from commit 6ebb5bd)
…etention (vllm-project#59146) Signed-off-by: Jared Wen <jaredwen@inferact.ai> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com> (cherry picked from commit d882bdd)
) Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com> Co-authored-by: Tai An <antai12232931@outlook.com> Co-authored-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 765872e)
…llm-project#59286) Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 105a4e0)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com> (cherry picked from commit 678baf5)
…cy kernel (vllm-project#59235) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 2798f66)
…m-project#59036) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 90e13fc)
vllm-project#58268) Signed-off-by: Harshal Adhav <harshal.adhav@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Li, Jiang <jiang1.li@intel.com> (cherry picked from commit 5faf81a)
…lm-project#58946) Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com> (cherry picked from commit 866fa13)
…up MLA (vllm-project#57652) (cherry picked from commit 5463fe4)
…pletion (vllm-project#59007) Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com> (cherry picked from commit ff1b87c)
…vllm-project#59282) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> (cherry picked from commit 3eb6cec)
…cy metrics (vllm-project#58725) Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 73c7cae)
…mpt-end eviction in align mode (vllm-project#59175) Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com> Co-authored-by: Nick Hill <nickhill123@gmail.com> (cherry picked from commit 7e583e6)
…ine (vllm-project#58542) Signed-off-by: LostFox11 <wangziyue17@huawei.com> Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: LostFox11 <wangziyue17@huawei.com> Co-authored-by: Nick Hill <nickhill123@gmail.com> (cherry picked from commit 7314360)
…lm-project#54024) Signed-off-by: R <Ganesh.R@amd.com> Signed-off-by: Ganesh R <Ganesh.R@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> (cherry picked from commit e12291d)
Signed-off-by: Robert Shaw <robshaw@redhat.com> Co-authored-by: Robert Shaw <robshaw@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Nick Hill <nickhill123@gmail.com> (cherry picked from commit e5e38ba)
…oject#59494) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 3a69636)
…es (vllm-project#59450) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com> (cherry picked from commit f7999d2)
… a saturated GPU pool (vllm-project#59309) Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com> (cherry picked from commit 4056c8a)
…ject#59309 backport Resolving the vllm-project#59309 cherry-pick conflict in tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in test_scheduled_prefix_hit_publishes_adopted_copies from vllm-project#57930, which is not on this branch; it imports _allocate_scheduled from tests.v1.core.test_prefix_caching and fails on every platform. After this change the file matches the upstream vllm-project#59309 diff. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
…formers v5.18 (vllm-project#59613) Signed-off-by: Isotr0py <Isotr0py@outlook.com> (cherry picked from commit b558f16)
…ject#59614) Signed-off-by: Isotr0py <Isotr0py@outlook.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> (cherry picked from commit 58b3298)
Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam: b0a14f2 #28, 76a3e6d, 6d4c756 #29) into one commit and replays it onto upstream v0.31.0 (db9527a, the commit vllm/vllm-openai:v0.31.0 is built from) as a 3-way merge against the real base. Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups) replaced the allocator the seam patched: one backing allocation, per-layer [B, H, N, C] views placed by the engine, page geometry read off the spec, and AttentionBackend.customize_spec applied to every layer's spec by both model runners. The seam is re-expressed on that and shrinks from 49 files to 27. Kept (re-applied on upstream's structure): - plugin KV-cache dtype registry; --kv-cache-dtype choices/type - TURBO_ATTN backend slot (now a distinct enum value: two None members made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default for plugin dtypes (#29), CUDA candidate once registered - lifecycle hooks on_model_loaded / on_draft_model_loaded / on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every backend in use - MLA wrapping in the selector; MLA chunked-context _get_gather_op - _tq_layer_idx injection for tkv layers - aggregated_layer_count: fused (composite) pages, now one shared page per fused set in upstream's single-allocation planner - per-group KV pool for hybrid models, re-expressed on v0.31's single backing allocation: the O(1) Mamba/GDN state groups get their own BlockPool sized for max_num_seqs, a region of the allocation after the attention pool; their pages are not unified with attention pages; per-pool admission, events and usage; CoW copies and warmup/profiling block ids per pool; VLLM_SAMPLER_RESERVE_MIB headroom Moved to upstream's extension points (turbo-attn plugin side): - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment) -> AttentionBackend.customize_spec - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar - get_kv_cache_shape(kv_cache_spec=...) passthrough and the backend-managed-dtype shape coercion -> spec-driven views Dropped: - spec_decode_warmup.py: upstream registers the same kernels with its JIT warmup registry (aed894c, vllm#56323) - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache - padded-page block fill (the split pool removes the page unification it worked around), drain hook / on_kv_manager_created, KVBlockZeroer clamp: no turbo-attn consumer - rotary fast-path registry: upstream guards the import (1f60771, vllm#42679) - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE, Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn workflow, draft-backend inheritance: unused or superseded - FA2 varlen paged split-K patch: never reached the overlay image (the image installs no compiled _vllm_fa2_C) and the FA pin moved PROTOCOL.md rewritten for the v0.31.0 seam. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Restores the per-group KV pool that #30 dropped. That drop was a mistake: #30 called the pools "conflicting with upstream's single-pool design". But the shared pool is exactly what wasted the KV budget on hybrid models, and why the seam had its own pools. turbo-attn ATTRIBUTION §7.3 measured Qwen3.5-0.8B at a 10 GB budget: 5,728 tokens with a shared pool, 355,696 with per-group pools.
Shape
The branch is v0.31.0 + one carry commit (
e54368700: #30's carry with the pool restored), thengit merge -s ours origin/mainso it merges without force-pushing main. The merge leaves the tree untouched (verified equal). Diff vs main: 16 files, +949/−141. The fork now changes 20vllm/files vs v0.31.0. #30 changed 12.How the pool is re-expressed on v0.31
v0.31 allocates one backing buffer and places a
[B, H, N, C]view per layer (vllm#51718). On v0.28 the pools were separate tensors; that code can't be replayed. The capability is rebuilt on upstream's structures instead._mamba_pool_eligible): Mamba cache modenone/align; no HiSparse, KV connector, DCP/PCP or hidden-state layers. Those features address blocks across groups by one id space, so they keep upstream's shared pool.max_num_seqs × Σ per-request blocks per group + 1blocks.VLLM_SAMPLER_RESERVE_MIB) comes off the budget before attention is planned. So auto-fit,num_gpu_blocks_override, the admission check and the per-rank re-plan all plan the attention pool only._layout_kv_cache_tensors, factored out unchanged). The regions never alias.KVCacheConfig.mamba_pool_num_blocks/mamba_pool_offset/group_num_blocks(), andKVCacheTensor.num_blocks.BlockPoolfor the state groups.get_num_blocks_to_allocate_per_pool,KVCacheManager._fits). Watermark and reservations count main-pool blocks.KVCacheBlock.poolupstream; a test confirms it.num_blocks, andcopy_kv_cache_blocks_inplaceapplies them only to views inside the state region.set_dummy_context) and both warmups number each pool's blocks separately.Per-hunk rows (changes vs #30)
v1/core/kv_cache_utils.pyget_kv_cache_configsfixed-cost accountingv1/kv_cache_interface.pyKVCacheConfig/KVCacheTensorv1/core/kv_cache_coordinator.pyBlockPool, per-pool block needsv1/core/kv_cache_manager.pyv1/worker/utils.py,gpu_model_runner.py,gpu/model_runner.pyv1/worker/gpu/input_batch.py,gpu/warmup.pyenvs.pyVLLM_SAMPLER_RESERVE_MIBPROTOCOL.mdNot restored, with the reason:
_fill_padded_page_block_size): it only worked around state/attention page unification, which the split removes.on_kv_manager_createdand the zeroer clamp: no turbo-attn consumer.Composite (smart-mix) layout
aggregated_layer_countis unchanged from #30: every fused layer's view is the same dense page region of the attention pool.test_mamba_pool.py::test_composite_attention_shares_one_dense_page_beside_the_state_poolchecks it.The plugin side adds one fix (turbo-attn PR, linked below). It puts each composite layer's region view into
runner.kv_cachesin place of the composite view. Otherwise a prefix-cache CoW copy, which works by page index, would move other layers' bytes in the layer-major composite.Verification
CPU only, in
vllm/vllm-openai:v0.31.0with the 20 files overlaid, on 10.1.0.202.tests/v1/core/test_mamba_pool.py(new): 12 passed. It covers:test_kv_seam_invariants,test_aggregated_layer_count,test_mla_wrapper_selection, and the hook tests intest_gpu_worker.tests/v1/core,tests/v1/workerandtests/engine/test_arg_utils.py, stock v0.31.0 vs overlay. The overlay shows no new failures and 18 more passes (the seam's own tests). The remaining failures are CPU-container device-inference errors, identical on stock. The one per-run difference (test_prefix_cache_default) fails identically on both when run alone.Not yet run
GPU validation of the engine image. Hybrid capacity needs measuring on a real boot (Qwen3.5-0.8B tkv: blocks reported vs v0.28 image).
🤖 Generated with Claude Code