Skip to content

Forward-port the tkv/turbo-attn vLLM seam onto upstream v0.31.0 - #30

Merged
arbi-dev merged 2084 commits into
mainfrom
upgrade/v0.31.0
Oct 7, 2026
Merged

arbi-dev merged 2084 commits into
mainfrom
upgrade/v0.31.0

Conversation

@arbi-dev

@arbi-dev arbi-dev commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Forward-ports the turbo-attn vLLM seam from upstream v0.28.0 to v0.31.0 (db9527a46, the latest stable release; vllm/vllm-openai:v0.31.0 is built from exactly this commit: image label org.opencontainers.image.revision=db9527a46873…, index digest sha256:c1c9f6fd5c10…).

Shape

Same method as #28: the whole carry (git diff v0.28.0 origin/main at 969c125e9) squashed into one commit on v0.28.0 and cherry-picked onto v0.31.0 as a 3-way merge against the real base. Result: v0.31.0 + 1 carry commit.

Carry on fork main (nothing else has merged since #28; no other forward-port branch exists):

commit what folded in
b0a14f2eb #28, the v0.28.0 forward-port yes
76a3e6d72 fix(kv-seam): hash_block_size on the per-group pool view (post-port) yes, then dropped with the per-group pools (below)
6d4c75611 / merge 969c125e9 #29 auto-default --attention-backend TURBO_ATTN for plugin kv dtypes (post-port) yes, kept

49 files / +4160 → 18 files / +1289 (12 of them under vllm/). Most of the shrink is upstream's v0.29–v0.31 KV-cache layout refactor (vllm#51718 and follow-ups): one backing allocation for every layer, per-layer [B, H, N, C] views placed by the engine from the spec's geometry (num_heads, state_content_size_bytes), and AttentionBackend.customize_spec applied to every layer's spec by both model runners. get_kv_cache_shape is no longer called at all. The seam's allocator patches were written against the code that refactor deleted, so they were re-expressed on it or dropped, and the spec-class hook moved to upstream's customize_spec.

This needs the matching turbo-attn plugin change: arbicity/turbo-attn deps/engine-forks-v0.31.0-v0.5.21 @ a2c4bd7d6 (customize_spec, layouts, [B,1,N,C]→3-D views, padded-slot review off on placed views). A v0.31 image without it will not serve tkv.

Conflicts (18 files) and resolutions

file conflict resolution
config/cache.py upstream added layout helpers + device_memory_utilization around our registry and str field upstream's code verbatim; registry + cache_dtype: str re-applied on top; skip-layers-dtype field dropped (below)
engine/arg_utils.py upstream's XPU batch-invariant default sits where #29's auto-default was both, upstream's first; AttentionBackendEnum now imported at module level so our local import went
envs.py upstream widened VLLM_KV_CACHE_LAYOUT to the new layouts upstream's; our two env vars dropped (below)
platforms/cuda.py upstream added TRITON_FLASHINFER/TRITON_FLASH_ATTN mm-prefix candidates upstream's lists; TURBO_ATTN append re-applied
model_executor/warmup/kernel_warmup.py import block churn upstream verbatim; our two warmups dropped (below)
v1/worker/gpu_worker.py upstream added set_torch_threads_for_runtime() at the end of load_model hooks run after it (serving thread count); CuTeDSL-cache init and post-capture spec warmup removed
v1/attention/selector.py upstream dropped get_required_kv_cache_layout + set_kv_cache_layout (layout is now negotiated by the engine core) and the module logger MLA wrapping re-applied; the layout block stays gone; logger re-added for the wrap log line
model_executor/layers/attention/mla_attention.py upstream's spec now adds is_index_group_leader and TMA-row block_stride_alignment after construction upstream's; our get_kv_cache_spec_class("mla") short-circuit dropped (it skipped both). The plugin's MLA wrapper now packs through customize_spec. _get_gather_op merged cleanly onto upstream's new CPU branch
v1/kv_cache_interface.py KVCacheConfig grew transfer/hisparse fields where per_group_num_blocks was upstream's; only aggregated_layer_count re-applied
v1/core/kv_cache_utils.py (9 hunks) upstream rewrote planning: single backing allocation, KVCacheTensor(layers, layer_stride, block_stride, offset), layout-aware strides upstream's file; fused-page support re-applied as get_tensor_slots in _get_kv_cache_bytes_per_block and the tensor emission loop
v1/core/kv_cache_coordinator.py, kv_cache_manager.py (2+5) per-group BlockPools vs upstream's shared pool upstream's (dropped, below)
v1/core/single_type_kv_cache_manager.py registry lookup rewritten (roles) upstream's (get_manager_class obsolete, below)
v1/worker/gpu_model_runner.py, v1/worker/gpu/attn_utils.py _reshape_kv_cache* deleted upstream (moved into allocate_kv_cache/create_kv_cache_views) upstream's; both of our hunks lived in the deleted functions
v1/worker/utils.py KVBlockZeroer rewritten for kernel-block ratios upstream's
tests/v1/worker/test_gpu_worker.py upstream added memory-accounting tests in the same tail upstream's file + our hook tests appended
tests/kernels/attention/test_flash_attn.py upstream's (FA2 patch dropped)

Auto-merged files were re-read against v0.31.0 too; attention.py, backend.py, platforms/interface.py, llm_base_proposer.py changed under the review below.

Per-hunk review against v0.31.0

file / hunk what it does verdict reason
config/cache.py registry, str field, validator plugin --kv-cache-dtype values still needed upstream has no plugin dtype registry; a Literal field makes plugin dtypes unconstructible
engine/arg_utils.py choices/type --kv-cache-dtype lists plugin dtypes still needed same
engine/arg_utils.py TURBO_ATTN auto-default (#29) named backend for plugin dtypes still needed hooks dispatch before attention groups exist
config/attention.py -→_ --attention-backend turbo-attn still needed documented spelling in turbo-attn
backends/registry.py TURBO_ATTN dedicated slot still needed, fixed was = None like CUSTOM, so CUSTOM was an Enum alias of TURBO_ATTN; now the plugin class path (distinct value, still is_overridden()-gated)
platforms/cuda.py append TURBO_ATTN as auto candidate once registered still needed covers selection that bypasses EngineArgs (draft models); also replaces the proposer hunk below
v1/attention/selector.py dtype check validate_cache_dtype instead of get_args(CacheDType) still needed
v1/attention/selector.py MLA wrapping wrap a stock MLA backend still needed no upstream equivalent
backend.py on_model_loaded, on_draft_model_loaded, on_kv_cache_initialized, adjust_kv_budget, wraps_mla_backend, resolve_user_selected_backend lifecycle hooks still needed all consumed by turbo-attn; no upstream worker hooks
gpu_worker.py dispatch fail-loud, every backend in use still needed
mla_attention.py _get_gather_op packed-latent gather in chunked prefill still needed consumed by TkvMLAImpl; no upstream hook
attention.py _tq_layer_idx per-layer bit widths at impl init still needed
kv_cache_interface.py + kv_cache_utils.py aggregated_layer_count fused (composite) pages conflicting → re-expressed v0.31's planner sums one page per layer; fused layers now share one page slot and alias one region
backend.py/attention.py/mla_attention.py/platforms/interface.py get_kv_cache_spec_class backend-declared spec class obsolete → moved to plugin upstream AttentionBackend.customize_spec, applied by both runners and by _align_hybrid_block_size
kv_cache_interface.py + single_type_kv_cache_manager.py get_manager_class spec picks its manager obsolete KVCacheSpecRegistry.get_manager_class walks the MRO; TKV specs subclass Full/SlidingWindow specs
backend.py get_supported_kv_cache_dtypes dynamic dtype list obsolete the wrapper sets the supported_kv_cache_dtypes ClassVar (plugin)
backend.py uses_per_group_kv_pool, get_gather_op (backend-level), on_kv_manager_created obsolete no consumer in turbo-attn
kv_cache_coordinator.py, kv_cache_manager.py, kv_cache_utils.py per-group BlockPools, _MultiGroupPoolView (+76a3e6d72), sampler reserve, padded-page block fill, drain hook O(1) mamba pool split conflicting → dropped v0.31 overlays every group on one pool by design (KVCacheConfig.num_blocks_of, single allocate_kv_cache buffer); re-expressing split pools means re-doing upstream's allocator. Hybrid models take upstream's alignment, which sizes the TKV page through customize_spec
worker/utils.py zeroer clamp per-group block counts obsolete only existed for split pools
gpu_model_runner.py/gpu/attn_utils.py kv_cache_spec passthrough, backend-managed dtype coercion get_kv_cache_shape args obsolete upstream no longer calls get_kv_cache_shape; views come from the spec (vllm#51718)
spec_decode_warmup.py + gpu_worker call warm EAGLE prep / rejection expand kernels obsolete upstream registers the same four kernels with its JIT warmup registry: aed894c19 (vllm#56323)
turbo_attn_warmup.py, utils/cutedsl_cache.py + calls dummy-run prefill warmup, persistent CuTeDSL cache obsolete (plugin) turbo-attn warms every served prefill kernel in on_kv_cache_initialized (tkv.engine.prewarm) and arms its own CuTeDSL cache in tkv/kernels/cute_dsl_cache.py
rotary_embedding/common.py registry survive FA4's namespace flash_attn obsolete upstream imports under suppress(ModuleNotFoundError): 1f60771c7 (vllm#42679)
llm_base_proposer.py draft inherits backend draft finds TURBO_ATTN obsolete the CUDA candidate append already makes it selectable
--kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE, VLLM_SAMPLER_RESERVE_MIB dropped unused by turbo-attn; skip-dtype also conflicts with upstream's skip_page_size_padded (assumes auto)
_custom_ops.py Triton fused fp8 GEMM dropped the gemm package is not in the image; dead
FA2 varlen split-K (cmake/apply_patches.cmake, patch, vllm_flash_attn.cmake, flash_attn_interface.py, test) dropped never reached the image (overlay ships no compiled _vllm_fa2_C; turbo-attn excluded the .py), and the FA pin moved f3e1a4f→9cd61de
notify-turbo-attn.yml dropped failing since 2026-08-03; turbo-attn polls the fork instead (poll-engine-forks.yml, vllm-project#741)
tests/evals/gsm8k startup waits dropped unrelated to the seam

Tests: kept and rewritten for v0.31: test_kv_seam_invariants.py (now the plugin dtype registry), test_aggregated_layer_count.py (one aliased page per fused set; equal-byte capacity), test_mla_wrapper_selection.py, test_gpu_worker.py hook tests, test_arg_utils.py (#29). Dropped with their code: test_kv_drain_hook.py, test_per_group_blockpool.py, test_tkv_hybrid_block_align.py, test_cutedsl_cache.py.

PROTOCOL.md is rewritten for this seam (it also fixes #28's drift: _get_gather_op returns the stock op, not None, and serves chunked-context prefill, not decode).

Known gaps

  • Hybrid capacity: dropping the per-group pools puts hybrid (mamba/GDN + attention) models on upstream's shared pool, so their reported KV capacity can be lower than on the v0.28 fork. Correctness is upstream's path.
  • Smart-mix composite on padded pages: if a hybrid model's attention page is ever padded (upstream pads it only when the mamba page sets the maximum and the ratio is not integral), the composite region view refuses the non-contiguous page instead of serving.
  • get_kv_cache_spec_class no longer exists, so a plugin that implemented only it (not customize_spec) silently gets stock specs. turbo-attn implements both.
  • test_plugin_kv_cache_dtype_auto_selects_turbo_attn_backend needs a device for create_engine_config; it errors in a CPU-only container exactly like upstream's own test_arg_utils.py cases do.

Verification

CPU only, in vllm/vllm-openai:v0.31.0 with this branch's 12 vllm/ files overlaid (as the image build does), on 10.1.0.202, no GPU visible:

  • seam tests: 37 passed, 9 skipped (need the CUDA allocator), 1 env error (above).
  • upstream tests/v1/core/test_kv_cache_utils.py, test_prefix_caching.py, test_single_type_kv_cache_manager.py, tests/engine/test_arg_utils.py, tests/v1/worker/test_gpu_worker.py: no new failures vs stock v0.31.0; the remaining failures are device-inference errors that fail identically on stock when run in isolation.
  • turbo-attn tests/vllm_plugin, tests/integrations/vllm, tests/test_vllm_*.py with the plugin branch: 2098 passed, 260 skipped; 2 failures are stale on turbo-attn main (_backend_resolution._MODEL_LOADED gone; logger test assumes no vLLM), 7 errors are the GPU e2e fixtures.
  • compileall clean, no conflict markers.

Not yet runtime-tested on a GPU: that is the engine-image build and GPU validation in turbo-attn on the new base. Merge after that is green.

🤖 Generated with Claude Code

Yejing-Lai and others added 30 commits September 24, 2026 02:05
…c value (vllm-project#54514)

Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…-project#54508)

Signed-off-by: priyansh jain <priyansh.jain2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
…vllm-project#58378)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
…ting Sentence Transformers from applying chat template (vllm-project#58117)

Signed-off-by: RyanMa29 <ziyang.ma@intel.com>
Signed-off-by: R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
…ect#57528)

Signed-off-by: Shrey Gajjar <shreygajjar007@gmail.com>
…sor (vllm-project#58460)

Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
… wo_a as a grouped FP8 GEMM (vllm-project#58456)

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…al tests (vllm-project#57347)

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
…ast` (vllm-project#58550)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
…eneration-config (vllm-project#50769)

Signed-off-by: Chenglun Hu <chenglunhu@gmail.com>
Signed-off-by: hclsys <chenglunhu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

Signed-off-by: Thomas Ortner <boh@zurich.ibm.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
…on (vllm-project#58593)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi <noreply@moonshot.cn>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Kimi Code <noreply@moonshot.cn>
…online shorthands (vllm-project#53585)

Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…mori-io (vllm-project#51681)

Signed-off-by: Vincent Cave <vincent.cave@amd.com>
Signed-off-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Shiksha Patel <shikpate@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
…oken_ids (vllm-project#58216)

Signed-off-by: Matt Mastracci <matthew@mastracci.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
…roject#48521)

Signed-off-by: Aaron Kang <aaron.h.kang@icloud.com>
Co-authored-by: Misha Goin <mgoin64@gmail.com>
zxd1997066 and others added 28 commits September 29, 2026 17:33
…t#58307)

Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
(cherry picked from commit 09c47db)
…vllm-project#59068)

Signed-off-by: Ren Yuzhou <54501155+yuzhouo7@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit ec5e0c3)
…-project#59137)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 6ebb5bd)
…etention (vllm-project#59146)

Signed-off-by: Jared Wen <jaredwen@inferact.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
(cherry picked from commit d882bdd)
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
Co-authored-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 765872e)
…llm-project#59286)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 105a4e0)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
(cherry picked from commit 678baf5)
…cy kernel (vllm-project#59235)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 2798f66)
…m-project#59036)

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 90e13fc)
vllm-project#58268)

Signed-off-by: Harshal Adhav <harshal.adhav@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
(cherry picked from commit 5faf81a)
…lm-project#58946)

Signed-off-by: Andreas Karatzas <andreas.karatzas@protonmail.com>
(cherry picked from commit 866fa13)
…pletion (vllm-project#59007)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit ff1b87c)
…vllm-project#59282)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit 3eb6cec)
…cy metrics (vllm-project#58725)

Signed-off-by: Yueh-Ting Chen <yueh.ting.chen@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 73c7cae)
…mpt-end eviction in align mode (vllm-project#59175)

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7e583e6)
…ct#59335)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit c055c1e)
…ine (vllm-project#58542)

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 7314360)
…lm-project#54024)

Signed-off-by: R <Ganesh.R@amd.com>
Signed-off-by: Ganesh R <Ganesh.R@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
(cherry picked from commit e12291d)
Signed-off-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Robert Shaw <robshaw@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit e5e38ba)
…oject#59494)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3a69636)
…es (vllm-project#59450)

Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
(cherry picked from commit f7999d2)
… a saturated GPU pool (vllm-project#59309)

Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit 4056c8a)
…ject#59309 backport

Resolving the vllm-project#59309 cherry-pick conflict in
tests/v1/kv_connector/unit/test_hisparse_connector.py pulled in
test_scheduled_prefix_hit_publishes_adopted_copies from vllm-project#57930, which is
not on this branch; it imports _allocate_scheduled from
tests.v1.core.test_prefix_caching and fails on every platform. After this
change the file matches the upstream vllm-project#59309 diff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: khluu <khluu000@gmail.com>
…formers v5.18 (vllm-project#59613)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
(cherry picked from commit b558f16)
…ject#59614)

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
(cherry picked from commit 58b3298)
Squashes arbicity/vllm-turbo main @969c125e9 (upstream v0.28.0 + the seam:
b0a14f2 #28, 76a3e6d, 6d4c756 #29) into one commit and replays it onto
upstream v0.31.0 (db9527a, the commit vllm/vllm-openai:v0.31.0 is built
from) as a 3-way merge against the real base.

Upstream's v0.29-v0.31 KV-cache layout refactor (vllm#51718 and follow-ups)
replaced the allocator the seam patched: one backing allocation, per-layer
[B, H, N, C] views placed by the engine, page geometry read off the spec, and
AttentionBackend.customize_spec applied to every layer's spec by both model
runners. The seam is re-expressed on that and shrinks from 49 files to 18.

Kept (re-applied on upstream's structure):
  - plugin KV-cache dtype registry; --kv-cache-dtype choices/type
  - TURBO_ATTN backend slot (now a distinct enum value: two None members
    made CUSTOM an alias of TURBO_ATTN), turbo-attn spelling, auto-default
    for plugin dtypes (#29), CUDA candidate once registered
  - lifecycle hooks on_model_loaded / on_draft_model_loaded /
    on_kv_cache_initialized / adjust_kv_budget, fail-loud dispatch to every
    backend in use
  - MLA wrapping in the selector; MLA chunked-context _get_gather_op
  - _tq_layer_idx injection for tkv layers
  - aggregated_layer_count: fused (composite) pages, now one shared page
    per fused set in upstream's single-allocation planner

Moved to upstream's extension points (turbo-attn plugin side):
  - get_kv_cache_spec_class (Attention, MLAAttention, hybrid alignment)
    -> AttentionBackend.customize_spec
  - KVCacheSpec.get_manager_class -> KVCacheSpecRegistry MRO lookup
  - get_supported_kv_cache_dtypes -> supported_kv_cache_dtypes ClassVar
  - get_kv_cache_shape(kv_cache_spec=...) passthrough and the
    backend-managed-dtype shape coercion -> spec-driven views

Dropped:
  - spec_decode_warmup.py: upstream registers the same kernels with its JIT
    warmup registry (aed894c, vllm#56323)
  - turbo_attn_warmup.py, utils/cutedsl_cache.py: superseded by turbo-attn's
    own prefill prewarm (on_kv_cache_initialized) and CuTeDSL cache
  - per-group BlockPools, sampler reserve, padded-page block fill, drain
    hook / on_kv_manager_created, KVBlockZeroer clamp: conflict with
    upstream's single-pool allocator; no turbo-attn consumer for the hooks
  - rotary fast-path registry: upstream guards the import (1f60771,
    vllm#42679)
  - --kv-cache-dtype-skip-layers-dtype, VLLM_KV_CACHE_SKIP_LAYERS_DTYPE,
    Triton fused fp8 GEMM hook, gsm8k startup waits, notify-turbo-attn
    workflow, draft-backend inheritance: unused or superseded
  - FA2 varlen paged split-K patch: never reached the overlay image (the
    image installs no compiled _vllm_fa2_C) and the FA pin moved

PROTOCOL.md rewritten for the v0.31.0 seam.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An ours-merge: the tree is exactly v0.31.0 + the carry commit. It only
records the old fork main as an ancestor so the PR merges without a
force-push of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.