Skip to content

fix(vllm): use engine-reported logical KV block sizes - #14590

Closed
JulienDarve wants to merge 4 commits into
mainfrom
jdarve/vllm-legacy-logical-kv-block-size
Closed

JulienDarve wants to merge 4 commits into
mainfrom
jdarve/vllm-legacy-logical-kv-block-size

Conversation

@JulienDarve

Copy link
Copy Markdown
Contributor

Summary

Use vLLM's initialized logical block size throughout the legacy Python backend. With physical block size 16 and DCP=2, full KV events contain 32 tokens; configuring Dynamo with 16 causes those events to be rejected.

  • Fetch typed metadata through AsyncLLM.get_kv_cache_group_metadata() and cache the main-attention logical size for the publisher, base-model registration, and LoRA registration.
  • Reject unavailable, invalid, or conflicting main-attention sizes. Preserve the no-cache case without deriving DCP or PCP multipliers.
  • Add compact metadata coverage and one four-H100 DCP event-indexing regression using Qwen2.5-3B.

Validation

  • 233 Python tests passed, including existing worker-factory and LoRA registration coverage.
  • Existing TinyLlama KV-router end-to-end test passed on one RTX 6000 Ada: 84.77 seconds.
  • Required pre-commit hooks passed.
  • TP=4/DCP=2 configuration validated through vLLM. Multi-GPU execution, negative control, VRAM profiling, and nightly-test promotion for PR CI remain pending.

Tests used Python source from JulienDarve/vllm#3 at 20905fbeda0eb9760e1782d6ea5a0f96d4fe2457, with existing precompiled extensions.

Dependencies

Draft until the metadata API is available across supported CUDA, CPU, and XPU runtimes. The current vLLM 0.28.0 CUDA/CPU and 0.27.1 XPU pins do not provide this contract.

Related Issues

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
@github-actions github-actions Bot added fix backend::vllm Relates to the vllm backend labels Sep 10, 2026
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
@JulienDarve

Copy link
Copy Markdown
Contributor Author

Being implemented in the version bump PR: #15182

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend fix size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant