Skip to content

[Feature][Frontend] Expose effective attention block size for DCP - #3

Open
JulienDarve wants to merge 97 commits into
mainfrom
jdarve/kv-cache-group-metadata
Open

JulienDarve wants to merge 97 commits into
mainfrom
jdarve/kv-cache-group-metadata

Conversation

@JulienDarve

@JulienDarve JulienDarve commented Sep 9, 2026 •

Copy link
Copy Markdown
Owner

Purpose

With decode context parallelism (DCP), a physical KV cache block of 16 tokens can represent 64 tokens across four ranks. Clients need that effective size to interpret full-attention cache events.

This PR exposes the value as CacheConfig.effective_attention_block_size, populated from the initialized cache managers. Python clients read it from their config, and the engine startup response carries it to Rust and gRPC Control.GetServerInfo.

Existing physical block-size fields remain unchanged. Without DCP, the effective size matches the full-attention physical block size. The new field is optional and unavailable when the engines cannot report a common full-attention block size.

Duplicate checks found no open PR exposing this DCP-adjusted value through config, startup metadata, and gRPC.

AI assistance: OpenAI Codex assisted with implementation and validation.

Test Plan

Python:

PYTHONPATH=. .venv/bin/python -m pytest \
  tests/config/test_config_utils.py \
  tests/v1/core/test_kv_cache_utils.py \
  tests/v1/engine/test_engine_core_client.py \
  -k 'effective_attention_block_size or apply_ready_response or cache_config_hash' \
  -q --tb=short

From rust/:

cargo nextest run -p vllm-engine-core-client -p vllm-server \
  -E 'test(python_msgpack_fixtures) | test(grpc::tests::control_)'
cargo clippy -p vllm-engine-core-client -p vllm-server --all-targets -- -D warnings

From the repository root:

.venv/bin/pre-commit run mypy-3.12 --all-files --hook-stage manual
cargo fmt --manifest-path rust/Cargo.toml --all -- --check

Test Result

  • Python: 9 passed, covering DCP=1/4 cache-manager values, startup synchronization, missing fields, rank disagreement, and compilation hash stability.
  • Rust: 7 passed, covering MessagePack compatibility and gRPC control behavior.
  • GPU smoke: hmellor/tiny-random-LlamaForCausalLM on RTX 6000 Ada generated four tokens. With DCP=1, the Python config and actual startup payload both reported 16 tokens; physical block size remained 16. The run used precompiled extensions and native sampling.
  • Multi-GPU DCP inference was not run; DCP=4 was covered at the cache-manager level.
  • Clippy, pre-commit hooks including manual mypy checks, protobuf lint, and protobuf backward-compatibility checks against the PR base passed.
  • Backward compatibility: the base Python decoder accepted the captured new startup payload; the updated decoder accepted the base payload with the new field unavailable.
  • Full-repository Python 3.12 type checking and Rust formatting passed.
  • Fork CI checks are skipped because the workflows run only for vllm-project; the validation results above were obtained locally.

Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
@JulienDarve JulienDarve changed the title [Feature][Frontend] Expose effective KV cache group block sizes [Feature][Frontend] Expose effective attention block size Sep 11, 2026
@JulienDarve JulienDarve changed the title [Feature][Frontend] Expose effective attention block size [Feature][Frontend] Expose effective attention block size for DCP Sep 11, 2026
JulienDarve and others added 23 commits September 11, 2026 17:03
Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
…lm-project#51104)

Signed-off-by: xiaguan <751080330@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…am (vllm-project#56382)

Signed-off-by: zhaoguochun1995 <zhaoguochun1995@163.com>
…ct#56687)

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
)

Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
… cannot be used (vllm-project#56560)

Signed-off-by: JohnQinAMD <yanyuan.qin@amd.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on (vllm-project#56844)

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Bhushan Asati <bhushanasati25@gmail.com>
Co-authored-by: Megha Agarwal <19240983+meghaagr13@users.noreply.github.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
…ection of a late fetch (vllm-project#53453)

Signed-off-by: Liran Schour <lirans@il.ibm.com>
…llm-project#56472)

Signed-off-by: 子华 <huaxi.shx@alibaba-inc.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
…vllm-project#53444)

Signed-off-by: wuhangxian <1391938827@qq.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
…llm-project#56893)

Signed-off-by: Yongye Zhu <yongye@inferact.ai>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Yongye Zhu <yongye@inferact.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ct#56640)

Signed-off-by: Yizheng Jiao <jyizheng@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
…h combine (vllm-project#52781)

Signed-off-by: Mario <404mario@users.noreply.github.com>
Co-authored-by: Mario <404mario@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Sherlock <sherlock@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
njhill and others added 30 commits September 15, 2026 20:36
Signed-off-by: Nick Hill <nickhill123@gmail.com>
…llm-project#56254)

Signed-off-by: Jared Wen <w13431838023@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ct#56346)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
…l push prefill (vllm-project#50494)

Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
vllm-project#56545)

Signed-off-by: Shang Wang <samshang.wang@mail.utoronto.ca>
Co-authored-by: Codex <noreply@openai.com>
…yer MTP kv caches during P/D (vllm-project#55055)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
…#56505)

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
…4 Checkpoint (vllm-project#56176)

Signed-off-by: Colin Zeng <Colin.Zeng@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
…3940)

Signed-off-by: ppalanga <ppalanga@amd.com>
Signed-off-by: Poovaiah Palangappa <poovaiah.palangappa@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Poovaiah Palangappa <poovaiah.palangappa@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com>
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
Signed-off-by: Francisco Javier Arceo <farceo@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
…roject#57033)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
…3674)

Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
…57046)

Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: mayuyuace <qiming1.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…#56096)

Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
vllm-project#51026)

Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: Islam <islam.almersawi@openinnovation.ai>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
…55885)

Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
…guage model only) (vllm-project#56231)

Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
vllm-project#56441)

Signed-off-by: 0z5a <0z5a@users.noreply.github.com>
Signed-off-by: Yongye Zhu <yongye@inferact.ai>
Co-authored-by: 0z5a <0z5a@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…56904)

Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: Sherlock <sherlock@raft.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Codex <noreply@openai.com>
…m-project#56709)

Signed-off-by: jacklin78911-collab <jacklin78911@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.