Skip to content

[Bugfix] FlashInfer: read the KV cache layout from attention metadata - #54834

Open
ptorsten wants to merge 2 commits into
vllm-project:mainfrom
ptorsten:fix-dflash-draft-kv-layout
Open

ptorsten wants to merge 2 commits into
vllm-project:mainfrom
ptorsten:fix-dflash-draft-kv-layout

Conversation

@ptorsten

@ptorsten ptorsten commented Sep 1, 2026 •

Copy link
Copy Markdown

Purpose

FlashInferImpl reads kv_cache_layout from a cache_config captured at construction. #51718 resolves the layout after model load and records it on the worker's CacheConfig. A drafter with a kv_cache_dtype override (DFlash, EAGLE, DSpark alike) is built under a replace()d config, so its impls hold a copy that never sees the resolution, and the first draft forward during profiling raises "KV cache layout has not been resolved yet".

Carry the layout on FlashInferMetadata instead. The builder already reads it from the worker's config to plan the wrappers, and every impl-side read is in forward(), so the impl needs no config of its own. FlashInfer is the only backend whose impl read the layout this way.

Replaces the first revision, which propagated the layout into draft config copies from worker_base.py. No open PR touches drafter layout resolution.

Test

pytest tests/v1/attention/test_attention_backends.py -k "flashinfer or causal_backend_correctness" - 110 passed on GB10 (sm_121).

2x DGX Spark (GB10), TP=2, FLASHINFER, target Qwen3.8-27B-NVFP4 with KV auto, DFlash2 drafter z-lab/Qwen3.8-27B-DFlash2 with kv_cache_dtype: fp8, --linear-backend flashinfer_cutlass, 8 speculative tokens. main at 8f816a3 fails in memory profiling with the error above from FlashInferImpl.kv_cache_layout. With this commit it reaches READY (34 piecewise + 2 full target graphs, 12 DFlash2 full graphs) and serves the 12-scenario tool-calling suite 12/12 at 154 tok/s mean decode, 72% draft acceptance. A W4A16 drafter does not load on main for an unrelated reason (_build_context_kv_buffers reads qkv_proj.weight after Marlin repacks it), hence the bf16 drafter.

AI assistance was used; every line reviewed and the tests above run by me.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@mergify mergify Bot added speculative-decoding dflash mrv2 Model Runner V2 specific labels Sep 1, 2026
@mergify mergify Bot added the bug Something isn't working label Sep 1, 2026

@Manny7717 Manny7717 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified locally on head 779e4183 (worktree + standalone probes; no DFlash hardware here, so mechanism-level verification).

Root cause confirmed by execution: kv_cache_layout is field(default=None, init=False) in CacheConfig (vllm/config/cache.py:89) and load_dflash_model builds the draft config via dataclasses.replace(vllm_config, ...) — which drops init=False fields. Executed probe: replace(cc, cache_dtype='auto') yields a draft whose kv_cache_layout is None even after the target resolved LHBNC. The engine's later set_kv_cache_layout collective RPC (core.py:294 → executor collective_rpc → worker_base.py:112) only ever touches the target config, so the draft's first forward dies in get_resolved_kv_cache_layout with exactly the reported "KV cache layout has not been resolved yet" (reproduced the raise).

Fix mechanics verified by probes:

  • Mirror step: draft_cache.kv_cache_layout = vllm_config.cache_config.kv_cache_layout carries an already-resolved value; guarded by draft_cache is not vllm_config.cache_config so the passthrough (no kv_cache_dtype override) case stays untouched.
  • Registration: _draft_cache_configs sibling list on the target config is consumed by the worker's set_kv_cache_layout, which now propagates to every registered draft copy — verified the full chain (mirror → register → record → both configs end at LHBNC).
  • Fail-closed preserved: record_kv_cache_layout still raises on a conflicting layout (already resolved to LHBNC; cannot change it to LBHNC) — the propagation can't silently flip an inconsistent draft.
  • get_resolved_kv_cache_layout raise on a stale (None-layout) draft still reproduces the pre-fix failure mode, confirming the propagation is what fixes it.

No regressions: dflash unit suites (test_dflash_causality, test_dflash_prepare_inputs, test_dflash_lookahead): head 2 failed / 14 passed / 16 errors vs base 73723b7 identical 2f/14p/16e — the failures/errors are CUDA-less torch.accelerator fixture noise, byte-identical on both. ruff (pinned 0.14.0, CI version) clean on both changed files. Author ran the real hardware path (2x DGX Spark, TP=2, Qwen3.8-27B + DFlash2 drafter) before/after.

Non-blocking nit: no automated test added — a config-level unit test pinning that load_dflash_model's replace path carries kv_cache_layout (or that set_kv_cache_layout reaches registered drafts) would guard this against future dataclass reordering; the author's e2e evidence covers the live path today.

@ptorsten

ptorsten commented Sep 2, 2026

Copy link
Copy Markdown
Author

Added a test, thanks for the nit.

@ptorsten
ptorsten force-pushed the fix-dflash-draft-kv-layout branch from 4d0fbbd to f8dc114 Compare September 2, 2026 06:28

@AndreasKaratzas AndreasKaratzas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM -- probably a second pair of eyes would be good for the worker base change

@@ -0,0 +1,58 @@
# SPDX-License-Identifier: Apache-2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if this test file is needed though.

@AndreasKaratzas AndreasKaratzas added ready ONLY add when PR is ready to merge/full CI is needed verified Run pre-commit for new contributors without triggering other tests labels Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

✅ @ptorsten, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@mergify

mergify Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Hi @ptorsten, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

FlashInferImpl captured cache_config from the current vllm config at
construction and read kv_cache_layout from it lazily, since the engine
core resolves the layout after model load (vllm-project#51718) and records it on the
worker's CacheConfig. A draft model with a kv_cache_dtype override is
built under a replace()d config whose CacheConfig is a copy, so its impls
never see the resolution and the first draft forward, during memory
profiling, raises "KV cache layout has not been resolved yet".

Put the resolved layout on FlashInferMetadata. The builder reads it from
the worker's config at build time and every impl-side read happens in
forward() with the metadata in hand, so the impl needs no config of its
own. This covers the DFlash, EAGLE and DSpark drafters alike, which all
derive their config the same way.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Patrik Torstensson <patrik.torstensson@gmail.com>
@ptorsten
ptorsten force-pushed the fix-dflash-draft-kv-layout branch from b0a9782 to 44aac86 Compare September 4, 2026 08:30
@ptorsten ptorsten changed the title [Bugfix] Propagate resolved KV cache layout to draft-model cache configs [Bugfix] FlashInfer: read the KV cache layout from attention metadata Sep 4, 2026
@coderabbitai

coderabbitai Bot commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Team

Run ID: f1289525-70d0-4f77-b7db-37e2b6080de6

📥 Commits

Reviewing files that changed from the base of the PR and between 9cd956c and 44aac86.

📒 Files selected for processing (1)
  • vllm/v1/attention/backends/flashinfer.py

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Refactor
    • Improved FlashInfer attention metadata handling by resolving the KV cache layout earlier and passing it through attention metadata.
    • Maintained consistent KV cache layout behavior across attention operations and decoding paths.

Walkthrough

FlashInfer now resolves the KV cache layout in FlashInferMetadataBuilder, stores it in FlashInferMetadata, and uses it in FlashInferImpl.forward() for stride ordering, layout checks, and decode kernel calls.

Changes

FlashInfer KV cache layout

Layer / File(s) Summary
Metadata layout contract
vllm/v1/attention/backends/flashinfer.py
FlashInferMetadata stores kv_cache_layout. The metadata builder populates the field from its resolved layout.
Forward layout consumption
vllm/v1/attention/backends/flashinfer.py
forward() uses the metadata layout for stride ordering, HND assertions, and XQA and TRTLLM decode kernel arguments.

Estimated code review effort: 2 (Simple) | ~10 minutes

Merge Risk: ⚪ Minimal · up to f1ab6

This change carries the resolved KV-cache layout through FlashInfer metadata so draft-model attention uses the correct layout during execution. The updated paths consistently use the propagated value, with no current merge-blocking risk identified.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 4 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly describes the main change: FlashInfer now reads the KV cache layout from attention metadata.
Description check ✅ Passed The description explains the KV cache layout resolution bug, the metadata-based fix, affected drafter scenarios, and test results. It is directly related to the changeset.
  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@LucasWilkinson

Copy link
Copy Markdown
Contributor

I opened #55384 as a draft alternative for design comparison, not as a second fix intended to merge alongside this PR. Instead of carrying the resolved layout through FlashInfer metadata, it makes CacheConfig copies share an explicit ResolvedKVCacheLayout state object. That addresses the same late-resolution issue generically across DFlash, EAGLE, DSpark, and other draft configs that override cache dtype, at the cost of a broader CacheConfig API change. Patrik is credited as a co-author.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working dflash mrv2 Model Runner V2 specific nvidia ready ONLY add when PR is ready to merge/full CI is needed speculative-decoding verified Run pre-commit for new contributors without triggering other tests

Projects

Status: No status
Status: Backlog

Development

Successfully merging this pull request may close these issues.

4 participants