Skip to content

[Bugfix][Kimi K3] Skip absent metadata during CUDA graph profiling - #53581

Merged
njhill merged 1 commit into
vllm-project:mainfrom
khluu:fix/kimi-k3-cudagraph-profiling
Aug 24, 2026
Merged

njhill merged 1 commit into
vllm-project:mainfrom
khluu:fix/kimi-k3-cudagraph-profiling

Conversation

@khluu

@khluu khluu commented Aug 24, 2026

Copy link
Copy Markdown
Member

Summary

PR #53306 makes CUDA graph memory profiling omit Mamba-family attention groups from dummy-run metadata. The generic GDN implementations were updated to treat a missing per-layer metadata entry as a warmup path, but the dedicated Kimi K3 NVIDIA and AMD KDA implementations still indexed the entry directly.

That raises KeyError: 'language_model.model.layers.0.self_attn' while initializing Kimi K3. Main CI reproduced it twice in the exact H200 Extra Initialization shard:

Use the same missing-entry guard as the generic GDN paths and add a regression test for the empty-metadata warmup case.

Duplicate-work check

Searched open issues and PRs for the exact Kimi K3 KeyError, missing attention metadata, and CUDA graph profiling fixes. No existing fix addresses this path.

Validation

  • uv run --no-project --python 3.12 python -m py_compile vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.py
  • uvx pre-commit run ruff-check --files vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.py
  • uvx pre-commit run ruff-format --files vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.py

The exact GPU regression job will be run in Buildkite before merge. No model eval is needed because this changes only the dummy profiling/warmup path and preserves real attention metadata execution.

AI assistance

AI assistance was used to analyze the CI failures, identify the missed KDA implementations, and prepare the patch. The submitter reviewed every changed line and the validation results.

Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the bug Something isn't working label Aug 24, 2026
@khluu

khluu commented Aug 24, 2026 •

Copy link
Copy Markdown
Member Author

Sherlock here, Kevin Luu's CI-monitoring agent. Exact H200 Extra Initialization validation ran in #85323 at literal branch head a891d1222c422e87de6c2e5b94d75e642b6bf2a8 with PR metadata and only basic-models-tests-extra-initialization selected.

The merge-gate shard 1 passed 92/92 (3 deselected, 0 failed) in 45m55s on physical h200-ci-4. It explicitly ran test_can_initialize_large_subset[KimiK3ForConditionalGeneration]; the full CUDA-graph memory-profiling and FlashInfer-autotune path completed without the prior KeyError.

This PR merged at 15:16 UTC as 4c56e62c85cea8fc2251efc25159836c214402aa. Natural webhook main #85345 left the exact shard blocked/not run, so I launched targeted post-merge main #85350 at the literal merge commit with only basic-models-tests-extra-initialization selected. The incident remains open until #85350 shard 1 passes.

Administrative note: #85322 used a mistyped full SHA and failed bootstrap before any test. It is not evidence and is superseded by #85323.

Resolved on main: targeted post-merge main #85350 shard 1 at merged commit 4c56e62c85ce finished exit 0 on physical h200-ci-4: 92 passed / 3 deselected / 0 failed in 49m30s. Exact test_can_initialize_large_subset[KimiK3ForConditionalGeneration] passed after completing CUDA-graph profiling and FlashInfer autotune without the prior missing-metadata KeyError. This satisfies the merge + exact post-main resolution gate.

@njhill njhill left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @khluu, sorry about this

@github-project-automation github-project-automation Bot moved this to Ready in NVIDIA Aug 24, 2026
@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 24, 2026
@njhill

njhill commented Aug 24, 2026

Copy link
Copy Markdown
Member

/ci run

@njhill
njhill enabled auto-merge (squash) August 24, 2026 14:16
@github-actions

Copy link
Copy Markdown

✅ CI is already running for this commit: https://buildkite.com/vllm/ci/builds/85323

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the work!

So sad that my PR just behind several minutes
#53584

Comment on lines +63 to +75
def test_kda_warmup_skips_missing_metadata(monkeypatch):
monkeypatch.setattr(
nvidia_kda,
"get_forward_context",
lambda: SimpleNamespace(attn_metadata={}),
)
layer = object.__new__(nvidia_kda.KimiK3DeltaAttention)
object.__setattr__(layer, "prefix", "language_model.model.layers.0.self_attn")
empty = torch.empty(0, device=DEVICE)

assert layer._forward(empty, empty, empty, empty, empty) is None


Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
def test_kda_warmup_skips_missing_metadata(monkeypatch):
monkeypatch.setattr(
nvidia_kda,
"get_forward_context",
lambda: SimpleNamespace(attn_metadata={}),
)
layer = object.__new__(nvidia_kda.KimiK3DeltaAttention)
object.__setattr__(layer, "prefix", "language_model.model.layers.0.self_attn")
empty = torch.empty(0, device=DEVICE)
assert layer._forward(empty, empty, empty, empty, empty) is None

I don't think we need a specific unit test for this update though

@njhill
njhill merged commit 4c56e62 into vllm-project:main Aug 24, 2026
26 checks passed
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Aug 24, 2026
@njhill

njhill commented Aug 24, 2026

Copy link
Copy Markdown
Member

@khluu's bot is too fast!

BTW I found one other similar regression: #53593

am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
…llm-project#53581)

Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
mikeshawcode pushed a commit to mikeshawcode/vllm that referenced this pull request Sep 1, 2026
…llm-project#53581)

Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working k3 kimi nvidia ready ONLY add when PR is ready to merge/full CI is needed

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants