[Bugfix][Kimi K3] Skip absent metadata during CUDA graph profiling - #53581
Conversation
Co-authored-by: Codex <codex@openai.com> Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
|
Sherlock here, Kevin Luu's CI-monitoring agent. Exact H200 Extra Initialization validation ran in #85323 at literal branch head The merge-gate shard 1 passed 92/92 (3 deselected, 0 failed) in 45m55s on physical This PR merged at 15:16 UTC as Administrative note: #85322 used a mistyped full SHA and failed bootstrap before any test. It is not evidence and is superseded by #85323. Resolved on main: targeted post-merge main #85350 shard 1 at merged commit |
|
/ci run |
|
✅ CI is already running for this commit: https://buildkite.com/vllm/ci/builds/85323 |
yewentao256
left a comment
There was a problem hiding this comment.
Thanks for the work!
So sad that my PR just behind several minutes
#53584
| def test_kda_warmup_skips_missing_metadata(monkeypatch): | ||
| monkeypatch.setattr( | ||
| nvidia_kda, | ||
| "get_forward_context", | ||
| lambda: SimpleNamespace(attn_metadata={}), | ||
| ) | ||
| layer = object.__new__(nvidia_kda.KimiK3DeltaAttention) | ||
| object.__setattr__(layer, "prefix", "language_model.model.layers.0.self_attn") | ||
| empty = torch.empty(0, device=DEVICE) | ||
|
|
||
| assert layer._forward(empty, empty, empty, empty, empty) is None | ||
|
|
||
|
|
There was a problem hiding this comment.
| def test_kda_warmup_skips_missing_metadata(monkeypatch): | |
| monkeypatch.setattr( | |
| nvidia_kda, | |
| "get_forward_context", | |
| lambda: SimpleNamespace(attn_metadata={}), | |
| ) | |
| layer = object.__new__(nvidia_kda.KimiK3DeltaAttention) | |
| object.__setattr__(layer, "prefix", "language_model.model.layers.0.self_attn") | |
| empty = torch.empty(0, device=DEVICE) | |
| assert layer._forward(empty, empty, empty, empty, empty) is None |
I don't think we need a specific unit test for this update though
…llm-project#53581) Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
…llm-project#53581) Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
Summary
PR #53306 makes CUDA graph memory profiling omit Mamba-family attention groups from dummy-run metadata. The generic GDN implementations were updated to treat a missing per-layer metadata entry as a warmup path, but the dedicated Kimi K3 NVIDIA and AMD KDA implementations still indexed the entry directly.
That raises
KeyError: 'language_model.model.layers.0.self_attn'while initializing Kimi K3. Main CI reproduced it twice in the exact H200 Extra Initialization shard:Use the same missing-entry guard as the generic GDN paths and add a regression test for the empty-metadata warmup case.
Duplicate-work check
Searched open issues and PRs for the exact Kimi K3
KeyError, missing attention metadata, and CUDA graph profiling fixes. No existing fix addresses this path.Validation
uv run --no-project --python 3.12 python -m py_compile vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.pyuvx pre-commit run ruff-check --files vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.pyuvx pre-commit run ruff-format --files vllm/models/kimi_k3/nvidia/kda.py vllm/models/kimi_k3/amd/kda.py tests/models/kimi_k3/test_kda.pyThe exact GPU regression job will be run in Buildkite before merge. No model eval is needed because this changes only the dummy profiling/warmup path and preserves real attention metadata execution.
AI assistance
AI assistance was used to analyze the CI failures, identify the missed KDA implementations, and prepare the patch. The submitter reviewed every changed line and the validation results.