update kernel config - #501
Merged
Merged
Conversation
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Contributor
There was a problem hiding this comment.
Pull request overview
This PR updates the XPU attention kernel configuration documentation and default paged-decode kernel preset to cover additional model families, and includes a small correctness/performance tweak in the Python reference attention path.
Changes:
- Update
paged_decode_default.confand docs to reflect expanded default coverage (e.g., Starcoder2, Phi, VLM2Vec). - Add model-family guidance and example parameter rows in
KERNEL_CONFIGURATION.md. - Ensure the reference paged-attention mask tensor is created on the same device as the attention tensor.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
vllm_xpu_kernels/flash_attn_interface.py |
Places the reference mask tensor on the attention device to avoid device mismatch/copies. |
KERNEL_CONFIGURATION.md |
Updates documentation around the default paged-decode preset coverage and adds model guidance/examples. |
csrc/xpu/attn/kernel_configs/README.md |
Updates kernel-config README table to match the expanded default preset coverage. |
csrc/xpu/attn/kernel_configs/paged_decode_default.conf |
Expands the default paged-decode config entries and updates header comments. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| v = (v.to(torch.float32) * v_descale).to(dtype) | ||
| attn = torch.einsum("qhd,khd->hqk", q, k).float() | ||
| empty_mask = torch.ones(query_len, kv_len) | ||
| empty_mask = torch.ones(query_len, kv_len).to(attn.device) |
Comment on lines
+15
to
16
| # Most models: causal=true, local=false, sink=false (standard autoregressive) | ||
| # |
jikunshang
enabled auto-merge (squash)
August 3, 2026 03:05
zufangzhu
approved these changes
Aug 3, 2026
YizhouZ
approved these changes
Aug 4, 2026
jinyouzhi
pushed a commit
to jinyouzhi/vllm-xpu-kernels
that referenced
this pull request
Aug 12, 2026
update config Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS ABOVE HAVE BEEN CONSIDERED.
Purpose
0.1.12.1 bump up PR: vllm-project/vllm#50441
CI https://buildkite.com/vllm/intel-ci/builds/7877#019fb598-d583-40ec-a16f-4f3f293e5ef5
Test Plan
Test Result
(Optional) Documentation Update
BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)