Skip to content

update kernel config - #501

Merged
jikunshang merged 1 commit into
vllm-project:mainfrom
jikunshang:kunshang/conf_from_ci
Aug 4, 2026
Merged

jikunshang merged 1 commit into
vllm-project:mainfrom
jikunshang:kunshang/conf_from_ci

Conversation

@jikunshang

Copy link
Copy Markdown
Member

Essential Elements of an Effective PR Description Checklist

  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS ABOVE HAVE BEEN CONSIDERED.

Purpose

0.1.12.1 bump up PR: vllm-project/vllm#50441
CI https://buildkite.com/vllm/intel-ci/builds/7877#019fb598-d583-40ec-a16f-4f3f293e5ef5

Test Plan

Test Result

(Optional) Documentation Update

BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)

Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Copilot AI review requested due to automatic review settings July 31, 2026 05:16
@jikunshang jikunshang changed the title update config update kernel config Jul 31, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the XPU attention kernel configuration documentation and default paged-decode kernel preset to cover additional model families, and includes a small correctness/performance tweak in the Python reference attention path.

Changes:

  • Update paged_decode_default.conf and docs to reflect expanded default coverage (e.g., Starcoder2, Phi, VLM2Vec).
  • Add model-family guidance and example parameter rows in KERNEL_CONFIGURATION.md.
  • Ensure the reference paged-attention mask tensor is created on the same device as the attention tensor.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.

File Description
vllm_xpu_kernels/flash_attn_interface.py Places the reference mask tensor on the attention device to avoid device mismatch/copies.
KERNEL_CONFIGURATION.md Updates documentation around the default paged-decode preset coverage and adds model guidance/examples.
csrc/xpu/attn/kernel_configs/README.md Updates kernel-config README table to match the expanded default preset coverage.
csrc/xpu/attn/kernel_configs/paged_decode_default.conf Expands the default paged-decode config entries and updates header comments.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

v = (v.to(torch.float32) * v_descale).to(dtype)
attn = torch.einsum("qhd,khd->hqk", q, k).float()
empty_mask = torch.ones(query_len, kv_len)
empty_mask = torch.ones(query_len, kv_len).to(attn.device)
Comment on lines +15 to 16
# Most models: causal=true, local=false, sink=false (standard autoregressive)
#
@jikunshang
jikunshang enabled auto-merge (squash) August 3, 2026 03:05
@jikunshang
jikunshang merged commit 4509e9f into vllm-project:main Aug 4, 2026
10 checks passed
jinyouzhi pushed a commit to jinyouzhi/vllm-xpu-kernels that referenced this pull request Aug 12, 2026
update config

Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants