Skip to content

Vulkan - Xe Flash Attn / GEMM improvements from main PR#24408 - #2

Draft
XanderSC wants to merge 12 commits into
masterfrom
intel-gpu-feature
Draft

Vulkan - Xe Flash Attn / GEMM improvements from main PR#24408#2
XanderSC wants to merge 12 commits into
masterfrom
intel-gpu-feature

Conversation

@XanderSC

@XanderSC XanderSC commented Aug 4, 2026

Copy link
Copy Markdown
Owner

Overview

Attempting merge of master PR24408 into forked llama.cpp for enhanced Vulkan performance, specifically from Xe flash attention + including GEMM optimizations on Xe2/Xe3.

Additional information

Requirements

fish-jiang and others added 4 commits June 10, 2026 17:25
…PG Plus (1/3, Xe1-ARLH)

Co-authored-by: Xia, Jie <jie.xia@intel.com>
Co-authored-by: Liu, Russell <russell.liu@intel.com>
…G Plus/Xe2/Xe3)

Co-authored-by: Xia, Jie <jie.xia@intel.com>
Co-authored-by: Liu, Russell <russell.liu@intel.com>
…n for Intel MoE path (3/3, Xe-LPG Plus/Xe2/Xe3)

Co-authored-by: Xia, Jie <jie.xia@intel.com>
Co-authored-by: Liu, Russell <russell.liu@intel.com>
…, refine Vulkan host code for temp buffer reuse and thread lock
@XanderSC XanderSC changed the title Intel gpu feature Vulkan - Xe Flash Attn / GEMM improvements from main PR#24408 Aug 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants