Skip to content

[CUDA] Auto-tune small-M decode MatMul (cuBLAS vs small-N GEMV) and speed up MatMulNBits GEMV - #32876

Merged
Tianlei Wu (tianleiwu) merged 3 commits into
mainfrom
tlwu/cuda-small-m-decode-gemv
Sep 29, 2026
Merged

Tianlei Wu (tianleiwu) merged 3 commits into
mainfrom
tlwu/cuda-small-m-decode-gemv

perf(cuda): auto-tune small-M MatMul between cuBLAS and small-N GEMV

a552f5f
Select commit
Loading
Failed to load commit list.

Select a check to view from the sidebar