Skip to content

opencl: optimize mean and sum_row kernels - #19614

Merged
lhez merged 3 commits into
ggml-org:masterfrom
qualcomm:sq/opencl-mean-sum_rows-opt
Feb 17, 2026
Merged

opencl: optimize mean and sum_row kernels#19614
lhez merged 3 commits into
ggml-org:masterfrom
qualcomm:sq/opencl-mean-sum_rows-opt

Conversation

@shaofeiqi

Copy link
Copy Markdown
Contributor

This PR optimizes the mean op and sum_rows op for the OpenCL backend.

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Feb 14, 2026
@lhez

lhez commented Feb 17, 2026

Copy link
Copy Markdown
Contributor

gemma-3n-E4B-Q8_0 on X Elite -

master,

common_perf_print: prompt eval time =    2205.01 ms /   235 tokens (    9.38 ms per token,   106.58 tokens per second)
common_perf_print:        eval time =   26895.76 ms /   256 runs   (  105.06 ms per token,     9.52 tokens per second)

this PR,

common_perf_print: prompt eval time =    2009.66 ms /   235 tokens (    8.55 ms per token,   116.94 tokens per second)
common_perf_print:        eval time =   21228.97 ms /   256 runs   (   82.93 ms per token,    12.06 tokens per second)

@lhez
lhez force-pushed the sq/opencl-mean-sum_rows-opt branch from bbdab67 to 598c7aa Compare February 17, 2026 07:06
@lhez
lhez marked this pull request as ready for review February 17, 2026 18:50
@lhez

lhez commented Feb 17, 2026

Copy link
Copy Markdown
Contributor

Workgroup size calculation can be further refined, but current form should be good for now.

@lhez
lhez merged commit 983559d into ggml-org:master Feb 17, 2026
78 checks passed
liparetejas pushed a commit to liparetejas/llama.cpp that referenced this pull request Feb 23, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
bartowski1182 pushed a commit to bartowski1182/llama.cpp that referenced this pull request Mar 2, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
ArberSephirotheca pushed a commit to ArberSephirotheca/llama.cpp that referenced this pull request Mar 3, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
ljubomirj pushed a commit to ljubomirj/llama.cpp that referenced this pull request May 6, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
fewtarius pushed a commit to fewtarius/CachyLLama that referenced this pull request May 30, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* opencl: optimize mean and sum_row kernels

* opencl: add comment for max subgroups

* opencl: format

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants