Skip to content

HIP: MMQ Dispatch config modification - separation of RDNA3, 3.5 from 4 and tune 4. - #26199

Merged
pwilkin merged 1 commit into
ggml-org:masterfrom
Geramy:mmq-rdna3-configs
Jul 29, 2026
Merged

HIP: MMQ Dispatch config modification - separation of RDNA3, 3.5 from 4 and tune 4.#26199
pwilkin merged 1 commit into
ggml-org:masterfrom
Geramy:mmq-rdna3-configs

Conversation

@Geramy

@Geramy Geramy commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Overview

This modification adds mmq configurations for host and device for RDNA3 and 3.5 separating them from RDNA4 as well as reconfiguring RDNA4 to be more optimal on real hardware.

I have also found that stream_k true helps a lot for Dense models and hurts MoE models, we should find a way to make MoE vs Dense have different mmq dispatch configurations to increase overall performance per type which I have found greatly improve performance as implied above.

Additional information

Also suggested to do the same in PR: #25940 which I volunteered to do because I already was working on this as it was something I had found as a good starting project.

Keep in mind my RDNA4 tests are on a MinisForum MS-S1 over TB5 with a R9700 inside a razor eGPU TB5 enclosure.
Performance may vary but more likely than not better on a board that has pcie 4.0 / 5.0 x16 vs the pcie 4.0 x4 I have over TB5.

RDNA 3.5 (Radeon 8060S iGPU, Ryzen AI MAX+ 395, gfx1151)

Model Microbatch size Test t/s master t/s mmq-rdna3-configs Speedup
nemotron_h_moe 31B.A3.5B Q4_K - Medium 16 pp2048 285.37 301.16 1.06
nemotron_h_moe 31B.A3.5B Q4_K - Medium 32 pp2048 204.43 218.77 1.07
nemotron_h_moe 31B.A3.5B Q4_K - Medium 64 pp2048 449.64 454.17 1.01
nemotron_h_moe 31B.A3.5B Q4_K - Medium 128 pp2048 415.14 409.74 0.99
nemotron_h_moe 31B.A3.5B Q4_K - Medium 256 pp2048 598.31 599.47 1.00
nemotron_h_moe 31B.A3.5B Q4_K - Medium 512 pp2048 869.75 861.07 0.99
nemotron_h_moe 31B.A3.5B Q4_K - Medium 1024 pp2048 1176.54 1153.26 0.98
nemotron_h_moe 31B.A3.5B Q4_K - Medium 2048 pp2048 1396.03 1370.54 0.98
qwen3 4B Q4_0 16 pp2048 753.19 823.78 1.09
qwen3 4B Q4_0 32 pp2048 826.36 817.00 0.99
qwen3 4B Q4_0 64 pp2048 1751.51 1759.21 1.00
qwen3 4B Q4_0 128 pp2048 2076.57 2084.85 1.00
qwen3 4B Q4_0 256 pp2048 2387.45 2367.47 0.99
qwen3 4B Q4_0 512 pp2048 2409.80 2399.34 1.00
qwen3 4B Q4_0 1024 pp2048 2356.71 2352.62 1.00
qwen3 4B Q4_0 2048 pp2048 2222.85 2219.39 1.00
qwen3 8B Q8_0 16 pp2048 349.31 367.08 1.05
qwen3 8B Q8_0 32 pp2048 508.57 520.80 1.02
qwen3 8B Q8_0 64 pp2048 896.66 890.80 0.99
qwen3 8B Q8_0 128 pp2048 1110.88 1106.22 1.00
qwen3 8B Q8_0 256 pp2048 1252.04 1251.09 1.00
qwen3 8B Q8_0 512 pp2048 1267.06 1267.72 1.00
qwen3 8B Q8_0 1024 pp2048 1286.55 1285.40 1.00
qwen3 8B Q8_0 2048 pp2048 1249.52 1255.84 1.01
qwen35 27B Q4_0 16 pp2048 132.63 136.68 1.03
qwen35 27B Q4_0 32 pp2048 140.19 153.57 1.10
qwen35 27B Q4_0 64 pp2048 286.07 287.16 1.00
qwen35 27B Q4_0 128 pp2048 336.75 337.62 1.00
qwen35 27B Q4_0 256 pp2048 363.48 365.50 1.01
qwen35 27B Q4_0 512 pp2048 357.52 355.71 0.99
qwen35 27B Q4_0 1024 pp2048 325.44 324.88 1.00
qwen35 27B Q4_0 2048 pp2048 300.38 300.60 1.00
qwen35 27B Q4_K - Medium 16 pp2048 114.09 121.81 1.07
qwen35 27B Q4_K - Medium 32 pp2048 181.06 203.24 1.12
qwen35 27B Q4_K - Medium 64 pp2048 250.48 250.64 1.00
qwen35 27B Q4_K - Medium 128 pp2048 303.64 303.69 1.00
qwen35 27B Q4_K - Medium 256 pp2048 327.49 327.73 1.00
qwen35 27B Q4_K - Medium 512 pp2048 336.56 331.88 0.99
qwen35 27B Q4_K - Medium 1024 pp2048 310.49 308.11 0.99
qwen35 27B Q4_K - Medium 2048 pp2048 285.88 291.43 1.02

RDNA 4 (Radeon AI PRO R9700, gfx1201)

Model Microbatch size Test t/s master t/s mmq-rdna3-configs Speedup
nemotron_h_moe 31B.A3.5B Q4_K - Medium 16 pp2048 497.99 518.36 1.04
nemotron_h_moe 31B.A3.5B Q4_K - Medium 32 pp2048 678.30 743.93 1.10
nemotron_h_moe 31B.A3.5B Q4_K - Medium 64 pp2048 935.93 991.05 1.06
nemotron_h_moe 31B.A3.5B Q4_K - Medium 128 pp2048 1081.18 1078.77 1.00
nemotron_h_moe 31B.A3.5B Q4_K - Medium 256 pp2048 1759.84 1757.94 1.00
nemotron_h_moe 31B.A3.5B Q4_K - Medium 512 pp2048 2715.48 2716.25 1.00
nemotron_h_moe 31B.A3.5B Q4_K - Medium 1024 pp2048 3763.76 3749.64 1.00
nemotron_h_moe 31B.A3.5B Q4_K - Medium 2048 pp2048 4482.67 4484.64 1.00
qwen3 4B Q4_0 16 pp2048 1061.06 1195.57 1.13
qwen3 4B Q4_0 32 pp2048 1525.66 1676.35 1.10
qwen3 4B Q4_0 64 pp2048 2242.85 2422.62 1.08
qwen3 4B Q4_0 128 pp2048 3699.20 3734.17 1.01
qwen3 4B Q4_0 256 pp2048 5438.61 5454.37 1.00
qwen3 4B Q4_0 512 pp2048 6620.78 6594.45 1.00
qwen3 4B Q4_0 1024 pp2048 7189.12 7167.10 1.00
qwen3 4B Q4_0 2048 pp2048 6857.16 6832.46 1.00
qwen3 8B Q8_0 16 pp2048 722.85 771.65 1.07
qwen3 8B Q8_0 32 pp2048 1022.69 1105.85 1.08
qwen3 8B Q8_0 64 pp2048 1475.42 1564.39 1.06
qwen3 8B Q8_0 128 pp2048 2750.42 2765.58 1.01
qwen3 8B Q8_0 256 pp2048 3962.44 3956.41 1.00
qwen3 8B Q8_0 512 pp2048 4486.76 4483.32 1.00
qwen3 8B Q8_0 1024 pp2048 4744.09 4742.62 1.00
qwen3 8B Q8_0 2048 pp2048 4645.29 4645.06 1.00
qwen35 27B Q4_0 16 pp2048 229.41 240.77 1.05
qwen35 27B Q4_0 32 pp2048 376.83 407.59 1.08
qwen35 27B Q4_0 64 pp2048 535.51 567.05 1.06
qwen35 27B Q4_0 128 pp2048 745.27 737.52 0.99
qwen35 27B Q4_0 256 pp2048 1012.16 1005.89 0.99
qwen35 27B Q4_0 512 pp2048 1150.41 1145.22 1.00
qwen35 27B Q4_0 1024 pp2048 1208.30 1204.81 1.00
qwen35 27B Q4_0 2048 pp2048 1206.24 1204.89 1.00
qwen35 27B Q4_K - Medium 16 pp2048 226.83 237.66 1.05
qwen35 27B Q4_K - Medium 32 pp2048 362.78 400.74 1.10
qwen35 27B Q4_K - Medium 64 pp2048 498.09 544.36 1.09
qwen35 27B Q4_K - Medium 128 pp2048 658.45 658.61 1.00
qwen35 27B Q4_K - Medium 256 pp2048 896.79 893.90 1.00
qwen35 27B Q4_K - Medium 512 pp2048 1013.67 1012.89 1.00
qwen35 27B Q4_K - Medium 1024 pp2048 1060.31 1056.47 1.00
qwen35 27B Q4_K - Medium 2048 pp2048 1061.01 1060.54 1.00

Requirements

  • I have read and agree with the contributing guidelines: Yes
  • AI usage disclosure: YES - only for research to find holes in areas of the project that do not have optimizations for rocm specifically. We have come across a list of roughly 6 items and this is one that seemed to be the easiest to get involved in the project in a constructive way.

@Geramy
Geramy requested a review from a team as a code owner July 27, 2026 21:41
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Jul 27, 2026
@Geramy Geramy changed the title add rdna3.5, and 3 to mmq configs so they can be tuned independently. HIP: MMQ Add separate configuration for RDNA3, 3.5 and tune 4. Jul 27, 2026
@Geramy Geramy changed the title HIP: MMQ Add separate configuration for RDNA3, 3.5 and tune 4. HIP: MMQ Dispatch config modification - separation of RDNA3, 3.5 from 4 and tune 4. Jul 27, 2026

@pwilkin pwilkin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pwilkin
pwilkin merged commit 60bccc3 into ggml-org:master Jul 29, 2026
15 of 20 checks passed
justinappler added a commit to justinappler/llama.cpp-strix-halo that referenced this pull request Aug 3, 2026
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants