Skip to content

[ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode - #58584

Merged
dllehr-amd merged 20 commits into
vllm-project:mainfrom
simondanielsson:feat/glm53flash-sparse-decode
Oct 8, 2026
Merged

dllehr-amd merged 20 commits into
vllm-project:mainfrom
simondanielsson:feat/glm53flash-sparse-decode

Conversation

@simondanielsson

@simondanielsson simondanielsson commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Decode rows in the bf16 sparse MLA path go through the ragged prefill kernel, whose grid is (q_len, heads_blocks). GLM-5.3-Flash at TP8 has one head block, so during decode with concurrency 4 launches only 4 workgroups on 256 CUs. Clearly underutilized.

This new kernel instead uses splitK+reduce over the context dim, i.e. with a grid (q_len, num_kv_splits, heads_blocks). DSv4 has a similar one.

Enabled if topk>=512, as microbenchmarks suggest (GLM uses topk=2048).

Perf: e2e -5-15% ITL. Microbenchmarks: 6x speedup on topk=2048

Test Result

Tested on MI350 unless mentioned otherwise.

VLLM_ROCM_USE_AITER=1 \
VLLM_USE_BREAKABLE_CUDAGRAPH=1 \
vllm serve zai-org/GLM-5.3-Flash --tensor-parallel-size 8 --max-num-seqs 512 --gpu-memory-utilization 0.8 --attention-backend ROCM_AITER_MLA_SPARSE --max-num-batched-tokens 1638

gsm8k

MI350

This branch:

Tasks Version Filter n-shot Metric Value Stderr
gsm8k 3 flexible-extract 5 exact_match ↑ 0.9196 ± 0.0075
strict-match 5 exact_match ↑ 0.9174 ± 0.0076

Main:

Tasks Version Filter n-shot Metric Value Stderr
gsm8k 3 flexible-extract 5 exact_match ↑ 0.9219 ± 0.0074
strict-match 5 exact_match ↑ 0.9196 ± 0.0075

MI300

This branch:

Tasks Version Filter n-shot Metric Value Stderr
gsm8k 3 flexible-extract 5 exact_match ↑ 0.9181 ± 0.0076
strict-match 5 exact_match ↑ 0.9174 ± 0.0076

Main:

Tasks Version Filter n-shot Metric Value Stderr
gsm8k 3 flexible-extract 5 exact_match ↑ 0.9136 ± 0.0077
strict-match 5 exact_match ↑ 0.9136 ± 0.0077

gpqa

Note no thinking mode enabled. I think low score on main is related to #59413.

lm_eval \
  --model local-completions \
  --model_args model=zai-org/GLM-5.3-Flash,base_url=http://localhost:8000/v1/completions,tokenized_requests=False,trust_remote_code=True,num_concurrent=32 \
  --tasks gpqa_diamond_zeroshot   \
  --seed 1234 --output_path /tmp/lm_eval_gsm8k

This branch

Tasks Version Filter n-shot Metric Value Stderr
gpqa_diamond_zeroshot 2.2 none 0 acc ↑ 0.4798 ± 0.0356
none 0 acc_norm ↑ 0.4798 ± 0.0356

Main

Tasks Version Filter n-shot Metric Value Stderr
gpqa_diamond_zeroshot 2.2 none 0 acc ↑ 0.4747 ± 0.0356
none 0 acc_norm ↑ 0.4747 ± 0.0356

e2e

MI350

ISL / OSL Conc Metric Nightly This branch Delta
1K / 1K 32 TTFT (ms) 561 560 -0.18%
1K / 1K 32 TPOT (ms) 14.35 12.88 -10.24%
1K / 1K 32 ITL (ms) 14.06 12.66 -9.96%
1K / 1K 32 Output tput (tok/s) 2083.73 2328.81 +11.76%
1K / 1K 128 TTFT (ms) 842 843 +0.12%
1K / 1K 128 TPOT (ms) 23.39 22.39 -4.28%
1K / 1K 128 ITL (ms) 20.18 19.05 -5.60%
1K / 1K 128 Output tput (tok/s) 5345.28 5639.88 +5.51%
64K / 1K 4 TTFT (ms) 2391 2403 +0.50%
64K / 1K 4 TPOT (ms) 15.57 13.32 -14.45%
64K / 1K 4 ITL (ms) 11.89 9.92 -16.57%
64K / 1K 4 Output tput (tok/s) 223.33 249.90 +11.90%
256K / 1K 4 TTFT (ms) 9423 9448 +0.27%
256K / 1K 4 TPOT (ms) 30.84 27.98 -9.27%
256K / 1K 4 ITL (ms) 12.38 10.36 -16.32%
256K / 1K 4 Output tput (tok/s) 99.81 107.59 +7.79%
512K / 1K 4 TTFT (ms) 22477 22836 +1.60%
512K / 1K 4 TPOT (ms) 59.46 57.13 -3.92%
512K / 1K 4 ITL (ms) 12.78 10.85 -15.10%
512K / 1K 4 Output tput (tok/s) 49.29 50.40 +2.25%

MI300, limited bench just for completeness. TPOT improves

8k/1k, conc 32

This branch

============ Serving Benchmark Result ============
Successful requests:                     320
Failed requests:                         0
Maximum request concurrency:             32
Benchmark duration (s):                  266.96
Total input tokens:                      2621440
Total generated tokens:                  327680
Request throughput (req/s):              1.20
Output token throughput (tok/s):         1227.47
Peak output token throughput (tok/s):    1984.00
Peak concurrent requests:                40.00
Total token throughput (tok/s):          11047.20
---------------Time to First Token----------------
Mean TTFT (ms):                          2428.67
Median TTFT (ms):                        2408.58
P99 TTFT (ms):                           9288.33
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          23.71
Median TPOT (ms):                        23.90
P99 TPOT (ms):                           25.81
---------------Inter-token Latency----------------
Mean ITL (ms):                           23.71
Median ITL (ms):                         16.77
P99 ITL (ms):                            598.27
==================================================

Main

============ Serving Benchmark Result ============
Successful requests:                     320
Failed requests:                         0
Maximum request concurrency:             32
Benchmark duration (s):                  291.54
Total input tokens:                      2621440
Total generated tokens:                  327680
Request throughput (req/s):              1.10
Output token throughput (tok/s):         1123.95
Peak output token throughput (tok/s):    1728.00
Peak concurrent requests:                41.00
Total token throughput (tok/s):          10115.55
---------------Time to First Token----------------
Mean TTFT (ms):                          2428.37
Median TTFT (ms):                        2396.20
P99 TTFT (ms):                           9227.89
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          26.11
Median TPOT (ms):                        26.40
P99 TPOT (ms):                           28.18
---------------Inter-token Latency----------------
Mean ITL (ms):                           26.11
Median ITL (ms):                         19.21
P99 ITL (ms):                            596.43
==================================================

Microbenchmark, MI350

python3 benchmarks/kernels/benchmark_sparse_mla_decode_bf16.py --rows 1 2 4 8 16 64 128 256 --lens 128 512 2048 --splits auto

topk 128. Here spitK loses, hence we don't use it in this domain

rows auto splits ragged us split-K us delta us
1 1 23.96 32.78 +8.82
2 1 24.22 32.81 +8.58
4 1 24.22 32.95 +8.73
8 1 23.86 33.06 +9.20
16 1 24.30 32.70 +8.39
64 1 24.79 32.96 +8.17
128 1 25.57 33.32 +7.75
256 1 26.66 30.15 +3.49

topk 512, splitk wins. Hence _SPARSE_DECODE_BF16_MIN_SPLIT_LEN = 512

rows auto splits ragged us split-K us speedup
1 16 61.09 32.62 1.87x
2 16 60.93 32.87 1.85x
4 16 61.15 32.85 1.86x
8 16 61.37 32.82 1.87x
16 16 62.04 32.86 1.89x
64 8 63.15 28.93 2.18x
128 4 63.55 31.70 2.00x
256 2 62.52 44.61 1.40x

topk 2048 (GLM-5.3-Flash)

rows auto splits ragged us split-K us speedup
1 32 210.00 33.05 6.35x
2 32 210.49 33.00 6.38x
4 32 211.41 32.81 6.44x
8 32 212.84 32.72 6.51x
16 16 213.46 30.15 7.08x
64 8 207.53 44.48 4.67x
128 4 201.11 68.77 2.92x
256 2 200.91 116.05 1.73x

Comparing this vs VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA added in #53492

vllm bench serve \
  --backend vllm --model zai-org/GLM-5.3-Flash  \
  --dataset_name random --random_input_len 64000 --random_output_len 1024 \
  --max_concurrency 4 --num_prompts 40 --num_warmups 8 \
  --host localhost --port 8000 --ignore_eos --seed 56552121

Testing this branch vs Oct 6 nightly vllm/vllm-openai-rocm:nightly-bb87d227d4b964abb2a966cdf9194f3d376d9bbe.

64K/1K, conc 4:

  • this branch: 2632ms TTFT, 11.92ms TPOT, 8.68ms ITL
  • VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 2291ms TTFT, 11.47ms TPOT, 8.64ms ITL

1k/1k. conc 32:

  • this branch: 713ms TTFT, 11.48ms TPOT, 11.40ms ITL
  • VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 666ms TTFT, 11.38ms TPOT, 11.26ms ITL

=> very similar performance, whereas the added splitk kernel is enabled by default and the aiter kernel is opt-in at the moment.

Traces

Before: 191us
image

After: 17us
image


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

@mergify mergify Bot added glm rocm Related to AMD ROCm labels Sep 24, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 24, 2026
@simondanielsson simondanielsson changed the title [ROCm][Perf][GLM-5.3-Flash] BF16 splitk sparse decode [ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode Sep 25, 2026
@simondanielsson simondanielsson changed the title [ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode [ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode (10x kernel improvement) Sep 25, 2026
@simondanielsson simondanielsson changed the title [ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode (10x kernel improvement) [ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode Sep 25, 2026
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
@simondanielsson
simondanielsson force-pushed the feat/glm53flash-sparse-decode branch from f813949 to 50a3ade Compare September 25, 2026 09:19
@mergify mergify Bot added the performance Performance-related issues label Sep 25, 2026
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
simondanielsson and others added 2 commits September 25, 2026 17:06
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
@simondanielsson
simondanielsson marked this pull request as ready for review September 25, 2026 15:08

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@simondanielsson

Copy link
Copy Markdown
Contributor Author

@claude review

@dllehr-amd
dllehr-amd self-requested a review September 30, 2026 14:53
Comment thread tests/kernels/attention/test_rocm_triton_attn_dsv4.py Outdated
sparse_len,
32,
)
if num_splits == 1:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you see if we can check num_splits in the original decode check to avoid having this function call appear twice in the same file?

@dllehr-amd

dllehr-amd commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Can you also double check with expert parallelism/dcp etc. Just to be sure we aren't exposed anywhere

@mergify mergify Bot added the needs-rebase label Oct 2, 2026
…rse-decode

Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
def test_rocm_sparse_triton_decode_routes_on_num_splits(
monkeypatch, num_splits, num_decode_tokens, expected
):
"""Split-K needs both a pure-decode batch and a heuristic asking for splits."""

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please let me know if you think these pure "is the kernel we expect to be called really called" tests are not useful, happy to remove to trim

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @simondanielsson, I just took another look at this, this test does seem close to trivial. If there's no good way to bolster its coverage, we can probably remove it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for having a look, I'll remove it

@dllehr-amd

Copy link
Copy Markdown
Contributor

/ci run

@dllehr-amd dllehr-amd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. Thanks @simondanielsson

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93362 for commit c882afc7e6c1.

simondanielsson and others added 2 commits October 7, 2026 12:05
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
@simondanielsson

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93367 for commit a175d19039ed.

@vllm-agent

Copy link
Copy Markdown
Contributor

CI selector (shadow): 15 test steps (26 jobs) instead of 75 (99 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 75 (99) 15 (26) 66 (82) 6 (9)
AMD mirrors 72 (94) 160 (224) 0 (0) 88 (130)
Selector would run (15)
  • amd-fp8-moe-kernels-mi355
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • amd-native-quantization-kernels-mi355
  • amd-qwen3-next-mtp-async-eplb-accuracy
  • cpu-kernel-tests ×2
  • deepseek-v4-kernel-test-h100
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd
  • kernels-attention-test ×7
  • kernels-b200 ×3
  • kernels-root-misc-test-b200
  • v1-attention-b200 ×2
  • v1-attention-h100-mi300 ×2
  • v1-executor-worker
Would skip (today's rules run them) (66)
  • ascend-npu-test
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multimodal-config
  • cpu-params-env-tokenizers-parser
  • cpu-reasoning-renderers
  • cpu-tool-parsers
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • metrics-tracing-2-gpus
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • regression
  • rl-entrypoints-tests
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-speculators
  • v1-core
  • v1-kv-connectors ×4
  • v1-kv-offload
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
Would add (today's rules do not run them) (6)
  • amd-fp8-moe-kernels-mi355 (code map)
  • amd-kernels-mi355 (code map)
  • amd-native-quantization-kernels-mi355 (code map)
  • cpu-kernel-tests ×2 (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • kernels-b200 ×3 (code map)
AMD mirrors: would skip (0)

none

AMD mirrors: would add (88)
  • basic-models-tests-extra-initialization ×14 (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • cudagraph (code map)
  • deepseek-v4-kernel-test-b200 (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • distributed-comm-ops (code map)
  • distributed-compile-comm-4-gpus (code map)
  • distributed-compile-rpc-tests-2-gpus (code map)
  • distributed-compile-unit-tests-2xh100 (code map)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-tests-8xh100 (code map)
  • distributed-torchrun-examples-4-gpus (code map)
  • distributed-torchrun-shutdown-tests-2-gpus (code map)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • engine (code map)
  • engine-1-gpu (code map)
  • entrypoints-unit-tests (code map)
  • eplb-algorithm (code map)
  • eplb-execution (code map)
  • examples (code map)
  • extract-hidden-states-integration (code map)
  • fault-tolerance-e2e-2xh100 (code map)
  • fusion-e2e-config-sweep-h100 (code map)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (code map)
  • gemm-rs-ar-2xb200 (code map)
  • glm5next-unit-tests (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (code map)
  • inkling-unit-tests-b200 (code map)
  • jit-monitor-no-runtime-jit (code map)
  • kernels-b200 ×3 (code map)
  • kernels-core-operation-test ×3 (code map)
  • kernels-flashmla-test-h100 (code map)
  • kernels-fp4-moe-test-b200 (code map)
  • kernels-fusedmoe-layer-test-2-b200s (code map)
  • kernels-fusedmoe-layer-test-2-h100s (code map)
  • kernels-helion-test ×5 (code map)
  • kernels-mamba-test (code map)
  • kernels-minimax-reduce-rms-test-2-gpus (code map)
  • kernels-moe-test ×5 (code map)
  • kernels-quantization-test ×6 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (code map)
  • kv-offload-medium (code map)
  • kv-offload-small (code map)
  • lm-eval-dspark-watermark-2xh100 (code map)
  • lm-eval-turboquant-k3v4nc (code map)
  • lm-eval-turboquant-k8v4 (code map)
  • lm-eval-turboquant-t3nc (code map)
  • lm-eval-turboquant-t4nc (code map)
  • lm-eval-watermarking (code map)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (code map)
  • model-executor (code map)
  • mooncake-ec-tcp-e2e-2-gpus (code map)
  • multi-modal-accuracy-eval-small-models (code map)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (code map)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (code map)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (code map)
  • pytorch-nightly-dependency-override-check (code map)
  • quantization ×4 (code map)
  • quantized-fusions (code map)
  • quantized-models-test (code map)
  • qwen4-exp-unit-tests (code map)
  • ray-dependency-compatibility-check (code map)
  • rayexecutorv2-4-gpus (code map)
  • rust-frontend-core-correctness (code map)
  • rust-frontend-distributed (code map)
  • rust-frontend-openai-coverage (code map)
  • rust-frontend-serve-admin-coverage (code map)
  • rust-frontend-tool-use (code map)
  • samplers-multimodal-beam-search (code map)
  • samplers-test (code map)
  • scale-out-ec-e2e-2-gpus (code map)
  • sharded-rdt-weight-transfer (code map)
  • spec-decode-draft-model ×4 (code map)
  • spec-decode-eagle-1-deepseek-qwen (code map)
  • spec-decode-eagle-2-llama3-qwen-vl-other (code map)
  • spec-decode-ngram-suffix (code map)
  • torch-stable-abi-audit (code map)
  • vllm-ir-tests (code map)

4 changed files · base a8260cbc5c · head a175d19039 · Python record: build 93039 at 1e5d0ea888 · kernel record: table 43b4aae (build 93244), map 43b4aae · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 65 optional steps the selector would also run

@simondanielsson

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93432 for commit d5de4fe57ee1.

@dllehr-amd
dllehr-amd merged commit 87d9996 into vllm-project:main Oct 8, 2026
193 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

glm performance Performance-related issues ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants