Skip to content

[ROCm][Attention] Add an mxfp4_mla KV cache for rope-free sparse MLA (read/write path) - #60321

Open
amd-dlimpus wants to merge 8 commits into
vllm-project:mainfrom
amd-dlimpus:dlimpus/mxfp4-mla-kv-rw
Open

amd-dlimpus wants to merge 8 commits into
vllm-project:mainfrom
amd-dlimpus:dlimpus/mxfp4-mla-kv-rw

Conversation

@amd-dlimpus

@amd-dlimpus amd-dlimpus commented Oct 7, 2026 •

Copy link
Copy Markdown

Overview

Adds an mxfp4_mla KV cache dtype for the ROCm AITER sparse MLA backend on rope-free (NoPE) latents such as GLM-5.3-Flash. Each 512-wide latent is stored as 256 bytes of E2M1 codes plus 16 inline E8M0 scales (one per 32 values): 272 bytes per token against 1024 for bf16, or 3.76× more tokens per GiB of KV cache.

This is the first of the smaller PRs split out of #59739, as requested there. It contains only the cache format, the write path and a Triton read path. Everything is Triton or Python: no HIP or C++ sources, and no new environment variables.

Claims

  • KV capacity: 3.76× more tokens per GiB than bf16 (272 vs 1024 bytes per token).
  • Correctness: attention over the packed cache is bit-identical to attention over the same dequantized values stored as bf16.
  • No change for other KV dtypes. The only shared-kernel change is a branch on the cache pointer's element type, which Triton resolves at compile time, so the bf16 and FP8 specializations compile to the same code as before.
  • No speed claim. The Triton read path unpacks to bf16 before both dots, so it doesn't make attention faster: at 16k tokens it takes about 1.13× as long as bf16 per layer. The gain is capacity. Faster arithmetic over this cache comes in a follow-up.

Validation

Unit tests (MI355X / gfx950):

pytest tests/v1/attention/test_mxfp4_mla*.py tests/v1/attention/test_rocm_glm5next_sparse.py
# 99 passed

The new tests cover:

  • the format math, checked bit-for-bit against AITER's per_1x32_f4_quant;
  • the store kernel, including padding slots and launches past 2^31 source elements;
  • the Triton read path, checked against a dequantize-then-attend reference, with both the hardware and software unpack;
  • the startup plumbing (KV spec bytes, hybrid block sizing, config validation);
  • the routing that sends both prefill and decode through the ragged kernel.

Accuracy (GLM-5.3-Flash, TP=4 on MI355X, seed 42; bf16 and MXFP4 served from the same image):

Benchmark bf16 mxfp4_mla, this PR
GSM8K (5-shot, 1319 questions, greedy) 97.19% (1282) 96.89% (1278)
GPQA-Diamond (198 questions, temperature 1.0, top_p 0.95), with #59412 90.91% (180) 89.90% (178)

Neither gap is significant. On GSM8K the two arms disagree on 12 of 1319 questions (4 right only with MXFP4, 8 right only with bf16; exact McNemar p = 0.39), and a repeat of the same bf16 configuration has differed by 0.76 points. On GPQA-Diamond they disagree on 16 of 198 (7 vs 9; p = 0.80), and both stay in the 88.9–91.9% range of earlier bf16 and MXFP4 runs without the indexer regression. For reference, the native FP8 × FP4 kernel over this same cache format scored 97.42% vs 96.97% for bf16 on GSM8K in #59739.

Long-context accuracy needs the indexer fix in #59412. On ROCm, main currently has a sparse-indexer page-addressing bug on GLM-5.3-Flash (#58858) that corrupts top-k selection beyond 2048 tokens, independently of the KV cache dtype. Without a fix, GPQA-Diamond fell to 80.8–85.4% for bf16 and MXFP4 alike across 8 runs on main ed3f6d1, with 13–18 of 198 generations running to the token limit. With #59412 applied, the runs above score 90.91% (bf16) and 89.90% (MXFP4), with 5 and 3 generations at the limit. GSM8K generations mostly stay below 2048 tokens and aren't affected. The GPQA row above is therefore measured with #59412 applied to both arms; this PR does not include or depend on it.

Details

Format and write path. The mxfp4_mla row is 256 bytes of packed E2M1 (low nibble holds the lower index) followed by 16 E8M0 scales. Scales use the round-up exponent, so no value clamps. MLAAttention.get_kv_cache_spec reports the 272-byte row through state_content_bytes, and Platform._align_hybrid_block_size uses the same size so the mamba state still fits one attention page on hybrid models. concat_and_cache_mla can't express per-group scales (its scale is a single float), so a Triton store kernel behind do_kv_cache_update quantizes and writes the row.

Read path. The ragged Triton sparse-attention kernel, which already serves both prefill and decode for rope-free models, gains a branch for a uint8 (packed MXFP4) cache, with no new kernel parameters. It unpacks each gathered tile to bf16 once and feeds both existing tl.dot calls:

  • On gfx950 the unpack uses the hardware FP4 converter (v_cvt_scalef32_pk_bf16_fp4) via inline asm.
  • On other GPUs it uses a software unpack. The two are tested bit-identical on gfx950, but I haven't run the software path on gfx942.

The decode-only kernels used by DeepSeek-V4 have no MXFP4 branch, so they refuse a packed cache rather than misread it.

Follow-ups (stacked on this PR)

Separately, two decode-overhead trims that apply to every KV dtype (host-side paged_kv_indptr for pure-decode steps, and skipping the NoPE q_concat copy) will go up as their own PR against main.


Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

New KV cache dtype `mxfp4_mla` for the ROCm AITER sparse MLA backend on
rope-free (NoPE) latents such as GLM-5.3-Flash: each 512-wide latent is stored
as 256 bytes of E2M1 codes plus 16 inline E8M0 scales per group of 32 (272
bytes per token vs 1024 for bf16, 3.76x more tokens per GiB).

- Write: dtype registration, 272-byte state_content_bytes, and a Triton
  store kernel behind MLAAttentionImpl.do_kv_cache_update.
- Read: the ragged Triton sparse-attention kernel unpacks MXFP4 rows in
  registers (KV_IS_MXFP4), with the gfx950 FP4 converter when available and
  a software unpack elsewhere. The decode-only kernels, which have no MXFP4
  branch, refuse a packed cache instead of misreading it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the rocm Related to AMD ROCm label Oct 7, 2026
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

Comment thread vllm/utils/torch_utils.py Outdated
"ultraquant_4bit": torch.uint8,
"nvfp4": torch.uint8,
"nvfp4_4over6": torch.uint8,
# MXFP4 MLA: 272-byte rows (256 packed E2M1 + 16 inline E8M0 scales) for a

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't need a comment on this one. Let the commit message handle it

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed in 8116f75.

if cache.dtype != torch.uint8 or cache.ndim == 0:
return
if cache.shape[-1] == row_bytes(nope_head_dim):
raise NotImplementedError(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lets cut down on the NotImplementedError here. Try to be as concise as you can

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the guard entirely in 8116f75. rocm_sparse_attn_decode is only reached from DeepSeek-V4, whose backend doesn't accept mxfp4_mla, so it could never fire.

other=0.0,
)
if KV_IS_MXFP4:
# Packed MXFP4: the row is KV_ROW_BYTES of uint8, not head_dim

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

less here if you don't mind

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cut to one line in 8116f75.

scale,
HAS_ATTN_SINK: tl.constexpr,
OUT_DV: tl.constexpr,
KV_IS_MXFP4: tl.constexpr,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you check if we can calculate this inside the kernel instead of having to add 5 new params? Or if we can't lets just create something like mxfp4_kv_metadata or some such, so we reduce the amount of additional variables

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8116f75, with no new params. The kernel checks kv_ptr.dtype.element_ty == tl.uint8 at compile time, and load_mxfp4_rows picks the gfx950 hardware unpack or the software unpack internally. The row pitch comes from the existing kv_stride_n.

@AndreasKaratzas AndreasKaratzas added the verified Run pre-commit for new contributors without triggering other tests label Oct 7, 2026
amd-dlimpus and others added 2 commits October 7, 2026 15:00
Drop the five MXFP4 constexprs from _sparse_attn_prefill_ragged_kernel.
The kernel now branches on kv_ptr's uint8 element type, and
load_mxfp4_rows picks the gfx950 hardware or software unpack internally,
taking the row pitch from kv_stride_n.

Remove the decode-path guard: rocm_sparse_attn_decode is only reached
from DeepSeek-V4, whose backend does not accept mxfp4_mla. Trim comments.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>
The kernel picks hardware unpack on gfx950 at compile time, so force the
software path in a fresh process and check it against the dequantized
bf16 cache.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>

@simondanielsson simondanielsson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the work! I think we can likely reconcile a lot of this work with the ultraquant development and general mxfp4 support we already have

Suggestion: Can we please check in both the existing mxfp4_utils and all of the ultraquant files for things we can re-use? Many of these ops already exist, see for instance _unpack_nibbles_last_dim

Suggestion: Can we gather some perf numbers to get a sense of the baseline performance of this here? On both gfx950/942 to see the effect of the software convert.

Comment thread vllm/v1/attention/backend.py Outdated
return
from vllm import _custom_ops as ops

if kv_cache_dtype == "mxfp4_mla":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: Can we move this rocm_aiter_sparse_mla instead (i.e. override do_kv_cache_update)?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. ROCMAiterMLASparseImpl.do_kv_cache_update owns the mxfp4 store. backend.py do_kv_cache_update is the shared concat_and_cache_mla path again.

mask=valid[:, None] & dim_mask[None, :],
other=0.0,
)
if kv_ptr.dtype.element_ty == tl.uint8:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: This looks a bit brittle, perhaps we can instead pass in the kv cache dtype into the kernel wrapper and add a IS_MXFP4 constexpr we can specialize on?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. The wrapper takes the kv cache dtype and passes one IS_MXFP4 constexpr. Both the ragged prefill kernel and the split-K partial kernel branch on IS_MXFP4, not on tl.uint8 (ragged in e130cd6, split-K in 908b18d).

q.shape[1],
).reshape(-1)
# A packed MXFP4 row is a byte pitch, not head_dim elements.
from vllm.v1.attention.ops.mxfp4_mla import (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: can we put this in the top of the file?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. row_bytes is imported at the top of backends/mla/rocm_aiter_mla_sparse.py.

Comment thread vllm/v1/attention/ops/mxfp4_mla_read.py Outdated
def _e2m1_codes_to_f32(codes):
"""4-bit E2M1 codes (as integers) -> signed fp32 values.

Deliberately transcendental-free. An earlier version used ``tl.exp2`` for

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: We can remove the reference to the "earlier version" of this 👍

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. The earlier-version / tl.exp2 history is gone from that docstring.

if _HW_UNPACK:
words = tl.load(
cache_ptr.to(tl.pointer_type(tl.int32))
+ base // 4

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: Should we assert somewhere that the row pitch is divisible by 4? Otherwise I think this will give the wrong value

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The pitch check is at startup, in ROCMAiterMLASparseImpl.__init__ (e130cd6): mxfp4_mla is rejected unless kv_lora_rank is a power of two and at least 128. row_bytes(rank) = rank/2 + rank/32 is then a multiple of 4, which is what the gfx950 path needs when it divides the pitch by 4 for the int32 load. No per-launch assert.

state_content_bytes={
"fp8_ds_mla": 656,
"nvfp4_ds_mla": 352,
"mxfp4_mla": self.head_size // 2 + self.head_size // 32,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: can we use row_bytes() here?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. Both mla_attention.py (get_kv_cache_spec) and platforms/interface.py (_align_hybrid_block_size) call row_bytes() for the mxfp4 page size.

``*scale``, so a single per-tensor scale is the only thing it can express, and
MXFP4 needs one E8M0 byte per group of 32.

Layout written (see :mod:`mxfp4_mla`): one 272-byte row per slot, 256 bytes of

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: can we trim these module docstrings a bit? Also, this specific one I think is just an examples as 272 row pitch will only be how head_dim=512, but models might have a different head dim than that

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trimmed in e130cd6. 0c705b9 states row_bytes(d) == d // 2 + d // 32 in the mxfp4_mla, store, and read module docstrings. 272 is only the head_dim=512 example on row_bytes().

_M5 = tl.constexpr(3.5)
_M6 = tl.constexpr(5.0)

assert _E2M1_MAX.value == E2M1_MAX

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: Could these asserts be in a test instead?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moved to test_triton_constexprs_match_reference in 0c705b9. The constexprs are still built from E2M1_MAX, E8M0_BIAS (127), and the midpoints of E2M1_VALUES; the test checks those against the reference. It does not need a GPU.

Comment thread vllm/v1/attention/ops/mxfp4_mla.py Outdated
# Midpoints between consecutive magnitudes, for round-to-nearest via bucketize.
_E2M1_MIDPOINTS: tuple[float, ...] = (0.25, 0.75, 1.25, 1.75, 2.5, 3.5, 5.0)

GROUP_SIZE = 32

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: Could we take this etc from ocp_mx_utils.OCP_MX_BLOCK_SIZE?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. GROUP_SIZE = ocp_mx_utils.OCP_MX_BLOCK_SIZE.

Comment thread vllm/platforms/interface.py Outdated
# Must match MLAAttention.get_kv_cache_spec, else the mamba
# state no longer fits the real (smaller) MLA page.
state_content_bytes=(
model_config.get_head_size() // 2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here, can we re-use the utility we have?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e130cd6. platforms/interface.py and mla_attention.py both call row_bytes() for the mxfp4 page size, instead of inlining head_size / 2 + head_size / 32.

@mergify

mergify Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @amd-dlimpus.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 8, 2026
amd-dlimpus and others added 2 commits October 8, 2026 13:14
…v-rw

Signed-off-by: Limpus, David <dlimpus@amd.com>

# Conflicts:
#	vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py
- Select the MXFP4 read with one IS_MXFP4 constexpr, set from the
  kv_cache_dtype passed to rocm_sparse_attn_prefill.
- Reject mxfp4_mla at init unless kv_lora_rank is a power of two >= 128,
  which the int32 row loads and the single-tile read require.
- Move the store dispatch into ROCMAiterMLASparseImpl.do_kv_cache_update.
- Keep mxfp4_mla decode off the bf16 split-K kernel.
- Reuse OCP_MX_BLOCK_SIZE, ultraquant's nibble helpers and row_bytes();
  derive the store kernel constants from the reference.
- Trim module docstrings.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>
@mergify

mergify Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Hi @amd-dlimpus, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

@mergify mergify Bot removed the needs-rebase label Oct 8, 2026
amd-dlimpus and others added 2 commits October 8, 2026 14:59
…v-rw

Signed-off-by: Limpus, David <dlimpus@amd.com>
The split-K partial kernel takes the same IS_MXFP4 path as the ragged kernel, and the backend no longer keeps a packed cache off that kernel. Annotate the routing-test capture so mypy accepts it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>
@amd-dlimpus

amd-dlimpus commented Oct 8, 2026 •

Copy link
Copy Markdown
Author

This replaces the earlier gfx942 kernel note with one end-to-end serving comparison: gfx942 (MI300X), one run, GLM-5.3-Flash, tensor parallel 4, 2048 input tokens, 128 output tokens, 16 prompts, concurrency 8, all 16 requests completed in both arms; both arms decode with the split-K kernel that landed in main (#58584, Simon Danielsson), the MXFP4 arm that same kernel reading the packed mxfp4_mla cache and the bf16 arm that kernel reading a bf16 cache.

Output tokens/s Time to first token Time per output token
Split-K bf16 348 1.18 s 13.9 ms
Split-K MXFP4 360 1.04 s 14.2 ms

Note: The MI300X (gfx942) does not have a hardware instruction that unpacks FP4 values. The MI355X (gfx950) has this instruction (v_cvt_scalef32_pk_bf16_fp4). On gfx942, the split-K kernel unpacks the MXFP4 cache in software. This software unpack adds more instructions to the kernel. Because of this, the time per output token does not decrease on gfx942.

@amd-dlimpus

amd-dlimpus commented Oct 8, 2026 •

Copy link
Copy Markdown
Author

This replaces the earlier gfx950 kernel note with one end-to-end serving comparison: gfx950 (MI355X), one run, GLM-5.3-Flash, tensor parallel 4, 2048 input tokens, 128 output tokens, 16 prompts, concurrency 8, all 16 requests completed in both arms; both arms decode with the split-K kernel that landed in main (#58584, Simon Danielsson), the MXFP4 arm that same kernel reading the packed mxfp4_mla cache and the bf16 arm that kernel reading a bf16 cache.

Output tokens/s Time to first token Time per output token
Split-K bf16 196 3.51 s 13.9 ms
Split-K MXFP4 299 1.75 s 10.6 ms

Simon asked to move the import-time constant asserts into a test. The constexprs stay derived from the reference; this test checks E2M1_MAX, the E8M0 bias, and the midpoints. The store and read module docs state the row-pitch formula.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Limpus, David <dlimpus@amd.com>
@amd-dlimpus

Copy link
Copy Markdown
Author

Reused ultraquant's _pack_nibbles_last_dim / _unpack_nibbles_last_dim for nibble pack and unpack, and GROUP_SIZE is ocp_mx_utils.OCP_MX_BLOCK_SIZE. The encoder and the in-kernel unpack stay ours: ultraquant uses a different scale rule (c=0.156 and Hadamard, zero-group scale byte 0 vs our 127), and its Triton dequant uses tl.exp2 into a full buffer, which measured 1.8x slower than bf16 inside the attention loop. mxfp4_utils dequant needs amd-quark and is PyTorch, so it cannot run in the kernel.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rocm Related to AMD ROCm verified Run pre-commit for new contributors without triggering other tests

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

4 participants