Skip to content

[ROCm] Enable shfl xor sync style reduce on amd - #3701

Closed
yiakwy-xpu-ml-framework-team wants to merge 11 commits into
sgl-project:mainfrom
yiakwy-xpu-ml-framework-team:enable_shfl_xor_sync_style_reduce_on_amd
Closed

yiakwy-xpu-ml-framework-team wants to merge 11 commits into
sgl-project:mainfrom
yiakwy-xpu-ml-framework-team:enable_shfl_xor_sync_style_reduce_on_amd

Conversation

@yiakwy-xpu-ml-framework-team

@yiakwy-xpu-ml-framework-team yiakwy-xpu-ml-framework-team commented Feb 19, 2025 •

Copy link
Copy Markdown
Contributor

Motivation

This is follow up of PR#3664

Modifications

  • enable shlf_xor_sync
  • enable flashinfer vec_t (should be removed once flashinfer-rocm [POC] fully finished)

Checklist

thanhhao98 pushed a commit to thanhhao98/sglang that referenced this pull request Sep 7, 2026
init_fi_a2a_workspace() rejected every device whose
is_mnnvl_fabric_supported() is False, i.e. every Blackwell box without an
IMEX/NVL72 fabric (x86 B200/B300), even though FlashInfer can serve that case:
since 0.6.16 (flashinfer sgl-project#3701) MnnvlMemory exports the workspace with POSIX
file-descriptor handles and exchanges them over AF_UNIX SCM_RIGHTS, which needs
no CAP_SYS_PTRACE. sglang pins FlashInfer 0.6.18, so the only remaining
requirement is that the DCP group is contained in one node.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants