Skip to content

[AMD][Bugfix] Fix vattn_asm HIP error 709 under CUDA graph capture on ROCm 10 - #39513

Merged
HaiShaw merged 5 commits into
sgl-project:mainfrom
chuyeh:amd/vattn-asm-rocm10-hip-lib
Sep 17, 2026
Merged

HaiShaw merged 5 commits into
sgl-project:mainfrom
chuyeh:amd/vattn-asm-rocm10-hip-lib

Conversation

@chuyeh

@chuyeh chuyeh commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

#37465 puts EAGLE verify / draft-extend / q_len-1 decode on the gfx950 vattn_asm kernel, launched through ctypes hipModuleLaunchKernel on the current torch stream. On ROCm 10 images, SGLANG_ASM_VERIFY_ATTN=1 fails during EAGLE verify HIP-graph capture:

RuntimeError: vattn_asm: hipModuleLaunchKernel failed: 709 (context is destroyed)

hipModuleLoad succeeded, so the kernel reports itself active, and the default stream (handle 0) still works. Only a torch stream fails — graph capture included, which is the path decode takes.

Root cause

ROCm 10 ships the HIP runtime (_rocm_sdk_core) and the toolchain (_rocm_sdk_devel) as separate wheels. Torch maps core's libamdhip64.so.7, which the unversioned CDLL("libamdhip64.so") does not match, so the loader takes devel's copy off LD_LIBRARY_PATH as a second HIP runtime: the module loads in devel while the stream belongs to core.

Unpatched, on MI355X:

Image torch HIP runtimes mapped CDLL("libamdhip64.so") binds Fix needed
rocm720 2.9.1 1 torch's, via the /opt/rocm symlink no
rocm724 2.11.0 1 torch's, torch/lib/libamdhip64.so no
rocm10 2.11.0 2 devel's, not torch's yes

So the trigger is the SDK split, not torch 2.11 — rocm724 has torch 2.11 and one runtime — and the change is a no-op wherever the soname already resolves to torch's object. Same failure mode #30870 fixed for CUDA, where TileLang's libcudart_stub.so could win over the real runtime.

Modifications

  • vattn_asm_gfx950/__init__.py: initialize CUDA so torch's HIP is mapped, then bind ctypes to it with the existing find_loaded_library() helper from Avoid TileLang CUDA runtime pollution #30870 rather than to the SONAME. Ten lines in one function; no new private API, no second /proc/self/maps parser.

No disambiguation between the two ROCm 10 copies is included, because none is needed: only torch's copy is ever mapped. After import torch, after CUDA init, and after importing aiter, unified_attention_3d_mtp and aiter_backend, /proc/self/maps holds exactly one libamdhip64:

maps after torch import: [.../_rocm_sdk_core/lib/libamdhip64.so.7]
maps after cuda init:    [.../_rocm_sdk_core/lib/libamdhip64.so.7]
maps after sglang+aiter: [.../_rocm_sdk_core/lib/libamdhip64.so.7]

The devel copy only enters the process if something explicitly dlopens it, which is precisely the bug being removed. Should a future image map both, find_loaded_library() is the single place to teach the preference, for CUDA and HIP at once.

Accuracy Tests

test/registered/amd/test_vattn_segplan_mi35x.py::TestVattnSegPlan::test_launches_on_a_side_stream_and_under_graph_capture (stage-b-test-1-gpu-small-amd-mi35x) launches the kernel on a torch side stream, then captures and replays it in a HIP graph, checking both against the file's existing fp32 reference. It lives with the other gfx950 asm kernel tests and reuses their fixtures.

Verified on MI355X (gfx950):

Image Unpatched This PR
v0.5.18-rocm10-mi35x-20260902 FAILED ... 709 (context is destroyed) 3 passed, 16 subtests passed
v0.5.19-rocm720-mi35x-20260913 passes 3 passed, 16 subtests passed
v0.5.18-rocm724-mi35x-20260825 passes 3 passed, 16 subtests passed

The ROCm 10 unpatched column is a negative control: reverting only _hip_lib() to CDLL("libamdhip64.so") makes the new test fail with exactly the reported 709, so the test guards the regression rather than merely passing. With the fix, ctypes binds /opt/venv/.../_rocm_sdk_core/lib/libamdhip64.so.7, the same runtime torch uses.

Speed Tests and Profiling

No speed change: one extra library-path lookup, once per process, at kernel load. This restores the assembly verify path under ROCm 10 HIP-graph capture; the kernel and launch geometry are untouched.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): 🚫 Run #35178338784
Latest PR Test (Extra): ❌ Run #35178338599
Latest PR Test (AMD ROCm 10): ❌ Run #35178338783

ROCm 10 wheels load HIP from _rocm_sdk_core while LD_LIBRARY_PATH points at
a second _rocm_sdk_devel copy. Opening the soname created a second HIP
runtime, so hipModuleLaunchKernel on a torch side stream or HIP graph
returned hipErrorContextIsDestroyed (709).

Co-authored-by: Cursor <cursoragent@cursor.com>
@chuyeh
chuyeh marked this pull request as ready for review September 15, 2026 04:09
@yichiche yichiche added the run-ci CI: run the baseline test suite on this PR label Sep 16, 2026
yichiche and others added 2 commits September 16, 2026 10:04
Reuse the helper sgl-project#30870 added for the same class of failure on CUDA instead
of parsing /proc/self/maps a second time, and drop the _rocm_sdk_core
preference: only torch's copy is ever mapped, so that branch never ran.

Move the regression test next to the other gfx950 asm kernel tests, where the
existing fixtures build the inputs, and revert the Triton wrapper's test file.

Co-authored-by: Cursor <cursoragent@cursor.com>
@chuyeh chuyeh changed the title [AMD] Bind vattn_asm ctypes launches to PyTorch's libamdhip64 [AMD] Bind vattn_asm ctypes launches to the libamdhip64 torch loaded Sep 16, 2026
The two copies do not differ by SONAME: torch maps _rocm_sdk_core's
libamdhip64.so.7, which an unversioned CDLL request never matches, so the
loader takes _rocm_sdk_devel's off LD_LIBRARY_PATH. That is also why the
single-runtime 7.2.0 and 7.2.4 images were never affected.

Co-authored-by: Cursor <cursoragent@cursor.com>
@yichiche

Copy link
Copy Markdown
Collaborator

@zijiecode Can you help review this PR?

@zijiecode

Copy link
Copy Markdown
Collaborator

Protection for ROCm 10, LGTM.

@yichiche yichiche left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed compatibility issue to support verify asm kernel on ROCm10. LGTM.

@chuyeh chuyeh changed the title [AMD] Bind vattn_asm ctypes launches to the libamdhip64 torch loaded [AMD][Bugfix] Fix vattn_asm HIP error 709 under CUDA graph capture on ROCm 10 Sep 17, 2026
@HaiShaw
HaiShaw merged commit 71ef869 into sgl-project:main Sep 17, 2026
229 of 259 checks passed
michaelzhang-ai added a commit that referenced this pull request Sep 17, 2026
#39513 fixed the ROCm 10 split-SDK failure by binding the HIP runtime already mapped by torch, but falling back to an unversioned SONAME can still map an unidentified second runtime when lookup is ambiguous.

Require exactly one mapped libamdhip64 path, bind that absolute path, and validate the binding before backend routing. This keeps GQA-8 requests on the existing fallback path if the runtime cannot be identified safely.

Add regression coverage for duplicate mappings and assert that the live launcher uses the process's sole mapped HIP runtime.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
michaelzhang-ai added a commit that referenced this pull request Sep 18, 2026
#39513 fixed the ROCm 10 split-SDK failure by binding the HIP runtime already mapped by torch, but falling back to an unversioned SONAME can still map an unidentified second runtime when lookup is ambiguous.

Require exactly one mapped libamdhip64 path, bind that absolute path, and validate the binding before backend routing. This keeps GQA-8 requests on the existing fallback path if the runtime cannot be identified safely.

Add regression coverage for duplicate mappings and assert that the live launcher uses the process's sole mapped HIP runtime.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
zijiecode added a commit to zijiecode/sglang that referenced this pull request Sep 20, 2026
…gl-project#39513)

ROCm 10 images ship a second libamdhip64 in _rocm_sdk_devel; the unversioned
CDLL("libamdhip64.so") picked it and every launch on a torch stream failed with
709 (context is destroyed). Initialize CUDA first and bind to torch's copy via
find_loaded_library(). Unit test passes on rocm720 and rocm10 MI35x images.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants