Skip to content

[PD] Validate Mooncake EFA allocator compatibility - #39973

Merged
ShangmingCai merged 3 commits into
sgl-project:mainfrom
nvpohanh:codex/mooncake-efa-allocator-validation
Sep 22, 2026
Merged

ShangmingCai merged 3 commits into
sgl-project:mainfrom
nvpohanh:codex/mooncake-efa-allocator-validation

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

[by Codex]

Motivation

Mooncake EFA transfers currently fail opaquely when their registered CUDA buffers come from a VMM-backed allocator. The current libfabric EFA provider applies pointer-level CU_POINTER_ATTRIBUTE_SYNC_MEMOPS to these buffers; CUDA VMM allocations do not support that attribute, so subsequent fi_read/fi_write calls fail with EINVAL.

The provider behavior and a sanitized two-node reproduction are tracked in ofiwg/libfabric#12853.

SGLANG_MOONCAKE_CUSTOM_MEM_POOL is intended for NVLink/MNNVL or BAREX, but it could previously be combined with MOONCAKE_PROTOCOL=efa. CUDA expandable segments can create the same incompatible VMM-backed transfer buffers.

Changes

  • Reject Mooncake custom memory pools when MOONCAKE_PROTOCOL=efa.
  • Reject expandable_segments:True from either PyTorch allocator environment variable under EFA.
  • Preserve the existing custom-pool behavior for non-EFA transports.
  • Add CPU unit coverage for supported custom pools, both allocator environment aliases, case-insensitive EFA selection, and the unaffected RDMA path.
  • Document how to replace SGLang's standard Mooncake CUDA 13 package with mooncake-transfer-engine-efa-cuda13.
  • Document the EFA wheel's host-runtime, architecture, and feature constraints, plus the current libfabric CUDA VMM limitation.

The validation runs before allocator initialization, replacing a later transfer-time failure with an actionable startup error. The error attributes the current limitation to the libfabric EFA provider rather than to EFA hardware.

Testing

  • Focused unit tests: 5 passed.
  • BLACK_NUM_WORKERS=1 SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failure.

CI States

Latest PR Test (Base): ✅ Run #35329691973
Latest PR Test (Extra): ❌ Run #35329691478
Latest PR Test (AMD ROCm 10): ❌ Run #35329691659

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 17, 2026

@ShangmingCai ShangmingCai left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Direction is right, the startup gate beats a transfer-time EINVAL. Three things before this lands.

if envs.MOONCAKE_PROTOCOL.get().lower() != "efa":
return

if enable_custom_mem_pool:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

enable_custom_mem_pool is True for INTRA_NODE_NVLINK, but that type never allocates a VMM pool: init_mooncake_custom_mem_pool bails out at L58 with (False, None, None), and disaggregation/utils.py:355 puts aux metadata on cuda through the default allocator precisely because custom_mem_pool is None. Those buffers are cudaMalloc-backed, EFA handles them fine.

We recommend that mode for A100/H20/H100 in pd_disaggregation.mdx:170 paired with MC_INTRANODE_NVLINK=true, which is exactly the p4d/p5 shape: NVLink inside the node, EFA across. So this turns a working config into a startup failure, with a reason that isn't true for it.

Gate on the type instead:

Suggested change
if enable_custom_mem_pool:
if enable_custom_mem_pool and custom_mem_pool_type in ("NVLINK", "BAREX"):

and drop INTRA_NODE_NVLINK from the loop in test_efa_rejects_custom_memory_pools.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] Addressed in f9b3372: EFA validation now rejects only NVLINK and BAREX, while INTRA_NODE_NVLINK remains allowed. The unit test and EFA documentation were updated accordingly.

enable_custom_mem_pool = False
custom_mem_pool_type = None

_validate_efa_allocator_compatibility(enable_custom_mem_pool, custom_mem_pool_type)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This only reaches the decode side when staging is on. conn.py:256 sits inside the PREFILL branch, and decode gets here only via _init_staging_allocator -> init_staging_allocator (staging_handler.py:756), with SGLANG_DISAGG_STAGING_BUFFER defaulting to False. The other decode path, maybe_init_custom_mem_pool, is gated on SGLANG_MOONCAKE_CUSTOM_MEM_POOL being set, which is the var EFA users are told here not to set.

So a default decode worker with expandable_segments:True gets no check, and the failure stays quiet: register_buffer_to_engine ignores the return of batch_register, which swallows the exception and logs at debug (mooncake_transfer_engine.py:166-181). That's the opaque fi_read/fi_write path this PR is meant to remove, still live on half the deployment.

Call it from somewhere both roles hit (CommonKVManager.__init__, or engine init) rather than from check_mooncake_custom_mem_pool_enabled.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] Addressed in f9b3372: allocator compatibility validation now runs at the start of MooncakeKVManager.init, which is shared by both prefill and decode roles, before engine initialization and buffer registration.

The current libfabric EFA provider cannot transfer CUDA VMM-backed buffers. Until [libfabric issue #12853](https://github.com/ofiwg/libfabric/issues/12853) is resolved, apply these settings to both prefill and decode workers:

- Do not set `SGLANG_MOONCAKE_CUSTOM_MEM_POOL`.
- Do not enable `expandable_segments:True` in `PYTORCH_CUDA_ALLOC_CONF` or `PYTORCH_ALLOC_CONF`.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Third VMM source that slips through: SGLANG_ENABLE_POST_CAPTURE_KV_SIZING=1 backs the KV cache with KvVmmBufferOwner (memory_pool.py:2364, cuMemCreate/cuMemMap). post_capture_kv_sizing_planned already opts out when SGLANG_MOONCAKE_CUSTOM_MEM_POOL is set (arg_groups/overrides.py:1822) but knows nothing about MOONCAKE_PROTOCOL, so EFA + post-capture sizing with neither flagged var set still registers VMM buffers and fails the same way.

Either add a bullet here, or add the protocol check to post_capture_kv_sizing_planned.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Codex] Addressed in f9b3372: post_capture_kv_sizing_planned now returns false for the Mooncake EFA backend, with unit-test coverage and documentation noting the automatic fallback.

@nvpohanh

Copy link
Copy Markdown
Collaborator Author

@ShangmingCai could you review again? Thanks!

@ShangmingCai ShangmingCai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ShangmingCai

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 21, 2026
@ShangmingCai

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@ShangmingCai
ShangmingCai merged commit 771c9d7 into sgl-project:main Sep 22, 2026
349 of 419 checks passed
iyastreb added a commit to iyastreb/sglang that referenced this pull request Sep 28, 2026
Conflicts:
- environ.py: keep the SGLANG_CUSTOM_MEM_POOL alias descriptor alongside
  upstream's new SGLANG_MOONCAKE_MAX_TRANSFER_BATCH_INDICES.
- test_mem_cache_utils.py: take upstream's removal of test_enabled_via_env
  (sgl-project#41297); only the rename in test_disabled_by_default remains.

Upstream code that still used the old name merged without textual
conflicts, so it is renamed here:
- test_mooncake_efa_allocator.py (sgl-project#39973) patched
  envs.SGLANG_MOONCAKE_CUSTOM_MEM_POOL, which no longer exists after the
  rename and would raise AttributeError.
- The EFA allocator error message and the EFA docs bullet now name
  SGLANG_CUSTOM_MEM_POOL.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants