Repository navigation
[PD] Validate Mooncake EFA allocator compatibility - #39973
ShangmingCai merged 3 commits into
Conversation
| if envs.MOONCAKE_PROTOCOL.get().lower() != "efa": | ||
| return | ||
|
|
||
| if enable_custom_mem_pool: |
There was a problem hiding this comment.
enable_custom_mem_pool is True for INTRA_NODE_NVLINK, but that type never allocates a VMM pool: init_mooncake_custom_mem_pool bails out at L58 with (False, None, None), and disaggregation/utils.py:355 puts aux metadata on cuda through the default allocator precisely because custom_mem_pool is None. Those buffers are cudaMalloc-backed, EFA handles them fine.
We recommend that mode for A100/H20/H100 in pd_disaggregation.mdx:170 paired with MC_INTRANODE_NVLINK=true, which is exactly the p4d/p5 shape: NVLink inside the node, EFA across. So this turns a working config into a startup failure, with a reason that isn't true for it.
Gate on the type instead:
| if enable_custom_mem_pool: | |
| if enable_custom_mem_pool and custom_mem_pool_type in ("NVLINK", "BAREX"): |
and drop INTRA_NODE_NVLINK from the loop in test_efa_rejects_custom_memory_pools.
There was a problem hiding this comment.
[by Codex] Addressed in f9b3372: EFA validation now rejects only NVLINK and BAREX, while INTRA_NODE_NVLINK remains allowed. The unit test and EFA documentation were updated accordingly.
| enable_custom_mem_pool = False | ||
| custom_mem_pool_type = None | ||
|
|
||
| _validate_efa_allocator_compatibility(enable_custom_mem_pool, custom_mem_pool_type) |
There was a problem hiding this comment.
This only reaches the decode side when staging is on. conn.py:256 sits inside the PREFILL branch, and decode gets here only via _init_staging_allocator -> init_staging_allocator (staging_handler.py:756), with SGLANG_DISAGG_STAGING_BUFFER defaulting to False. The other decode path, maybe_init_custom_mem_pool, is gated on SGLANG_MOONCAKE_CUSTOM_MEM_POOL being set, which is the var EFA users are told here not to set.
So a default decode worker with expandable_segments:True gets no check, and the failure stays quiet: register_buffer_to_engine ignores the return of batch_register, which swallows the exception and logs at debug (mooncake_transfer_engine.py:166-181). That's the opaque fi_read/fi_write path this PR is meant to remove, still live on half the deployment.
Call it from somewhere both roles hit (CommonKVManager.__init__, or engine init) rather than from check_mooncake_custom_mem_pool_enabled.
There was a problem hiding this comment.
[by Codex] Addressed in f9b3372: allocator compatibility validation now runs at the start of MooncakeKVManager.init, which is shared by both prefill and decode roles, before engine initialization and buffer registration.
| The current libfabric EFA provider cannot transfer CUDA VMM-backed buffers. Until [libfabric issue #12853](https://github.com/ofiwg/libfabric/issues/12853) is resolved, apply these settings to both prefill and decode workers: | ||
|
|
||
| - Do not set `SGLANG_MOONCAKE_CUSTOM_MEM_POOL`. | ||
| - Do not enable `expandable_segments:True` in `PYTORCH_CUDA_ALLOC_CONF` or `PYTORCH_ALLOC_CONF`. |
There was a problem hiding this comment.
Third VMM source that slips through: SGLANG_ENABLE_POST_CAPTURE_KV_SIZING=1 backs the KV cache with KvVmmBufferOwner (memory_pool.py:2364, cuMemCreate/cuMemMap). post_capture_kv_sizing_planned already opts out when SGLANG_MOONCAKE_CUSTOM_MEM_POOL is set (arg_groups/overrides.py:1822) but knows nothing about MOONCAKE_PROTOCOL, so EFA + post-capture sizing with neither flagged var set still registers VMM buffers and fails the same way.
Either add a bullet here, or add the protocol check to post_capture_kv_sizing_planned.
There was a problem hiding this comment.
[by Codex] Addressed in f9b3372: post_capture_kv_sizing_planned now returns false for the Mooncake EFA backend, with unit-test coverage and documentation noting the automatic fallback.
|
@ShangmingCai could you review again? Thanks! |
|
/tag-and-rerun-ci |
|
/tag-and-rerun-ci |
Conflicts: - environ.py: keep the SGLANG_CUSTOM_MEM_POOL alias descriptor alongside upstream's new SGLANG_MOONCAKE_MAX_TRANSFER_BATCH_INDICES. - test_mem_cache_utils.py: take upstream's removal of test_enabled_via_env (sgl-project#41297); only the rename in test_disabled_by_default remains. Upstream code that still used the old name merged without textual conflicts, so it is renamed here: - test_mooncake_efa_allocator.py (sgl-project#39973) patched envs.SGLANG_MOONCAKE_CUSTOM_MEM_POOL, which no longer exists after the rename and would raise AttributeError. - The EFA allocator error message and the EFA docs bullet now name SGLANG_CUSTOM_MEM_POOL. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[by Codex]
Motivation
Mooncake EFA transfers currently fail opaquely when their registered CUDA buffers come from a VMM-backed allocator. The current libfabric EFA provider applies pointer-level
CU_POINTER_ATTRIBUTE_SYNC_MEMOPSto these buffers; CUDA VMM allocations do not support that attribute, so subsequentfi_read/fi_writecalls fail withEINVAL.The provider behavior and a sanitized two-node reproduction are tracked in ofiwg/libfabric#12853.
SGLANG_MOONCAKE_CUSTOM_MEM_POOLis intended for NVLink/MNNVL or BAREX, but it could previously be combined withMOONCAKE_PROTOCOL=efa. CUDA expandable segments can create the same incompatible VMM-backed transfer buffers.Changes
MOONCAKE_PROTOCOL=efa.expandable_segments:Truefrom either PyTorch allocator environment variable under EFA.mooncake-transfer-engine-efa-cuda13.The validation runs before allocator initialization, replacing a later transfer-time failure with an actionable startup error. The error attributes the current limitation to the libfabric EFA provider rather than to EFA hardware.
Testing
BLACK_NUM_WORKERS=1 SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failure.CI States
Latest PR Test (Base): ✅ Run #35329691973
Latest PR Test (Extra): ❌ Run #35329691478
Latest PR Test (AMD ROCm 10): ❌ Run #35329691659