[Bugfix] Pin EPLB and MLA host-to-device transfer buffers - #56138
Conversation
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: khluu <khluu000@gmail.com>
LucasWilkinson
left a comment
There was a problem hiding this comment.
Overall LGTM but left a comment that I think could be a worthwhile improvement
| if max(context_lens, default=0) <= 0: | ||
| return None | ||
| if PIN_MEMORY and not context_lens_cpu.is_pinned(): | ||
| context_lens_cpu = context_lens_cpu.pin_memory() |
There was a problem hiding this comment.
Could we pin this at the sparse MLA source instead?
context_lens_cpu = torch.empty(
num_prefills, dtype=seq_lens_cpu.dtype, pin_memory=PIN_MEMORY
)
torch.subtract(
seq_lens_cpu[num_decodes : num_decodes + num_prefills],
prefill_query_lens_cpu,
out=context_lens_cpu,
)Then this helper-level fallback should not be necessary for production callers.
There was a problem hiding this comment.
Addressed in a9f8512. Sparse MLA now allocates context_lens_cpu with pin_memory=PIN_MEMORY and computes directly into it with torch.subtract(out=...), matching dense MLA. Removed the helper fallback. Updated the CUDA regression to exercise the sparse caller with mixed decode/prefill inputs under strict sync checking and verify the transferred lengths. Pre-commit (including mypy) passes; exact-revision attention CI is running at https://buildkite.com/vllm/ci/builds/88262. The PR description distinguishes this pending run from the earlier GPU results.
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
|
/ci run |
|
✅ Triggered Buildkite CI #88264 for commit |
…ct#56138) Signed-off-by: khluu <khluu000@gmail.com> Signed-off-by: Kevin H. Luu <khluu000@gmail.com> Co-authored-by: OpenAI Codex <noreply@openai.com>
The gloo staging regression from vllm-project#56138 mocks eplb_comm.P2POp with a positional lambda. This PR's build_ops now passes group/group_peer as keywords, so the mock raises TypeError. Mirror the full P2POp signature (peer, group, tag, group_peer) without weakening the pinned-memory sync check. Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Main CI build 87842 fails both Qwen3 Sync EPLB accuracy jobs and the TP1/PCP4 evaluation when strict GPU synchronization checks detect nonblocking H2D copies from pageable CPU memory.
Pin Gloo receive staging buffers and allocate sparse MLA context lengths directly in pinned CPU memory with torch.subtract(out=...), matching the dense MLA caller. Preserve the strict checks. Add focused CUDA regressions that exercise the actual transfers and verify the copied values.
Duplicate check: searched open PRs for EPLB/pinned, unpinned, context_lens_cpu, and the failure signatures. No existing PR fixes these two transfers. PCP still separately needs #55879 for autotuning memory and #55499 for FlashInfer ragged-prefill synchronization; this PR does not duplicate them.
Current revision (
a9f8512c0c), addressing Lucas Wilkinson's review:Earlier-revision validation (does not validate the review update above):
pre-commit run --files vllm/distributed/eplb/eplb_communicator.py vllm/model_executor/layers/attention/mla_attention.py tests/distributed/test_eplb_execute.py tests/v1/attention/test_mla_context_chunks.py: passed.pytest -v -s v1/attention/test_mla_context_chunks.py— 13 passed;pytest -v -s distributed/test_eplb_execute.py— 24 passed, 7 skipped. All threetest_eplb_spec_decode.pycases skipped under their existing conditions.AI assistance: implemented and validated with OpenAI Codex. The human submitter confirmed review of these changes and authorized PR submission. Test results above identify the automated validation performed.