Skip to content

[Bugfix] Pin EPLB and MLA host-to-device transfer buffers - #56138

Merged
khluu merged 4 commits into
vllm-project:mainfrom
khluu:codex/fix-ci-unpinned-transfers
Sep 11, 2026
Merged

khluu merged 4 commits into
vllm-project:mainfrom
khluu:codex/fix-ci-unpinned-transfers

Conversation

@khluu

@khluu khluu commented Sep 9, 2026 •

Copy link
Copy Markdown
Member

Main CI build 87842 fails both Qwen3 Sync EPLB accuracy jobs and the TP1/PCP4 evaluation when strict GPU synchronization checks detect nonblocking H2D copies from pageable CPU memory.

Pin Gloo receive staging buffers and allocate sparse MLA context lengths directly in pinned CPU memory with torch.subtract(out=...), matching the dense MLA caller. Preserve the strict checks. Add focused CUDA regressions that exercise the actual transfers and verify the copied values.

Duplicate check: searched open PRs for EPLB/pinned, unpinned, context_lens_cpu, and the failure signatures. No existing PR fixes these two transfers. PCP still separately needs #55879 for autotuning memory and #55499 for FlashInfer ragged-prefill synchronization; this PR does not duplicate them.

Current revision (a9f8512c0c), addressing Lucas Wilkinson's review:

  • Moved pinning to the sparse MLA source and removed the helper-level fallback.
  • Updated the CUDA regression to exercise the sparse caller with a decode followed by two prefills under strict GPU sync checking.
  • Pre-commit checks passed for all three revised files (including mypy). Local runtime tests could not run: this checkout's available local environment has no torch/CUDA.
  • Exact-revision attention CI passed: https://buildkite.com/vllm/ci/builds/88262 — image build and both H100 V1 Attention shards passed at a9f8512.

Earlier-revision validation (does not validate the review update above):

AI assistance: implemented and validated with OpenAI Codex. The human submitter confirmed review of these changes and authorized PR submission. Test results above identify the automated validation performed.

khluu and others added 3 commits September 9, 2026 03:09
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: khluu <khluu000@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the bug Something isn't working label Sep 9, 2026

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM but left a comment that I think could be a worthwhile improvement

if max(context_lens, default=0) <= 0:
return None
if PIN_MEMORY and not context_lens_cpu.is_pinned():
context_lens_cpu = context_lens_cpu.pin_memory()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we pin this at the sparse MLA source instead?

context_lens_cpu = torch.empty(
    num_prefills, dtype=seq_lens_cpu.dtype, pin_memory=PIN_MEMORY
)
torch.subtract(
    seq_lens_cpu[num_decodes : num_decodes + num_prefills],
    prefill_query_lens_cpu,
    out=context_lens_cpu,
)

Then this helper-level fallback should not be necessary for production callers.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in a9f8512. Sparse MLA now allocates context_lens_cpu with pin_memory=PIN_MEMORY and computes directly into it with torch.subtract(out=...), matching dense MLA. Removed the helper fallback. Updated the CUDA regression to exercise the sparse caller with mixed decode/prefill inputs under strict sync checking and verify the transferred lengths. Pre-commit (including mypy) passes; exact-revision attention CI is running at https://buildkite.com/vllm/ci/builds/88262. The PR description distinguishes this pending run from the earlier GPU results.

Co-authored-by: OpenAI Codex <noreply@openai.com>

Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
@khluu

khluu commented Sep 11, 2026

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #88264 for commit a9f8512c0cb6.

@khluu
khluu merged commit 1cf6555 into vllm-project:main Sep 11, 2026
150 of 159 checks passed
ItsRoy69 pushed a commit to ItsRoy69/vllm that referenced this pull request Sep 15, 2026
…ct#56138)

Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
LostFox11 pushed a commit to LostFox11/vllm that referenced this pull request Sep 24, 2026
The gloo staging regression from vllm-project#56138 mocks eplb_comm.P2POp with a
positional lambda. This PR's build_ops now passes group/group_peer as
keywords, so the mock raises TypeError. Mirror the full P2POp signature
(peer, group, tag, group_peer) without weakening the pinned-memory sync
check.

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants