Skip to content

[II] Bind lazy B12X DCP pools to graph channels - #337

Closed
voipmonitor wants to merge 1 commit into
agent/ii-b12x-graph-pool-capturefrom
agent/ii-b12x-lazy-graph-pools
Closed

voipmonitor wants to merge 1 commit into
agent/ii-b12x-graph-pool-capturefrom
agent/ii-b12x-lazy-graph-pools

Conversation

@voipmonitor

Copy link
Copy Markdown

Status

Qualified.

Behavior

B12X DCP pools initialized after a CUDA graph context opens join that graph's semantic channel before their first operation. LSE reduce-scatter and query-head gather select the channel associated with their process group; eager execution continues to use vllm:eager:dcp.

Technical reason

Projection geometry can create a B12X pool after the graph-owner context has already entered. Binding only the pools present at context entry leaves such a pool on the eager channel, which violates B12X graph ownership and prevents safe graph replay.

Compatibility

The change is generic to B12X DCP collectives and does not alter model geometry, supported world sizes, tensor formats, numerical operations, or NCCL fallback behavior. It depends on the graph-owner integration in #324.

Validation

  • Ruff formatting and checks pass.
  • git diff --check passes.
  • 37 relevant tests from tests/distributed/test_dcp_a2a.py pass in the CUDA 13.3/PyTorch 2.13 Kimi-K3 image.
  • Four focused tests cover pool creation inside graph context, active channel selection for LSE reduction and query gather, existing-pool capture, and eager-channel preservation.

The remaining full-file test categories are outside this change: packed-LSE dtype expectations, NCCL tests requiring a larger Docker shared-memory allocation, and a test for the removed B12X _load_extension symbol.

Record each active B12X DCP graph channel by process-group identity. A pool initialized after the graph context opens now prepares and enters that channel before its first operation, and collective calls select the active channel instead of the eager scheduler channel.

Eager execution is unchanged. The change depends on the graph-owner integration in vLLM #324.

Validation: Ruff format/check and git diff --check pass. Thirty-seven relevant DCP tests pass in the CUDA 13.3/PyTorch 2.13 Kimi-K3 image; four focused eager/graph channel tests pass.
@coderabbitai

coderabbitai Bot commented Aug 15, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5ef7b4b4-a68f-4d14-b084-d6089958a42d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@voipmonitor

Copy link
Copy Markdown
Author

The behavior implemented by this pull request is preserved unchanged in vLLM #384 as commit 11e9d87433eb. Stable Git patch IDs match. Review and merge #384; this pull request is closed to avoid duplicate review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant