Skip to content

refactor(multimodal): route RDMA through the shared tensor payload resolver - #1915

Merged
slin1237 merged 1 commit into
mainfrom
refactor/mm-rdma-payload-path
Jul 13, 2026
Merged

slin1237 merged 1 commit into
mainfrom
refactor/mm-rdma-payload-path

Conversation

@slin1237

Copy link
Copy Markdown
Member

Description

Problem

After #1908 extracted the NIXL mechanics into smg-mm-rdma, the RDMA transport was still bolted onto the TokenSpeed path: export ran via TokenSpeedTensor::try_export_nixl_remote / try_export_encoder_inputs_nixl_remote, the engine-neutral resolver resolve_mm_tensor_payload had no Remote variant, and the Python puller lived in the tokenspeed/ package importing the TokenSpeed runtime. That coupling blocks RDMA from becoming a first-class transport (like /dev/shm) and blocks vLLM support.

Solution

Move the RDMA decision into the shared, engine-neutral payload resolver and hoist the Python puller into an engine-neutral module, so RDMA sits alongside inline/SHM for every backend. vLLM emit stays gated off until its puller + capability gate land (next PR); a remote payload sent to a pre-support vLLM worker would hard-fail, so the guard is load-bearing.

Changes

Rust

  • resolve_mm_tensor_payload takes an optional MmRdmaExport and resolves RDMA → SHM → inline; on an export Err(bytes) the bytes fall through so the payload is never dropped. Added MmTensorPayload::Remote + Remote arms to the vLLM/TokenSpeed payload mappers.
  • The general TokenSpeed export moves into into_proto (fresh random slot key, encoder_input only — never the small side tensors). EPD stages explicitly with the load-bearing bootstrap_room via stage_tokenspeed_tensor_rdma, then calls into_proto(false).
  • Deleted try_export_nixl_remote and try_export_encoder_inputs_nixl_remote.
  • VllmMultimodalData carries rdma_enabled (held false; vLLM cannot pull RDMA payloads yet).

Python

  • Moved tokenspeed/rdma_pixel.py → smg_grpc_servicer/mm_rdma.py, engine-neutral: gateway_agent_name is a constructor param and feature_from_remote returns the feature tensor, with the hash_feature + ShmTensorHandle.publish moved into the TokenSpeed caller (dropping both tokenspeed.runtime imports from the puller). Updated the encoder + scheduler servicers.

Test Plan

Behavior-preserving; the SMGRDMA1 descriptor wire format is unchanged.

  • cargo check -p smg (default) and --features mm-rdma both compile.
  • cargo test -p smg --lib routers::grpc::proto_wrapper — payload tests pass, incl. the remote-encoder-input round-trip.
  • cargo clippy clean: -p smg default and --features mm-rdma -- -D warnings.
  • cargo +nightly fmt --all --check clean; codespell clean.
  • Python: py_compile + ruff check + ruff format --check clean on the moved module and both servicers.

vLLM RDMA is wired but emits nothing (rdma_enabled=false); the live NIXL/vLLM validation lands with the emit + capability gate in the next PR.

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Jul 13, 2026
@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 24 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: fcfd5e48-58f7-47e2-989d-2d1d27a84b4d

📥 Commits

Reviewing files that changed from the base of the PR and between 1d33392 and e5bc230.

📒 Files selected for processing (7)
  • grpc_servicer/smg_grpc_servicer/mm_rdma.py
  • grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py
  • grpc_servicer/smg_grpc_servicer/tokenspeed/servicer.py
  • model_gateway/src/routers/grpc/client.rs
  • model_gateway/src/routers/grpc/epd_encode.rs
  • model_gateway/src/routers/grpc/multimodal/assemble.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch refactor/mm-rdma-payload-path

Comment @coderabbitai help to get the list of available commands.

…solver

Make RDMA a first-class transport in the engine-neutral tensor payload path
(alongside inline/SHM), decouple it from TokenSpeed, and hoist the Python
puller out of the tokenspeed package. vLLM emit stays gated off until its
puller lands.

Rust:
- `resolve_mm_tensor_payload` takes an optional `MmRdmaExport` and resolves
  RDMA -> SHM -> inline; on export failure the bytes fall through so the
  payload is never dropped. Add `MmTensorPayload::Remote` + `Remote` arms to
  the vLLM/TokenSpeed payload mappers.
- The general TokenSpeed export moves into `into_proto` (random slot key,
  encoder_input only); EPD stages explicitly with `bootstrap_room` via
  `stage_tokenspeed_tensor_rdma`. Delete `try_export_nixl_remote` and
  `try_export_encoder_inputs_nixl_remote`.
- `VllmMultimodalData` carries `rdma_enabled` (held `false`: vLLM cannot pull
  RDMA payloads yet).

Python:
- Move `tokenspeed/rdma_pixel.py` to `smg_grpc_servicer/mm_rdma.py`,
  engine-neutral: `gateway_agent_name` is a constructor param and
  `feature_from_remote` returns the tensor, with the SHM publish moved to the
  TokenSpeed caller (dropping the `tokenspeed.runtime` imports).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the refactor/mm-rdma-payload-path branch from 59c14a8 to e5bc230 Compare July 13, 2026 16:13

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the RDMA (NIXL) pixel-payload pulling and exporting mechanism to be more engine-neutral and centralized. Key changes include making the gateway agent name configurable in RdmaPixelPuller, removing SHM publishing logic from feature_from_remote (shifting it to the calling servicer), and integrating RDMA transport decisions directly into the model gateway's payload resolution flow (resolve_mm_tensor_payload). Additionally, support for staging multimodal tensors over RDMA has been structured for both TokenSpeed and vLLM. There are no review comments provided, so I have no further feedback to offer.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 59c14a8267

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".


return (tensor if copied else tensor.clone()), None
# Copy out of the landing slot before it returns to the free ring.
return tensor if copied else tensor.clone()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid cloning RDMA tensors before SHM publishing

When EPD_PIXEL_SHM is left at its default true and the remote tensor already has the model dtype, this return clones the full RDMA landing buffer before TokenSpeedEncoderServicer._items_from_proto immediately hashes it and calls ShmTensorHandle.publish(feature). For the large encoder-input tensors that use RDMA, that adds an extra full-size copy and temporary allocation on every EPD encode item; keep the publish/hash inside the landing-slot lifetime (for example via a caller-supplied materializer) so the SHM path can copy directly from the landing view.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit e8a9034 into main Jul 13, 2026
81 of 83 checks passed
@slin1237
slin1237 deleted the refactor/mm-rdma-payload-path branch July 13, 2026 20:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant