Repository navigation
refactor(multimodal): route RDMA through the shared tensor payload resolver - #1915
Conversation
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 24 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
…solver Make RDMA a first-class transport in the engine-neutral tensor payload path (alongside inline/SHM), decouple it from TokenSpeed, and hoist the Python puller out of the tokenspeed package. vLLM emit stays gated off until its puller lands. Rust: - `resolve_mm_tensor_payload` takes an optional `MmRdmaExport` and resolves RDMA -> SHM -> inline; on export failure the bytes fall through so the payload is never dropped. Add `MmTensorPayload::Remote` + `Remote` arms to the vLLM/TokenSpeed payload mappers. - The general TokenSpeed export moves into `into_proto` (random slot key, encoder_input only); EPD stages explicitly with `bootstrap_room` via `stage_tokenspeed_tensor_rdma`. Delete `try_export_nixl_remote` and `try_export_encoder_inputs_nixl_remote`. - `VllmMultimodalData` carries `rdma_enabled` (held `false`: vLLM cannot pull RDMA payloads yet). Python: - Move `tokenspeed/rdma_pixel.py` to `smg_grpc_servicer/mm_rdma.py`, engine-neutral: `gateway_agent_name` is a constructor param and `feature_from_remote` returns the tensor, with the SHM publish moved to the TokenSpeed caller (dropping the `tokenspeed.runtime` imports). Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
59c14a8 to
e5bc230
Compare
There was a problem hiding this comment.
Code Review
This pull request refactors the RDMA (NIXL) pixel-payload pulling and exporting mechanism to be more engine-neutral and centralized. Key changes include making the gateway agent name configurable in RdmaPixelPuller, removing SHM publishing logic from feature_from_remote (shifting it to the calling servicer), and integrating RDMA transport decisions directly into the model gateway's payload resolution flow (resolve_mm_tensor_payload). Additionally, support for staging multimodal tensors over RDMA has been structured for both TokenSpeed and vLLM. There are no review comments provided, so I have no further feedback to offer.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 59c14a8267
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
|
||
| return (tensor if copied else tensor.clone()), None | ||
| # Copy out of the landing slot before it returns to the free ring. | ||
| return tensor if copied else tensor.clone() |
There was a problem hiding this comment.
Avoid cloning RDMA tensors before SHM publishing
When EPD_PIXEL_SHM is left at its default true and the remote tensor already has the model dtype, this return clones the full RDMA landing buffer before TokenSpeedEncoderServicer._items_from_proto immediately hashes it and calls ShmTensorHandle.publish(feature). For the large encoder-input tensors that use RDMA, that adds an extra full-size copy and temporary allocation on every EPD encode item; keep the publish/hash inside the landing-slot lifetime (for example via a caller-supplied materializer) so the SHM path can copy directly from the landing view.
Useful? React with 👍 / 👎.
Description
Problem
After #1908 extracted the NIXL mechanics into
smg-mm-rdma, the RDMA transport was still bolted onto the TokenSpeed path: export ran viaTokenSpeedTensor::try_export_nixl_remote/try_export_encoder_inputs_nixl_remote, the engine-neutral resolverresolve_mm_tensor_payloadhad noRemotevariant, and the Python puller lived in thetokenspeed/package importing the TokenSpeed runtime. That coupling blocks RDMA from becoming a first-class transport (like/dev/shm) and blocks vLLM support.Solution
Move the RDMA decision into the shared, engine-neutral payload resolver and hoist the Python puller into an engine-neutral module, so RDMA sits alongside inline/SHM for every backend. vLLM emit stays gated off until its puller + capability gate land (next PR); a
remotepayload sent to a pre-support vLLM worker would hard-fail, so the guard is load-bearing.Changes
Rust
resolve_mm_tensor_payloadtakes an optionalMmRdmaExportand resolves RDMA → SHM → inline; on an exportErr(bytes)the bytes fall through so the payload is never dropped. AddedMmTensorPayload::Remote+Remotearms to the vLLM/TokenSpeed payload mappers.into_proto(fresh random slot key,encoder_inputonly — never the small side tensors). EPD stages explicitly with the load-bearingbootstrap_roomviastage_tokenspeed_tensor_rdma, then callsinto_proto(false).try_export_nixl_remoteandtry_export_encoder_inputs_nixl_remote.VllmMultimodalDatacarriesrdma_enabled(heldfalse; vLLM cannot pull RDMA payloads yet).Python
tokenspeed/rdma_pixel.py→smg_grpc_servicer/mm_rdma.py, engine-neutral:gateway_agent_nameis a constructor param andfeature_from_remotereturns the feature tensor, with thehash_feature+ShmTensorHandle.publishmoved into the TokenSpeed caller (dropping bothtokenspeed.runtimeimports from the puller). Updated the encoder + scheduler servicers.Test Plan
Behavior-preserving; the
SMGRDMA1descriptor wire format is unchanged.cargo check -p smg(default) and--features mm-rdmaboth compile.cargo test -p smg --lib routers::grpc::proto_wrapper— payload tests pass, incl. the remote-encoder-input round-trip.cargo clippyclean:-p smgdefault and--features mm-rdma -- -D warnings.cargo +nightly fmt --all --checkclean;codespellclean.py_compile+ruff check+ruff format --checkclean on the moved module and both servicers.vLLM RDMA is wired but emits nothing (
rdma_enabled=false); the live NIXL/vLLM validation lands with the emit + capability gate in the next PR.Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspasses