Repository navigation
feat(vllm): add external encoder handoff - #13293
Conversation
This comment has been minimized.
This comment has been minimized.
6b0b89f to
bcd4d65
Compare
bcd4d65 to
0653bf1
Compare
0653bf1 to
51b27ed
Compare
f42d16b to
8e8332e
Compare
8e8332e to
24b3084
Compare
24b3084 to
c4dbac5
Compare
c4dbac5 to
7f5b03c
Compare
7f5b03c to
1a0f101
Compare
b64bc29 to
e656d71
Compare
Signed-off-by: furionw <qiwa@nvidia.com>
e656d71 to
01332e4
Compare
Signed-off-by: furionw <qiwa@nvidia.com>
Signed-off-by: furionw <qiwa@nvidia.com>
Signed-off-by: furionw <qiwa@nvidia.com>
Signed-off-by: furionw <qiwa@nvidia.com>
Signed-off-by: furionw <qiwa@nvidia.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. WalkthroughThe change adds an external vision encoder handoff for aggregated vLLM requests. It serializes image embeddings into ChangesExternal Encoder Handoff
Priority: ⬆️ High Estimated code review effort: 4 (Complex) | ~45 minutes Merge Risk: 🔵 Low · up to The external encoder flow lacks an end-to-end assertion for forwarding its reconstructed prompt to generation. Add that focused integration test before relying on this coverage for future changes. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: b99cd151-207c-4d05-b606-932f43329d98
📒 Files selected for processing (9)
components/src/dynamo/common/backend/engine.pycomponents/src/dynamo/vllm/handlers.pycomponents/src/dynamo/vllm/multimodal_utils/custom_encoder/__init__.pycomponents/src/dynamo/vllm/multimodal_utils/custom_encoder/adapter/linear.pycomponents/src/dynamo/vllm/multimodal_utils/custom_encoder/backend/base.pycomponents/src/dynamo/vllm/multimodal_utils/custom_encoder/external.pycomponents/src/dynamo/vllm/multimodal_utils/custom_encoder/handoff.pycomponents/src/dynamo/vllm/tests/multimodal_utils/custom_encoder/test_vllm_external_handoff.pycomponents/src/dynamo/vllm/tests/test_vllm_external_encoder.py
Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review.
Signed-off-by: furionw <qiwa@nvidia.com>
Why
Dynamo already maintains the
VisionEncoderBackendframework, but applications that place the encoder outside the aggregated vLLM worker still had to reimplement image extraction, encoder driving, tensor packing, request sanitization, and the decoder-side prompt reconstruction contract.This PR makes that handoff a vLLM-owned contract while keeping
GenerateRequest.encoder_resultopaque indynamo.common.User contract
Users implement
VisionEncoderBackend; Dynamo owns the rest of the external handoff.What changed
ExternalEncoderHandofflifecycle and request-preparation API.dynamo.vllm.Version 0 remains intentionally narrow: image inputs, ordered two-dimensional CPU linear embeddings, MsgPack request-plane transport, token-in/token-out, and a text-only aggregated vLLM worker with
--enable-prompt-embeds.The thin application-owned example is #14859.
Test plan
Full PR sequence
The three open PRs form GitHub stack #15572. The first two remain in completed stack #15197.
Summary by CodeRabbit