feat: add --kv-transfer-config NixlConnector to disagg scripts and recipes - #6560
Conversation
…cipes Signed-off-by: alec-flowers <aflowers@nvidia.com>
…sfer Signed-off-by: alec-flowers <aflowers@nvidia.com> # Conflicts: # examples/backends/vllm/launch/disagg_multimodal_epd.sh
Signed-off-by: alec-flowers <aflowers@nvidia.com>
WalkthroughAdds Changes
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
examples/backends/vllm/launch/disagg_multimodal_epd.sh (1)
69-79: Optional: deduplicate repeated JSON config literals.You can reduce drift risk by defining shared config variables once and reusing them across workers.
♻️ Proposed refactor
EXTRA_ARGS="" +KV_TRANSFER_CONFIG='{"kv_connector":"NixlConnector","kv_role":"kv_both"}' +KV_EVENTS_CONFIG_ENCODE='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20080"}' +KV_EVENTS_CONFIG_PREFILL='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20081"}' +KV_EVENTS_CONFIG_DECODE='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20082"}' @@ -VLLM_NIXL_SIDE_CHANNEL_PORT=20097 CUDA_VISIBLE_DEVICES=$DYN_ENCODE_WORKER_GPU python -m dynamo.vllm --multimodal-encode-worker --enable-multimodal --model $MODEL_NAME --gpu-memory-utilization $DYN_ENCODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20080"}' & +VLLM_NIXL_SIDE_CHANNEL_PORT=20097 CUDA_VISIBLE_DEVICES=$DYN_ENCODE_WORKER_GPU python -m dynamo.vllm --multimodal-encode-worker --enable-multimodal --model $MODEL_NAME --gpu-memory-utilization $DYN_ENCODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_ENCODE" & @@ -CUDA_VISIBLE_DEVICES=$DYN_PREFILL_WORKER_GPU python -m dynamo.vllm --multimodal-worker --route-to-encoder --disaggregation-mode prefill --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_PREFILL_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20081"}' & +CUDA_VISIBLE_DEVICES=$DYN_PREFILL_WORKER_GPU python -m dynamo.vllm --multimodal-worker --route-to-encoder --disaggregation-mode prefill --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_PREFILL_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_PREFILL" & @@ -CUDA_VISIBLE_DEVICES=$DYN_DECODE_WORKER_GPU python -m dynamo.vllm --multimodal-decode-worker --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_DECODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20082"}' & +CUDA_VISIBLE_DEVICES=$DYN_DECODE_WORKER_GPU python -m dynamo.vllm --multimodal-decode-worker --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_DECODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_DECODE" &🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@examples/backends/vllm/launch/disagg_multimodal_epd.sh` around lines 69 - 79, The three worker launch lines repeat the same JSON literals for --kv-transfer-config and similar --kv-events-config patterns; create reusable shell variables (e.g. KV_TRANSFER_CONFIG='--kv-transfer-config '\''{"kv_connector":"NixlConnector","kv_role":"kv_both"}'\'' ' and a base KV_EVENTS_CONFIG_PREFIX='--kv-events-config '\''{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:' and then build per-worker KV_EVENTS_CONFIGs by appending ports like 20080/20081/20082) and replace the inline JSON literals in the VLLM_NIXL_SIDE_CHANNEL_PORT/CUDA_VISIBLE_DEVICES ... python -m dynamo.vllm invocations with those variables so the same payload is defined once and reused for the multimodal-encode-worker, --multimodal-worker (prefill) and --multimodal-decode-worker lines.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Nitpick comments:
In `@examples/backends/vllm/launch/disagg_multimodal_epd.sh`:
- Around line 69-79: The three worker launch lines repeat the same JSON literals
for --kv-transfer-config and similar --kv-events-config patterns; create
reusable shell variables (e.g. KV_TRANSFER_CONFIG='--kv-transfer-config
'\''{"kv_connector":"NixlConnector","kv_role":"kv_both"}'\'' ' and a base
KV_EVENTS_CONFIG_PREFIX='--kv-events-config
'\''{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:' and then build
per-worker KV_EVENTS_CONFIGs by appending ports like 20080/20081/20082) and
replace the inline JSON literals in the
VLLM_NIXL_SIDE_CHANNEL_PORT/CUDA_VISIBLE_DEVICES ... python -m dynamo.vllm
invocations with those variables so the same payload is defined once and reused
for the multimodal-encode-worker, --multimodal-worker (prefill) and
--multimodal-decode-worker lines.
ℹ️ Review info
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (9)
examples/backends/vllm/launch/disagg.shexamples/backends/vllm/launch/disagg_multimodal_epd.shexamples/backends/vllm/launch/disagg_multimodal_llama.shexamples/backends/vllm/launch/disagg_router.shexamples/backends/vllm/launch/disagg_same_gpu.shrecipes/deepseek-r1/vllm/disagg/deploy_hopper_16gpu.yamlrecipes/llama-3-70b/vllm/disagg-multi-node/deploy.yamlrecipes/llama-3-70b/vllm/disagg-single-node/deploy.yamlrecipes/qwen3-32b/vllm/disagg-kv-router/deploy.yaml
…cipes (ai-dynamo#6560) Signed-off-by: alec-flowers <aflowers@nvidia.com>
Summary
--kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}'to all vLLM disaggregated serving launch scripts and recipes that were missing it--connectorin favor of the upstream vLLM--kv-transfer-configflagFiles changed
Launch scripts:
disagg.sh- decode + prefill workersdisagg_router.sh- 2 decode + 2 prefill workersdisagg_same_gpu.sh- decode + prefill workersdisagg_multimodal_epd.sh- encode + prefill + decode workersdisagg_multimodal_llama.sh- prefill + decode workersRecipes:
recipes/deepseek-r1/vllm/disagg/deploy_hopper_16gpu.yaml- decode + prefillrecipes/llama-3-70b/vllm/disagg-single-node/deploy.yaml- prefill + decoderecipes/llama-3-70b/vllm/disagg-multi-node/deploy.yaml- prefill + decoderecipes/qwen3-32b/vllm/disagg-kv-router/deploy.yaml- decode + prefillTest plan
--kv-transfer-config🤖 Generated with Claude Code
Summary by CodeRabbit