Skip to content

feat: add --kv-transfer-config NixlConnector to disagg scripts and recipes - #6560

Merged
alec-flowers merged 3 commits into
mainfrom
feat/add-nixl-kv-transfer
Feb 25, 2026
Merged

alec-flowers merged 3 commits into
mainfrom
feat/add-nixl-kv-transfer

Conversation

@alec-flowers

@alec-flowers alec-flowers commented Feb 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Adds explicit --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' to all vLLM disaggregated serving launch scripts and recipes that were missing it
  • This is needed as part of deprecating --connector in favor of the upstream vLLM --kv-transfer-config flag

Files changed

Launch scripts:

  • disagg.sh - decode + prefill workers
  • disagg_router.sh - 2 decode + 2 prefill workers
  • disagg_same_gpu.sh - decode + prefill workers
  • disagg_multimodal_epd.sh - encode + prefill + decode workers
  • disagg_multimodal_llama.sh - prefill + decode workers

Recipes:

  • recipes/deepseek-r1/vllm/disagg/deploy_hopper_16gpu.yaml - decode + prefill
  • recipes/llama-3-70b/vllm/disagg-single-node/deploy.yaml - prefill + decode
  • recipes/llama-3-70b/vllm/disagg-multi-node/deploy.yaml - prefill + decode
  • recipes/qwen3-32b/vllm/disagg-kv-router/deploy.yaml - decode + prefill

Test plan

  • Verify disagg scripts launch correctly with explicit --kv-transfer-config
  • Verify vLLM DSR1 recipe deploys successfully
  • Verify vLLM Llama 3 disagg recipes deploy successfully

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Introduced KV transfer configuration for disaggregated serving mode, enabling optimized cache handling across prefill and decode operations in multi-GPU and multi-node deployment setups.

…cipes

Signed-off-by: alec-flowers <aflowers@nvidia.com>
…sfer

Signed-off-by: alec-flowers <aflowers@nvidia.com>

# Conflicts:
#	examples/backends/vllm/launch/disagg_multimodal_epd.sh
@alec-flowers
alec-flowers requested a review from a team as a code owner February 25, 2026 01:49
@alec-flowers
alec-flowers requested a review from a team February 25, 2026 01:49
@alec-flowers
alec-flowers requested a review from a team as a code owner February 25, 2026 01:49
@alec-flowers
alec-flowers requested a review from a team February 25, 2026 01:49
@github-actions github-actions Bot added feat backend::vllm Relates to the vllm backend labels Feb 25, 2026
Signed-off-by: alec-flowers <aflowers@nvidia.com>
@coderabbitai

coderabbitai Bot commented Feb 25, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds --kv-transfer-config flag with NixlConnector configuration to vLLM disaggregation commands across multiple deployment scripts and recipes. The flag specifies {"kv_connector":"NixlConnector","kv_role":"kv_both"} for key-value transfer behavior in encode, prefill, and decode modes.

Changes

Cohort / File(s) Summary
Launch Scripts
examples/backends/vllm/launch/disagg.sh, disagg_multimodal_epd.sh, disagg_multimodal_llama.sh, disagg_router.sh, disagg_same_gpu.sh
Adds --kv-transfer-config flag with NixlConnector JSON payload to vLLM command invocations across disaggregation modes (encode, prefill, decode). Adjusts line continuations in some files to accommodate new argument placement.
Deployment Recipes
recipes/deepseek-r1/vllm/disagg/deploy_hopper_16gpu.yaml, recipes/llama-3-70b/vllm/disagg-multi-node/deploy.yaml, recipes/llama-3-70b/vllm/disagg-single-node/deploy.yaml, recipes/qwen3-32b/vllm/disagg-kv-router/deploy.yaml
Injects --kv-transfer-config with NixlConnector configuration into prefill and decode worker container args within deployment manifests.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Poem

🐰 Hoppy hops through launch scripts with glee,
Adding connectors for KV to flow free,
NixlConnector magic, both roles combined,
A transfer so smooth, perfectly aligned! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The PR title accurately and concisely describes the main change: adding --kv-transfer-config NixlConnector to disaggregation scripts and recipes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Description check ✅ Passed The PR description provides a clear summary, lists all changed files with context, explains the purpose of changes, and includes a test plan with verified checks.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
examples/backends/vllm/launch/disagg_multimodal_epd.sh (1)

69-79: Optional: deduplicate repeated JSON config literals.

You can reduce drift risk by defining shared config variables once and reusing them across workers.

♻️ Proposed refactor
 EXTRA_ARGS=""
+KV_TRANSFER_CONFIG='{"kv_connector":"NixlConnector","kv_role":"kv_both"}'
+KV_EVENTS_CONFIG_ENCODE='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20080"}'
+KV_EVENTS_CONFIG_PREFILL='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20081"}'
+KV_EVENTS_CONFIG_DECODE='{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20082"}'
@@
-VLLM_NIXL_SIDE_CHANNEL_PORT=20097 CUDA_VISIBLE_DEVICES=$DYN_ENCODE_WORKER_GPU python -m dynamo.vllm --multimodal-encode-worker --enable-multimodal --model $MODEL_NAME --gpu-memory-utilization $DYN_ENCODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20080"}' &
+VLLM_NIXL_SIDE_CHANNEL_PORT=20097 CUDA_VISIBLE_DEVICES=$DYN_ENCODE_WORKER_GPU python -m dynamo.vllm --multimodal-encode-worker --enable-multimodal --model $MODEL_NAME --gpu-memory-utilization $DYN_ENCODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_ENCODE" &
@@
-CUDA_VISIBLE_DEVICES=$DYN_PREFILL_WORKER_GPU python -m dynamo.vllm --multimodal-worker --route-to-encoder --disaggregation-mode prefill --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_PREFILL_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20081"}' &
+CUDA_VISIBLE_DEVICES=$DYN_PREFILL_WORKER_GPU python -m dynamo.vllm --multimodal-worker --route-to-encoder --disaggregation-mode prefill --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_PREFILL_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_PREFILL" &
@@
-CUDA_VISIBLE_DEVICES=$DYN_DECODE_WORKER_GPU python -m dynamo.vllm --multimodal-decode-worker --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_DECODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}' --kv-events-config '{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:20082"}' &
+CUDA_VISIBLE_DEVICES=$DYN_DECODE_WORKER_GPU python -m dynamo.vllm --multimodal-decode-worker --enable-multimodal --enable-mm-embeds --model $MODEL_NAME --gpu-memory-utilization $DYN_DECODE_GPU_MEM $EXTRA_ARGS --kv-transfer-config "$KV_TRANSFER_CONFIG" --kv-events-config "$KV_EVENTS_CONFIG_DECODE" &
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@examples/backends/vllm/launch/disagg_multimodal_epd.sh` around lines 69 - 79,
The three worker launch lines repeat the same JSON literals for
--kv-transfer-config and similar --kv-events-config patterns; create reusable
shell variables (e.g. KV_TRANSFER_CONFIG='--kv-transfer-config
'\''{"kv_connector":"NixlConnector","kv_role":"kv_both"}'\'' ' and a base
KV_EVENTS_CONFIG_PREFIX='--kv-events-config
'\''{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:' and then build
per-worker KV_EVENTS_CONFIGs by appending ports like 20080/20081/20082) and
replace the inline JSON literals in the
VLLM_NIXL_SIDE_CHANNEL_PORT/CUDA_VISIBLE_DEVICES ... python -m dynamo.vllm
invocations with those variables so the same payload is defined once and reused
for the multimodal-encode-worker, --multimodal-worker (prefill) and
--multimodal-decode-worker lines.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@examples/backends/vllm/launch/disagg_multimodal_epd.sh`:
- Around line 69-79: The three worker launch lines repeat the same JSON literals
for --kv-transfer-config and similar --kv-events-config patterns; create
reusable shell variables (e.g. KV_TRANSFER_CONFIG='--kv-transfer-config
'\''{"kv_connector":"NixlConnector","kv_role":"kv_both"}'\'' ' and a base
KV_EVENTS_CONFIG_PREFIX='--kv-events-config
'\''{"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:' and then build
per-worker KV_EVENTS_CONFIGs by appending ports like 20080/20081/20082) and
replace the inline JSON literals in the
VLLM_NIXL_SIDE_CHANNEL_PORT/CUDA_VISIBLE_DEVICES ... python -m dynamo.vllm
invocations with those variables so the same payload is defined once and reused
for the multimodal-encode-worker, --multimodal-worker (prefill) and
--multimodal-decode-worker lines.

ℹ️ Review info

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 5a4c96d and 5c87279.

📒 Files selected for processing (9)
  • examples/backends/vllm/launch/disagg.sh
  • examples/backends/vllm/launch/disagg_multimodal_epd.sh
  • examples/backends/vllm/launch/disagg_multimodal_llama.sh
  • examples/backends/vllm/launch/disagg_router.sh
  • examples/backends/vllm/launch/disagg_same_gpu.sh
  • recipes/deepseek-r1/vllm/disagg/deploy_hopper_16gpu.yaml
  • recipes/llama-3-70b/vllm/disagg-multi-node/deploy.yaml
  • recipes/llama-3-70b/vllm/disagg-single-node/deploy.yaml
  • recipes/qwen3-32b/vllm/disagg-kv-router/deploy.yaml

@biswapanda biswapanda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@alec-flowers
alec-flowers enabled auto-merge (squash) February 25, 2026 02:02
@alec-flowers
alec-flowers merged commit eac9432 into main Feb 25, 2026
59 of 60 checks passed
@alec-flowers
alec-flowers deleted the feat/add-nixl-kv-transfer branch February 25, 2026 02:26
yao531441 pushed a commit to yao531441/dynamo that referenced this pull request May 13, 2026
…cipes (ai-dynamo#6560)

Signed-off-by: alec-flowers <aflowers@nvidia.com>

This branch was previously deployed

1 inactive deployment
GITLAB — 82047662 Deployed Feb 25, 2026 by copy-pr-bot[bot] via Trigger CI Pipeline #18303
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend feat size/M

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants