Skip to content

feat(vllm): add Python RL serving parity for generate endpoint - #15179

Merged
biswapanda merged 18 commits into
ai-dynamo:mainfrom
biswapanda:feat/dynamo-vllm-rl-parity
Sep 28, 2026
Merged

biswapanda merged 18 commits into
ai-dynamo:mainfrom
biswapanda:feat/dynamo-vllm-rl-parity

Conversation

@biswapanda

@biswapanda biswapanda commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This brings the Python dynamo.vllm worker to parity with Dynamo's RL path for aggregated vLLM deployments. It adapts Rust-frontend token-in/token-out Generate envelopes, preserves trainer replay metadata, publishes the topology and capabilities needed by RL orchestration, and serializes standard vLLM LoRA updates with active generation.

  • adapt native Generate/TITO requests, including sampling controls, cache salts, and preprocessed image features
  • fold engine streams into unary Generate responses while preserving routed-expert tensors and sampling masks
  • expose RL administration routes, NCCL world size, native Generate support, and tower-LoRA capability metadata
  • reject unsupported data-parallel and external-launcher RL topologies instead of publishing an incorrect receiver count
  • keep versioned standard vLLM LoRA adapters loaded for the full request lifetime, draining active generation before hot-swap or unload
  • add focused Python and Rust coverage for request adaptation, discovery metadata, replay metadata, topology validation, and LoRA lifecycle behavior

area breakdown by code owner

Changed area in PR CODEOWNER team
components/src/dynamo/common/rl/admin.py @ai-dynamo/dynamo-rl-codeowners
Other components/src/dynamo/common/... tests/support @ai-dynamo/dynamo-runtime-codeowners
components/src/dynamo/vllm/... and its tests @ai-dynamo/dynamo-backend-vllm-codeowners
lib/llm/src/protocols/openai/generate.rs @ai-dynamo/dynamo-frontend-codeowners

Validation

  • repository-configured pre-commit hooks passed after the Omni split
  • cargo fmt --all -- --check
  • cargo test -p dynamo-llm --no-default-features protocols::openai::generate::tests -- --nocapture — 32 passed
  • test_vllm_legacy_lora.py — 17 passed with vLLM 0.29 after the split
  • test_vllm_worker_handler.py and test_vllm_engine_generate.py — 119 passed, 5 dependency-gated skips after the split
  • Qwen3-0.6B MathEnv dense NCCL run — 3/3 steps passed; mismatch KL 0.0004, 0.0004, 0.0003
  • Qwen3-0.6B MathEnv filesystem-LoRA run after the lifecycle fix — 3/3 steps passed; mismatch KL 0.0003, 0.0004, 0.0006; immutable adapters v0-v3 loaded and drained successfully

Summary by CodeRabbit

  • New Features

    • Added support for structured routed-expert metadata, including tensor payloads.
    • Generation responses can now include sampling-mask data.
    • Added native support for adapted vLLM generation requests, including multimodal prompts and routing metadata.
    • Route descriptions can report configured world size.
  • Bug Fixes

    • Improved validation for malformed generation metadata, token limits, routing hashes, and sampling masks.
    • LoRA loading and unloading now safely wait for active requests to finish.
    • Added validation to prevent unsupported reinforcement-learning deployment topologies.

Post-review validation

  • Exact vLLM 0.29 focused request-adaptation and LoRA suites: 28 passed.
  • Worker-handler suite: 108 passed, 5 dependency-gated skips.
  • Repository pre-commit hooks, Python compilation, and diff checks passed.

Where should the reviewer start?

  • components/src/dynamo/vllm/engine_generate.py for native Generate request adaptation.
  • components/src/dynamo/vllm/handlers.py and components/src/dynamo/vllm/lora_state.py for response metadata and LoRA request draining.
  • components/src/dynamo/common/rl/admin.py for RL topology publication.

Related Issues

  • None.

@biswapanda
biswapanda deployed to external_collaborator September 22, 2026 18:14 — with GitHub Actions Active
@biswapanda
biswapanda deployed to external_collaborator September 22, 2026 18:14 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi biswapanda! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added external-contribution Pull request is from an external contributor feat labels Sep 22, 2026
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test 6190cd3

@github-actions github-actions Bot added backend::vllm Relates to the vllm backend frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` labels Sep 22, 2026
@biswapanda
biswapanda deployed to external_collaborator September 22, 2026 19:39 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test e7b84bb

@biswapanda biswapanda removed the external-contribution Pull request is from an external contributor label Sep 22, 2026
@biswapanda
biswapanda force-pushed the feat/dynamo-vllm-rl-parity branch from e7b84bb to 23061d1 Compare September 22, 2026 20:56
@biswapanda
biswapanda deployed to external_collaborator September 22, 2026 20:56 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test 23061d1

@biswapanda
biswapanda deployed to external_collaborator September 22, 2026 21:01 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test 613fd66

@biswapanda biswapanda self-assigned this Sep 22, 2026
@biswapanda biswapanda changed the title feat(vllm): add Python RL generate parity feat(vllm): add Python RL generate api parity Sep 22, 2026
@biswapanda
biswapanda marked this pull request as ready for review September 22, 2026 23:09
@biswapanda
biswapanda requested review from a team as code owners September 22, 2026 23:09
@biswapanda
biswapanda enabled auto-merge (squash) September 22, 2026 23:09
Comment thread components/src/dynamo/vllm/omni/omni_handler.py Outdated
Comment thread components/src/dynamo/vllm/omni/omni_handler.py Outdated
Comment thread components/src/dynamo/vllm/omni/omni_handler.py Outdated
Comment thread components/src/dynamo/vllm/omni/omni_handler.py Outdated
Comment thread components/src/dynamo/vllm/tests/test_vllm_legacy_lora.py Outdated
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
@biswapanda
biswapanda force-pushed the feat/dynamo-vllm-rl-parity branch from c71affd to 404be27 Compare September 24, 2026 05:24
@biswapanda
biswapanda deployed to external_collaborator September 24, 2026 05:24 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test 404be27

@dynamo-review-agent dynamo-review-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Previously reported defects still present:

  • Original discussion: Still present: adapt_engine_generate_request rejects every sampling_params.stop request even though vLLM supports stop strings when detokenization is enabled, and test_tito_adapter_rejects_stop_strings_but_preserves_stop_token_ids now codifies that unsupported rejection instead of preserving stop-string handling.

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
@biswapanda
biswapanda deployed to external_collaborator September 24, 2026 06:31 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test eee5e44

Comment thread components/src/dynamo/vllm/engine_generate.py
Comment thread components/src/dynamo/vllm/lora_state.py
Comment thread components/src/dynamo/vllm/lora_state.py
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
@biswapanda
biswapanda deployed to external_collaborator September 24, 2026 19:28 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test 0ecc46b

@rmccorm4 rmccorm4 mentioned this pull request Sep 27, 2026
Comment thread components/src/dynamo/vllm/engine_generate.py
Comment thread components/src/dynamo/vllm/handlers.py Outdated
@biswapanda
biswapanda deployed to external_collaborator September 28, 2026 18:46 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test f168ae1

@biswapanda
biswapanda requested a review from a team as a code owner September 28, 2026 19:21
@biswapanda
biswapanda deployed to external_collaborator September 28, 2026 19:21 — with GitHub Actions Active
@dynamo-ops

Copy link
Copy Markdown
Contributor

/ok to test d9a2532

This branch was successfully deployed

1 active deployment
external_collaborator — d9a25329 Deployed Sep 28, 2026 by biswapanda via ok-to-test #23376
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend feat frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` multimodal size/XXL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants