Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 17 additions & 4 deletions container/templates/vllm_runtime.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -182,6 +182,9 @@ COPY --chmod=775 --chown=dynamo:0 --from=wheel_builder /opt/dynamo/dist/*.whl /o
{# Inline expression, not a block tag: render.py leaves trim_blocks off, so a tag
on its own line inside the RUN breaks the backslash continuation. #}
{% set vllm_rs_required = "1" if device == "cuda" else "0" %}
{# TODO: Remove this workaround once bundled vllm-rs accepts extra output fields. #}
{% set vllm_rs_allowlist = "1" if target not in ("dev", "local-dev") else "0" %}
{% set vllm_rs_plugins = "modelexpress" if context.vllm.enable_modelexpress == "true" else "" %}

# The vLLM 0.28.0 release images resolve the unbounded `transformers>=5.5.3`
# requirement to 5.15.1, but vLLM-Omni 0.28.0rc1 caps Transformers below 5.15.
Expand Down Expand Up @@ -531,18 +534,28 @@ if actual != expected:
raise RuntimeError(f"expected transformers {expected}, found {actual}")
PY

# `vllm-rs` ships inside the installed `vllm` package, not as a console script;
# linking it keeps the binary at that package's vLLM revision. Fatal on cuda only.
# Use the packaged binary to match the installed vLLM version.
RUN set -eu; \
pkg="$({{ python_executable }} -c 'import os, vllm; print(os.path.dirname(vllm.__file__))')"; \
if [ -f "${pkg}/vllm-rs" ] && [ -x "${pkg}/vllm-rs" ]; then \
ln -sf "${pkg}/vllm-rs" {{ vllm_rs_link }}; \
if [ "{{ vllm_rs_allowlist }}" = "1" ]; then \
printf '%s\n' \
'#!/bin/sh' \
'# Keep Omni from changing the EngineCore output schema.' \
'VLLM_PLUGINS="${VLLM_PLUGINS-{{ vllm_rs_plugins }}}"' \
'export VLLM_PLUGINS' \
"exec \"${pkg}/vllm-rs\" \"\$@\"" \
> {{ vllm_rs_link }}; \
chmod 755 {{ vllm_rs_link }}; \
else \
ln -sf "${pkg}/vllm-rs" {{ vllm_rs_link }}; \
fi; \
vllm-rs --help >/dev/null; \
elif [ "{{ vllm_rs_required }}" = "1" ]; then \
echo "ERROR: installed vllm package (${pkg}) ships no executable vllm-rs" >&2; \
exit 1; \
else \
echo "WARNING: installed vllm package (${pkg}) ships no executable vllm-rs; not linking it onto PATH" >&2; \
echo "WARNING: installed vllm package (${pkg}) ships no executable vllm-rs; not putting it onto PATH" >&2; \
Comment thread
glamr-agent marked this conversation as resolved.
fi

USER dynamo
Expand Down
15 changes: 13 additions & 2 deletions lib/sidecar/vllm/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,15 @@ The official `Qwen/Qwen3-ASR-1.7B` repository currently needs Rust-frontend-comp

### Runtime compatibility

The Python `vllm` package and `vllm-rs` must expose compatible EngineCore and gRPC contracts. Prefer artifacts built from the same vLLM source revision; do not combine a Python wheel from one nightly with a `vllm-rs` binary from another. The sidecar's vendored gRPC source revisions are recorded in [`proto/README.md`](proto/README.md).
The Python `vllm` package and `vllm-rs` must come from compatible vLLM revisions. Do not combine a wheel from one nightly with a binary from another. The sidecar's vendored gRPC source revisions are recorded in [`proto/README.md`](proto/README.md).

vLLM-Omni changes the engine response format, causing `vllm-rs` to reject
responses. The Dynamo vLLM runtime image provides a `vllm-rs` wrapper that
disables Omni by default. Use `vllm-rs` from `PATH` when starting the engine.

The wrapper enables only ModelExpress when installed; otherwise it disables
all plugins. An exported `VLLM_PLUGINS` overrides this default. The `dev` and
`local-dev` images do not install Omni and retain normal plugin discovery.

Start vLLM with its gRPC listener:

Expand Down Expand Up @@ -180,7 +188,10 @@ command.

The sidecar waits for both the Control and Inference services through the standard gRPC health API before registering the worker. The deployment manifests retain lightweight socket probes for container lifecycle monitoring. The engine image must include a `vllm-rs` build compatible with the vendored protocol.

The Dynamo vLLM CUDA runtime image exposes `vllm-rs` on `PATH`, linked from the `vllm` package that image installs, so `vllm-rs serve` runs there by name; that image fails to build if its `vllm` package ever stops shipping the binary. The XPU and CPU variants build from separate upstream vLLM distributions and link it the same way when it is present, but only warn when it is not, so check `command -v vllm-rs` before relying on it there. A stock upstream engine image ships the same binary inside the package but leaves it off `PATH`; `deploy/agg.yaml` and `deploy/disagg.yaml` run that image and resolve the path out of the package themselves.
The Dynamo vLLM runtime image exposes `vllm-rs` through the
[wrapper described above](#runtime-compatibility). On CPU and XPU, check that
the binary is available with `command -v vllm-rs`. The example manifests use
upstream vLLM images and locate the binary inside the Python package.

### Prerequisites

Expand Down
Loading