diff --git a/README.md b/README.md index 2cd6a574a48..434c197cdc3 100644 --- a/README.md +++ b/README.md @@ -16,6 +16,7 @@ Easy, fast, and cheap omni-modality model serving for everyone *Latest News* 🔥 +- [2026/09] We released [0.30.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.30.0), rebased onto vLLM 0.30.0, featuring a unified full-duplex serving framework around engine-owned sessions for [MiniCPM-o 4.5](recipes/OpenBMB/MiniCPM-o-4_5.md) and AURA, native cross-stage KV and multimodal payload transfer (Mooncake AR-to-DiT handoff and NIXL connectors), interactive world-model serving with [LingBot World](recipes/Robbyant/LingBot-World-2.0.md), and realtime [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) Turbo inference on NVIDIA Blackwell and Ascend 950. - [2026/08] We released [0.28.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.28.0), featuring production-ready [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) serving on GPU and NPU, a unified AR/DiT paged KV cache runtime, and enhanced realtime full-duplex serving for the [MiniCPM-o series](recipes/OpenBMB/MiniCPM-o-4_5.md). - [2026/08] [VeRL-Omni](https://github.com/verl-project/verl-omni) `v0.2.0` is released: faster diffusion RL powered by vLLM-Omni (request-level/step-wise batching with FA3), rebuilt Qwen3-Omni multimodal training (DPO & GSPO), plus LTX-2.3, Qwen-Image-Edit support and more. See the [release notes](https://github.com/verl-project/verl-omni/releases/tag/v0.2.0). - [2026/08] We released [0.26.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.26.0) - aligned with the vLLM 0.26 release line, featuring [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) joint video/audio generation, an experimental full-duplex realtime runtime for [MiniCPM-o 4.5](recipes/OpenBMB/MiniCPM-o-4_5.md), distributed layerwise diffusion offload, and broader model, hardware, streaming, TTS, and quantization support. @@ -58,9 +59,9 @@ vLLM-Omni is flexible and easy to use with: vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including: - **Omni-modality models** (e.g. Qwen3-Omni, MiniCPM-o 4.5, Cosmos3, HunyuanImage, BAGEL) -- **TTS models** (e.g. Qwen3-TTS, IndexTTS 2.5, CosyVoice3) -- **Diffusion models** — image, video, and audio generation (e.g. MiniMax H3, LTX-2.5, SANA-Video, Wan2.2) -- **Robot-policy and action models** (e.g. π0, GR00T-N1.7, DreamZero-DROID, InternVLA-A1) +- **TTS models** (e.g. Qwen3-TTS, Tencent AuK, Breeze-TTS-2, CosyVoice3) +- **Diffusion models** — image, video, and audio generation (e.g. MiniMax H3, LingBot World, MAGI-2, LTX-2.5, Wan2.2) +- **Robot-policy and action models** (e.g. π0.5, GR00T-N1.7, DreamZero-DROID, InternVLA-A1) ## Getting Started diff --git a/docs/README.md b/docs/README.md index 3733a9d2d59..92df8da11b6 100644 --- a/docs/README.md +++ b/docs/README.md @@ -56,9 +56,9 @@ vLLM-Omni is flexible and easy to use with: vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including: - Omni-modality models (e.g. Qwen3-Omni, MiniCPM-o 4.5, Cosmos3, HunyuanImage, BAGEL) -- TTS models (e.g. Qwen3-TTS, IndexTTS 2.5, CosyVoice3) -- Diffusion models — image, video, and audio generation (e.g. MiniMax H3, LTX-2.5, SANA-Video, Wan2.2) -- Robot-policy and action models (e.g. π0, GR00T-N1.7, DreamZero-DROID, InternVLA-A1) +- TTS models (e.g. Qwen3-TTS, Tencent AuK, Breeze-TTS-2, CosyVoice3) +- Diffusion models — image, video, and audio generation (e.g. MiniMax H3, LingBot World, MAGI-2, LTX-2.5, Wan2.2) +- Robot-policy and action models (e.g. π0.5, GR00T-N1.7, DreamZero-DROID, InternVLA-A1) For more information, checkout the following: diff --git a/docs/configuration/README.md b/docs/configuration/README.md index 3d065f9a1f7..fc12c26636f 100644 --- a/docs/configuration/README.md +++ b/docs/configuration/README.md @@ -2,7 +2,7 @@ This section lists the most common options for running vLLM-Omni. -For options within a vLLM Engine, please refer to the [vLLM 0.29 configuration guide](https://docs.vllm.ai/en/v0.29.0/configuration/index.html). +For options within a vLLM Engine, please refer to the [vLLM 0.30 configuration guide](https://docs.vllm.ai/en/v0.30.0/configuration/index.html). Each model defines fixed topology in a registered `PipelineConfig` and runtime overrides in a deploy YAML. diff --git a/docs/configuration/environment_variables.md b/docs/configuration/environment_variables.md index eaefb1015af..e8a56b58128 100644 --- a/docs/configuration/environment_variables.md +++ b/docs/configuration/environment_variables.md @@ -162,7 +162,7 @@ their keys only. ## Inherited vLLM variables vLLM-Omni also reads variables through its aligned vLLM dependency. Refer to -the [vLLM 0.29 environment-variable reference](https://docs.vllm.ai/en/v0.29.0/configuration/env_vars.html) +the [vLLM 0.30 environment-variable reference](https://docs.vllm.ai/en/v0.30.0/configuration/env_vars.html) for their definitions. This includes vLLM launch, cache, logging, plugin, ROCm, XPU, ModelScope, and FlashInfer workspace settings. diff --git a/docs/getting_started/installation/README.md b/docs/getting_started/installation/README.md index 077b35329ab..507a0bd96ec 100644 --- a/docs/getting_started/installation/README.md +++ b/docs/getting_started/installation/README.md @@ -1,7 +1,7 @@ # Installation !!! important - vLLM-Omni is released against the matching upstream vLLM major/minor version. The 0.30 development line requires vLLM 0.30.x; the source-install instructions pin vLLM 0.30.0. Published vLLM-Omni 0.28.0 wheels and images still require vLLM 0.28.x. + vLLM-Omni is released against the matching upstream vLLM major/minor version. The vLLM-Omni 0.30.x release line uses vLLM 0.30.x; the source-install instructions pin vLLM 0.30.0. vLLM-Omni supports the following hardware platforms: diff --git a/docs/getting_started/installation/gpu.md b/docs/getting_started/installation/gpu.md index 0dc58faf9aa..260e3fe7832 100644 --- a/docs/getting_started/installation/gpu.md +++ b/docs/getting_started/installation/gpu.md @@ -34,9 +34,7 @@ vLLM-Omni is a Python library that supports the following GPU variants. The libr ### Pre-built wheels -The source-install instructions target vLLM 0.30.0. Until vLLM-Omni 0.30.0 artifacts are published, the CUDA/ROCm pre-built examples below remain on the matching 0.28.0 release pair. - -Note: Pre-built wheels are currently available for vLLM-Omni 0.11.0rc1, 0.12.0rc1, 0.14.0rc1, 0.14.0, 0.16.0, 0.18.0, 0.20.0, 0.21.0, 0.22.0, 0.23.0, 0.24.0, 0.25.0, 0.26.0, and 0.28.0. If you need a newer unreleased revision, please [build from source](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/installation/gpu/#build-wheel-from-source). +Note: Pre-built wheels are currently available for vLLM-Omni 0.11.0rc1, 0.12.0rc1, 0.14.0rc1, 0.14.0, 0.16.0, 0.18.0, 0.20.0, 0.21.0, 0.22.0, 0.23.0, 0.24.0, 0.25.0, 0.26.0, 0.28.0, and 0.30.0. If you need a newer unreleased revision, please [build from source](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/installation/gpu/#build-wheel-from-source). === "NVIDIA CUDA" diff --git a/docs/getting_started/installation/gpu/cuda.inc.md b/docs/getting_started/installation/gpu/cuda.inc.md index c02efc7cfa4..d3704b0b10f 100644 --- a/docs/getting_started/installation/gpu/cuda.inc.md +++ b/docs/getting_started/installation/gpu/cuda.inc.md @@ -5,7 +5,7 @@ # --8<-- [end:requirements] # --8<-- [start:set-up-using-python] -vLLM-Omni depends on the matching major/minor release of vLLM. The 0.30 development line uses vLLM 0.30.x. Published 0.28.0 wheels and images use vLLM 0.28.x. +vLLM-Omni depends on the matching major/minor release of vLLM. The vLLM-Omni 0.30.x release line uses vLLM 0.30.x. !!! note PyTorch installed via `conda` will statically link `NCCL` library, which can cause issues when vLLM tries to use `NCCL`. See for more details. @@ -18,22 +18,22 @@ Therefore, it is recommended to install vLLM and vLLM-Omni with a **fresh new** #### Installation of vLLM -These pre-built wheel instructions install the published vLLM-Omni 0.28.0 release. For the 0.30 development line, use the source-install instructions below. +These pre-built wheel instructions install the published vLLM-Omni 0.30.0 release. vLLM-Omni is built based on vLLM. Please install it with command below. ```bash -uv pip install vllm==0.28.0 --torch-backend=auto +uv pip install vllm==0.30.0 --torch-backend=auto ``` #### Installation of vLLM-Omni ```bash -uv pip install vllm-omni==0.28.0 +uv pip install vllm-omni==0.30.0 ``` To run Gradio demos, also install the optional extras: ```bash -uv pip install 'vllm-omni[demo]==0.28.0' +uv pip install 'vllm-omni[demo]==0.30.0' ``` # --8<-- [end:pre-built-wheels] @@ -111,7 +111,7 @@ docker run --runtime nvidia --gpus 2 \ --env "HF_TOKEN=$HF_TOKEN" \ -p 8091:8091 \ --ipc=host \ - vllm/vllm-omni:v0.28.0 \ + vllm/vllm-omni:v0.30.0 \ vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni --port 8091 ``` diff --git a/docs/getting_started/installation/gpu/rocm.inc.md b/docs/getting_started/installation/gpu/rocm.inc.md index adf6985a3eb..63675ea94e1 100644 --- a/docs/getting_started/installation/gpu/rocm.inc.md +++ b/docs/getting_started/installation/gpu/rocm.inc.md @@ -10,9 +10,8 @@ For ROCm, vLLM-Omni currently recommends the setup steps through Docker Images. -vLLM-Omni depends on the matching major/minor release of vLLM. The 0.30 -development line uses vLLM 0.30.x. Published 0.28.0 wheels and images use vLLM -0.28.x. +vLLM-Omni depends on the matching major/minor release of vLLM. The vLLM-Omni +0.30.x release line uses vLLM 0.30.x. The Dockerfile's `BASE_IMAGE` pin applies only to Docker builds. The `vllm-omni` package does not install vLLM as a dependency, so non-Docker source @@ -23,12 +22,12 @@ installing vLLM-Omni, as shown below. #### Installation of vLLM -These pre-built wheel instructions install the published vLLM-Omni 0.28.0 release. For the 0.30 development line, use the source-install instructions below. +These pre-built wheel instructions install the published vLLM-Omni 0.30.0 release. vLLM-Omni is built based on vLLM. Please install it with command below. ```bash -uv pip install vllm==0.28.0+rocm723 --extra-index-url https://wheels.vllm.ai/rocm/0.28.0/rocm723 +uv pip install vllm==0.30.0+rocm723 --extra-index-url https://wheels.vllm.ai/rocm/0.30.0/rocm723 ``` #### Installation of vLLM-Omni @@ -37,7 +36,7 @@ uv pip install vllm==0.28.0+rocm723 --extra-index-url https://wheels.vllm.ai/roc # we need to add --no-build-isolation as the torch # is not obtained from pypi, we have to install using the # torch installed in our environment -uv pip install vllm-omni==0.28.0 +uv pip install vllm-omni==0.30.0 # Optional if want to run Qwen3 TTS uv pip uninstall onnxruntime # should be removed before we can install onnxruntime-rocm @@ -150,6 +149,8 @@ vllm-omni-rocm vLLM-Omni offers an official docker image for deployment. These images are built on top of vLLM docker images and available on Docker Hub as [vllm/vllm-omni-rocm](https://hub.docker.com/r/vllm/vllm-omni-rocm/tags). The version of vLLM-Omni indicates which release of vLLM it is based on. +The prebuilt ROCm image is published separately from the release pipeline: `v0.28.0` is the latest tag currently available on Docker Hub, and newer-tag availability is tracked in [#7405](https://github.com/vllm-project/vllm-omni/issues/7405). + #### Launch vLLM-Omni Server Here's an example deployment command that has been verified on 2 x MI300's: diff --git a/docs/getting_started/installation/npu/npu.inc.md b/docs/getting_started/installation/npu/npu.inc.md index 021c59638d3..ad29c844738 100644 --- a/docs/getting_started/installation/npu/npu.inc.md +++ b/docs/getting_started/installation/npu/npu.inc.md @@ -53,7 +53,7 @@ docker run --rm \ # Inside the container, install vLLM-Omni from source cd /vllm-workspace -git clone -b v0.30.0rc1 https://github.com/vllm-project/vllm-omni.git +git clone -b v0.30.0 https://github.com/vllm-project/vllm-omni.git cd vllm-omni pip install -v -e . --no-build-isolation # or VLLM_OMNI_TARGET_DEVICE=npu pip install -v -e .