Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ Easy, fast, and cheap omni-modality model serving for everyone

*Latest News* 🔥

- [2026/09] We released [0.30.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.30.0), rebased onto vLLM 0.30.0, featuring a unified full-duplex serving framework around engine-owned sessions for [MiniCPM-o 4.5](recipes/OpenBMB/MiniCPM-o-4_5.md) and AURA, native cross-stage KV and multimodal payload transfer (Mooncake AR-to-DiT handoff and NIXL connectors), interactive world-model serving with [LingBot World](recipes/Robbyant/LingBot-World-2.0.md), and realtime [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) Turbo inference on NVIDIA Blackwell and Ascend 950.
- [2026/08] We released [0.28.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.28.0), featuring production-ready [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) serving on GPU and NPU, a unified AR/DiT paged KV cache runtime, and enhanced realtime full-duplex serving for the [MiniCPM-o series](recipes/OpenBMB/MiniCPM-o-4_5.md).
- [2026/08] [VeRL-Omni](https://github.com/verl-project/verl-omni) `v0.2.0` is released: faster diffusion RL powered by vLLM-Omni (request-level/step-wise batching with FA3), rebuilt Qwen3-Omni multimodal training (DPO & GSPO), plus LTX-2.3, Qwen-Image-Edit support and more. See the [release notes](https://github.com/verl-project/verl-omni/releases/tag/v0.2.0).
- [2026/08] We released [0.26.0](https://github.com/vllm-project/vllm-omni/releases/tag/v0.26.0) - aligned with the vLLM 0.26 release line, featuring [MiniMax H3](recipes/MiniMaxAI/MiniMax-H3.md) joint video/audio generation, an experimental full-duplex realtime runtime for [MiniCPM-o 4.5](recipes/OpenBMB/MiniCPM-o-4_5.md), distributed layerwise diffusion offload, and broader model, hardware, streaming, TTS, and quantization support.
Expand Down Expand Up @@ -58,9 +59,9 @@ vLLM-Omni is flexible and easy to use with:
vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including:

- **Omni-modality models** (e.g. Qwen3-Omni, MiniCPM-o 4.5, Cosmos3, HunyuanImage, BAGEL)
- **TTS models** (e.g. Qwen3-TTS, IndexTTS 2.5, CosyVoice3)
- **Diffusion models** — image, video, and audio generation (e.g. MiniMax H3, LTX-2.5, SANA-Video, Wan2.2)
- **Robot-policy and action models** (e.g. π0, GR00T-N1.7, DreamZero-DROID, InternVLA-A1)
- **TTS models** (e.g. Qwen3-TTS, Tencent AuK, Breeze-TTS-2, CosyVoice3)
- **Diffusion models** — image, video, and audio generation (e.g. MiniMax H3, LingBot World, MAGI-2, LTX-2.5, Wan2.2)
- **Robot-policy and action models** (e.g. π0.5, GR00T-N1.7, DreamZero-DROID, InternVLA-A1)

## Getting Started

Expand Down
6 changes: 3 additions & 3 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,9 +56,9 @@ vLLM-Omni is flexible and easy to use with:
vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including:

- Omni-modality models (e.g. Qwen3-Omni, MiniCPM-o 4.5, Cosmos3, HunyuanImage, BAGEL)
- TTS models (e.g. Qwen3-TTS, IndexTTS 2.5, CosyVoice3)
- Diffusion models — image, video, and audio generation (e.g. MiniMax H3, LTX-2.5, SANA-Video, Wan2.2)
- Robot-policy and action models (e.g. π0, GR00T-N1.7, DreamZero-DROID, InternVLA-A1)
- TTS models (e.g. Qwen3-TTS, Tencent AuK, Breeze-TTS-2, CosyVoice3)
- Diffusion models — image, video, and audio generation (e.g. MiniMax H3, LingBot World, MAGI-2, LTX-2.5, Wan2.2)
- Robot-policy and action models (e.g. π0.5, GR00T-N1.7, DreamZero-DROID, InternVLA-A1)

For more information, checkout the following:

Expand Down
2 changes: 1 addition & 1 deletion docs/configuration/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

This section lists the most common options for running vLLM-Omni.

For options within a vLLM Engine, please refer to the [vLLM 0.29 configuration guide](https://docs.vllm.ai/en/v0.29.0/configuration/index.html).
For options within a vLLM Engine, please refer to the [vLLM 0.30 configuration guide](https://docs.vllm.ai/en/v0.30.0/configuration/index.html).

Each model defines fixed topology in a registered `PipelineConfig` and runtime
overrides in a deploy YAML.
Expand Down
2 changes: 1 addition & 1 deletion docs/configuration/environment_variables.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ their keys only.
## Inherited vLLM variables

vLLM-Omni also reads variables through its aligned vLLM dependency. Refer to
the [vLLM 0.29 environment-variable reference](https://docs.vllm.ai/en/v0.29.0/configuration/env_vars.html)
the [vLLM 0.30 environment-variable reference](https://docs.vllm.ai/en/v0.30.0/configuration/env_vars.html)
for their definitions. This includes vLLM launch, cache, logging, plugin, ROCm,
XPU, ModelScope, and FlashInfer workspace settings.

Expand Down
2 changes: 1 addition & 1 deletion docs/getting_started/installation/README.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Installation

!!! important
vLLM-Omni is released against the matching upstream vLLM major/minor version. The 0.30 development line requires vLLM 0.30.x; the source-install instructions pin vLLM 0.30.0. Published vLLM-Omni 0.28.0 wheels and images still require vLLM 0.28.x.
vLLM-Omni is released against the matching upstream vLLM major/minor version. The vLLM-Omni 0.30.x release line uses vLLM 0.30.x; the source-install instructions pin vLLM 0.30.0.

vLLM-Omni supports the following hardware platforms:

Expand Down
4 changes: 1 addition & 3 deletions docs/getting_started/installation/gpu.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,9 +34,7 @@ vLLM-Omni is a Python library that supports the following GPU variants. The libr

### Pre-built wheels

The source-install instructions target vLLM 0.30.0. Until vLLM-Omni 0.30.0 artifacts are published, the CUDA/ROCm pre-built examples below remain on the matching 0.28.0 release pair.

Note: Pre-built wheels are currently available for vLLM-Omni 0.11.0rc1, 0.12.0rc1, 0.14.0rc1, 0.14.0, 0.16.0, 0.18.0, 0.20.0, 0.21.0, 0.22.0, 0.23.0, 0.24.0, 0.25.0, 0.26.0, and 0.28.0. If you need a newer unreleased revision, please [build from source](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/installation/gpu/#build-wheel-from-source).
Note: Pre-built wheels are currently available for vLLM-Omni 0.11.0rc1, 0.12.0rc1, 0.14.0rc1, 0.14.0, 0.16.0, 0.18.0, 0.20.0, 0.21.0, 0.22.0, 0.23.0, 0.24.0, 0.25.0, 0.26.0, 0.28.0, and 0.30.0. If you need a newer unreleased revision, please [build from source](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/installation/gpu/#build-wheel-from-source).

=== "NVIDIA CUDA"

Expand Down
12 changes: 6 additions & 6 deletions docs/getting_started/installation/gpu/cuda.inc.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
# --8<-- [end:requirements]
# --8<-- [start:set-up-using-python]

vLLM-Omni depends on the matching major/minor release of vLLM. The 0.30 development line uses vLLM 0.30.x. Published 0.28.0 wheels and images use vLLM 0.28.x.
vLLM-Omni depends on the matching major/minor release of vLLM. The vLLM-Omni 0.30.x release line uses vLLM 0.30.x.

!!! note
PyTorch installed via `conda` will statically link `NCCL` library, which can cause issues when vLLM tries to use `NCCL`. See <gh-issue:8420> for more details.
Expand All @@ -18,22 +18,22 @@ Therefore, it is recommended to install vLLM and vLLM-Omni with a **fresh new**

#### Installation of vLLM

These pre-built wheel instructions install the published vLLM-Omni 0.28.0 release. For the 0.30 development line, use the source-install instructions below.
These pre-built wheel instructions install the published vLLM-Omni 0.30.0 release.

vLLM-Omni is built based on vLLM. Please install it with command below.
```bash
uv pip install vllm==0.28.0 --torch-backend=auto
uv pip install vllm==0.30.0 --torch-backend=auto
```

#### Installation of vLLM-Omni

```bash
uv pip install vllm-omni==0.28.0
uv pip install vllm-omni==0.30.0
```

To run Gradio demos, also install the optional extras:
```bash
uv pip install 'vllm-omni[demo]==0.28.0'
uv pip install 'vllm-omni[demo]==0.30.0'
```

# --8<-- [end:pre-built-wheels]
Expand Down Expand Up @@ -111,7 +111,7 @@ docker run --runtime nvidia --gpus 2 \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8091:8091 \
--ipc=host \
vllm/vllm-omni:v0.28.0 \
vllm/vllm-omni:v0.30.0 \
vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni --port 8091
```

Expand Down
13 changes: 7 additions & 6 deletions docs/getting_started/installation/gpu/rocm.inc.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,8 @@

For ROCm, vLLM-Omni currently recommends the setup steps through Docker Images.

vLLM-Omni depends on the matching major/minor release of vLLM. The 0.30
development line uses vLLM 0.30.x. Published 0.28.0 wheels and images use vLLM
0.28.x.
vLLM-Omni depends on the matching major/minor release of vLLM. The vLLM-Omni
0.30.x release line uses vLLM 0.30.x.

The Dockerfile's `BASE_IMAGE` pin applies only to Docker builds. The
`vllm-omni` package does not install vLLM as a dependency, so non-Docker source
Expand All @@ -23,12 +22,12 @@ installing vLLM-Omni, as shown below.

#### Installation of vLLM

These pre-built wheel instructions install the published vLLM-Omni 0.28.0 release. For the 0.30 development line, use the source-install instructions below.
These pre-built wheel instructions install the published vLLM-Omni 0.30.0 release.

vLLM-Omni is built based on vLLM. Please install it with command below.

```bash
uv pip install vllm==0.28.0+rocm723 --extra-index-url https://wheels.vllm.ai/rocm/0.28.0/rocm723
uv pip install vllm==0.30.0+rocm723 --extra-index-url https://wheels.vllm.ai/rocm/0.30.0/rocm723
```

#### Installation of vLLM-Omni
Expand All @@ -37,7 +36,7 @@ uv pip install vllm==0.28.0+rocm723 --extra-index-url https://wheels.vllm.ai/roc
# we need to add --no-build-isolation as the torch
# is not obtained from pypi, we have to install using the
# torch installed in our environment
uv pip install vllm-omni==0.28.0
uv pip install vllm-omni==0.30.0

# Optional if want to run Qwen3 TTS
uv pip uninstall onnxruntime # should be removed before we can install onnxruntime-rocm
Expand Down Expand Up @@ -150,6 +149,8 @@ vllm-omni-rocm

vLLM-Omni offers an official docker image for deployment. These images are built on top of vLLM docker images and available on Docker Hub as [vllm/vllm-omni-rocm](https://hub.docker.com/r/vllm/vllm-omni-rocm/tags). The version of vLLM-Omni indicates which release of vLLM it is based on.

The prebuilt ROCm image is published separately from the release pipeline: `v0.28.0` is the latest tag currently available on Docker Hub, and newer-tag availability is tracked in [#7405](https://github.com/vllm-project/vllm-omni/issues/7405).

#### Launch vLLM-Omni Server

Here's an example deployment command that has been verified on 2 x MI300's:
Expand Down
2 changes: 1 addition & 1 deletion docs/getting_started/installation/npu/npu.inc.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ docker run --rm \

# Inside the container, install vLLM-Omni from source
cd /vllm-workspace
git clone -b v0.30.0rc1 https://github.com/vllm-project/vllm-omni.git
git clone -b v0.30.0 https://github.com/vllm-project/vllm-omni.git
cd vllm-omni
pip install -v -e . --no-build-isolation
# or VLLM_OMNI_TARGET_DEVICE=npu pip install -v -e .
Expand Down
Loading