[Model]: Support Helios from ByteDance by princepride · Pull Request #1604 · vllm-project/vllm-omni

princepride · 2026-03-02T09:42:21Z

Purpose

Add full support for the Helios video generation model family in vLLM-Omni, covering 3 model variants (Helios-Base, Helios-Mid, Helios-Distilled) across 3 generation tasks (T2V, I2V, V2V) — a complete 3×3 support matrix.

Variant	Pipeline Class	Key Features
Helios-Base	`HeliosPipeline`	Stage 1 single-stage denoising, `guidance_scale=5.0`
Helios-Mid	`HeliosPyramidPipeline`	Stage 2 pyramid denoising, CFG-Zero*
Helios-Distilled	`HeliosPyramidPipeline`	Stage 2 pyramid + DMD few-step inference, `guidance_scale=1.0`

Test Plan

Manual end-to-end testing across all supported model×task combinations, NPU is supported as well. Run from the repository root:

T2V (Text-to-Video):

# Helios-Base (Stage 1)
python3 examples/offline_inference/helios/end2end.py \
--sample-type t2v \
--model BestWishYsh/Helios-Base \
--prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \
--num-frames 600 \
--seed 42 \
--output helios_t2v_base.mp4

# Helios-Mid (Stage 2 + CFG-Zero*)
python examples/offline_inference/helios/end2end.py \
    --model BestWishYsh/Helios-Base --sample-type t2v \
    --prompt "A dynamic time-lapse video showing the rapidly moving scenery from the window of a speeding train." \
    --guidance-scale 5.0 --is-enable-stage2 \
    --pyramid-num-inference-steps-list 20 20 20 \
    --use-cfg-zero-star --use-zero-init --zero-steps 1 \
    --output helios_t2v_mid.mp4

# Helios-Distilled (Stage 2 + DMD)
python examples/offline_inference/helios/end2end.py \
    --model BestWishYsh/Helios-Base --sample-type t2v \
    --prompt "A dynamic time-lapse video showing the rapidly moving scenery from the window of a speeding train." \
    --num-frames 240 --guidance-scale 1.0 --is-enable-stage2 \
    --pyramid-num-inference-steps-list 2 2 2 \
    --is-amplify-first-chunk --output helios_t2v_distilled.mp4

I2V (Image-to-Video):

# Helios-Base
python examples/offline_inference/helios/end2end.py \
    --model ./Helios-Base --sample-type i2v \
    --image-path wave.jpg \
    --prompt "A towering emerald wave surges forward, its crest curling with raw power and energy." \
    --guidance-scale 5.0 --output helios_i2v_base.mp4

V2V (Video-to-Video):

# Helios-Base
python examples/offline_inference/helios/end2end.py \
    --model ./Helios-Base --sample-type v2v \
    --video-path car.mp4 \
    --prompt "A bright yellow Lamborghini speeds along a curving mountain road." \
    --guidance-scale 5.0 --output helios_v2v_base.mp4

Test Result

Helios-Base, T2V

Details

helios_output_1.mp4

Helios-Mid, T2V

Details

helios_t2v_mid.mp4

Helios-Distilled, T2V

Details

helios_t2v_distilled.mp4

Helios-Base, I2V

Details

helios_i2v_base.mp4

Helios-Base, V2V

Details

helios_v2v_base.mp4

Signed-off-by: princepride <wangzhipeng628@gmail.com>

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 69497c6179

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Signed-off-by: princepride <wangzhipeng628@gmail.com>

hsliuustc0106 · 2026-03-02T13:48:05Z

vllm-omni-reviewer Code Review

Found 3 critical issues:

1. Documentation formatting bug breaks table rendering

Line 10 has an orphaned HeliosPyramidPipeline outside the markdown table that breaks rendering. This appears to be a copy-paste error - the pipeline name is already in the table row above.

vllm-omni/docs/models/supported_models.md

Lines 9 to 11 in 83ce087

    
           ## List of Supported Models for Nvidia GPU / AMD GPU 
        
           <style>

Fix: Delete the orphaned line.

2. Missing test coverage for new model

This PR adds 3362 lines of new model code (transformer, pipeline, scheduler) but includes no test files. Per conventions, new model implementations require correctness validation tests.

https://github.com/vllm-project/vllm-omni/blob/83ce0872685aa373966665f8658b91f23d4a50c5/vllm_omni/diffusion/models/helios/

Fix: Add tests validating model loading and basic inference for the Helios variants.

3. Memory/hardware requirements not documented

The PR adds 3 model variants (Base, Mid, Distilled) with different memory profiles but doesn't document minimum VRAM, recommended GPUs, or memory configuration settings. Users have no guidance on hardware requirements.

vllm-omni/docs/models/supported_models.md

Line 9 in 83ce087

## List of Supported Models for Nvidia GPU / AMD GPU

Fix: Add memory requirements documentation for each variant (e.g., "Helios-Base: X GB VRAM minimum").

🤖 Generated with vllm-omni-reviewer

hsliuustc0106

vllm-omni-reviewer Code Review

Found 2 issues requiring fixes:

Markdown table broken (line 10) - orphaned HeliosPyramidPipeline text needs deletion
Missing memory requirements (line 9) - document VRAM needs for 3 model variants

See inline comments below.

Signed-off-by: 汪志鹏 <wangzhipeng628@gmail.com>

Gaohan123 · 2026-03-02T16:11:04Z

The link https://github.com/BestWishYsh/Helios cannot open

princepride · 2026-03-03T00:13:42Z

The model still didn't open source, please wait

Signed-off-by: gcanlin <canlinguosdu@gmail.com>

Support NPU

gcanlin · 2026-03-04T05:59:50Z

python3 examples/offline_inference/helios/end2end.py \
--sample-type t2v \
--model ./Helios-Base \
--prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \
--num-frames 600 \
--seed 42 \
--output helios_t2v_base.mp4 \
--tensor-parallel-size 2

NPU needs to enable TP=2. And end-to-end can run successfully.

Sy0307 · 2026-03-04T11:30:54Z

If Helios has any plans for optimization on vllm-omni in the future, I would be happy to participate.

Signed-off-by: princepride <wangzhipeng628@gmail.com> Signed-off-by: 汪志鹏 <wangzhipeng628@gmail.com> Signed-off-by: gcanlin <canlinguosdu@gmail.com> Co-authored-by: gcanlin <canlinguosdu@gmail.com>

### vllm-omni-api - Source: [PR #1724](vllm-project/vllm-omni#1724) - Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" - Changes: - New feature: Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" ### vllm-omni-contrib - Source: [PR #1724](vllm-project/vllm-omni#1724) - Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" - Changes: - New feature: Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" ### vllm-omni-api - Source: [PR #1716](vllm-project/vllm-omni#1716) - [Feature]: Add vae-patch-parallel CLI argument in online serving - Changes: - New feature: [Feature]: Add vae-patch-parallel CLI argument in online serving ### vllm-omni-contrib - Source: [PR #1716](vllm-project/vllm-omni#1716) - [Feature]: Add vae-patch-parallel CLI argument in online serving - Changes: - New feature: [Feature]: Add vae-patch-parallel CLI argument in online serving ### vllm-omni-contrib - Source: [PR #1693](vllm-project/vllm-omni#1693) - [skip CI][Docs] Add TTS model developer guide - Changes: - New feature: [skip CI][Docs] Add TTS model developer guide ### vllm-omni-audio-tts - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-distributed - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-perf - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-perf - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-distributed - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-api - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Additions: - `/v1/audio/speech` ### vllm-omni-quantization - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-cicd - Source: [PR #1683](vllm-project/vllm-omni#1683) - [CI] Remove high concurrency tests before issue #1374 fixed. - Changes: - Bug fix: [CI] Remove high concurrency tests before issue #1374 fixed. ### vllm-omni-audio-tts - Source: [PR #1678](vllm-project/vllm-omni#1678) - Add non-async chunk support for Qwen3-TTS - Changes: - New feature: Add non-async chunk support for Qwen3-TTS ### vllm-omni-cicd - Source: [PR #1678](vllm-project/vllm-omni#1678) - Add non-async chunk support for Qwen3-TTS - Changes: - New feature: Add non-async chunk support for Qwen3-TTS ### vllm-omni-cicd - Source: [PR #1677](vllm-project/vllm-omni#1677) - Replace hard-coded cuda generator with current_omni_platform.device_type ### vllm-omni-perf - Source: [PR #1677](vllm-project/vllm-omni#1677) - Replace hard-coded cuda generator with current_omni_platform.device_type ### vllm-omni-serving - Source: [PR #1675](vllm-project/vllm-omni#1675) - [Misc] remove logits_processor_pattern this field, because vllm have … ### vllm-omni-cicd - Source: [PR #1666](vllm-project/vllm-omni#1666) - [Cleanup] Move cosyvoice3 tests to model subdirectory ### vllm-omni-audio-tts - Source: [PR #1664](vllm-project/vllm-omni#1664) - [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder - Changes: - Bug fix: [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder ### vllm-omni-cicd - Source: [PR #1664](vllm-project/vllm-omni#1664) - [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder - Changes: - Bug fix: [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder ### vllm-omni-distributed - Source: [PR #1656](vllm-project/vllm-omni#1656) - [Optimize][Qwen3-Omni] Reduce inter-packet latency in async chunk ### vllm-omni-contrib - Source: [PR #1656](vllm-project/vllm-omni#1656) - [Optimize][Qwen3-Omni] Reduce inter-packet latency in async chunk ### vllm-omni-quantization - Source: [PR #1652](vllm-project/vllm-omni#1652) - [UX] Add progress bar for diffusion models - Changes: - New feature: [UX] Add progress bar for diffusion models ### vllm-omni-perf - Source: [PR #1652](vllm-project/vllm-omni#1652) - [UX] Add progress bar for diffusion models - Changes: - New feature: [UX] Add progress bar for diffusion models ### vllm-omni-distributed - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-quantization - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-perf - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-contrib - Source: [PR #1649](vllm-project/vllm-omni#1649) - [Misc] update wechat ### vllm-omni-perf - Source: [PR #1642](vllm-project/vllm-omni#1642) - [chore] add _repeated_blocks for regional compilation support - Changes: - New feature: [chore] add _repeated_blocks for regional compilation support ### vllm-omni-api - Source: [PR #1641](vllm-project/vllm-omni#1641) - [Bugfix] Add TTS request validation to prevent engine crashes - Changes: - New feature: [Bugfix] Add TTS request validation to prevent engine crashes ### vllm-omni-cicd - Source: [PR #1641](vllm-project/vllm-omni#1641) - [Bugfix] Add TTS request validation to prevent engine crashes - Changes: - New feature: [Bugfix] Add TTS request validation to prevent engine crashes ### vllm-omni-image-gen - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer - Additions: - text-to-image - Text-to-Image - Flux ### vllm-omni-quantization - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer - Additions: - FP8 support or improvements ### vllm-omni-contrib - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer ### vllm-omni-perf - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer ### vllm-omni-contrib - Source: [PR #1631](vllm-project/vllm-omni#1631) - [BugFix] Fix LongCat Sequence Parallelism / Small Cleanup - Changes: - Bug fix: [BugFix] Fix LongCat Sequence Parallelism / Small Cleanup ### vllm-omni-cicd - Source: [PR #1628](vllm-project/vllm-omni#1628) - [Test][Qwen3-Omni]Modify Qwen3-Omni benchmark test cases ### vllm-omni-perf - Source: [PR #1628](vllm-project/vllm-omni#1628) - [Test][Qwen3-Omni]Modify Qwen3-Omni benchmark test cases ### vllm-omni-perf - Source: [PR #1619](vllm-project/vllm-omni#1619) - [Bugfix] Fix Qwen3-TTS code predictor crash due to missing vLLM config context - Changes: - Bug fix: [Bugfix] Fix Qwen3-TTS code predictor crash due to missing vLLM config context ### vllm-omni-perf - Source: [PR #1617](vllm-project/vllm-omni#1617) - [Refactor][Perf] Qwen3-TTS: re-prefill Code Predictor with torch.compile + enable Code2Wav decoder CUDA Graph - Changes: - Performance improvement: [Refactor][Perf] Qwen3-TTS: re-prefill Code Predictor with torch.compile + enable Code2Wav decoder CUDA Graph ### vllm-omni-contrib - Source: [PR #1615](vllm-project/vllm-omni#1615) - [Doc] Fix links in the configuration doc - Changes: - Bug fix: [Doc] Fix links in the configuration doc ### vllm-omni-audio-tts - Source: [PR #1614](vllm-project/vllm-omni#1614) - perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor - Changes: - Performance improvement: perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor ### vllm-omni-perf - Source: [PR #1614](vllm-project/vllm-omni#1614) - perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor - Changes: - Performance improvement: perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor ### vllm-omni-image-gen - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Additions: - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image ### vllm-omni-api - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation ### vllm-omni-perf - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation ### vllm-omni-contrib - Source: [PR #1604](vllm-project/vllm-omni#1604) - [Model]: support Helios from ByteDance ### vllm-omni-perf - Source: [PR #1604](vllm-project/vllm-omni#1604) - [Model]: support Helios from ByteDance ### vllm-omni-serving - Source: [PR #1602](vllm-project/vllm-omni#1602) - [Bugfix] fix kernel error for qwen3-omni - Changes: - Bug fix: [Bugfix] fix kernel error for qwen3-omni ### vllm-omni-distributed - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-image-gen - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Additions: - HunyuanImage3 - HunyuanImage3Pipeline - HunyuanImage3 - HunyuanImage-3 - HunyuanImage-3 - HunyuanImage-3 - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage-3 ### vllm-omni-quantization - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-perf - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-audio-tts - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-api - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-cicd - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-contrib - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-api - Source: [PR #1579](vllm-project/vllm-omni#1579) - [1/N][Refactor] Clean up dead code in output processor ### vllm-omni-serving - Source: [PR #1579](vllm-project/vllm-omni#1579) - [1/N][Refactor] Clean up dead code in output processor ### vllm-omni-distributed - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-cicd - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-perf - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-contrib - Source: [PR #1576](vllm-project/vllm-omni#1576) - 0.16.0 release ### vllm-omni-audio-tts - Source: [PR #1570](vllm-project/vllm-omni#1570) - [bugfix] Fix unexpected argument 'is_finished' in function llm2code2wav_async_chunk of mimo-audio - Changes: - Bug fix: [bugfix] Fix unexpected argument 'is_finished' in function llm2code2wav_async_chunk of mimo-audio ### vllm-omni-api - Source: [PR #1566](vllm-project/vllm-omni#1566) - [Bugfix] Import InputPreprocessor into Renderer - Changes: - Bug fix: [Bugfix] Import InputPreprocessor into Renderer ### vllm-omni-distributed - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-quantization - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-perf - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-image-gen - Source: [PR #1537](vllm-project/vllm-omni#1537) - [NPU] [Features] [Bugfix] Support mindiesd adaln - Changes: - New feature: [NPU] [Features] [Bugfix] Support mindiesd adaln - Additions: - mindiesd - mindiesd - Qwen-Image-Edit-2509 - mindiesd - mindiesd - mindiesd - mindiesd ### vllm-omni-perf - Source: [PR #1537](vllm-project/vllm-omni#1537) - [NPU] [Features] [Bugfix] Support mindiesd adaln - Changes: - New feature: [NPU] [Features] [Bugfix] Support mindiesd adaln ### vllm-omni-serving - Source: [PR #1536](vllm-project/vllm-omni#1536) - [Bugfix] Fix transformers 5.x compat issues in online TTS serving - Changes: - Bug fix: [Bugfix] Fix transformers 5.x compat issues in online TTS serving ### vllm-omni-perf - Source: [PR #1536](vllm-project/vllm-omni#1536) - [Bugfix] Fix transformers 5.x compat issues in online TTS serving - Changes: - Bug fix: [Bugfix] Fix transformers 5.x compat issues in online TTS serving

princepride added 8 commits March 1, 2026 08:11

init helios support

79badb1

Signed-off-by: princepride <wangzhipeng628@gmail.com>

add zero*

b86d8d3

Signed-off-by: princepride <wangzhipeng628@gmail.com>

add HeliosPyramidPipeline

4694b1b

Signed-off-by: princepride <wangzhipeng628@gmail.com>

add I2V and V2V

b491fa1

Signed-off-by: princepride <wangzhipeng628@gmail.com>

add end2end example and readme.md

beed79e

Signed-off-by: princepride <wangzhipeng628@gmail.com>

fix some bug

02bead2

Signed-off-by: princepride <wangzhipeng628@gmail.com>

fix some bug

b3e86fd

Signed-off-by: princepride <wangzhipeng628@gmail.com>

adjust end2end and README

d059e94

Signed-off-by: princepride <wangzhipeng628@gmail.com>

princepride requested a review from hsliuustc0106 as a code owner March 2, 2026 09:42

adjust supported models

69497c6

Signed-off-by: princepride <wangzhipeng628@gmail.com>

chatgpt-codex-connector Bot reviewed Mar 2, 2026

View reviewed changes

Comment thread vllm_omni/diffusion/models/helios/pipeline_helios.py Outdated

Comment thread vllm_omni/diffusion/models/helios/pipeline_helios.py

support load model from huggingface id

83ce087

Signed-off-by: princepride <wangzhipeng628@gmail.com>

hsliuustc0106 reviewed Mar 2, 2026

View reviewed changes

Comment thread docs/models/supported_models.md Outdated

Comment thread docs/models/supported_models.md Outdated

Fix formatting of HeliosPipeline entry in supported models

0d9523f

Signed-off-by: 汪志鹏 <wangzhipeng628@gmail.com>

hsliuustc0106 added the ready label to trigger buildkite CI label Mar 2, 2026

princepride enabled auto-merge (squash) March 4, 2026 05:02

princepride and others added 3 commits March 4, 2026 13:03

Merge branch 'main' into helios-support

9d63c17

Support NPU

758c456

Signed-off-by: gcanlin <canlinguosdu@gmail.com>

Merge pull request #7 from gcanlin/helios

259ac69

Support NPU

hsliuustc0106 disabled auto-merge March 4, 2026 06:17

hsliuustc0106 merged commit f9ec07f into vllm-project:main Mar 4, 2026
6 of 7 checks passed

hsliuustc0106 changed the title ~~[Model]: Helios support~~ [Model]: support Helios from ByteDance Mar 4, 2026

david6666666 mentioned this pull request Mar 5, 2026

vLLM-Omni Model Support #808

Open

63 tasks

Sy0307 mentioned this pull request Mar 5, 2026

[RFC] Multi-Stage Abort / Barge-in for Omni Models #1480

Closed

wtomin mentioned this pull request Mar 12, 2026

[RFC]: Continuous Diffusion Model Acceleration Support #1217

Open

1 task

princepride changed the title ~~[Model]: support Helios from ByteDance~~ [Model]: Support Helios from ByteDance Mar 17, 2026

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Model]: Support Helios from ByteDance#1604

[Model]: Support Helios from ByteDance#1604
hsliuustc0106 merged 14 commits intovllm-project:mainfrom
princepride:helios-support

princepride commented Mar 2, 2026 •

edited

Loading

Uh oh!

chatgpt-codex-connector Bot left a comment

Uh oh!

Uh oh!

Uh oh!

hsliuustc0106 commented Mar 2, 2026

Uh oh!

hsliuustc0106 left a comment

Uh oh!

Uh oh!

Uh oh!

Gaohan123 commented Mar 2, 2026

Uh oh!

princepride commented Mar 3, 2026

Uh oh!

gcanlin commented Mar 4, 2026 •

edited

Loading

Uh oh!

Uh oh!

Sy0307 commented Mar 4, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

Conversation

princepride commented Mar 2, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Purpose

Test Plan

Test Result

Uh oh!

chatgpt-codex-connector Bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

Uh oh!

Uh oh!

hsliuustc0106 commented Mar 2, 2026

vllm-omni-reviewer Code Review

1. Documentation formatting bug breaks table rendering

2. Missing test coverage for new model

3. Memory/hardware requirements not documented

Uh oh!

hsliuustc0106 left a comment

Choose a reason for hiding this comment

vllm-omni-reviewer Code Review

Uh oh!

Uh oh!

Uh oh!

Gaohan123 commented Mar 2, 2026

Uh oh!

princepride commented Mar 3, 2026

Uh oh!

gcanlin commented Mar 4, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

Sy0307 commented Mar 4, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

princepride commented Mar 2, 2026 •

edited

Loading

gcanlin commented Mar 4, 2026 •

edited

Loading