Skip to content

[Model] Add LingBot World Ulysses sequence parallelism - #6841

Merged
tzhouam merged 19 commits into
vllm-project:mainfrom
wtz2333:codex/lingbot-world-sp-vae-parallel
Sep 11, 2026
Merged

tzhouam merged 19 commits into
vllm-project:mainfrom
wtz2333:codex/lingbot-world-sp-vae-parallel

Conversation

@wtz2333

@wtz2333 wtz2333 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

The core computational workload of LingBot World lies in its DiT component, and compute bound is the primary bottleneck affecting performance. Currently, the single-GPU performance of LingBot World fails to meet real-time requirements. Moreover, the main branch does not natively support sequence parallelism, so its DiT workload cannot be distributed across sequence shards to reduce request latency.

This PR adds pure Ulysses SP for direct and AR-Diffusion execution. Hidden tokens, camera features, timestep conditioning and RoPE remain aligned, while attention and paged/text KV use local head shards. Historical KV stays on each head shard, keeping exchange payloads proportional to the current AR block. Ring, AllGather-KV and VAE parallelism remain unsupported for this model.

What Changes

File Change
vllm_omni/diffusion/models/lingbot_world/transformer.py Token-aligned SP plan, Ulysses attention exchanges, local KV heads and output gather.
vllm_omni/diffusion/models/lingbot_world/pipeline.py Pure Ulysses validation, rejection of unsupported ulysses_a2a_permute, and SP-local paged/text KV cache geometry.
tests/diffusion/models/lingbot_world/test_lingbot_world_attention.py SP import stubs for isolated attention tests.
tests/diffusion/models/lingbot_world/test_lingbot_world_transformer.py Token-aligned SP plan coverage.
tests/diffusion/models/lingbot_world/test_pipeline_lingbot_world.py Supported/rejected parallel configurations and cache geometry coverage.
docs/design/feature/realtime_ar_diffusion.md Ulysses execution contract.
examples/offline_inference/diffusion/lingbot_world_v2.py Expose --ulysses-degree and validate positive degrees.
tests/examples/offline_inference/test_lingbot_world_v2.py Validate the default, explicit and invalid values of the Ulysses CLI option.

Test Plan

  • vLLM Version: 0.28.0.
  • vLLM-Omni Commit: 5263a564c for both arms; rebase base 34aa1d2af.
  • PyTorch 2.13.0+cu130, CUDA 13.0, Transformers 5.14.0, Diffusers 0.40.0.
  • NVIDIA H200 devices.
  • robbyant/lingbot-world-v2-14b-causal-fast-diffusers, 480x832, 81 frames, seed 42, BF16, eager, TP1, concurrency/batch 1; four DMD steps per AR block.
  • Additional latent/video comparisons: 81 frames with seed 43 and 117 frames with seed 44, distinct prompts/camera scripts on the same input image.
  • Separate profiling: same head and 81-frame SP2 workload; one warmup and one control before each diagnostic request. Two-rank Nsight Systems 2026.1 and a separate Torch shape trace; synchronized stage timers run separately from both traces and the benchmark.

GPU regression commands:

python -m examples.offline_inference.diffusion.lingbot_world_v2 \
  --model robbyant/lingbot-world-v2-14b-causal-fast-diffusers \
  --image /path/to/first_frame.jpg --action-dir /path/to/actions/forward \
  --prompt "The camera moves slowly forward through the scene." \
  --height 480 --width 832 --num-frames 81 --seed 42 --fps 16 \
  --tensor-parallel-size 1 --ulysses-degree "$sp" --output "lingbot_sp${sp}.mp4"

CPU regression commands:

pytest -q tests/diffusion/models/lingbot_world tests/diffusion/ar_diffusion \
  -m "core_model and cpu"
pytest -q tests/examples/offline_inference/test_lingbot_world_v2.py -m "core_model and cpu"

Test Result

Current pipeline and offline CLI CPU regression suites: 121 passed at 4327313eb.

Latency and memory

torch compile enabled

Metric SP1 SP2 SP4
Measured requests 10 5 5
Video-request median (ms) 17735.48 12362.91 8868.13
Video-request nearest-rank p95 (ms) 17798.36 12436.05 8959.37
Full-window AR-block GPU median (ms) 2397.47 1344.50 724.99
Peak allocated memory, maximum rank (GiB) 69.28 69.74 59.58
Peak reserved memory, maximum rank (GiB) 72.12 72.46 62.28

Output consistency

The following results hold across SP1, SP2 and SP4:

torch.compile enabled

Frames / seed Pair Latent max abs / RMSE SSIM mean / worst Worst PSNR (dB)
81 / 42 SP1 / SP2 1.34756 / 0.0314614 0.975202 / 0.954628 30.8642
81 / 42 SP1 / SP4 1.08444 / 0.048722 0.958006 / 0.887530 25.8456
81 / 43 SP1 / SP2 1.58044 / 0.0632767 0.939152 / 0.823248 24.0161
81 / 43 SP1 / SP4 1.79564 / 0.0468043 0.964307 / 0.894646 26.5263
117 / 44 SP1 / SP2 0.948059 / 0.0357667 0.973674 / 0.908727 26.8081
117 / 44 SP1 / SP4 0.662083 / 0.0234112 0.984706 / 0.964608 32.3986

torch.compile disabled

Frames Seed Camera Latent max abs / RMSE Decoded SSIM mean / worst Decoded PSNR MP4 hashes
81 42 Forward 0 / 0 1.0 / 1.0 ∞ Identical
81 43 Left 0 / 0 1.0 / 1.0 ∞ Identical
117 44 Forward 0 / 0 1.0 / 1.0 ∞ Identical

SP2 profiling

Traces show 35 DiT forwards (four DMD probes plus one clean-KV commit per block) and 8,400 SendRecv kernels per rank. Each block has 4,680 tokens; SP2 begins with 2,340 tokens per rank and exchanges sequence shards for 20 local heads.

Activity Exposed interval union (ms) Share of trace span
NCCL 532.44 4.31%
NCCL + all-to-all layout copies (inclusive) 939.67 7.60%

Exposure is the union across both ranks with concurrent unrelated GPU work removed, divided by the 12363.33 ms trace span. NCCL includes peer-arrival waiting; these values are not pure transfer time or guaranteed optimization savings.

Separate stage timing identifies DiT execution as the main time consumer:

Stage (rank 0) Total (ms) Attribution
prepare_encode 1948.99 Includes 1728.14 ms of unsharded VAE encoding.
28 denoise_step calls 7983.74 Four DMD probes per block.
7 post_decode calls 2010.75 Mainly clean-KV forwards plus next-block preparation; no pixel decoding.
step_scheduler 5.21 Scheduler updates.

Communication alone does not explain the gap to ideal two-GPU scaling. Stage timing and communication exposure come from separate runs and must not be added together; they do not isolate per-stage SP1→SP2 savings. All diagnostic latents match the reference.

@wtz2333 wtz2333 changed the title [world-model][feature] Add LingBot Ulysses and spatial VAE parallel decode [Lingbot World][feature] Add LingBot Ulysses and spatial VAE parallel decode Aug 31, 2026
@wtz2333
wtz2333 force-pushed the codex/lingbot-world-sp-vae-parallel branch from 28458b0 to 49f456c Compare August 31, 2026 03:30
@hsliuustc0106 hsliuustc0106 added world model diffusion codes related to diffusion models labels Aug 31, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches tests/diffusion/, vllm_omni/diffusion/, examples/offline_inference/, docs/design/, vllm_omni/experimental/ (14 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@yenuo26 @NickCao @Bounty-hunter

Could one of you take a look when you get a chance? Thanks!

@wtz2333
wtz2333 force-pushed the codex/lingbot-world-sp-vae-parallel branch from 49f456c to 313fd1c Compare September 1, 2026 07:33
@wtz2333 wtz2333 changed the title [Lingbot World][feature] Add LingBot Ulysses and spatial VAE parallel decode [LingBot World][Feature] Add LingBot Ulysses sequence parallelism Sep 1, 2026
Signed-off-by: wtz2333 <2955110911@qq.com>
@wtz2333
wtz2333 force-pushed the codex/lingbot-world-sp-vae-parallel branch from 313fd1c to 5263a56 Compare September 7, 2026 06:34
@wtz2333 wtz2333 changed the title [LingBot World][Feature] Add LingBot Ulysses sequence parallelism [Model] Add LingBot World Ulysses sequence parallelism Sep 7, 2026
@wtz2333
wtz2333 marked this pull request as ready for review September 7, 2026 06:41
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/parallelism.md.

Module owners: @princepride @yuanheng-zhao @xuechendi @Bounty-hunter

@wtz2333, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: three questions on the performance claim

@wtz2333 this PR reads as a performance or value claim:

  • claim: On the rebased head, SP2 reduces the median 81-frame latent-request latency from 19557.39 ms to 12233.09 ms (1.60×).
  • claim: SP2 reduces median 81-frame latent-request latency by 37.45%, from 19,557.39 ms to 12,233.09 ms, a 1.60× speedup.

Before the full evidence checklist, three short questions:

  1. Bottleneck — what is the current bottleneck, and which profile, trace or per-stage measurement shows it?
  2. Value — what does the change buy the user or the system (latency, throughput, memory, cost), and at which workload?
  3. A/B or ablation — is there a same-workload, same-head/config comparison that isolates each main claim on its own? For stacked optimizations, one number per item rather than a blended delta.

When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution.

@wtz2333
wtz2333 requested a review from ywang96 as a code owner September 9, 2026 08:34
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 629a3de262c2 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: wtz2333 <2955110911@qq.com>
@wtz2333
wtz2333 force-pushed the codex/lingbot-world-sp-vae-parallel branch from 68246b5 to e1a9b81 Compare September 9, 2026 08:43
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
@wtz2333
wtz2333 requested a review from congw729 as a code owner September 10, 2026 09:06
@wtz2333

wtz2333 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Reviewed the SP implementation in depth. Summary: the sharding math is correct as far as I can tell, and the speedup is real (17.7s -> 12.4s -> 8.9s). My concerns are a memory regression that also lands on the existing non-SP path, one config flag that is accepted and then breaks, and the absence of any test that executes a real all-to-all.

What I verified as correct, so nobody re-derives it:

  • shard_kv_heads indexes ulysses_rank * num_sp_heads, which matches the head range all_to_all_4D actually delivers to rank r, and sp_shard hands rank r the r-th contiguous chunk. These line up only because the new validator pins sequence_parallel_size == ulysses_degree with ring/allgather at 1 - good call making that explicit.
  • RoPE is applied before the exchange on the local shard, with correspondingly sharded (S, head_dim) tables matching expected_dims=2.
  • Token order is frame-major in patch_embedding, _LingBotCameraPatchEmbedding and _rotary_embedding alike, so the timestep expansion aligns.
  • current_start / sink_tokens / query_len stay in global token coordinates and are computed before sharding, which is right given attention sees the full sequence post-exchange.
  • No double head division: DiffusionPagedAttentionLayerAdapter divides by ulysses_degree only when skip_sequence_parallel is False, and this model sets it True and pre-divides itself.

TP is not broken. I loaded the pre-PR and post-PR LingBotAttentionBlock side by side with identical weights and inputs: output is bitwise equal (max abs diff 0.0), including the cross-attention cache tensors, so the per-frame -> per-token modulation refactor is a pure representation change. The head-count substitution in pipeline.py:568 is also identical for every TP size the model accepts, and get_sp_group() adds no initialization precondition since _SP is created unconditionally alongside _TP in initialize_model_parallel.

Two TP-adjacent points that have no diff line to attach to:

  • test_ar_diffusion_capability_uses_fixed_tp_local_lingbot_geometry (test_pipeline_lingbot_world.py:444) still does module.get_tensor_model_parallel_world_size = lambda: 1, but this PR removed that symbol from pipeline.py entirely. The line is now inert and the test no longer exercises what its name claims - the head count comes from the _RecordingTransformer stub's hardcoded num_sp_heads=2. Worth removing the stub or re-pointing the test.
  • TP > 1 combined with ulysses > 1 is newly permitted (num_attention_heads // tp // ulysses) and has no coverage at all: every benchmark here is tp=1 and every unit test stubs tp=1.

Inline comments below, roughly in severity order.

Thanks for the detailed review. I pushed the follow-up changes in ea2dcca5d and the preceding commits:

  • Non-SP execution now retains frame-broadcast timestep modulation; token expansion runs only with Ulysses enabled.
  • Configuration validation rejects advanced_uaa, defaults a missing ulysses_degree to 1, and tests the Ring/AllGather clauses independently.
  • The transformer checks that all SP inputs were actually sharded before entering attention blocks.
  • The output head now projects local tokens before the gather, reducing the 14B gathered width from 5120 to 64. The head uses global token offsets to handle shards that split a frame.
  • Text K/V head shards now own compact storage. I evaluated removing the cross-attention exchanges, but retained the head-sharded implementation to keep this SP integration within the existing shared cache contract. The final diff does not extend the AR cache framework.
  • Cache head counts are derived from model and parallel configuration; the module-tree dependency and its test stub are removed.
  • I reverted the unrelated whole-method torch.compiler.disable change. This PR no longer claims to solve direct-cache recompilation.
  • Added real NCCL regression tests for SP2 FP32 direct execution, SP4 BF16 direct execution, and TP2×SP2 BF16 paged execution.They compare outputs and text K/V against SP1 across repeated denoise calls and cache-window updates.

@wtz2333

wtz2333 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author

SP4 chunk latency: 1 warmup + 3 measured requests

Full-window chunk median: 730.06 ms. Each chunk contains 3 latent frames, with four denoise steps plus one KV commit (five transformer forwards).

Correction: the previously highlighted 8,739.08 ms was the full 81-frame video-request median, not one chunk. The chunk results below come directly from the per-block CUDA events recorded during that same run; they are not the request latency divided by seven. No additional benchmark run was performed for this correction.

Commit: 3ce6cfb39e1ea1b5dba655ed486cb47e4805e371. Ordinary offline video generation, SP4/TP1, batch=1 and concurrency=1; torch.compile requested (enforce_eager=False, regional blocks, dynamic=True).

  • Hardware: 4 × NVIDIA L20X; allocated GPU IDs 0,1,3,4; driver 570.133.20.
  • Model: robbyant/lingbot-world-v2-14b-causal-fast-diffusers, checkpoint revision 59cccf49f2d2dd27418ae7a04b82b10868d455c2; dtype torch.bfloat16.
  • Workload: 480×832, 81 output frames, 16 FPS export, seed 42, forward camera trajectory, text length limit 512, flow shift 5.0. Four DMD denoise steps per AR block plus cache commit; seven AR blocks per request.
  • Attention: self=FLASH_ATTN, cross=FLASH_ATTN.
  • Software: torch=2.13.0, vllm=0.28.0, diffusers=0.40.0, transformers=5.14.0.
  • Chunk timing: CUDA events around pipeline._generate_block, taking the maximum rank duration for each chunk. Includes latent initialization, four DMD denoise steps and one KV commit. Excludes text encoding/input preprocessing outside this call, VAE encode/decode, and video export. No profiler was enabled.

An 81-frame request contains seven AR chunks. To keep growing-context and full-window costs separate, here are all chunk positions from the three measured requests (warmup excluded):

Chunk Start latent frame Test 1 (ms) Test 2 (ms) Test 3 (ms) Median (ms)
1 0 545.64 552.90 547.49 547.49
2 3 550.47 550.89 552.22 550.89
3 6 593.02 589.40 593.13 593.02
4 9 643.86 644.77 646.01 644.77
5 12 685.21 691.52 687.66 687.66
6 15 730.89 733.00 729.24 730.89
7 18 726.31 733.75 728.00 728.00

The configured attention window is 18 latent frames. Chunks 6 and 7 reach that window size; pooling these two chunk positions across the three measured requests gives six full-window chunk observations:

  • Median: 730.06 ms/chunk.
  • Mean: 730.20 ms/chunk.
  • Range: 726.31–733.75 ms/chunk.
  • If comparing exactly one fixed chunk position, the final chunk (start latent frame 18) has a three-run median of 728.00 ms.

These are chunk-generation GPU event durations, not decoded-video streaming delivery latency. The six observations come from three requests and should not be treated as six independent full-request trials.

No new Dynamo graphs were observed in the measured requests and no compile-limit/fallback message was detected. Repeated measured latents and returned video arrays were bit-identical; the final 81-frame MP4 decoded successfully. This validates repeatability, not SP1 parity.

Whole-request timing, memory, and reproduction details (secondary metrics)
Request Video-request latency (ms)
Warmup (excluded) 121910.55
Test 1 8764.11
Test 2 8739.08
Test 3 8713.89
Measured statistic Value
Mean 8739.03 ms
Median 8739.08 ms
Min / max 8713.89 / 8764.11 ms
Sample standard deviation / CV 25.11 ms / 0.29%
Serial throughput 0.1144 requests/s; 9.27 output frames/s

Peak allocator memory across the three measured requests, per rank (GiB):

Rank Allocated Reserved
0 59.219 62.277
1 59.602 62.209
2 59.602 62.209
3 59.602 62.209

No new Dynamo graphs were observed during the three measured requests, and no compile-limit/fallback message was detected.

All 3 measured requests returned finite video arrays of the expected shape. Repeated latent tensors were bit-identical: True; repeated returned video arrays were bit-identical: True. The final MP4 decoded successfully as 81 frames at 832×480 and 16 FPS. This checks repeatability, not SP1 parity.

This is a three-sample SP4 rerun, not a fresh SP1/SP4 A/B or a tail-latency study. No speedup ratio is inferred from older measurements. The retained implementation uses compact head-sharded cross-attention K/V; the earlier no-exchange microbenchmark is not the implementation measured here.

Workload command and repetition protocol
python -m examples.offline_inference.diffusion.lingbot_world_v2 \
  --model /path/to/lingbot-world-v2-14b-causal-fast-diffusers \
  --image /path/to/first_frame.jpg --action-dir /path/to/actions/forward81 \
  --prompt "The camera moves slowly forward through the scene." \
  --height 480 --width 832 --num-frames 81 --fps 16 --seed 42 \
  --flow-shift 5.0 --tensor-parallel-size 1 --ulysses-degree 4 \
  --output sp4.mp4

An external measurement wrapper invokes this committed entry point once and repeats its real Omni.generate call on the same engine: exactly one warmup followed by three measured requests, each with fresh sampling state and the same seed. There are no benchmark flags or benchmark changes in the PR. Allocator peaks are reset before each request; rank statistics are collected after the timed call.

@tzhouam tzhouam added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 11, 2026

@tzhouam tzhouam left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
@tzhouam

tzhouam commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

@vllm-omni-review-bot

@wtz2333
wtz2333 force-pushed the codex/lingbot-world-sp-vae-parallel branch from 3ce6cfb to c019972 Compare September 11, 2026 07:16
@tzhouam tzhouam added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 11, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict under experiment vllm-omni-strict-5050-20260829.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Scan:

Category Result
Tests / verification 9 finding(s) below
Security no finding reported
Docs / comments no finding reported
Behavior / compatibility 1 finding(s) below
Correctness 1 finding(s) below

Validated:

  • [resolved] advanced_uaa fail-closed: pipeline.py:218/222 ulysses_mode != "strict" — residual: ulysses_a2a_permute still unvalidated.
  • [resolved] Real collectives: test_sp2/sp4/tp2_sp2_* use native init+load+rollout; residual=enforce_eager=True, not merge-gated after c019972.
  • [resolved] tzhouam no real all-to-all: GPU tests exist at test_lingbot_world_transformer.py:978-993; residual is CI placement not absence
  • [resolved] Real collectives coverage: test_sp2_direct_matches_sp1_fp32 / sp4 / tp2_sp2_paged (transformer tests:978-993) — residual: hardware_test L4 lanes not in merge-gate grep of .buildkite.
  • [resolved] non-SP timestep stays frame-sized (_LingBotSPPrepare:156); residual=compile/eager evidence mismatch (kept).
  • [claim-verified] Ring/AllGather/VAE SP unsupported: sequence_parallel removed from size bans; hybrid/allgather/ring/advanced_uaa covered by test_unsupported_sp_config.

Keep the primary fail-closed and KV-geometry defects (heads/token divisibility early gates; config-vs-live num_kv_heads; reject ignored ulysses_a2a_permute), plus the CI placement test-integrity issue for multi-GPU SP tests and a nit on compile-labeled results vs eager-only SP parity. Drop secondary doc/example polish and merge the duplicate permute-gate comments onto one survivor.

Verdict: REQUEST CHANGES

Findings

  • **[P1] This PR switches ar_diffusion_kv_cache_spec from live TP to od_config.parall…** — vllm_omni/diffusion/models/lingbot_world/pipeline.pyThis PR switchesar_diffusion_kv_cache_specfrom live TP tood_config.parallel_config TP×Ulysses (pipeline.py:574-598: num_kv_heads=num_tp_heads // ulysses_degree), while allocate_cachestill uses liveblocks[0].self_attn.num_sp_heads (transformer.py:1010) and the runner sizes the paged pool from spec.num_kv_heads (runner.py:138). Attention freezes num_sp_headsfromget_sp_group() at construct (transformer.py:194-200). Config/live skew—including test_ar_diffusion_capability_uses_config_local_head_geometrysettingulysses_degree=2without patchingget_sp_group—can advertise fewer KV heads than post-A2A attention uses. DreamZero still uses live heads (pipeline_dreamzero.py:143). Prefer num_kv_heads=int(self.transformer.blocks[0].self_attn.num_sp_heads)` and make the capability test patch SP/TP so both sources agree.

Evidence: pipeline.py:575-598 tp_size = getattr(parallel_config, "tensor_parallel_size", 1) / ulysses_degree = getattr(parallel_config, "ulysses_degree", 1) or 1 / num_kv_heads=num_tp_heads // ulysses_degree,; transformer.py:194-200 self.ulysses_world_size, self.ulysses_rank, self.ulysses_group = _ulysses_state() / self.num_sp_heads = self.num_local_heads // self.ulysses_world_size; transformer.py:1010 num_local_heads = self.blocks[0].self_attn.num_sp_heads; unchanged by this diff, present in the PR-time tree: runner.py:138 num_kv_heads=spec.num_kv_heads,; unchanged by this diff, present in the PR-time tree: dreamzero/pipeline_dreamzero.py:143 num_kv_heads=int(transformer.blocks[0].self_attn.tp_num_heads),; test_pipeline_lingbot_world.py:447-453 sets ulysses_degree = sp_size and asserts spec.num_kv_heads == 8 // tp_size // sp_size with no get_sp_group patch.

Suggestion: return ARDiffusionKVCacheSpec(
num_layers=int(self.transformer.config.num_layers),
num_kv_heads=int(self.transformer.blocks[0].self_attn.num_sp_heads),
head_size=int(self.transformer.config.attention_head_dim),

  • **[P2] This PR's pure-Ulysses gate (pipeline.py:217–222) rejects Ring/AllGather/advan…** — vllm_omni/diffusion/models/lingbot_world/pipeline.py This PR's pure-Ulysses gate (pipeline.py:217–222) rejects Ring/AllGather/advanced_uaabut still acceptsulysses_a2a_permute=True. LingBot builds Attention with skip_sequence_parallel=True(so the shared Ulysses strategy that reads that knob never runs) and performs Q/K/V exchange via manualSeqAllToAll4D(no a2a_permute path). Extend the same fail-closed conjunction—andtest_unsupported_sp_config_fails_before_component_loading—to reject ulysses_a2a_permute`.

Evidence: pipeline.py:217-222 ulysses_mode = getattr(parallel_config, "ulysses_mode", "strict") / if ( sequence_parallel_size != ulysses_degree or ring_degree != 1 or allgather_degree != 1 or ulysses_mode != "strict" ): — no ulysses_a2a_permute check. Unchanged by this diff, present in the PR-time tree: data.py:228-229 ulysses_a2a_permute: bool = False / "Use fused permute-free all-to-all for eligible strict Ulysses exchanges.". transformer.py:222 skip_sequence_parallel=True,. Unchanged by this diff, present in the PR-time tree: layer.py:283-284 if self.skip_sequence_parallel: return self._no_parallel_strategy. transformer.py:332-334 query = SeqAllToAll4D.apply(self.ulysses_group, query, 2, 1, False) (and same for key/value) — manual A2A, not the ulysses.py path that takes ulysses_a2a_permute. test_pipeline_lingbot_world.py:642-648 override matrix lists ring/allgather/advanced_uaa/None degree but not ulysses_a2a_permute.

Suggestion: ulysses_a2a_permute = bool(getattr(parallel_config, "ulysses_a2a_permute", False))
if (
sequence_parallel_size != ulysses_degree
or ring_degree != 1
or allgather_degree != 1
or ulysses_mode != "strict"
or ulysses_a2a_permute
):

  • [P2] Minor residual only: test_sp2/sp4/tp2_sp2_ already use native NCCL init + load…* — ``
    Minor residual only: test_sp2/sp4/tp2_sp2_* already use native NCCL init + load_weights + _rollout in _worker; OmniDiffusionConfig still hard-codes enforce_eager=True (line 924), so graph-mode SP parity is not covered and is not merge-gated.

Evidence: test_lingbot_world_transformer.py:907-909 torch.distributed.init_process_group(\n "nccl", init_method=f"file://{rendezvous}", world_size=world_size, rank=rank, timeout=timedelta(seconds=90)\n ); :915-920 initialize_model_parallel(... ulysses_degree=sp, tensor_parallel_size=tp); :921-924 config = OmniDiffusionConfig(\n model=str(Path(rendezvous).parent),\n dtype=dtype,\n enforce_eager=True,; :938-940 model.load_weights(iter(weights.items()))\n _apply_sequence_parallel_if_enabled(...)\n result, cross = _rollout(model, mode, dtype, batch); :980-993 def test_sp2_direct_matches_sp1_fp32 / test_sp4_direct_matches_sp1_bf16 / test_tp2_sp2_paged_matches_sp1_bf16 all call _run(...).

  • [P2] Real multi-GPU Ulysses SP coverage exists at test_lingbot_world_transformer.py:… — test_lingbot_world_transformer.py:978
    Real multi-GPU Ulysses SP coverage exists at test_lingbot_world_transformer.py:978-993 (_worker imports live CausalLingBotWorldTransformer3DModel + NCCL SP); identity SeqAllToAll4D is only the CPU stub path. Residual is CI placement of these @parallel/@hardware_test cases, not absence of real all-to-all tests.

Evidence: tests/diffusion/models/lingbot_world/test_lingbot_world_transformer.py:978-982 @pytest.mark.parallel / @hardware_test(res={"cuda": "L4"}, num_cards=2) / def test_sp2_direct_matches_sp1_fp32(tmp_path): / _run(tmp_path, 2, 1, "direct", torch.float32, 2) — multi-GPU SP regression present. tests/diffusion/models/lingbot_world/test_lingbot_world_transformer.py:773-774 from vllm_omni.diffusion.models.lingbot_world.transformer import CausalLingBotWorldTransformer3DModel — GPU worker loads real transformer, not stubbed attention module. unchanged by this diff, present in the PR-time tree: vllm_omni/diffusion/distributed/comm.py:117 return all_to_all_4D(input, scatter_idx, gather_idx, group=group, use_sync=use_sync) and :51 dist.all_to_all_single(output, input_t, group=group) — production SeqAllToAll4D is a real collective. Identity stub only in tests/diffusion/models/lingbot_world/test_lingbot_world_attention.py:65-68 _SeqAllToAll4D / return value for CPU unit stubs.

  • [P2] Document (or extend recipe/deploy beyond devices:"0") that LingBot USP rides sh… — async_omni_engine.py:1067
    Document (or extend recipe/deploy beyond devices:"0") that LingBot USP rides shared DiffusionParallelConfig via serve --usp → ulysses_degree and async_omni_engine sequence_parallel_size←ulysses_degree (× ring when unset); LingBot rejects non-pure Ulysses. Recipe/deploy still pins one GPU and serving/e2e stepwise tests are num_cards=1 — only the LingBot offline/GPU SP matrix was validated.

Evidence: unchanged by this diff, present in the PR-time tree: async_omni_engine.py:1067-1068 if sequence_parallel_size is None: / sequence_parallel_size = allgather_degree if allgather_degree > 1 else ulysses_degree * ring_degree; unchanged: serve.py:526-528 --usp / dest="ulysses_degree"; unchanged: lingbot_world_v2_stepwise.yaml:17 devices: "0"; in-diff: pipeline.py:218-224 sequence_parallel_size != ulysses_degree … raises NotImplementedError pure Ulysses only; unchanged e2e: test_lingbot_world_v2_stepwise.py (online) @hardware_test(..., num_cards=1).

  • [P2] Strip or qualify the PR What Changes rows that claim an optional repeated-video… — lingbot_world_v2.py:102
    Strip or qualify the PR What Changes rows that claim an optional repeated-video benchmark (per-rank memory/frame hashes) and example-test warmup/statistics coverage: at this head the example only adds --ulysses-degree (lingbot_world_v2.py:102, :171) and test_lingbot_world_v2.py only asserts that flag (author already notes the harness is pending).

Evidence: examples/offline_inference/diffusion/lingbot_world_v2.py:102 parser.add_argument("--ulysses-degree", type=int, default=1, help="Pure Ulysses sequence parallel degree."); examples/offline_inference/diffusion/lingbot_world_v2.py:171 "ulysses_degree": args.ulysses_degree, — kwargs contain only that SP flag, no benchmark/warmup path; tests/examples/offline_inference/test_lingbot_world_v2.py:13-24 sole test test_offline_ulysses_argument asserts CLI/kwargs for ulysses_degree only; PR #6841 What Changes (unchanged by this tree) still claims lingbot_world_v2.py “optional repeated video benchmark…” and the example test “warmup exclusion, statistics…”, with author note that those are a pending local extension.

allgather_degree = getattr(parallel_config, "allgather_degree", 1) or 1
ulysses_mode = getattr(parallel_config, "ulysses_mode", "strict")
if (
sequence_parallel_size != ulysses_degree

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] _validate_parallel_config (pipeline.py:212–228) only allowlists pure Ulysses…

_validate_parallel_config (pipeline.py:212–228) only allowlists pure Ulysses (sequence_parallel_size == ulysses_degree, ring/allgather=1, ulysses_mode="strict"). Pipeline __init__ then prefetches components (482) before building the transformer (512), so illegal degrees still start load work. Official 40 heads fail only later in LingBotSelfAttention when num_local_heads % ulysses_world_size (transformer.py:195–199)—e.g. ulysses=3. Separately, request H/W is only checked for VAE×patch alignment (pipeline.py:742–745); a legal 16-aligned 240×240 yields 675 tokens/block (3×15×15) and fails only at global_tokens % sp_size (transformer.py:1247–1248) for SP=2. Mirror MAGI’s early heads % (TP*SP) gate (pipeline_magi2.py:275–279) in _validate_parallel_config, and reject SP-indivisible block token counts beside the existing height/width divisor checks—do not add Wan-style auto_pad.

Evidence: pipeline.py:218-222 sequence_parallel_size != ulysses_degree or ring_degree != 1 or allgather_degree != 1 or ulysses_mode != "strict" — pure-Ulysses allowlist only, no heads%(TP×SP). pipeline.py:460 then :482 _validate_parallel_config(od_config) then prefetch_subfolders(...) — illegal SP still prefetches. transformer.py:786 num_attention_heads: int = 40 and :195-198 if self.num_local_heads % self.ulysses_world_size: raise ValueError(...) — head reject deferred to attention ctor. pipeline.py:742-745 if height % height_divisor / if width % width_divisor — no SP token-shard gate. transformer.py:1247-1248 if global_tokens % sp_size: raise ValueError("LingBot Ulysses requires the token count to be divisible by its degree."). unchanged by this diff, present in the PR-time tree: magi2/pipeline_magi2.py:275-279 attention_shards = tp_size * sp_size / if MAGI2_PREVIEW_CONFIG.num_heads_q % attention_shards: raise ValueError(...).

Suggestion: sequence_parallel_size != ulysses_degree
or ring_degree != 1
or allgather_degree != 1
or ulysses_mode != "strict"
):
raise NotImplementedError(
"LingBot World sequence parallelism currently supports pure Ulysses only: "
f"sequence_parallel_size={sequence_parallel_size}, ulysses_degree={ulysses_degree}, "
f"ring_degree={ring_degree}, allgather_degree={allgather_degree}, ulysses_mode={ulysses_mode!r}."
)
tp_size = getattr(parallel_config, "tensor_parallel_size", 1) or 1
if 40 % (tp_size * ulysses_degree):
raise NotImplementedError(
"LingBot World v2 attention heads (40) must be divisible by "
f"tensor_parallel_size * ulysses_degree; got tp={tp_size}, ulysses={ulysses_degree}."
)

@pytest.mark.parallel
@hardware_test(res={"cuda": "L4"}, num_cards=2)
def test_sp2_direct_matches_sp1_fp32(tmp_path):
_run(tmp_path, 2, 1, "direct", torch.float32, 2)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Change test_sp2_direct_matches_sp1_fp32 to batch=1 (or add a batch=1 SP2 arm):…

Change test_sp2_direct_matches_sp1_fp32 to batch=1 (or add a batch=1 SP2 arm): with batch=2, LingBotSelfAttention takes the else branch at transformer.py:394–395 (self.attn) because query.shape[0]!=1, so SP2 never hits the CUDA flash path that SP4 batch=1 and paged TP2SP2 already cover.

Evidence: tests/diffusion/models/lingbot_world/test_lingbot_world_transformer.py:981 _run(tmp_path, 2, 1, "direct", torch.float32, 2) — SP2 direct uses batch=2; :987 _run(tmp_path, 4, 1, "direct", torch.bfloat16, 1) and :993 _run(tmp_path, 2, 2, "paged", torch.bfloat16, 1) use batch=1. Unchanged-by-diff consumer in PR-time tree: vllm_omni/diffusion/models/lingbot_world/transformer.py:359 if query.is_cuda and query.shape[0] == 1: then flash via ar_diffusion_paged_attention; :394-395 else: output = self.attn(query, visible_key, visible_value). Worker forces SDPA backend at test file:931 diffusion_attention_config=AttentionConfig(default=AttentionSpec(backend="TORCH_SDPA")).

@tzhouam tzhouam added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 11, 2026
@tzhouam
tzhouam merged commit 02aaa34 into vllm-project:main Sep 11, 2026
6 of 9 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models ready label to trigger buildkite CI world model

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants