Skip to content

[Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution - #6408

Merged
linyueqian merged 17 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/omnivoice_batch
Sep 9, 2026
Merged

linyueqian merged 17 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/omnivoice_batch

Conversation

@sphinxkkkbc

@sphinxkkkbc sphinxkkkbc commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Support request batching and step execution for OmniVoice by implementing model-specific helpers following the existing execute_stepwise() lifecycle.

Bit-exact parity with outputs generated before this PR is not guaranteed because the attention backend and packed execution layout have changed. WER results are listed below. The OmniVoice decoder is unchanged from main, and each request is decoded separately using exact-length codec tokens.

An earlier version of this PR implemented padded Decoder batching, but it was later reverted. With the packed varlen Generator, generated codec tokens are now kept at exact per-request lengths, so batching Decoder execution would require padding them again to the longest request. The current implementation therefore decodes completed requests individually using exact-length codec tokens. I have not benchmarked padded Decoder batching against per-request exact-length decoding, so it is unclear which approach is faster.

1. Add packed variable-length attention to the OmniVoice generator

The previous Generator layout stored conditional and unconditional inputs as separate rows. Since their sequence lengths differ, unconditional inputs had to be padded to the conditional sequence length. In a request batch, shorter
requests were also padded to the longest request, introducing substantial attention and Transformer overhead.

This PR changes Generator inputs to a packed, token-major 2D layout:

[cond0, uncond0, cond1, uncond1, ...]

with tensor shapes:

input_ids:     [total_seq_len, num_codebooks]
audio_mask:    [total_seq_len]
hidden_states: [total_seq_len, hidden_size]
logits:        [num_codebooks, total_seq_len, audio_vocab_size]

For B requests, CFG produces 2B independent sequences. Their cumulative offsets are represented by cu_seqs, allowing PyTorch varlen attention to keep requests and cond/uncond branches mutually invisible while retaining full bidirectional attention within each sequence.

Position IDs are reset independently for every packed sequence:

[0..cond0_len-1, 0..uncond0_len-1, 0..cond1_len-1, 0..uncond1_len-1, ...]

_position_ids_from_cu_seqs() derives these sequence-local positions directly from cu_seqs.

Per-request random generators are retained so one request does not consume or change another request's Gumbel sampling stream.

2. Make packed varlen attention compatible with CUDA Graphs

OmniVoice already supported CUDA Graphs, but the previous graph inputs and keys were based on padded [batch, codebook, sequence] tensors.

CUDA Graph keys are now: (request_batch_size, token_bucket)

Captured static inputs use:

input_ids:  [token_bucket, num_codebooks]
audio_mask: [token_bucket]
cu_seqs:    [2 * request_batch_size + 2]

The 2B+2 cu_seqs layout represents:

  • 2B real cond/uncond sequences;
  • the initial zero offset;
  • one additional tail sequence used for Graph bucket padding.

In eager mode, the tail has zero length: [..., real_total, real_total]

In Graph mode, the last offset is changed to the selected bucket: [..., real_total, token_bucket]

This keeps the metadata shape static while isolating padded tail tokens from all real requests.

max_q and max_k use the fixed packed token bucket during graph capture, so the real maximum sequence length does not need to be part of the graph key.

Common request batch sizes derived from default config are pre-captured. Other batch sizes and oversized token shapes are captured lazily.

3. Support request-batch and step-execution lifecycles

Request-batch mode packs the input_ids and audio_mask values produced by _prepare_request_input() along the token dimension. cond_lens and target_lens are used to construct cu_seqs once before Transformer execution and to locate each request's target logits and generated codec tokens.

Step-execution stores each request's persistent latents as: [request_seq_len, num_codebooks]

InputBatch.make_batch() therefore concatenates active requests directly along dimension 0 without model-specific repadding. denoise_step() constructs the packed metadata, runs one varlen Generator step, and returns the same 2D layout so the Runner can scatter rows back to individual request states.

The standard lifecycle methods are used:

  • prepare_encode() initializes request-local state and sampling metadata;
  • denoise_step() executes one packed Generator step;
  • step_scheduler() persists the request slice and advances its step;
  • post_decode() decodes the completed request.

Generated codec tokens are maintained at exact per-request lengths.

Test Plan

Unit Tests

pytest -v \
  tests/entrypoints/openai_api/test_serving_speech.py \
  tests/model_executor/models/omnivoice/test_cuda_graph_generator.py \
  tests/model_executor/models/omnivoice/test_pipeline_batching.py \
  tests/model_executor/models/omnivoice/test_mask_dtype.py \
  tests/model_executor/models/omnivoice/test_fused_projection_load.py

The CUDA Graph tests continue the validation coverage from the earlier implementation for packed Generator. Additional tests were added for the varlen layout to verify dynamic cu_seqs updates across replays of the same graph key, and eager/Graph parity.

E2E Tests

pytest -v \
  tests/e2e/online_serving/test_omnivoice_expansion.py \
  tests/e2e/online_serving/test_omnivoice_parity.py

The E2E tests verify:

  • request-batch execution correctness test;
  • step execution generates valid audio without an error response, batch step-execution correctness test;
  • B=1 step execution produces the same seeded WAV bytes as B=1 request-mode execution.

Perf and Acc

  • request-rate=inf: batch vs per-req
  • request-rate=1: batch vs step-execution

Serve Command:

vllm serve k2-fsa/OmniVoice \
  --omni \
  --port 8091 \
  --trust-remote-code \
  --max-num-seqs 8 \
  --request-batch-max-wait-ms 50 \
  --step-execution

Remove --step-execution for benchmarking request-batch.

Benchmark Command:

vllm bench serve --omni \
  --host 127.0.0.1 \
  --port 8091 \
  --model k2-fsa/OmniVoice \
  --backend openai-audio-speech \
  --endpoint /v1/audio/speech \
  --dataset-name seed-tts-text \
  --dataset-path /path/to/seedtts_testset \
  --seed-tts-locale en \
  --num-prompts 32 \
  --num-warmups 2 \
  --extra-body '{
    "language": "English",
    "seed": 42,
    "extra_params": {
      "num_inference_steps": 32,
      "guidance_scale": 2.0
    }
  }' \
  --max-concurrency 4 \
  --seed-tts-wer-eval \
  --request-rate inf \
  --percentile-metrics e2el,audio_rtf,audio_ttfp,audio_duration \
  --save-result \
  --result-dir ./omnivoice_results

Environment:

  • vLLM Version: 0.28.0
  • vLLM-Omni Commit: c4197ffe163c23017f179bf2f766c67766aca249

Test Results

Unit Test: 1 failed, 302 passed. The single failure is unrelated to this PR and is not addressed here.

FAILED tests/entrypoints/openai_api/test_serving_speech.py::TestTTSAsyncOffloading::test_prepare_speech_generation_awaits_voxtral_async - ValueError: Invalid speaker 'test'. Supported: bar

E2E : 14 passed

Performance

1. Batch vs Per-Request

C=4

Metrics C=4 Batch C=4 per-req ∆
Benchmark duration (s) 15.96 24.04 -33.61%
Request throughput (req/s) 2.00 1.33 +50.38%
Mean E2EL (ms) 1993.46 2852.22 -30.11%
Median E2EL (ms) 2257.81 2920.93 -22.70%
P99 E2EL (ms) 2290.35 3400.14 -32.64%
Audio throughput (audio duration/s) 7.35 4.88 +50.61%
Total Token throughput (tok/s) 26.12 17.35 +50.55%
Mean AUDIO_RTF 0.58 0.84 -30.95%
Median AUDIO_RTF 0.56 0.78 -28.21%
P99 AUDIO_RTF 0.95 1.57 -39.49%
Mean AUDIO_TTFP (ms) 1991.95 2851.77 -30.15%
Median AUDIO_TTFP (ms) 2256.20 2920.65 -22.75%
P99 AUDIO_TTFP (ms) 2288.25 3399.69 -32.69%

C=8

Metrics C=8 Batch C=8 per-req ∆
Benchmark duration (s) 14.75 24.44 -39.65%
Request throughput (req/s) 2.17 1.31 +65.65%
Mean E2EL (ms) 3579.96 5416.61 -33.91%
Median E2EL (ms) 3575.87 6025.53 -40.65%
P99 E2EL (ms) 4092.70 6495.25 -36.99%
Audio throughput (audio duration/s) 7.96 4.80 +65.83%
Total Token throughput (tok/s) 28.28 17.06 +65.77%
Mean AUDIO_RTF 1.08 1.60 -32.50%
Median AUDIO_RTF 1.02 1.54 -33.77%
P99 AUDIO_RTF 1.98 3.42 -42.11%
Mean AUDIO_TTFP (ms) 3577.48 5416.21 -33.95%
Median AUDIO_TTFP (ms) 3573.22 6025.31 -40.70%
P99 AUDIO_TTFP (ms) 4091.90 6494.96 -37.00%

2. Batch vs Step-Execution (C=4, request-rate=1)

Metrics Step-Execution Request-batch ∆
Benchmark duration (s) 32.73 32.70 -0.09%
Request throughput (req/s) 0.98 0.98 0.00%
Mean E2EL (ms) 981.70 943.41 -3.90%
Median E2EL (ms) 934.53 785.74 -15.92%
P99 E2EL (ms) 1764.82 1974.45 +11.88%
Audio throughput (audio duration/s) 3.59 3.59 0.00%
Total Token throughput (tok/s) 12.74 12.75 +0.08%
Mean AUDIO_RTF 0.28 0.27 -3.57%
Median AUDIO_RTF 0.24 0.21 -12.50%
P99 AUDIO_RTF 0.84 0.98 +16.67%
Mean AUDIO_TTFP (ms) 981.21 942.79 -3.92%
Median AUDIO_TTFP (ms) 934.14 785.35 -15.93%
P99 AUDIO_TTFP (ms) 1764.36 1974.01 +11.88%

Accuracy

All request passed for each benchmark with 32 requests. All benchmark result with Median WER=0.0000

Mode Mean WER
Baseline (Original) 0.0074
Request-batch + varlen (C=4, inf) 0.0074
Request-batch + varlen (C=8, inf) 0.0074
Step-execution (C=4, rate=1) 0.0048
Request-batch (C=4, rate=1) 0.0096

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.

(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md.

Module owners: @tzhouam @gcanlin

@sphinxkkkbc, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added tts code related to tts models enhancement New feature or request labels Aug 21, 2026
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/omnivoice_batch branch 2 times, most recently from 8117d69 to f7cf964 Compare August 23, 2026 12:21
@sphinxkkkbc sphinxkkkbc changed the title [Feat][OmniVoice]Support Request-Batch and Step-Execution [Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution Aug 23, 2026
@sphinxkkkbc
sphinxkkkbc force-pushed the feat/omnivoice_batch branch from aaf8a26 to 7a096c3 Compare August 23, 2026 18:35
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian could you help review this PR? Thanks!

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the writeup, the packed varlen layout and the 2B+2 cu_seqs trick with the padded tail sequence are a clean way to keep the graph metadata static. Numbers on the request-batch path look convincing.

Three blocking items and a few smaller ones, all inline.

Blocking

  1. pipeline_omnivoice.py:584: declaring supports_request_batch = True routes through execute_model_batch, where allow_single_output=False. Returning a bare DiffusionOutput on a per-request validation error now raises RuntimeError for every batch size including B=1, where main returned a clean user error.
  2. pipeline_omnivoice.py:376: in step mode the same error return is discarded by the runner, state.latents stays None, and _prepare_latents raises outside the per-request try, failing every concurrent request in the batch.
  3. omni_base.py:182: defaulting diffusion_batch_size to max_num_seqs changes startup behavior for all 51 diffusion pipelines. 43 of them lack supports_request_batch and will now fail to start under --max-num-seqs > 1 where they previously ran serially. This is unrelated to the PR title and has no test.

Test coverage

test_cuda_graph_generator.py is the only core_model test, and its _SyntheticGenerator never calls _varlen_attn, so it covers the padding, slicing and replay bookkeeping but not the attention change itself. The e2e tests are slow plus tts plus L4, so they only run in the weekly sweep, which does pick up the new test_omnivoice_parity.py automatically through the marker-driven collection in .buildkite/cuda/test-weekly.yml:93. test_omnivoice_parity.py also sets OMNIVOICE_CUDA_GRAPH=0 on both sides, so the CUDA-graph plus varlen combination, which is the riskiest part of this change, has no automated coverage at all. Could you add a core_model test that drives a small real OmniVoiceGenerator through _varlen_attn in both eager and graph mode, including a replay with a nonzero padded tail?

Reviewed at 7a096c39668eb3295c016f792a16d4b3ea21f577. Static review only, no GPU run on my side.

Comment thread vllm_omni/entrypoints/omni_base.py Outdated
diffusion_batch_size: int = kwargs.pop("diffusion_batch_size", 1)
diffusion_batch_size = kwargs.pop(
"diffusion_batch_size",
kwargs.get("max_num_seqs", 1),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] This changes the diffusion batch default for every pipeline, not just OmniVoice, and the PR body does not mention it.

diffusion_batch_size is force-assigned onto od_config.max_num_seqs at stage_init_utils.py:1444 and stage_engine_startup.py:1561, overwriting whatever the stage config resolved. With the old default of 1, --max-num-seqs N was silently ignored by diffusion stages. Now it lands, which is what makes the --max-num-seqs 8 e2e test in this PR work, and it does match docs/user_guide/diffusion/execution_modes.md:232.

The side effect: only 8 of 51 diffusion pipelines set supports_request_batch = True, and diffusion_engine.py:229 raises at startup when max_num_seqs > 1 without it. Serve commands that run today (serially) will fail to start after this. It also still clobbers a per-stage engine_args.max_num_seqs with the global CLI value.

Please split this into its own PR with a unit test in tests/engine/test_async_omni_engine_stage_init.py, or at minimum call it out in the body and in the execution-modes doc.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change was removed in d5f0bfb. #6405 is already opened, so we can support that PR.

Also, #6525 resolved the issue where max_num_seqs was being overridden to 1. Now we can configure this in deploy.yaml. I've set the default max_num_seqs to 8 in 3354dbf.

extra = request.sampling_params.extra_args or {}
prepared = self._prepare_request_input(prompt, extra)
if isinstance(prepared, DiffusionOutput):
return prepared

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] Returning a bare DiffusionOutput here now raises RuntimeError instead of surfacing the user error.

Declaring supports_request_batch = True routes OmniVoice through execute_model_batch, which calls _normalize_pipeline_outputs(..., allow_single_output=False) (diffusion_model_runner.py:650). That helper raises

RuntimeError: OmniVoicePipeline.forward returned a single DiffusionOutput;
request-batch forward must return list[DiffusionOutput].

at diffusion_model_runner.py:84, for every batch size including B=1. On main OmniVoice went through execute_model with allow_single_output=True, so the same prompt produced a clean per-request error.

Trigger: offline Omni.generate with two clips in multi_modal_data["audio"] (line 258 above), or a dict prompt with no text key (line 276).

Fix direction: allocate outputs = [None] * len(req.requests) up front, place the error DiffusionOutput in that request's slot, and keep the remaining requests running.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved in d5f0bfb, added

outputs = [None] * len(req.requests)
prepared_indices: list[int] = []

to handle this and preserve the original request indices to prevent slot mismatch

Also added regression test in tests/model_executor/models/omnivoice/test_pipeline_batching.py for changes in pipeline_omnivoice.py

extra = state.sampling.extra_args or {}
prepared = self._prepare_request_input(prompt, extra)
if isinstance(prepared, DiffusionOutput):
return prepared

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] In step mode this error return is discarded, and the malformed request takes down the whole step batch.

diffusion_model_runner.py:714 calls prepare_encode(state) for its side effects only and drops the return value. When _prepare_request_input returns an error, state.latents is never assigned, so InputBatch.make_batch hits _prepare_latents and raises ValueError("All requests must have 'latents' initialized.") (input_batch.py:347). That happens at execute_stepwise line 776, outside the per-request try at line 810, so one bad request fails every concurrent request in the batch.

The annotation on line 371 also matches neither the interface (interface.py:60 declares -> StepRequestState) nor the success path, which returns None. Suggest recording the error on the state and letting the runner emit it per request, then fixing the annotation.

@sphinxkkkbc sphinxkkkbc Aug 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was initially resolved in d5f0bfb, but after merging with the latest main, I found that #5810 already fixed it a few days ago. I've reverted it in f78c290 and now directly using the handling from main.

Annotation fixed.

# Default bucket count is 10; 16 gives modest headroom for edge cases
# (seq_len > max bucket or non-CFG batch) without unbounded GPU growth.
_MAX_LAZY_GRAPHS: int = 16
_DEFAULT_CAPTURE_BATCH_SIZES: tuple[int, ...] = (1, 2, 3, 4)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Medium] The pre-warm set does not cover the concurrency this PR targets, so the hot path falls into exact-size lazy capture.

The graph key is now (request_batch_size, packed_token_total), but this tuple is hard-coded to (1, 2, 3, 4) and the largest bucket is 1024 (configs/omnivoice.py:88).

Worked example with the parity-test prompt: RuleDurationEstimator gives target_len about 80, so the packed total per request is about 200 (cond_len about 120 plus target_len 80).

  • B=4 gives about 800, bucket 1024, fine.
  • B=8 gives about 1600, past every bucket. _find_bucket returns None and line 629 sets bucket = seq_len exactly.

So every distinct combination of request lengths becomes its own lazy capture of a 28-layer graph on the hot path, thrashing the 16-entry cache. B in {5..8} is never pre-warmed at all, and tests/e2e/online_serving/test_omnivoice_expansion.py:34 runs this model with --max-num-seqs 8.

Fix direction: derive the capture batch sizes from max_num_seqs, and round oversized totals up to a capped ladder instead of capturing at the exact packed length.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved in d5f0bfb6ae982bd21f73d3e4756aad458202f637, now _OmniVoiceCUDAGraphForward will read od_config and derive buckets and batches to be captured. I configured the buckets in a ladder pattern (shown below), which properly activates the LRU cache. With MAX_LAZY_GRAPH = 16, lazy-captured buckets are rounded up to multiples of 128. In benchmarking (using the command described), most scenarios do not trigger lazy capture, except for three specific (batch, seq_len) combinations: (7, 1574), (8, 1511), and (7, 1511). I believe the current capture strategy is reasonable.

This PR measures ~7 GiB peak memory after pre-capture, versus ~5 GiB on the main branch.

Long-text generation issues persist as noted in #6333. #6409 proposes a chunk-based solution with a per-request limit of ~15 seconds for ~375 tokens in a single forward pass, so capturing extremely long sequences is unnecessary.

B=1: [128, 192, 256, 320, 384, 448, 512]
B=2: [256, 320, 384, 448, 512, 640, 768]
B=3: [384, 448, 512, 640, 768, 896, 1024]
B=4: [512, 640, 768, 896, 1024]
B=5: [640, 768, 896, 1024]
B=6: [768, 896, 1024]
B=7: [896, 1024]
B=8: [1024]

Regression test added in tests/model_executor/models/omnivoice/test_cuda_graph_generator.py for LRU.

graph = torch.cuda.CUDAGraph()
with torch.no_grad():
with torch.cuda.graph(graph, pool=self._pool_handle):
with torch.cuda.graph(graph, pool=current_platform.get_global_graph_pool()):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Medium] This reverses a deliberate decision on main, and the class docstring above still describes the old behavior.

main used a per-instance torch.cuda.graph_pool_handle() with an explicit comment:

Lazy-init per-instance pool handle: isolates OmniVoice CUDA Graph memory from other vllm modules (unlike get_global_graph_pool which shares a single pool across all captured graphs and can cause memory aliasing when two graphs replay concurrently).

That comment is gone, but the class docstring at lines 486 to 490 still says graphs "share a single per-instance pool handle, which isolates OmniVoice CUDA memory from other vllm modules". Please either restore the per-instance pool, or state why the aliasing concern no longer applies and update the docstring to match.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reverted in d5f0bfb. This change was introduced while debugging an issue during testing, but there were two potential variables involved, and this change turned out to be one of them. After double-checking, I found no reason to keep it. I'll comment below with more details.

uncond_start = cond_end

# Extract logits for target region; upcast only the slices we actually consume.
c_logits = batch_logits[:, cond_end - t_len : cond_end, :].unsqueeze(0).to(torch.float32)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Medium] Lines 488 to 551 duplicate omnivoice_generator.py:956-1016 nearly verbatim.

The CFG fuse, the log_probs[..., mask_id] = -inf mask, the layer penalty, the Gumbel position noise, the top-k select and the cond/uncond mirror write are copied line for line, comments included. The only differences are guidance_scales[i] vs guidance_scale, generators[i] vs request_generator, and where sample_tokens comes from. Any future correctness fix has to land in both places.

Extracting one _unmask_one_request(...) helper owned by the generator would also remove denoise_step's reach into generator._prepare_embeddings, _transformer_forward, _get_logits and _cuda_graph_fwd, which are all private.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved in d5f0bfb

state.extra["tokens"] = tokens

def denoise_step(self, input_batch: InputBatch, *, states: Sequence[StepRequestState] | None = None, **kwargs: Any):
use_cuda_graph = self.generator._cuda_graph_fwd is not None

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor] Two small things in this method.

  1. This line has no input_ids.is_cuda guard, unlike the request-mode path at omnivoice_generator.py:929. Worth keeping the two consistent.
  2. Line 450 reads state.extra.get("guidance", self.guidance_scale), but prepare_encode sets state.guidance (line 420), not state.extra["guidance"]. The .get always falls through, so the per-request guidance plumbing is dead. Either write state.extra["guidance"] in prepare_encode or read state.guidance here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both resolved in d5f0bfb, regression test for Q2 added in tests/model_executor/models/omnivoice/test_pipeline_batching.py

entry["graph"].replay()

output = entry["static_output"]
if bucket is not None and bucket != seq_len:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor] bucket is not None cannot be false: line 629 assigns bucket = seq_len when _find_bucket returns None, so the check is dead and only bucket != seq_len matters.

Also, _lazy_graphs is documented as LRU but behaves as FIFO: the hit path on line 639 never calls move_to_end, so a frequently used key is still evicted by popitem(last=False).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both issues have been fixed.

Detailed comments on the _lazy_graphs FIFO issue have been inlined in the bucket issue. This was an existing unresolved issue on main.

if "num_inference_steps" in extra:
sampling.num_inference_steps = int(extra["num_inference_steps"])

if "guidance_scale" in extra:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor] This block is out of scope for OmniVoice and unvalidated.

OmniVoice ignores both fields: prepare_encode reads self.num_step and self.guidance_scale from the model config, never state.sampling. So this changes behavior only for other diffusion TTS pipelines, with no test in this PR.

Two details if it stays:

  • guidance_scale_provided = True is redundant. OmniDiffusionRequest.__post_init__ already sets it whenever guidance_scale is not None (vllm_omni/diffusion/request.py:66).
  • int(...) and float(...) run on unvalidated client input. extra_params: {"num_inference_steps": "abc"} raises ValueError and surfaces as a 500 rather than a 400.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These two changes are primarily required for scheduler correctness:

  1. StepScheduler._get_total_steps() reads sampling_params.num_inference_steps. If the value remains only in extra_args and is not promoted to the standard sampling field, the scheduler receives None and fails when evaluating int(None).

  2. Request-batch compatibility is determined from the standard sampling fields, including guidance_scale. If the request value remains only in extra_args, requests with different guidance scales may receive the same compatibility key and be incorrectly admitted to the same request batch.

Regression coverage was added in tests/entrypoints/openai_api/test_serving_speech.py

import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.nn.attention.varlen import _varlen_attn

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor] _varlen_attn is a private torch API imported at module scope, so it becomes a hard load-time requirement for OmniVoice on every platform.

What is the minimum torch version, and is torch.nn.attention.varlen present on the npu, rocm and xpu images? A short note in the module docstring plus the minimum version in the model doc would help, since there is no other use of this API in the repo.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved in 9432df3. Now using the in-tree class diffusion.attention.layer.Attention with an explicit attn_mask fallback when varlen_attn is not supported.

As for the version question, the torch version I used is the one installed with vLLM(2.13), and the minimum version that supports varlen_attn is torch 2.10.

@sphinxkkkbc

sphinxkkkbc commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Inline comments are replied. Perf number still hold.

Test coverage updates:

  1. test_omnivoice_parity.py now verifies both graph and eager modes.
  2. test_cuda_graph_generator.py now uses the small real OmniVoiceGenerator in all cases rather than the _SyntheticGenerator. The results and commands have also been updated with descriptions.

@sphinxkkkbc
sphinxkkkbc requested a review from linyueqian August 29, 2026 01:12
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian PTAL, thanks!

@linyueqian

Copy link
Copy Markdown
Collaborator

One practical follow-up to the approval, because it is approved but still cannot merge.

Auto-merge is armed, so this will land by itself once the required checks are satisfied. The blocker is that buildkite/vllm-omni is a required check and it is still showing build 14666, which is red on the mooncake test_concurrent_put_get_threaded_both_sides MD5 mismatch. That failure is not yours, as covered in the review, but GitHub does not care whose it is.

I just tried clearing it without troubling you: cycling the ready label re-fired the NPU lane, but the general lane stayed pinned to the old failed build. The general pipeline triggers on push, so a label cycle cannot replace a stale general result. Only a new commit will.

So when you have a moment, merging current main into the branch will give you a fresh general run, and the flake is intermittent so it should come back clean. That is also worth doing on its own terms, since main has moved since 70a2b19 and you have already been bitten twice by tests on main that call APIs this branch changes. I would rather you spend the push on a real rebase than on an empty commit.

Nothing else is outstanding from my side.

…ised yet)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…tion

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…_args

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…data

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
auto-merge was automatically disabled September 8, 2026 01:21

Head branch was pushed to by a user without write access

@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

Rebased cleanly with no conflicts, ran the unit tests locally with no new errors raised. The ready label needs to be cycled.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 9, 2026
@linyueqian
linyueqian merged commit cdf0278 into vllm-project:main Sep 9, 2026
6 of 9 checks passed
haic0 pushed a commit to haic0/vllm-omni that referenced this pull request Sep 9, 2026
vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
wenj-yan pushed a commit to wenj-yan/vllm-omni that referenced this pull request Sep 9, 2026
vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: wenjie.yan <wenjyan@outlook.com>
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants