Repository navigation
[Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution - #6408
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/model_integration.md. Module owners: @tzhouam @gcanlin @sphinxkkkbc, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
8117d69 to
f7cf964
Compare
aaf8a26 to
7a096c3
Compare
|
@linyueqian could you help review this PR? Thanks! |
linyueqian
left a comment
There was a problem hiding this comment.
Thanks for the writeup, the packed varlen layout and the 2B+2 cu_seqs trick with the padded tail sequence are a clean way to keep the graph metadata static. Numbers on the request-batch path look convincing.
Three blocking items and a few smaller ones, all inline.
Blocking
pipeline_omnivoice.py:584: declaringsupports_request_batch = Trueroutes throughexecute_model_batch, whereallow_single_output=False. Returning a bareDiffusionOutputon a per-request validation error now raisesRuntimeErrorfor every batch size including B=1, wheremainreturned a clean user error.pipeline_omnivoice.py:376: in step mode the same error return is discarded by the runner,state.latentsstaysNone, and_prepare_latentsraises outside the per-requesttry, failing every concurrent request in the batch.omni_base.py:182: defaultingdiffusion_batch_sizetomax_num_seqschanges startup behavior for all 51 diffusion pipelines. 43 of them lacksupports_request_batchand will now fail to start under--max-num-seqs > 1where they previously ran serially. This is unrelated to the PR title and has no test.
Test coverage
test_cuda_graph_generator.py is the only core_model test, and its _SyntheticGenerator never calls _varlen_attn, so it covers the padding, slicing and replay bookkeeping but not the attention change itself. The e2e tests are slow plus tts plus L4, so they only run in the weekly sweep, which does pick up the new test_omnivoice_parity.py automatically through the marker-driven collection in .buildkite/cuda/test-weekly.yml:93. test_omnivoice_parity.py also sets OMNIVOICE_CUDA_GRAPH=0 on both sides, so the CUDA-graph plus varlen combination, which is the riskiest part of this change, has no automated coverage at all. Could you add a core_model test that drives a small real OmniVoiceGenerator through _varlen_attn in both eager and graph mode, including a replay with a nonzero padded tail?
Reviewed at 7a096c39668eb3295c016f792a16d4b3ea21f577. Static review only, no GPU run on my side.
| diffusion_batch_size: int = kwargs.pop("diffusion_batch_size", 1) | ||
| diffusion_batch_size = kwargs.pop( | ||
| "diffusion_batch_size", | ||
| kwargs.get("max_num_seqs", 1), |
There was a problem hiding this comment.
[High] This changes the diffusion batch default for every pipeline, not just OmniVoice, and the PR body does not mention it.
diffusion_batch_size is force-assigned onto od_config.max_num_seqs at stage_init_utils.py:1444 and stage_engine_startup.py:1561, overwriting whatever the stage config resolved. With the old default of 1, --max-num-seqs N was silently ignored by diffusion stages. Now it lands, which is what makes the --max-num-seqs 8 e2e test in this PR work, and it does match docs/user_guide/diffusion/execution_modes.md:232.
The side effect: only 8 of 51 diffusion pipelines set supports_request_batch = True, and diffusion_engine.py:229 raises at startup when max_num_seqs > 1 without it. Serve commands that run today (serially) will fail to start after this. It also still clobbers a per-stage engine_args.max_num_seqs with the global CLI value.
Please split this into its own PR with a unit test in tests/engine/test_async_omni_engine_stage_init.py, or at minimum call it out in the body and in the execution-modes doc.
| extra = request.sampling_params.extra_args or {} | ||
| prepared = self._prepare_request_input(prompt, extra) | ||
| if isinstance(prepared, DiffusionOutput): | ||
| return prepared |
There was a problem hiding this comment.
[High] Returning a bare DiffusionOutput here now raises RuntimeError instead of surfacing the user error.
Declaring supports_request_batch = True routes OmniVoice through execute_model_batch, which calls _normalize_pipeline_outputs(..., allow_single_output=False) (diffusion_model_runner.py:650). That helper raises
RuntimeError: OmniVoicePipeline.forward returned a single DiffusionOutput;
request-batch forward must return list[DiffusionOutput].
at diffusion_model_runner.py:84, for every batch size including B=1. On main OmniVoice went through execute_model with allow_single_output=True, so the same prompt produced a clean per-request error.
Trigger: offline Omni.generate with two clips in multi_modal_data["audio"] (line 258 above), or a dict prompt with no text key (line 276).
Fix direction: allocate outputs = [None] * len(req.requests) up front, place the error DiffusionOutput in that request's slot, and keep the remaining requests running.
There was a problem hiding this comment.
Resolved in d5f0bfb, added
outputs = [None] * len(req.requests)
prepared_indices: list[int] = []
to handle this and preserve the original request indices to prevent slot mismatch
Also added regression test in tests/model_executor/models/omnivoice/test_pipeline_batching.py for changes in pipeline_omnivoice.py
| extra = state.sampling.extra_args or {} | ||
| prepared = self._prepare_request_input(prompt, extra) | ||
| if isinstance(prepared, DiffusionOutput): | ||
| return prepared |
There was a problem hiding this comment.
[High] In step mode this error return is discarded, and the malformed request takes down the whole step batch.
diffusion_model_runner.py:714 calls prepare_encode(state) for its side effects only and drops the return value. When _prepare_request_input returns an error, state.latents is never assigned, so InputBatch.make_batch hits _prepare_latents and raises ValueError("All requests must have 'latents' initialized.") (input_batch.py:347). That happens at execute_stepwise line 776, outside the per-request try at line 810, so one bad request fails every concurrent request in the batch.
The annotation on line 371 also matches neither the interface (interface.py:60 declares -> StepRequestState) nor the success path, which returns None. Suggest recording the error on the state and letting the runner emit it per request, then fixing the annotation.
| # Default bucket count is 10; 16 gives modest headroom for edge cases | ||
| # (seq_len > max bucket or non-CFG batch) without unbounded GPU growth. | ||
| _MAX_LAZY_GRAPHS: int = 16 | ||
| _DEFAULT_CAPTURE_BATCH_SIZES: tuple[int, ...] = (1, 2, 3, 4) |
There was a problem hiding this comment.
[Medium] The pre-warm set does not cover the concurrency this PR targets, so the hot path falls into exact-size lazy capture.
The graph key is now (request_batch_size, packed_token_total), but this tuple is hard-coded to (1, 2, 3, 4) and the largest bucket is 1024 (configs/omnivoice.py:88).
Worked example with the parity-test prompt: RuleDurationEstimator gives target_len about 80, so the packed total per request is about 200 (cond_len about 120 plus target_len 80).
- B=4 gives about 800, bucket 1024, fine.
- B=8 gives about 1600, past every bucket.
_find_bucketreturnsNoneand line 629 setsbucket = seq_lenexactly.
So every distinct combination of request lengths becomes its own lazy capture of a 28-layer graph on the hot path, thrashing the 16-entry cache. B in {5..8} is never pre-warmed at all, and tests/e2e/online_serving/test_omnivoice_expansion.py:34 runs this model with --max-num-seqs 8.
Fix direction: derive the capture batch sizes from max_num_seqs, and round oversized totals up to a capped ladder instead of capturing at the exact packed length.
There was a problem hiding this comment.
Resolved in d5f0bfb6ae982bd21f73d3e4756aad458202f637, now _OmniVoiceCUDAGraphForward will read od_config and derive buckets and batches to be captured. I configured the buckets in a ladder pattern (shown below), which properly activates the LRU cache. With MAX_LAZY_GRAPH = 16, lazy-captured buckets are rounded up to multiples of 128. In benchmarking (using the command described), most scenarios do not trigger lazy capture, except for three specific (batch, seq_len) combinations: (7, 1574), (8, 1511), and (7, 1511). I believe the current capture strategy is reasonable.
This PR measures ~7 GiB peak memory after pre-capture, versus ~5 GiB on the main branch.
Long-text generation issues persist as noted in #6333. #6409 proposes a chunk-based solution with a per-request limit of ~15 seconds for ~375 tokens in a single forward pass, so capturing extremely long sequences is unnecessary.
B=1: [128, 192, 256, 320, 384, 448, 512]
B=2: [256, 320, 384, 448, 512, 640, 768]
B=3: [384, 448, 512, 640, 768, 896, 1024]
B=4: [512, 640, 768, 896, 1024]
B=5: [640, 768, 896, 1024]
B=6: [768, 896, 1024]
B=7: [896, 1024]
B=8: [1024]
Regression test added in tests/model_executor/models/omnivoice/test_cuda_graph_generator.py for LRU.
| graph = torch.cuda.CUDAGraph() | ||
| with torch.no_grad(): | ||
| with torch.cuda.graph(graph, pool=self._pool_handle): | ||
| with torch.cuda.graph(graph, pool=current_platform.get_global_graph_pool()): |
There was a problem hiding this comment.
[Medium] This reverses a deliberate decision on main, and the class docstring above still describes the old behavior.
main used a per-instance torch.cuda.graph_pool_handle() with an explicit comment:
Lazy-init per-instance pool handle: isolates OmniVoice CUDA Graph memory from other vllm modules (unlike get_global_graph_pool which shares a single pool across all captured graphs and can cause memory aliasing when two graphs replay concurrently).
That comment is gone, but the class docstring at lines 486 to 490 still says graphs "share a single per-instance pool handle, which isolates OmniVoice CUDA memory from other vllm modules". Please either restore the per-instance pool, or state why the aliasing concern no longer applies and update the docstring to match.
There was a problem hiding this comment.
Reverted in d5f0bfb. This change was introduced while debugging an issue during testing, but there were two potential variables involved, and this change turned out to be one of them. After double-checking, I found no reason to keep it. I'll comment below with more details.
| uncond_start = cond_end | ||
|
|
||
| # Extract logits for target region; upcast only the slices we actually consume. | ||
| c_logits = batch_logits[:, cond_end - t_len : cond_end, :].unsqueeze(0).to(torch.float32) |
There was a problem hiding this comment.
[Medium] Lines 488 to 551 duplicate omnivoice_generator.py:956-1016 nearly verbatim.
The CFG fuse, the log_probs[..., mask_id] = -inf mask, the layer penalty, the Gumbel position noise, the top-k select and the cond/uncond mirror write are copied line for line, comments included. The only differences are guidance_scales[i] vs guidance_scale, generators[i] vs request_generator, and where sample_tokens comes from. Any future correctness fix has to land in both places.
Extracting one _unmask_one_request(...) helper owned by the generator would also remove denoise_step's reach into generator._prepare_embeddings, _transformer_forward, _get_logits and _cuda_graph_fwd, which are all private.
| state.extra["tokens"] = tokens | ||
|
|
||
| def denoise_step(self, input_batch: InputBatch, *, states: Sequence[StepRequestState] | None = None, **kwargs: Any): | ||
| use_cuda_graph = self.generator._cuda_graph_fwd is not None |
There was a problem hiding this comment.
[Minor] Two small things in this method.
- This line has no
input_ids.is_cudaguard, unlike the request-mode path atomnivoice_generator.py:929. Worth keeping the two consistent. - Line 450 reads
state.extra.get("guidance", self.guidance_scale), butprepare_encodesetsstate.guidance(line 420), notstate.extra["guidance"]. The.getalways falls through, so the per-request guidance plumbing is dead. Either writestate.extra["guidance"]inprepare_encodeor readstate.guidancehere.
There was a problem hiding this comment.
Both resolved in d5f0bfb, regression test for Q2 added in tests/model_executor/models/omnivoice/test_pipeline_batching.py
| entry["graph"].replay() | ||
|
|
||
| output = entry["static_output"] | ||
| if bucket is not None and bucket != seq_len: |
There was a problem hiding this comment.
[Minor] bucket is not None cannot be false: line 629 assigns bucket = seq_len when _find_bucket returns None, so the check is dead and only bucket != seq_len matters.
Also, _lazy_graphs is documented as LRU but behaves as FIFO: the hit path on line 639 never calls move_to_end, so a frequently used key is still evicted by popitem(last=False).
There was a problem hiding this comment.
Both issues have been fixed.
Detailed comments on the _lazy_graphs FIFO issue have been inlined in the bucket issue. This was an existing unresolved issue on main.
| if "num_inference_steps" in extra: | ||
| sampling.num_inference_steps = int(extra["num_inference_steps"]) | ||
|
|
||
| if "guidance_scale" in extra: |
There was a problem hiding this comment.
[Minor] This block is out of scope for OmniVoice and unvalidated.
OmniVoice ignores both fields: prepare_encode reads self.num_step and self.guidance_scale from the model config, never state.sampling. So this changes behavior only for other diffusion TTS pipelines, with no test in this PR.
Two details if it stays:
guidance_scale_provided = Trueis redundant.OmniDiffusionRequest.__post_init__already sets it wheneverguidance_scale is not None(vllm_omni/diffusion/request.py:66).int(...)andfloat(...)run on unvalidated client input.extra_params: {"num_inference_steps": "abc"}raisesValueErrorand surfaces as a 500 rather than a 400.
There was a problem hiding this comment.
These two changes are primarily required for scheduler correctness:
-
StepScheduler._get_total_steps()readssampling_params.num_inference_steps. If the value remains only in extra_args and is not promoted to the standard sampling field, the scheduler receives None and fails when evaluating int(None). -
Request-batch compatibility is determined from the standard sampling fields, including
guidance_scale. If the request value remains only inextra_args, requests with different guidance scales may receive the same compatibility key and be incorrectly admitted to the same request batch.
Regression coverage was added in tests/entrypoints/openai_api/test_serving_speech.py
| import torch | ||
| import torch.nn as nn | ||
| import torch.nn.functional as F | ||
| from torch.nn.attention.varlen import _varlen_attn |
There was a problem hiding this comment.
[Minor] _varlen_attn is a private torch API imported at module scope, so it becomes a hard load-time requirement for OmniVoice on every platform.
What is the minimum torch version, and is torch.nn.attention.varlen present on the npu, rocm and xpu images? A short note in the module docstring plus the minimum version in the model doc would help, since there is no other use of this API in the repo.
There was a problem hiding this comment.
Resolved in 9432df3. Now using the in-tree class diffusion.attention.layer.Attention with an explicit attn_mask fallback when varlen_attn is not supported.
As for the version question, the torch version I used is the one installed with vLLM(2.13), and the minimum version that supports varlen_attn is torch 2.10.
|
Inline comments are replied. Perf number still hold. Test coverage updates:
|
|
@linyueqian PTAL, thanks! |
|
One practical follow-up to the approval, because it is approved but still cannot merge. Auto-merge is armed, so this will land by itself once the required checks are satisfied. The blocker is that I just tried clearing it without troubling you: cycling the So when you have a moment, merging current main into the branch will give you a fresh general run, and the flake is intermittent so it should come back clean. That is also worth doing on its own terms, since main has moved since 70a2b19 and you have already been bitten twice by tests on main that call APIs this branch changes. I would rather you spend the push on a real rebase than on an empty commit. Nothing else is outstanding from my side. |
…ised yet) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…tion Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…_args Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
…data Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Head branch was pushed to by a user without write access
70a2b19 to
5557903
Compare
|
Rebased cleanly with no conflicts, ran the unit tests locally with no new errors raised. The ready label needs to be cycled. |
vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> Signed-off-by: wenjie.yan <wenjyan@outlook.com>
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763) Signed-off-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> * [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253) Signed-off-by: liangmengh <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794) Signed-off-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> * [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240) Signed-off-by: Tianyao Wu <rayroy31@gmail.com> * [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Doc] Add AI usage policy for contributions (vllm-project#7305) Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> * [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539) Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> * [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730) Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> * [Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution (vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> * [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> * [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051) Signed-off-by: zjli2013 <leezhengjiang@126.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003) Signed-off-by: ZenAlexa <zimingwang945@gmail.com> * [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> * [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076) Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> * [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896) Signed-off-by: xutianle <xutianle@fudan.edu.cn> * [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Core] Split Omni connector model runner mixin (vllm-project#6903) Signed-off-by: natureofnature <wzliu@connect.hku.hk> * [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231) Signed-off-by: mglyn <1203789601@qq.com> * [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299) Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> * [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289) Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> * [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308) Signed-off-by: mglyn <1203789601@qq.com> * [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185) Signed-off-by: NumberWan <wantszkin2003@gmail.com> * [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685) Signed-off-by: zouyizhou <zouyizhou@huawei.com> * [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328) Signed-off-by: ZhengWG <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200) Signed-off-by: kunkunblueberry <1833921874@qq.com> * [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741) Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Nick Cao <ncao@redhat.com> * [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007) Signed-off-by: mershi <mershi@tencent.com> Co-authored-by: mershi <mershi@tencent.com> * [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155) Signed-off-by: hyw <yuweih205@gmail.com> * [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202) Signed-off-by: Sy03 <1370724210@qq.com> * [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301) Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> * [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292) Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336) Signed-off-by: andyluo7 <andy.luo@amd.com> * [NPU][CI] Add A5 and 310P CI support (vllm-project#6875) Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> * [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350) Signed-off-by: mglyn <1203789601@qq.com> * [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344) Signed-off-by: Sy03 <1370724210@qq.com> * [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230) Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453) Signed-off-by: herotai214 <herotai214@gmail.com> * [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356) Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597) Signed-off-by: wangyu <410167048@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: Samit <285365963@qq.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: Samit <285365963@qq.com> * [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI] Isolate layerwise offload memory measurements (vllm-project#6938) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560) Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> * [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362) Signed-off-by: tly <2200895168@qq.com> * [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363) Signed-off-by: psv666 <2693925048@qq.com> * Cosmos3 action policy improvements (vllm-project#6460) Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561) Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> * [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281) Signed-off-by: david6666666 <530634352@qq.com> * [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930) Signed-off-by: yancaocn <yancaochn@163.com> Co-authored-by: yancaocn <yancaochn@163.com> * [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005) Signed-off-by: samithuang <285365963@qq.com> * [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559) Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172) Signed-off-by: hyw <yuweih205@gmail.com> * [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171) Signed-off-by: hyw <yuweih205@gmail.com> * [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167) Signed-off-by: hyw <yuweih205@gmail.com> * [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841) Signed-off-by: wtz2333 <2955110911@qq.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> * [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853) Signed-off-by: XIN GAO <1037396230@qq.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430) Signed-off-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Codex <noreply@openai.com> * [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Co-authored-by: Sy03 <1370724210@qq.com> * [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441) Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427) Signed-off-by: Maciej Bala <mbala@nvidia.com> * [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094) Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390) Signed-off-by: Guangjian <hiro20833@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167) Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> * [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889) Signed-off-by: psv666 <2693925048@qq.com> * [NPU] upgrade to v0.29.0 (vllm-project#7433) Signed-off-by: Weiming Liao <liaowm5@gmail.com> * [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128) Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> * [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907) Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876) Signed-off-by: gerayking <399geray@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982) Signed-off-by: Qihan Kang <rollykanggg@gmail.com> * [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> --------- Signed-off-by: Asthenia <asthenia0412@gmail.com> Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: liangmengh <liangmengh@nvidia.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Signed-off-by: KrystalRay <keeleiray@gmail.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com> Signed-off-by: specture724 <specture724@gmail.com> Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: zjli2013 <leezhengjiang@126.com> Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Signed-off-by: ZenAlexa <zimingwang945@gmail.com> Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Signed-off-by: Allen Wu <allenwu2795@gmail.com> Signed-off-by: xutianle <xutianle@fudan.edu.cn> Signed-off-by: wangyu <410167048@qq.com> Signed-off-by: natureofnature <wzliu@connect.hku.hk> Signed-off-by: mglyn <1203789601@qq.com> Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Signed-off-by: NumberWan <wantszkin2003@gmail.com> Signed-off-by: zouyizhou <zouyizhou@huawei.com> Signed-off-by: ZhengWG <zwg0606@gmail.com> Signed-off-by: kunkunblueberry <1833921874@qq.com> Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Signed-off-by: mershi <mershi@tencent.com> Signed-off-by: hyw <yuweih205@gmail.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> Signed-off-by: Guangjian <hiro20833@gmail.com> Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Signed-off-by: herotai214 <herotai214@gmail.com> Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Signed-off-by: Samit <285365963@qq.com> Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Signed-off-by: tly <2200895168@qq.com> Signed-off-by: psv666 <2693925048@qq.com> Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Signed-off-by: Nick Cao <ncao@redhat.com> Signed-off-by: david6666666 <530634352@qq.com> Signed-off-by: yancaocn <yancaochn@163.com> Signed-off-by: samithuang <285365963@qq.com> Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Signed-off-by: wtz2333 <2955110911@qq.com> Signed-off-by: XIN GAO <1037396230@qq.com> Signed-off-by: QI JIA <qi.jia@shengshu.ai> Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Signed-off-by: gerayking <399geray@gmail.com> Signed-off-by: Qihan Kang <rollykanggg@gmail.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Signed-off-by: José Carlos <jose@valendra.tech> Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com> Co-authored-by: liangmenghuang <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Lei Ke <1141466880@qq.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com> Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com> Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com> Co-authored-by: boatman <1930807094@qq.com> Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com> Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com> Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com> Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> Co-authored-by: xutianle <24210290017@m.fudan.edu.cn> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> Co-authored-by: NATURE <wzliu@connect.hku.hk> Co-authored-by: Mu GuanLin <1203789601@qq.com> Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com> Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: NumberWan <wantszkin2003@gmail.com> Co-authored-by: zyz111222 <zouyizhou@huawei.com> Co-authored-by: Zheng Wengang <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com> Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com> Co-authored-by: Nick Cao <ncao@redhat.com> Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com> Co-authored-by: mershi <mershi@tencent.com> Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com> Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com> Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com> Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: Samit <285365963@qq.com> Co-authored-by: wkutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com> Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com> Co-authored-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com> Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com> Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com> Co-authored-by: yancaocn <yancaochn@163.com> Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: wtz2333 <2955110911@qq.com> Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com> Co-authored-by: Qi Jia <kuafou@gmail.com> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: DanaerLee <mrdanaer@gmail.com> Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com> Co-authored-by: junpengw67-max <junpengw67@gmail.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: geray <48796550+gerayking@users.noreply.github.com> Co-authored-by: KANG Qihan <3149604185@qq.com>
vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Purpose
Support request batching and step execution for OmniVoice by implementing model-specific helpers following the existing
execute_stepwise()lifecycle.Bit-exact parity with outputs generated before this PR is not guaranteed because the attention backend and packed execution layout have changed. WER results are listed below. The OmniVoice decoder is unchanged from
main, and each request is decoded separately using exact-length codec tokens.An earlier version of this PR implemented padded Decoder batching, but it was later reverted. With the packed varlen Generator, generated codec tokens are now kept at exact per-request lengths, so batching Decoder execution would require padding them again to the longest request. The current implementation therefore decodes completed requests individually using exact-length codec tokens. I have not benchmarked padded Decoder batching against per-request exact-length decoding, so it is unclear which approach is faster.
1. Add packed variable-length attention to the OmniVoice generator
The previous Generator layout stored conditional and unconditional inputs as separate rows. Since their sequence lengths differ, unconditional inputs had to be padded to the conditional sequence length. In a request batch, shorter
requests were also padded to the longest request, introducing substantial attention and Transformer overhead.
This PR changes Generator inputs to a packed, token-major 2D layout:
with tensor shapes:
For
Brequests, CFG produces2Bindependent sequences. Their cumulative offsets are represented bycu_seqs, allowing PyTorch varlen attention to keep requests and cond/uncond branches mutually invisible while retaining full bidirectional attention within each sequence.Position IDs are reset independently for every packed sequence:
_position_ids_from_cu_seqs()derives these sequence-local positions directly fromcu_seqs.Per-request random generators are retained so one request does not consume or change another request's Gumbel sampling stream.
2. Make packed varlen attention compatible with CUDA Graphs
OmniVoice already supported CUDA Graphs, but the previous graph inputs and keys were based on padded
[batch, codebook, sequence]tensors.CUDA Graph keys are now:
(request_batch_size, token_bucket)Captured static inputs use:
The
2B+2cu_seqslayout represents:2Breal cond/uncond sequences;In eager mode, the tail has zero length:
[..., real_total, real_total]In Graph mode, the last offset is changed to the selected bucket:
[..., real_total, token_bucket]This keeps the metadata shape static while isolating padded tail tokens from all real requests.
max_qandmax_kuse the fixed packed token bucket during graph capture, so the real maximum sequence length does not need to be part of the graph key.Common request batch sizes derived from default config are pre-captured. Other batch sizes and oversized token shapes are captured lazily.
3. Support request-batch and step-execution lifecycles
Request-batch mode packs the
input_idsandaudio_maskvalues produced by_prepare_request_input()along the token dimension.cond_lensandtarget_lensare used to constructcu_seqsonce before Transformer execution and to locate each request's target logits and generated codec tokens.Step-execution stores each request's persistent latents as:
[request_seq_len, num_codebooks]InputBatch.make_batch()therefore concatenates active requests directly along dimension 0 without model-specific repadding.denoise_step()constructs the packed metadata, runs one varlen Generator step, and returns the same 2D layout so the Runner can scatter rows back to individual request states.The standard lifecycle methods are used:
prepare_encode()initializes request-local state and sampling metadata;denoise_step()executes one packed Generator step;step_scheduler()persists the request slice and advances its step;post_decode()decodes the completed request.Generated codec tokens are maintained at exact per-request lengths.
Test Plan
Unit Tests
The CUDA Graph tests continue the validation coverage from the earlier implementation for packed Generator. Additional tests were added for the varlen layout to verify dynamic
cu_seqsupdates across replays of the same graph key, and eager/Graph parity.E2E Tests
The E2E tests verify:
Perf and Acc
Serve Command:
Remove
--step-executionfor benchmarking request-batch.Benchmark Command:
Environment:
c4197ffe163c23017f179bf2f766c67766aca249Test Results
Unit Test:
1 failed, 302 passed. The single failure is unrelated to this PR and is not addressed here.E2E :
14 passedPerformance
1. Batch vs Per-Request
C=4
C=8
2. Batch vs Step-Execution (C=4, request-rate=1)
Accuracy
All request passed for each benchmark with 32 requests. All benchmark result with
Median WER=0.0000BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)