Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion vllm_omni/entrypoints/omni_base.py
Original file line number Diff line number Diff line change
Expand Up @@ -177,7 +177,12 @@ def __init__(
# override the deploy YAML's ``async_chunk: true`` default.
async_chunk = kwargs.get("async_chunk")
output_modalities = kwargs.pop("output_modalities", None)
diffusion_batch_size: int = kwargs.pop("diffusion_batch_size", 1)
# Stage init overwrites ``od_config.max_num_seqs`` with this value.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocking] This comment is false on current main. #6525 removed the batch_size -> od_config.max_num_seqs overwrite and made --max-num-seqs flow through stage runtime overrides into the diffusion scheduler directly. The legacy value changed here is only logged by the subprocess and inline clients, so this fallback cannot increase request concurrency. The intended fix has already shipped; please close this PR as obsolete rather than rebasing an ineffective assignment.

# ``vllm serve --omni --max-num-seqs N`` is the documented request-batch
# knob; use it when ``diffusion_batch_size`` is not passed explicitly.
explicit_batch = kwargs.pop("diffusion_batch_size", None)
max_num_seqs = kwargs.get("max_num_seqs") or 1
diffusion_batch_size = int(explicit_batch if explicit_batch is not None else max_num_seqs)
Comment on lines +180 to +185

if "log_requests" in kwargs:
raise TypeError("`log_requests` has been removed in Omni/AsyncOmni. Use `log_stats`.")
Expand Down
Loading