Skip to content

[Refactor] P0.2: Migrate API server helpers out of api_server - #5453

Merged
Gaohan123 merged 15 commits into
vllm-project:mainfrom
herotai214:hero/api-helper-latest-main
Sep 10, 2026
Merged

Gaohan123 merged 15 commits into
vllm-project:mainfrom
herotai214:hero/api-helper-latest-main

Conversation

@herotai214

Copy link
Copy Markdown
Contributor

[Refactor] P0.2: Migrate API server helpers out of api_server

cc @linyueqian
See if/how a P0.1: Add API server guardrails should be added prior merging this PR.
This PR tgt with P0.3 serves as the fundamental structure of our whole refactor project.

Summary

This is Refactor [P0.2] for #5227.

It moves non-route helpers and request models out of vllm_omni/entrypoints/openai/api_server.py into endpoint-owned or server-owned modules. The goal is to stop api_server.py from being the place other code imports helper implementation from.

This PR does not migrate endpoint route bodies yet. That is intentional: Phase 0 is split into:

  • P0.1: API server guardrails/tests
  • P0.2: helper migration, this PR
  • P0.3: endpoint route-owner extraction, immediately after this PR

What Changed

Moved API-server helpers into clearer owners:

  • openai/app_state.py for app-state accessors
  • openai/chat_template.py for chat-template bootstrap helpers
  • openai/diffusion.py for shared diffusion-stage generation/sampling helpers
  • openai/lora.py for OpenAI LoRA request parsing
  • openai/images/helpers.py for image-only helpers
  • openai/models/serving.py for diffusion-only /v1/models serving shim
  • serve/utils/errors.py for server-level engine exception handling
  • serve/utils/routes.py for app/router route-table mutation helpers
  • serve/profile/protocol.py and serve/profile/utils.py for profiler route support
  • serve/omni_control/protocol.py for sleep/wakeup request models

This follows upstream vLLM's current ownership direction where practical:

  • server/app mechanics live under entrypoints/serve/utils;
  • OpenAI-compatible response shaping stays under entrypoints/openai;
  • model-list serving helpers live near entrypoints/openai/models;
  • endpoint-family helpers move toward their owning OpenAI endpoint packages instead of remaining in the server composition root.

Updated tests and serving imports to use the new helper locations.

Why api_server.py Still Imports These Helpers

This PR intentionally does not move endpoint bodies yet, so api_server.py temporarily imports helpers from their new owners.

For example:

from vllm_omni.entrypoints.openai.images.helpers import (
    _check_max_generated_image_size,
    _extract_images_from_result,
    _load_input_images,
)
from vllm_omni.entrypoints.openai.lora import _parse_lora_request
from vllm_omni.entrypoints.serve.utils.errors import (
    _create_engine_error_json_response,
)

That looks a little inverted today because the route bodies are still in api_server.py. The next PR, Refactor [P0.3], will move those route bodies beside these helpers, so these temporary imports disappear or become local endpoint-package imports.

Out Of Scope

This PR does not move decorated endpoint bodies such as:

  • /v1/chat/completions
  • /v1/images/generations
  • /v1/images/edits
  • /v1/videos*
  • /v1/realtime
  • /v1/video/chat/stream
  • /v1/duplex
  • /v1/omni/sleep
  • /v1/omni/wakeup

It also does not move duplex_capability.py, because that helper was already outside api_server.py on latest main.

Relationship To Phase 0

This PR is the middle step of Phase 0:

P0.1 tests/guardrails
  -> P0.2 helper migration
  -> P0.3 endpoint router migration

The endpoint migration should start immediately after this PR. Keeping P0.2 focused on helpers keeps this review smaller and makes the P0.3 route-body move mostly mechanical.

Test Plan

Focused API-server pytest passed in a local development environment with matching vllm / vllm-omni versions:
vLLM Version: 0.25.0

vLLM-Omni Commit: d688aa8

pytest -q tests/entrypoints/openai_api/test_image_server.py \
  tests/entrypoints/openai_api/test_video_server.py \
  tests/entrypoints/openai_api/test_serving_speech.py \
  tests/entrypoints/openai_api/test_omni_sleep_wakeup.py

(Those 4 are only the focused regression set we ran for this PR, not the only tests in the repo.

They were chosen because they most directly hit the helpers moved out of api_server.py:

  • test_image_server.py — image helpers / diffusion / LoRA / models serving
  • test_video_server.py — video generation helpers / store / form parsing
  • test_serving_speech.py — speech error-response helpers
  • test_omni_sleep_wakeup.py — sleep/wakeup protocol + app-state init)

Result:

357 passed, 18 warnings in 13.22s

Additional checks run:

python - <<'PY'
from pathlib import Path

for path in [
    "vllm_omni/entrypoints/openai/api_server.py",
    "vllm_omni/entrypoints/openai/app_state.py",
    "vllm_omni/entrypoints/serve/utils/errors.py",
    "vllm_omni/entrypoints/serve/utils/routes.py",
]:
    compile(Path(path).read_text(), path, "exec")

print("compile ok")
PY

Also checked IDE lints on touched helper files; no linter errors were reported.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@herotai214

Copy link
Copy Markdown
Contributor Author

Additional Test Plan

pytest -q tests/entrypoints -m "core_model and cpu"

Result: 1049 passed, 2 skipped, 165 deselected (~31s)

Scope: all CPU core_model tests under tests/entrypoints/ (38 files / 1051 collected).

Covered modules

  • Audio: generate, speech, speech stream, audio format, decode URL
  • Images: image server
  • Video: generation server, video stream, video output stream, API utils, storage
  • Chat: sampling params, multistage, metrics, speaker
  • Duplex / realtime: handler, capability, protocol, session, async duplex, fence, realtime helpers
  • OpenPI: connection, serving
  • Serve / omni core: sleep/wakeup, profiler, stage utils, async omni, diffusion config, snapshot download, utils, modality validation

Files

38 files

openai_api/

  • test_audio_format.py, test_decode_audio_url.py
  • test_duplex_capability.py, test_duplex_handler.py
  • test_image_server.py, test_modality_validation.py
  • test_omni_sleep_wakeup.py
  • test_openpi_connection.py, test_openpi_serving.py
  • test_serving_audio_generate.py
  • test_serving_chat_metrics.py, test_serving_chat_multistage_generation.py, test_serving_chat_sampling_params.py, test_serving_chat_speaker.py
  • test_serving_speech.py, test_serving_speech_stream.py
  • test_serving_video_output_stream.py, test_serving_video_stream.py
  • test_storage_manager.py, test_video_api_utils.py, test_video_server.py

openai/

  • test_duplex_protocol.py, test_duplex_session_attachment.py

entrypoints/

  • test_async_omni.py, test_async_omni_diffusion_config.py, test_async_omni_duplex.py, test_async_omni_prompt_update.py
  • test_duplex_fence_propagation.py
  • test_omni_base_profiler.py, test_omni_entrypoints.py, test_omni_new_request_data.py, test_omni_snapshot_download.py
  • test_realtime_connection_helpers.py, test_resolve_dreamzero_config.py
  • test_serve.py, test_stage_utils.py, test_stream_finish_reason.py, test_utils.py
```

@hsliuustc0106 hsliuustc0106 added refactor refactoring for better code scalability and quality frontend code related to entrypoint labels Jul 28, 2026
@herotai214
herotai214 force-pushed the hero/api-helper-latest-main branch from 30d76e1 to 8775bd0 Compare July 28, 2026 02:17

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the diff and validated it end to end on an H100 node (merge-base vs head, vllm 0.26). The zero-behavior-change claim holds.

Runtime equivalence

  • Routes identical: 42 vs 42, empty set diff.
  • OpenAPI schema deep-equal after key sorting (373 definitions each).
  • Live endpoints match: /health, /ping, /version, /v1/models, /v1/audio/voices (same 9 voices), and POST /v1/audio/speech returning byte-identical audio (53804 B, 1.12 s, 24 kHz mono) on both trees.
  • All 10 error probes return identical status codes and byte-identical JSON bodies, which covers the moved exception handlers.
  • All 43 names that left api_server's namespace resolve from their new modules; nothing unaccounted for. Import time is flat (-0.5%, measured interleaved in one process).

Tests — tests/entrypoints/, two full runs per tree: head never regresses. The 8 differing tests all go base-fail to head-pass, and base disagrees with itself on exactly those 8 across its own two runs, so they are flaky under shared-GPU pressure rather than affected by this PR. Five failures are pre-existing on both trees.

Structure — the direction is right: none of the 13 new modules import the composition root, so the dependency graph stays acyclic, and keeping app_state / models.serving reachable as module aliases preserves existing monkeypatch targets. I also AST-compared every moved symbol against main: 46 of 47 are byte-identical, the only delta being a docstring (inline comment below).

One blocker, with a verified one-line fix. The ReadTheDocs build fails on this PR. It is green on 10 of 10 other currently-open PRs, so it is attributable here. Details and the fix are in the inline comment on serve/omni_control/protocol.py.

On your question about P0.1: I would land the guardrails before this. The route manifest and the OpenAPI snapshot are exactly what made this review conclusive, and having them in CI means later phases do not depend on someone re-running this by hand. Happy to review that PR as well.

@@ -0,0 +1,14 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""Request models for Omni server control routes."""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: this PR breaks the ReadTheDocs strict build. Reproduced locally on both trees — merge-base builds clean (exit 0), this head aborts:

WARNING - mkdocs_autorefs: api/README.md: Could not find cross-reference target
  'vllm_omni.entrypoints.serve.omni_control.protocol.OmniSleepRequest'
  'vllm_omni.entrypoints.serve.omni_control.protocol.OmniWakeupRequest'
  'vllm_omni.entrypoints.serve.profile.protocol.ProfileRequest'
Aborted with 3 warnings in strict mode!

Two things collide. docs/mkdocs/hooks/generate_api_readme.py AST-scans the tree and emits [dotted.path][] autorefs for every public class; its exclude list covers vllm_omni.entrypoints.openai but not vllm_omni.entrypoints.serve, so these three pydantic models get links. Meanwhile mkdocs.yml sets api-autonav: on_implicit_namespace_package: skip, and the 8 new package directories have no __init__.py, so autonav skips them, no API page is generated, the anchors never exist, and fail_on_warning: true in .readthedocs.yml turns the dangling references into a failed build.

Fix, verified: add __init__.py to the 8 new directories (serve/, serve/utils/, serve/profile/, serve/omni_control/, openai/images/, openai/models/, openai/video/, openai/video/generation/). With those files present, mkdocs build --strict exits 0 with zero warnings. Adding vllm_omni.entrypoints.serve to the hook's exclude list also silences it, but the __init__.py route is better: it matches the other 206 packages in the tree and gives these modules real API pages.

For the record, packaging is not affected either way — I built a wheel from this head and all 13 new modules are present, because [tool.setuptools.packages.find] defaults to namespaces = true. This is a docs-only failure.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree; I'll add the __init__.py for them.



def _register_omni_exception_handlers(app) -> None:
"""Override upstream vLLM exception handlers with Omni-aware versions."""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The move dropped the body of this docstring. On main it explained why Omni overrides the upstream handler at all:

The upstream engine_error_handler is designed for AsyncLLM (single EngineCore process). Omni uses a multi-stage orchestrator with different health semantics, so we register our own handlers that log multi-stage diagnostic info (orchestrator liveness, per-stage health) when an EngineDeadError is caught, call terminate_if_errored, and return an OpenAI-compatible error JSON response.

That paragraph is the rationale for this module existing, and it is more valuable here than it was in api_server.py. Worth restoring verbatim. This is the only content delta I found across all 47 moved symbols, the code itself is byte-identical.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch! I'll add the docstring body back.

ReferenceVideo,
)
from vllm_omni.entrypoints.openai.storage import STORAGE_MANAGER
from vllm_omni.entrypoints.openai.stores import VIDEO_STORE

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Worth a note for whoever touches the video tests next: VIDEO_STORE and STORAGE_MANAGER are now bound in two modules, here and in api_server.py, because the route bodies stay put until P0.3. Production behavior is unaffected since both are from ... import bindings of the same singletons in openai/stores.py and openai/storage.py, but a test that patches only one of the two targets will silently fail to isolate the other path, which is why test_video_server.py now has to patch both. P0.3 dissolves this once the routes move next to these helpers, so no change requested here.

@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 290476c0635d produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@herotai214

Copy link
Copy Markdown
Contributor Author

Omni ReviewBot triage note

Automated triage of commit c90a2805317e produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Refactor update base on v0.28.0 will be ready soon within these 1-2 days.

@Gaohan123 Gaohan123 modified the milestones: v0.28.0, v0.30.0 Sep 2, 2026
Signed-off-by: herotai214 <herotai214@gmail.com>
@herotai214

herotai214 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Resolved another conflict-minimal merge onto latest main. Only api_server.py conflicted; Incoming was #4728 (image/video benchmark metrics).
@Gaohan123 @hsliuustc0106 — Could you help land this sooner if the extract looks right? main still has these helpers in api_server.py, so every nearby video/image change rebases as a dump-vs-empty conflict and we have to re-forward the real edit frequently.

Resolved conflict details:

Same pattern as MAGI / Cosmos3: we kept this PR’s empty side after edit_images and forwarded the real #4728 edits into video/generation/helpers.py:

  • _unpack_video_generation_result (new)
  • _run_video_generation_job now unpacks the 4- or 5-tuple and stores metadata
  • _parse_video_form accepts return_stage_metrics
  • video_response_from_request now fills fps / num_frames / duration_s (this one was missed in the merge commit and forwarded in a follow-up so queued /v1/videos matches main)

Route bodies stay in api_server.py: generate_images / edit_images already auto-merged the metrics wiring; create_video / create_video_sync still call the helpers.

_build_image_generation_response and _build_image_response_metrics are still in api_server.py on purpose. They are internal helpers used only by the image endpoint functions (generate_images / edit_images), not public API and not a missing images/serving.py. P0.2 left the first one next to the route; #4728 added the second beside it. Both move in the next PR, P0.3, together with those endpoint bodies.

P0.2 scope is otherwise unchanged: no route-body moves in this PR. New request/job helpers should land in the owning package (openai/video/generation/helpers.py, openai/images/helpers.py, entrypoints/duplex/, …), not in api_server.py.

Signed-off-by: herotai214 <herotai214@gmail.com>
@herotai214

Copy link
Copy Markdown
Contributor Author

If possible, plz refrain from merging PRs touching the api_server for 1-2 days & merge this PR first....

Resolved another conflict-minimal merge onto latest main. Three files conflicted (api_server.py, test_serving_speech.py, check_forbidden_imports.py). Could you help land this sooner if the extract looks right? main still has these helpers in api_server.py, so every nearby change rebases as a dump-vs-empty conflict and we have to re-forward the real edit.

Same pattern as MAGI / Cosmos3 / #4728: Incoming dumped the already-extracted helper block after profiler_router. We kept this PR’s empty side and forwarded the real edits:

_build_image_generation_response and _build_image_response_metrics are still in api_server.py on purpose. They are internal helpers used only by the image endpoint functions, not public API. Both move in P0.3 together with those endpoint bodies.

P0.2 scope is otherwise unchanged: no route-body moves in this PR. New request/job helpers should land in the owning package, not in api_server.py.

Signed-off-by: herotai214 <herotai214@gmail.com>
@herotai214

Copy link
Copy Markdown
Contributor Author

AGAIN

Resolved another conflict-minimal merge onto latest main. Only api_server.py conflicted; Incoming was #6963 (image pixel limits on video refs). Could you help land this sooner if the extract looks right? main still has these helpers in api_server.py, so every nearby video change rebases as a dump-vs-empty conflict and we have to re-forward the real edit.

Same pattern as MAGI / Cosmos3 / #4728 / #5381: Incoming dumped the already-extracted helper block after edit_images. We kept this PR’s empty side and forwarded the real #6963 edits into video/generation/helpers.py:

  • _validate_minimax_h3_image_payload now runs _validate_image_pixel_limit and maps InvalidInputReferenceError / DecompressionBombError
  • _persist_uploaded_media_references decodes uploads via _decode_image_bytes

video_api_utils.py and the video tests already auto-merged. api_server.py does not import those decode helpers — helpers → utils only.

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI merge-test label to trigger buildkite merge test CI labels Sep 10, 2026
@Gaohan123

Copy link
Copy Markdown
Collaborator

Please resolve conflicts

Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: herotai214 <herotai214@gmail.com>
@herotai214

herotai214 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Resolved another conflict-minimal merge onto latest main. Only api_server.py conflicted; Incoming was #7230 (rebase to vLLM 0.29).

This one was not a helper dump. #7230 only moved vLLM import paths on api_server.py. We kept this PR’s extracted helper imports and took the 0.29 homes for names that file still uses:

  • RequestResponseMetadata ← generate.base.protocol
  • make_arg_parser ← launchers.cli_args
  • serve_http ← launchers.launcher
  • get_uvicorn_log_config ← launchers.utils.server_utils
  • ErrorResponse ← serve.engine.protocol

We did not pull Incoming’s ModelCard / terminate_if_errored / create_error_response / 0.28 try/except back into api_server.py — those already live in the extract.

Follow-up on this branch (same import rewrite #7230 could not see):

  • serve/utils/errors.py and video/generation/helpers.py — launchers.launcher.terminate_if_errored + from vllm.entrypoints.serve import create_error_response
  • models/serving.py — serve.engine.protocol for ModelCard / ModelList / ModelPermission
  • openai/errors.py — serve.engine.protocol for ErrorResponse

The real 0.29 behavior changes (video_api_utils pyav → opencv, AsyncOmni.finish_weight_update(weight_version=...)) already auto-merged. No new routes or helpers.


Full L1 set (guards + image/video/speech/audio/sleep/log-stats): 620 passed. (After below)

One local fallout from the pyav → opencv switch, not from the extract:

vLLM 0.29 dropped the pyav decoder; video_api_utils._decode_video_bytes now uses backend="opencv". opencv-python-headless still dlopens libxcb.so.1 (and libGL / libX11 / libXext / libSM / libICE). On a slim GPU container those packages were missing, so 14 test_video_server.py cases initially failed on my side. The user-visible error is a 400 InvalidInputReferenceError: ... not a valid video, with ImportError: libxcb.so.1: cannot open shared object file as the cause — easy to misread as bad bytes.

After installing libxcb1 libx11-6 libxext6 libgl1 libsm6 libice6, video tests went 14 failed → 111 passed.

CI (vllm/vllm-openai:v0.29.0) may already ship these. Worth a glance at #7230 / the image so other slim or custom boxes don’t hit the same wrap. Not something P0.2 should paper over in api_server.py.

P0.2 scope is otherwise unchanged: no route-body moves. New request/job helpers should land in the owning package, not in api_server.py.

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI merge-test label to trigger buildkite merge test CI labels Sep 10, 2026
@herotai214

herotai214 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author
image

The failing CIs buildkite with amd, intel and npu should be all due to vllm version mismatch after the rebase v0.29 commit; and not introduced by this PR or api server related files.

For example, the ImportError: cannot import name 'compute_layout_strides' from 'vllm.v1.kv_cache_interface' (/usr/local/lib/python3.12/dist-packages/vllm/v1/kv_cache_interface.py here is far from our api_server/entrypoints refactor

The buildkite/vllm-omni passed.

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123
Gaohan123 merged commit bbac4df into vllm-project:main Sep 10, 2026
6 of 9 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

frontend code related to entrypoint high priority high priority issue, needs to be done asap ready label to trigger buildkite CI refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants