Skip to content

[Bugfix] Fix host-memory leak on aborted /v1/images/generations (#6462) - #6561

Merged
hsliuustc0106 merged 3 commits into
vllm-project:mainfrom
zhang-keliang:fix/6462-frontend-abort-consumed-metrics-leak
Sep 11, 2026
Merged

hsliuustc0106 merged 3 commits into
vllm-project:mainfrom
zhang-keliang:fix/6462-frontend-abort-consumed-metrics-leak

Conversation

@zhang-keliang

Copy link
Copy Markdown
Contributor

Purpose

Fixes #6462 (host-memory leak in api_server on aborted /v1/images/generations).

Every request keeps self._consumed_metric_messages[request_id] — a per-request set[int] populated by _handle_output_message (omni_base.py) to de-dup metrics. The completion path (_log_summary_and_cleanup) pops both request_states[rid] and _consumed_metric_messages[rid]. The abort path — AsyncOmni._abort and the sync mirror Omni.abort — popped only request_states, so every aborted request retained one set[int] entry for the process lifetime.

Reporter's fingerprint matches: abort-only, diffusion-only (audio has its own dict with explicit .pop() sites in serving_speech.py), linear, retained after load stops. Related to #6413 / #6439 which fixed the executor-side abort leak (_completed_outputs); this covers the frontend side.

Fix: _abort (async) and Omni.abort (sync) now pop _consumed_metric_messages[req_id] in lockstep with request_states. Uses getattr(self, "_consumed_metric_messages", None) so existing subclasses / test doubles that don't initialise the dict still work.

Test Plan

Added tests/entrypoints/test_async_omni.py::test_abort_evicts_consumed_metric_messages_entry alongside the existing test_abort_handles_internal_request_mapping. Uses the pre-existing get_async_omni_instance scaffold, seeds both dicts, calls omni._abort([req_id]), asserts both dicts drop the key. Fails on main, passes on this branch.

CPU-only tracemalloc reproduction (200 abort cycles × 40 metrics messages per request):

  • Before fix: 71.1 KB retained (364 B/req); _consumed_metric_messages size = 200.
  • After fix: 3232 B retained (16 B/req, all tracemalloc bookkeeping); dict empty.
  • Completion control (both before and after): ~4 B/req, dict empty.

vLLM Version: N/A for this fix (pure frontend Python; no vLLM version bump)

vLLM-Omni Commit: off 3d035bf (main at the time of the fix)

Test Result

  • Regression test passes on this branch (verified in isolation via a standalone driver that extracts _abort source, since the full pytest suite has heavy runtime deps not present in my minimal env).
  • ruff check clean; ruff format --check clean (ruff==0.14.10 per .pre-commit-config.yaml).
  • No behavior change on the non-abort paths; no public API change.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/entrypoints.md.

Module owners: @alex-jw-brooks @linyueqian @NickCao

@zhang-keliang, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@NickCao

NickCao commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Could _consumed_metric_messages, instead of being a dict keyed by request_id, be an attr attached to the request object itself? That way it can never leak.

zhang-keliang added a commit to zhang-keliang/vllm-omni that referenced this pull request Aug 24, 2026
The AsyncOmni frontend keeps a per-request `set[int]` of already-consumed
metric message ids to guard against double-counting. Previously this was a
class-level `dict[str, set[int]]` on `OmniBase._consumed_metric_messages`,
populated in `_handle_output_message`. `_log_summary_and_cleanup` popped
its entry on completion, but the abort paths popped only `request_states`,
so every aborted request retained a `set[int]` entry for the process
lifetime (~210 KB/abort linear growth in `api_server` PID 1 on
`/v1/images/generations`, reported by the issue).

Rather than adding a matching `.pop()` to each abort path — sync `Omni.abort`,
async `AsyncOmni._abort`, and any future entry point — attach the set to
`ClientRequestState` (per @NickCao's review on vllm-project#6561). The dedup set is
now `req_state.consumed_metric_message_ids`, so it drops with the state
object via GC whenever `request_states.pop(rid)` runs, whichever path
called it. This removes the class-level dict, the setdefault helper, and
the explicit dict pop from `_log_summary_and_cleanup` — the leak becomes
structurally unreachable rather than defended per call site.

Regression: `tests/entrypoints/test_async_omni.py::
test_abort_drops_consumed_metric_message_ids_with_state` — asserts a
populated `consumed_metric_message_ids` set attached to a
`ClientRequestState` is dropped when the abort path pops
`request_states[rid]`, and that no class-level `_consumed_metric_messages`
dict re-appears (regression guard).

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
@zhang-keliang
zhang-keliang force-pushed the fix/6462-frontend-abort-consumed-metrics-leak branch from 94d5dd8 to 2868c44 Compare August 24, 2026 14:44
@zhang-keliang

Copy link
Copy Markdown
Contributor Author

Adopted, thanks — moved the dedup set to ClientRequestState.consumed_metric_message_ids. Dropping the class-level OmniBase._consumed_metric_messages dict + its setdefault helper + the explicit dict pop in _log_summary_and_cleanup; the set now drops with the state via GC on both completion and abort. Net -14 LoC. New head 2868c44.

@zhang-keliang

zhang-keliang commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Self-review:

  • Scope: 4 files, +35/-16, all Python host-side. No CUDA / kernel / config touched.
  • Root cause: _consumed_metric_messages[rid] populated on every stage metric emit but only popped in _log_summary_and_cleanup; abort skipped it → linear leak.
  • After @NickCao's suggestion the set now lives on ClientRequestState, so completion and abort share one cleanup site (request_states.pop(rid)); no per-callsite dict maintenance.
  • Regression test: test_abort_drops_consumed_metric_message_ids_with_state — populated set on the state → abort → state pops → set gone (also guards against re-introducing a class-level dict via hasattr).
  • Ran locally: pytest tests/entrypoints/test_async_omni.py tests/metrics/test_prometheus.py -m cpu → 13 passed. ruff check and ruff format --check on all 6 modified files → clean (only pre-existing UP038 warnings on unrelated lines).
  • No GPU / model-weight reproduction on this branch (host-side lifecycle only; leak measured by reporter with tracemalloc in api_server PID 1, and my CPU repro in the linked issue does the same). No perf benchmarks claimed.

@hsliuustc0106 hsliuustc0106 added the bug Something isn't working label Aug 25, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

resolve conflicts

@zhang-keliang
zhang-keliang force-pushed the fix/6462-frontend-abort-consumed-metrics-leak branch from 2868c44 to 5abddb2 Compare August 28, 2026 02:08
zhang-keliang added a commit to zhang-keliang/vllm-omni that referenced this pull request Aug 28, 2026
The AsyncOmni frontend keeps a per-request `set[int]` of already-consumed
metric message ids to guard against double-counting. Previously this was a
class-level `dict[str, set[int]]` on `OmniBase._consumed_metric_messages`,
populated in `_handle_output_message`. `_log_summary_and_cleanup` popped
its entry on completion, but the abort paths popped only `request_states`,
so every aborted request retained a `set[int]` entry for the process
lifetime (~210 KB/abort linear growth in `api_server` PID 1 on
`/v1/images/generations`, reported by the issue).

Rather than adding a matching `.pop()` to each abort path — sync `Omni.abort`,
async `AsyncOmni._abort`, and any future entry point — attach the set to
`ClientRequestState` (per @NickCao's review on vllm-project#6561). The dedup set is
now `req_state.consumed_metric_message_ids`, so it drops with the state
object via GC whenever `request_states.pop(rid)` runs, whichever path
called it. This removes the class-level dict, the setdefault helper, and
the explicit dict pop from `_log_summary_and_cleanup` — the leak becomes
structurally unreachable rather than defended per call site.

Regression: `tests/entrypoints/test_async_omni.py::
test_abort_drops_consumed_metric_message_ids_with_state` — asserts a
populated `consumed_metric_message_ids` set attached to a
`ClientRequestState` is dropped when the abort path pops
`request_states[rid]`, and that no class-level `_consumed_metric_messages`
dict re-appears (regression guard).

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
@zhang-keliang

Copy link
Copy Markdown
Contributor Author

Rebased onto latest main and resolved the conflict — main had moved prom_metrics initialization earlier in __init__ (post-0.28 rebase), and the new test_same_finished_image_message_is_observed_exactly_once in test_emit_calls.py stubbed the old class-level dict, so its fixture now seeds request_states with a per-request state instead. Branch is conflict-free (head 5abddb2); pytest tests/metrics/test_emit_calls.py tests/entrypoints/test_async_omni.py tests/metrics/test_prometheus.py -m cpu → 54 passed locally.

from vllm_omni.entrypoints.client_request_state import ClientRequestState

async def run():
omni = get_async_omni_instance()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Exercise the production leak path in this regression

This helper uses object.__new__, manually seeds the new field, and never runs OmniBase.__init__ or a metric-bearing output through _handle_output_message. Reintroducing and populating the old request-keyed map in production could therefore still pass this test. Drive a metric message through the production handler, abort the request, and assert that no independent request-keyed metric state survives.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in d9d2b69 — the regression now drives a metric-bearing OutputMessage through _handle_output_message (asserting the handler records the de-dup entry on the state and de-duplicates a replay of the same message object), aborts via the public abort(), and then sweeps vars(omni) so no mapping on the instance still keys anything by the request id. A reintroduced request-keyed map populated by the production handler fails either the sweep or the hasattr guard.

zhang-keliang added a commit to zhang-keliang/vllm-omni that referenced this pull request Aug 31, 2026
Per vllm-project#6561 review: the regression now drives a metric-bearing
OutputMessage through _handle_output_message (which records the
de-dup entry on ClientRequestState and de-duplicates a replay),
aborts the request, and sweeps the instance so no request-keyed
metric state survives -- a reintroduced class-level map populated
by the production handler fails the sweep.

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit c3d8baaf6503 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Aug 31, 2026
@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 1, 2026

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found one blocking integration issue in the regression test; see the inline comment. No additional high-confidence production-code findings.

Comment thread tests/entrypoints/test_async_omni.py Outdated
assert not hasattr(omni, "_consumed_metric_messages")

await omni.abort("req-1")
assert rid not in omni.request_states

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 — Please rebase this regression test onto current main's abort lifecycle. Since #6367, AsyncOmni._abort() keeps request_states registered, queues a terminal abort output, and relies on generate() cleanup to remove the state. Current main therefore conflicts with this test, and if this assertion is carried over it fails immediately after await omni.abort("req-1"). Please preserve the newer abort tests and assert that the per-request set is released after the normal generate()/finally cleanup, or update this test to the terminal-message contract.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done — rebased onto current main (973d20f) and moved the test to the #6367 contract. Both new abort tests from main (test_abort_enqueues_prefix_tokens_from_engine, test_abort_enqueues_synthetic_finished_when_engine_returns_empty) are preserved. The regression now: drives a metric-bearing OutputMessage through _handle_output_message → abort() → asserts the state is still registered and the terminal abort output is on its queue (with the de-dup set living only on that state) → runs generate()'s cleanup via _log_summary_and_cleanup → asserts the state and the set are released together, with a vars(omni) sweep for any surviving request-keyed mapping. pytest tests/entrypoints/test_async_omni.py tests/metrics/test_prometheus.py tests/metrics/test_emit_calls.py -m cpu → 58 passed.

The AsyncOmni frontend keeps a per-request `set[int]` of already-consumed
metric message ids to guard against double-counting. Previously this was a
class-level `dict[str, set[int]]` on `OmniBase._consumed_metric_messages`,
populated in `_handle_output_message`. `_log_summary_and_cleanup` popped
its entry on completion, but the abort paths popped only `request_states`,
so every aborted request retained a `set[int]` entry for the process
lifetime (~210 KB/abort linear growth in `api_server` PID 1 on
`/v1/images/generations`, reported by the issue).

Rather than adding a matching `.pop()` to each abort path — sync `Omni.abort`,
async `AsyncOmni._abort`, and any future entry point — attach the set to
`ClientRequestState` (per @NickCao's review on vllm-project#6561). The dedup set is
now `req_state.consumed_metric_message_ids`, so it drops with the state
object via GC whenever `request_states.pop(rid)` runs, whichever path
called it. This removes the class-level dict, the setdefault helper, and
the explicit dict pop from `_log_summary_and_cleanup` — the leak becomes
structurally unreachable rather than defended per call site.

Regression: `tests/entrypoints/test_async_omni.py::
test_abort_drops_consumed_metric_message_ids_with_state` — asserts a
populated `consumed_metric_message_ids` set attached to a
`ClientRequestState` is dropped when the abort path pops
`request_states[rid]`, and that no class-level `_consumed_metric_messages`
dict re-appears (regression guard).

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Per vllm-project#6561 review: the regression drives a metric-bearing OutputMessage
through _handle_output_message (which records the de-dup entry on
ClientRequestState and de-duplicates a replay), aborts the request, and
asserts that no request-keyed metric state exists anywhere on the
instance except on the still-registered state.

Follows the vllm-project#6367 abort lifecycle: abort() keeps request_states
registered and enqueues a terminal abort output; the test then runs
generate()'s cleanup (_log_summary_and_cleanup) and asserts the state
and the set attached to it are released together.

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
@zhang-keliang
zhang-keliang force-pushed the fix/6462-frontend-abort-consumed-metrics-leak branch from d9d2b69 to 973d20f Compare September 3, 2026 08:57
@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. and removed ready label to trigger buildkite CI labels Sep 10, 2026
@hsliuustc0106
hsliuustc0106 enabled auto-merge (squash) September 10, 2026 23:44

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@hsliuustc0106
hsliuustc0106 merged commit 5a386d7 into vllm-project:main Sep 11, 2026
6 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda-test Used to trigger vllm-omni cuda CI separately. high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Host-memory leak in api_server frontend on aborted /v1/images/generations (diffusion) requests

5 participants