Skip to content

[Bugfix] Fix video prewarm cache retention and cancel-restart delay - #7363

Merged
Gaohan123 merged 3 commits into
vllm-project:mainfrom
psv666:fix/7223-video-prewarm-eviction
Sep 10, 2026
Merged

Gaohan123 merged 3 commits into
vllm-project:mainfrom
psv666:fix/7223-video-prewarm-eviction

Conversation

@psv666

@psv666 psv666 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Addresses D2 and D4 in #7223 (P0.5) for the streaming video session handler.

  • D2 — evicted frames remain cached: a frame can leave the bounded frame buffer while its prewarm decode is still running. The completed task then re-inserts the PIL image into frame_pil_cache, where later buffer evictions no longer remove it. Check buffer membership after decoding before publishing either the PIL image or the failed-frame marker.
  • D4 — cancel-then-restart adds a fixed 100 ms delay: after cancelling and awaiting the previous query task and awaiting abort(), the session still calls asyncio.sleep(0.1). Remove that extra wait while preserving the existing cancellation and task-exit ordering. AsyncOmni already awaits the Orchestrator abort acknowledgment; elapsed wall time is not an additional completion signal.

The prewarm cache/sentinel annotations are also made consistent across the base and Qwen handlers, existing test-output attributes are declared, and the engine-client/base64-string assumptions are explicit. These changes resolve the seven existing mypy errors in the touched code. SPDX headers follow the repository hook.

The D4 change removes the fixed delay. It does not claim to fix or disprove a deeper EngineCore scheduler race: Orchestrator acknowledgment is not a guarantee that all GPU execution has stopped. Existing abort-error handling is unchanged.

Test Plan

  • D2 regression: delay frame A's real PIL decode until B is accepted. With max_frames=1, A's image must be collectible while the session is open; with max_frames=2, it stays cached until session teardown. The evicted case fails before the fix.
  • D4 regression: exercise the actual session handler with immediate and delayed engine abort acknowledgments. Each case interrupts three successive queries, verifies the old generator closes and abort completes before the next generation, checks distinct request IDs and the final response, and records requested sleeps instead of using a machine-dependent wall-clock limit. Both cases fail before removing the sleep, recording [0.1, 0.1, 0.1].
  • Focused regression: both regressions, frame acknowledgments/consumption, cancellation, and waiting for an in-flight query on video.done.
  • Full relevant suite: all tests in the eight files listed below, excluding two model-weight-dependent request-ID hardware tests (test_diffusion_generate_request_id and test_omni_generate_request_id).
  • Static checks: all applicable local pre-commit gates, including mypy for Python 3.10.

vLLM Version: 0.29.0
vLLM-Omni Commit: a44fb5fcb28e910e1d39a8a3dce2322733ec8c43 (merge base: bbac4df7d3392791ce64a574c4f78c18927e61fc)
Environment: Python 3.12.13, PyTorch 2.13.0+cu129, Transformers 5.14.1, pytest 9.1.1.

Prerequisites: the repository's matching vLLM and development environment. These checks use controlled engine outputs and require no model weights or deployed model server. On the shared CUDA host, pytest ran under gpu run --gpus 1 --nonblock --timeout 8m. No real-model scheduler stress test was performed. Logs retain the warning from stale installed Omni version metadata; imports resolve to the tested checkout.

Run from the repository root:

# D2 regression: the evicted case fails before the D2 fix.
.venv/bin/python -m pytest -q tests/entrypoints/openai_api/test_serving_video_stream.py \
  -k test_frame_prewarm_only_keeps_retained_images

# D4 regression: both cases fail with the fixed sleep still present.
.venv/bin/python -m pytest -q tests/entrypoints/openai_api/test_serving_video_stream.py \
  -k test_interrupted_queries_wait_for_abort_without_fixed_delay

# Combined focused regression.
.venv/bin/python -m pytest -q tests/entrypoints/openai_api/test_serving_video_stream.py \
  -k 'prewarm or frame_ack or frames_consumed or interrupted_queries or cancels_in_flight or waits_for_in_flight'

# Full relevant test suite.
.venv/bin/python -m pytest -q \
  tests/entrypoints/openai_api/test_serving_video_stream.py \
  tests/entrypoints/openai_api/test_video_stream_handler.py \
  tests/entrypoints/openai_api/test_video_stream_session.py \
  tests/entrypoints/openai_api/test_video_frame_filter.py \
  tests/entrypoints/openai_api/test_api_server_guards.py \
  tests/entrypoints/test_async_omni.py \
  tests/engine/test_async_omni_engine_abort_ack.py \
  tests/engine/test_orchestrator.py \
  -k 'not test_diffusion_generate_request_id and not test_omni_generate_request_id'

# CI-style L1 selection for the modified regression module.
.venv/bin/python -m pytest -q tests/entrypoints/openai_api/test_serving_video_stream.py \
  -m 'core_model and cpu' --run-level=core_model

# Static checks.
.venv/bin/python -m pre_commit run --files \
  vllm_omni/entrypoints/openai/video_stream_base.py \
  vllm_omni/entrypoints/openai/serving_video_stream.py \
  tests/entrypoints/openai_api/test_serving_video_stream.py

Test Result

Check Result
D2 pre-fix regression 1 failed, 1 passed: evicted image remained alive
D4 pre-fix regression 2 failed, 21 deselected: each restart requested a 100 ms sleep
Combined focused regression 10 passed, 13 deselected
Full relevant suite 198 passed, 2 deselected
Pre-commit, including mypy All applicable hooks passed

An additional controlled diagnostic using the real AsyncOmni.generate() and abort RPC transport passed 3/3 runs against the fixed source: withholding the acknowledgment for 150 ms prevented the next ADD; after acknowledgment, the previous frontend state was removed before the new ADD. The remote acknowledgment and model outputs were controlled test inputs, not a real-model scheduler run.

Author self-review: checked both prewarm publication paths, eviction/retention and session cleanup, consistent cache typing, cancellation/task-exit ordering, delayed acknowledgment, and repeated interrupts. No protocol event names or payload containers are changed.

Full captured logs follow; local checkout paths are replaced with <checkout>. The two red logs are from the respective pre-fix states. The focused/full/static logs are from the source committed as a44fb5fcb28e910e1d39a8a3dce2322733ec8c43.

D2 pre-fix regression (based on aff7d64)
Reserved 1 GPU(s): [7] for command execution (timeout: 0h 10m 0s)
Released automatically if no GPU usage is detected for 30m
Running 2 items in this shard
F.                                                                       [100%]
=================================== FAILURES ===================================
____________ test_frame_prewarm_only_keeps_retained_images[evicted] ____________

monkeypatch = <_pytest.monkeypatch.MonkeyPatch object at 0x7f876cc7d7c0>
max_frames = 1

    @pytest.mark.asyncio
    @pytest.mark.parametrize("max_frames", [1, 2], ids=["evicted", "retained"])
    async def test_frame_prewarm_only_keeps_retained_images(monkeypatch, max_frames):
        frame_a = _make_jpeg(255, 0, 0)
        frame_b = _make_jpeg(0, 255, 0)
        decode_started = asyncio.Event()
        release_decode = asyncio.Event()
        frame_b_accepted = asyncio.Event()
        decoded_images: dict[bytes, weakref.ReferenceType[Image.Image]] = {}
        prewarm_task: asyncio.Task | None = None
        original_to_thread = asyncio.to_thread
    
        async def controlled_to_thread(function, *args, **kwargs):
            nonlocal prewarm_task
            if function is video_stream_base._decode_frame_bytes:
                if args[0] == frame_a:
                    prewarm_task = asyncio.current_task()
                    decode_started.set()
                    await release_decode.wait()
                image = await original_to_thread(function, *args, **kwargs)
                decoded_images[args[0]] = weakref.ref(image)
                return image
            return await original_to_thread(function, *args, **kwargs)
    
        class AckWebSocket(TimedWebSocket):
            async def send_json(self, data):
                await super().send_json(data)
                if data.get("type") == "video.frame.ack" and data.get("frame_id") == "B":
                    frame_b_accepted.set()
    
        monkeypatch.setattr(video_stream_base.asyncio, "to_thread", controlled_to_thread)
        ws = AckWebSocket()
        handler = QwenOmniStreamingVideoHandler(chat_service=object(), idle_timeout=5.0)
        session_task = asyncio.create_task(handler.handle_session(ws))
        ws.put({"type": "session.config", "max_frames": max_frames, "enable_frame_filter": False})
        try:
            ws.put({"type": "video.frame", "frame_id": "A", "data": _b64(frame_a)})
            await asyncio.wait_for(decode_started.wait(), timeout=5.0)
            ws.put({"type": "video.frame", "frame_id": "B", "data": _b64(frame_b)})
            await asyncio.wait_for(frame_b_accepted.wait(), timeout=5.0)
            ack = next(message for message in ws.sent if message.get("frame_id") == "B")
            assert ack["accepted"] is True
            assert ack.get("dropped_frame_id") == ("A" if max_frames == 1 else None)
    
            # Finish A's real decode only after B has either evicted A or joined it.
            release_decode.set()
            assert prewarm_task is not None
            await asyncio.wait_for(asyncio.shield(prewarm_task), timeout=5.0)
            gc.collect()
            assert not session_task.done()
            # A finished task must not keep an evicted PIL image alive for the session.
>           assert (decoded_images[frame_a]() is not None) == (max_frames == 2)
E           AssertionError: assert (<PIL.Image.Image image mode=RGB size=64x64 at 0x7F876CAF7C80> is not None) == (1 == 2)
E            +  where <PIL.Image.Image image mode=RGB size=64x64 at 0x7F876CAF7C80> = <weakref at 0x7f876caf3ce0; to 'Image' at 0x7f876caf7c80>()

tests/entrypoints/openai_api/test_serving_video_stream.py:644: AssertionError
---------------------------- Captured stdout setup -----------------------------
INFO 09-10 15:04:22 [scheduler.py:277] Chunked prefill is enabled with max_num_batched_tokens=2048.
INFO 09-10 15:04:22 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
--- Running test: test_frame_prewarm_only_keeps_retained_images[evicted]
------------------------------ Captured log setup ------------------------------
INFO     vllm.config.scheduler:scheduler.py:277 Chunked prefill is enabled with max_num_batched_tokens=2048.
INFO     vllm.config.kernel:kernel.py:369 Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
=============================== warnings summary ===============================
.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  <checkout>/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

vllm_omni/version.py:55
  <checkout>/vllm_omni/version.py:55: RuntimeWarning: vLLM and vLLM-Omni appear to have mismatched major/minor versions:
   --> vLLM-Omni version 0.28.0rc2.dev75+g58cb8de68
   --> vLLM version 0.29.0
  This will likely cause compatibility issues.
    warn_if_misaligned_vllm_version()

.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24
  <checkout>/.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24: DeprecationWarning: `vllm.entrypoints.openai.api_server` is deprecated and will likely beunsupported in a future version. Use the corresponding function from `vllm.entrypoints.launchers` instead.
    warnings.warn(

vllm_omni/entrypoints/duplex/audio.py:15
  <checkout>/vllm_omni/entrypoints/duplex/audio.py:15: DeprecationWarning: 'audioop' is deprecated and slated for removal in Python 3.13
    from audioop import alaw2lin, lin2alaw, lin2ulaw, ulaw2lin

vllm_omni/entrypoints/openai/protocol/audio.py:412
  <checkout>/vllm_omni/entrypoints/openai/protocol/audio.py:412: PydanticDeprecatedSince20: Support for class-based `config` is deprecated, use ConfigDict instead. Deprecated in Pydantic V2.0 to be removed in V3.0. See Pydantic V2 Migration Guide at https://errors.pydantic.dev/2.13/migration/
    class CreateAudio(BaseModel):

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
--- Running Summary
=========================== short test summary info ============================
FAILED tests/entrypoints/openai_api/test_serving_video_stream.py::test_frame_prewarm_only_keeps_retained_images[evicted]
1 failed, 1 passed, 19 deselected, 18 warnings in 11.09s
D4 pre-fix regression (D2 present, fixed sleep still present)
Reserved 1 GPU(s): [7] for command execution (timeout: 0h 3m 0s)
Released automatically if no GPU usage is detected for 30m
Running 2 items in this shard
FF                                                                       [100%]
=================================== FAILURES ===================================
__ test_interrupted_queries_wait_for_abort_without_fixed_delay[immediate-ack] __

monkeypatch = <_pytest.monkeypatch.MonkeyPatch object at 0x7f405c1381a0>
delay_abort = False

    @pytest.mark.asyncio
    @pytest.mark.parametrize("delay_abort", [False, True], ids=["immediate-ack", "delayed-ack"])
    async def test_interrupted_queries_wait_for_abort_without_fixed_delay(monkeypatch, delay_abort):
        started: asyncio.Queue[str] = asyncio.Queue()
        abort_started: asyncio.Queue[str] = asyncio.Queue()
        allow_abort = asyncio.Event()
        if not delay_abort:
            allow_abort.set()
        request_ids: list[str] = []
        closed: list[str] = []
        aborted: list[str] = []
        delays: list[float] = []
        original_sleep = asyncio.sleep
    
        async def record_sleep(delay, result=None):
            # Observe requested delays without a machine-dependent latency limit.
            delays.append(delay)
            return await original_sleep(0, result)
    
        monkeypatch.setattr(video_stream_base.asyncio, "sleep", record_sleep)
    
        class BlockingEngine:
            async def generate(self, *, request_id, **kwargs):
                if request_ids:
                    assert aborted[-1] == request_ids[-1]
                request_ids.append(request_id)
                started.put_nowait(request_id)
                try:
                    if len(request_ids) <= 3:
                        yield _text_result("partial")
                        await asyncio.Event().wait()
                    else:
                        yield _text_result("final")
                finally:
                    closed.append(request_id)
    
            async def abort(self, request_id):
                assert request_id in closed
                abort_started.put_nowait(request_id)
                await allow_abort.wait()
                aborted.append(request_id)
    
        class PreprocessedHandler(QwenOmniStreamingVideoHandler):
            async def _preprocess_to_engine_prompt(self, request):
                return {"prompt_token_ids": [1]}
    
        ws = TimedWebSocket()
        handler = PreprocessedHandler(chat_service=object(), engine_client=BlockingEngine(), idle_timeout=5.0)
        task = asyncio.create_task(handler.handle_session(ws))
        try:
            ws.put({"type": "session.config", "modalities": ["text"], "enable_frame_filter": False})
            ws.put({"type": "video.frame", "data": _b64(_make_jpeg())})
            ws.put({"type": "video.query", "text": "first"})
            previous_id = await asyncio.wait_for(started.get(), timeout=2.0)
    
            for _ in range(3):
                ws.put({"type": "video.query", "text": "interrupt"})
                assert await asyncio.wait_for(abort_started.get(), timeout=2.0) == previous_id
                if delay_abort:
                    assert previous_id not in aborted
                    assert started.empty()
                    allow_abort.set()
                previous_id = await asyncio.wait_for(started.get(), timeout=2.0)
                if delay_abort:
                    allow_abort.clear()
    
            ws.put({"type": "video.done"})
            await asyncio.wait_for(task, timeout=2.0)
    
            assert aborted == request_ids[:-1]
            assert len(set(request_ids)) == 4
            assert [msg["text"] for msg in ws.sent if msg["type"] == "response.text.done"] == ["final"]
            assert "error" not in ws.sent_types()
            assert "session.done" in ws.sent_types()
>           assert not [delay for delay in delays if delay > 0], "Restart must not add a timed grace period after abort"
E           AssertionError: Restart must not add a timed grace period after abort
E           assert not [0.1, 0.1, 0.1]

tests/entrypoints/openai_api/test_serving_video_stream.py:592: AssertionError
---------------------------- Captured stdout setup -----------------------------
INFO 09-10 16:22:02 [scheduler.py:277] Chunked prefill is enabled with max_num_batched_tokens=2048.
INFO 09-10 16:22:02 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
--- Running test: test_interrupted_queries_wait_for_abort_without_fixed_delay[immediate-ack]
------------------------------ Captured log setup ------------------------------
INFO     vllm.config.scheduler:scheduler.py:277 Chunked prefill is enabled with max_num_batched_tokens=2048.
INFO     vllm.config.kernel:kernel.py:369 Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
----------------------------- Captured stdout call -----------------------------
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-d2f1740b7793
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-52d94c1c216d
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-bf3208731542
INFO 09-10 16:22:02 [video_stream_base.py:826] [TIMING] mode=on total=0.00s first_text=0.00s first_audio=-1.00s audio_chunks=0
------------------------------ Captured log call -------------------------------
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-d2f1740b7793
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-52d94c1c216d
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-bf3208731542
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:826 [TIMING] mode=on total=0.00s first_text=0.00s first_audio=-1.00s audio_chunks=0
___ test_interrupted_queries_wait_for_abort_without_fixed_delay[delayed-ack] ___

monkeypatch = <_pytest.monkeypatch.MonkeyPatch object at 0x7f405c1af9e0>
delay_abort = True

    @pytest.mark.asyncio
    @pytest.mark.parametrize("delay_abort", [False, True], ids=["immediate-ack", "delayed-ack"])
    async def test_interrupted_queries_wait_for_abort_without_fixed_delay(monkeypatch, delay_abort):
        started: asyncio.Queue[str] = asyncio.Queue()
        abort_started: asyncio.Queue[str] = asyncio.Queue()
        allow_abort = asyncio.Event()
        if not delay_abort:
            allow_abort.set()
        request_ids: list[str] = []
        closed: list[str] = []
        aborted: list[str] = []
        delays: list[float] = []
        original_sleep = asyncio.sleep
    
        async def record_sleep(delay, result=None):
            # Observe requested delays without a machine-dependent latency limit.
            delays.append(delay)
            return await original_sleep(0, result)
    
        monkeypatch.setattr(video_stream_base.asyncio, "sleep", record_sleep)
    
        class BlockingEngine:
            async def generate(self, *, request_id, **kwargs):
                if request_ids:
                    assert aborted[-1] == request_ids[-1]
                request_ids.append(request_id)
                started.put_nowait(request_id)
                try:
                    if len(request_ids) <= 3:
                        yield _text_result("partial")
                        await asyncio.Event().wait()
                    else:
                        yield _text_result("final")
                finally:
                    closed.append(request_id)
    
            async def abort(self, request_id):
                assert request_id in closed
                abort_started.put_nowait(request_id)
                await allow_abort.wait()
                aborted.append(request_id)
    
        class PreprocessedHandler(QwenOmniStreamingVideoHandler):
            async def _preprocess_to_engine_prompt(self, request):
                return {"prompt_token_ids": [1]}
    
        ws = TimedWebSocket()
        handler = PreprocessedHandler(chat_service=object(), engine_client=BlockingEngine(), idle_timeout=5.0)
        task = asyncio.create_task(handler.handle_session(ws))
        try:
            ws.put({"type": "session.config", "modalities": ["text"], "enable_frame_filter": False})
            ws.put({"type": "video.frame", "data": _b64(_make_jpeg())})
            ws.put({"type": "video.query", "text": "first"})
            previous_id = await asyncio.wait_for(started.get(), timeout=2.0)
    
            for _ in range(3):
                ws.put({"type": "video.query", "text": "interrupt"})
                assert await asyncio.wait_for(abort_started.get(), timeout=2.0) == previous_id
                if delay_abort:
                    assert previous_id not in aborted
                    assert started.empty()
                    allow_abort.set()
                previous_id = await asyncio.wait_for(started.get(), timeout=2.0)
                if delay_abort:
                    allow_abort.clear()
    
            ws.put({"type": "video.done"})
            await asyncio.wait_for(task, timeout=2.0)
    
            assert aborted == request_ids[:-1]
            assert len(set(request_ids)) == 4
            assert [msg["text"] for msg in ws.sent if msg["type"] == "response.text.done"] == ["final"]
            assert "error" not in ws.sent_types()
            assert "session.done" in ws.sent_types()
>           assert not [delay for delay in delays if delay > 0], "Restart must not add a timed grace period after abort"
E           AssertionError: Restart must not add a timed grace period after abort
E           assert not [0.1, 0.1, 0.1]

tests/entrypoints/openai_api/test_serving_video_stream.py:592: AssertionError
---------------------------- Captured stdout setup -----------------------------
--- Running test: test_interrupted_queries_wait_for_abort_without_fixed_delay[delayed-ack]
----------------------------- Captured stdout call -----------------------------
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-d13898e800e1
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-83635c293678
INFO 09-10 16:22:02 [video_stream_base.py:291] Interrupt signaled for video-c09c920fe120
INFO 09-10 16:22:02 [video_stream_base.py:826] [TIMING] mode=on total=0.00s first_text=0.00s first_audio=-1.00s audio_chunks=0
------------------------------ Captured log call -------------------------------
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-d13898e800e1
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-83635c293678
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:291 Interrupt signaled for video-c09c920fe120
INFO     vllm_omni.entrypoints.openai.video_stream_base:video_stream_base.py:826 [TIMING] mode=on total=0.00s first_text=0.00s first_audio=-1.00s audio_chunks=0
=============================== warnings summary ===============================
.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  <checkout>/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

vllm_omni/version.py:55
  <checkout>/vllm_omni/version.py:55: RuntimeWarning: vLLM and vLLM-Omni appear to have mismatched major/minor versions:
   --> vLLM-Omni version 0.28.0rc2.dev75+g58cb8de68
   --> vLLM version 0.29.0
  This will likely cause compatibility issues.
    warn_if_misaligned_vllm_version()

.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24
  <checkout>/.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24: DeprecationWarning: `vllm.entrypoints.openai.api_server` is deprecated and will likely beunsupported in a future version. Use the corresponding function from `vllm.entrypoints.launchers` instead.
    warnings.warn(

vllm_omni/entrypoints/duplex/audio.py:15
  <checkout>/vllm_omni/entrypoints/duplex/audio.py:15: DeprecationWarning: 'audioop' is deprecated and slated for removal in Python 3.13
    from audioop import alaw2lin, lin2alaw, lin2ulaw, ulaw2lin

vllm_omni/entrypoints/openai/protocol/audio.py:412
  <checkout>/vllm_omni/entrypoints/openai/protocol/audio.py:412: PydanticDeprecatedSince20: Support for class-based `config` is deprecated, use ConfigDict instead. Deprecated in Pydantic V2.0 to be removed in V3.0. See Pydantic V2 Migration Guide at https://errors.pydantic.dev/2.13/migration/
    class CreateAudio(BaseModel):

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
--- Running Summary
=========================== short test summary info ============================
FAILED tests/entrypoints/openai_api/test_serving_video_stream.py::test_interrupted_queries_wait_for_abort_without_fixed_delay[immediate-ack]
FAILED tests/entrypoints/openai_api/test_serving_video_stream.py::test_interrupted_queries_wait_for_abort_without_fixed_delay[delayed-ack]
2 failed, 21 deselected, 18 warnings in 5.96s
Combined focused regression
Running 10 items in this shard
..........                                                               [100%]
=============================== warnings summary ===============================
.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  <checkout>/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

vllm_omni/version.py:55
  <checkout>/vllm_omni/version.py:55: RuntimeWarning: vLLM and vLLM-Omni appear to have mismatched major/minor versions:
   --> vLLM-Omni version 0.28.0rc2.dev75+g58cb8de68
   --> vLLM version 0.29.0
  This will likely cause compatibility issues.
    warn_if_misaligned_vllm_version()

.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24
  <checkout>/.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24: DeprecationWarning: `vllm.entrypoints.openai.api_server` is deprecated and will likely beunsupported in a future version. Use the corresponding function from `vllm.entrypoints.launchers` instead.
    warnings.warn(

vllm_omni/entrypoints/duplex/audio.py:15
  <checkout>/vllm_omni/entrypoints/duplex/audio.py:15: DeprecationWarning: 'audioop' is deprecated and slated for removal in Python 3.13
    from audioop import alaw2lin, lin2alaw, lin2ulaw, ulaw2lin

vllm_omni/entrypoints/openai/protocol/audio.py:412
  <checkout>/vllm_omni/entrypoints/openai/protocol/audio.py:412: PydanticDeprecatedSince20: Support for class-based `config` is deprecated, use ConfigDict instead. Deprecated in Pydantic V2.0 to be removed in V3.0. See Pydantic V2 Migration Guide at https://errors.pydantic.dev/2.13/migration/
    class CreateAudio(BaseModel):

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
--- Running Summary
10 passed, 13 deselected, 18 warnings in 7.58s
Full relevant suite
Running 198 items in this shard
........................................................................ [ 36%]
........................................................................ [ 72%]
......................................................                   [100%]
=============================== warnings summary ===============================
.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  <checkout>/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

vllm_omni/version.py:55
  <checkout>/vllm_omni/version.py:55: RuntimeWarning: vLLM and vLLM-Omni appear to have mismatched major/minor versions:
   --> vLLM-Omni version 0.28.0rc2.dev75+g58cb8de68
   --> vLLM version 0.29.0
  This will likely cause compatibility issues.
    warn_if_misaligned_vllm_version()

.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24
  <checkout>/.venv/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py:24: DeprecationWarning: `vllm.entrypoints.openai.api_server` is deprecated and will likely beunsupported in a future version. Use the corresponding function from `vllm.entrypoints.launchers` instead.
    warnings.warn(

vllm_omni/entrypoints/duplex/audio.py:15
  <checkout>/vllm_omni/entrypoints/duplex/audio.py:15: DeprecationWarning: 'audioop' is deprecated and slated for removal in Python 3.13
    from audioop import alaw2lin, lin2alaw, lin2ulaw, ulaw2lin

vllm_omni/entrypoints/openai/protocol/audio.py:412
  <checkout>/vllm_omni/entrypoints/openai/protocol/audio.py:412: PydanticDeprecatedSince20: Support for class-based `config` is deprecated, use ConfigDict instead. Deprecated in Pydantic V2.0 to be removed in V3.0. See Pydantic V2 Migration Guide at https://errors.pydantic.dev/2.13/migration/
    class CreateAudio(BaseModel):

.venv/lib/python3.12/site-packages/fastapi/testclient.py:1
  <checkout>/.venv/lib/python3.12/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
    from starlette.testclient import TestClient as TestClient  # noqa

tests/entrypoints/openai_api/test_video_stream_session.py: 29 warnings
tests/entrypoints/openai_api/test_video_frame_filter.py: 5 warnings
  <checkout>/tests/entrypoints/openai_api/conftest_video.py:28: DeprecationWarning: 'mode' parameter is deprecated and will be removed in Pillow 13 (2026-10-15)
    img = Image.fromarray(arr, "RGB")

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
--- Running Summary
198 passed, 2 deselected, 53 warnings in 14.67s
Local pre-commit, including mypy
check yaml................................................(no files to check)Skipped
debug statements (python).....................................................Passed
fix end of files..............................................................Passed
mixed line ending.............................................................Passed
trim trailing whitespace......................................................Passed
ruff check....................................................................Passed
ruff format...................................................................Passed
typos.........................................................................Passed
Lint GitHub Actions workflow files........................(no files to check)Skipped
markdownlint-cli2.........................................(no files to check)Skipped
Run mypy for Python 3.10......................................................Passed
Ensure test files have CI marks (no direct SKU)...............................Passed
Ratchet TTS per-model branches out of serving_speech.py...(no files to check)Skipped
- hook id: check-tts-adapter-migration
Lint shell scripts........................................(no files to check)Skipped
Check SPDX headers............................................................Passed
Check forbidden imports.......................................................Passed
Prevent new torch.cuda API calls..............................................Passed
Validate Buildkite Pipelines..............................(no files to check)Skipped
Suggestion....................................................................Passed
- hook id: suggestion
- duration: 0s

To bypass all the pre-commit hooks, add --no-verify to git commit. To skip a specific hook, prefix the commit command with SKIP=<hook-id>.
Additional controlled AsyncOmni / abort RPC diagnostic on fixed source
<checkout>/vllm_omni/version.py:55: RuntimeWarning: vLLM and vLLM-Omni appear to have mismatched major/minor versions:
 --> vLLM-Omni version 0.28.0rc2.dev75+g58cb8de68
 --> vLLM version 0.29.0
This will likely cause compatibility issues.
  warn_if_misaligned_vllm_version()
INFO 09-10 16:25:18 [patch.py:252] NVFP4 W4A4 weight_scale NaN-clamp: installed.
INFO 09-10 16:25:19 [patch.py:503] inductor factorable-divisibility patch: installed.
INFO 09-10 16:25:19 [patch.py:568] [cumem-cuda] CuMemAllocator._python_free_callback patched: asleep guard extended to all platforms.
INFO 09-10 16:25:25 [video_stream_base.py:291] Interrupt signaled for video-7d8e2d221006
PASS: ack held for 150ms; second ADD absent; first frontend state retained
INFO 09-10 16:25:25 [async_omni.py:665] [AsyncOmni] Request video-7d8e2d221006-a5284375 aborted.
INFO 09-10 16:25:25 [video_stream_base.py:825] [TIMING] mode=on total=0.00s first_text=-1.00s first_audio=-1.00s audio_chunks=0
PASS: first frontend state removed before second ADD; ack -> ADD = 1.256ms; skip_sleep=False
INFO 09-10 16:25:25 [video_stream_base.py:291] Interrupt signaled for video-9e0605e16196
PASS: ack held for 150ms; second ADD absent; first frontend state retained
INFO 09-10 16:25:25 [async_omni.py:665] [AsyncOmni] Request video-9e0605e16196-84f51939 aborted.
INFO 09-10 16:25:25 [video_stream_base.py:825] [TIMING] mode=on total=0.00s first_text=-1.00s first_audio=-1.00s audio_chunks=0
PASS: first frontend state removed before second ADD; ack -> ADD = 0.741ms; skip_sleep=False
INFO 09-10 16:25:25 [video_stream_base.py:291] Interrupt signaled for video-bc326f030b82
PASS: ack held for 150ms; second ADD absent; first frontend state retained
INFO 09-10 16:25:26 [async_omni.py:665] [AsyncOmni] Request video-bc326f030b82-8e4bebe3 aborted.
INFO 09-10 16:25:26 [video_stream_base.py:825] [TIMING] mode=on total=0.00s first_text=-1.00s first_audio=-1.00s audio_chunks=0
PASS: first frontend state removed before second ADD; ack -> ADD = 0.738ms; skip_sleep=False
Runs=3; delays_ms=[1.256, 0.741, 0.738]

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/profiling.md, docs/design/module/entrypoints.md, docs/design/module/observability.md.

Module owners: @alex-jw-brooks @linyueqian @NickCao

Routing: @alex-jw-brooks via module of the changed files, semantic router, CODEOWNERS; @linyueqian via module of the changed files, semantic router, CODEOWNERS; @NickCao via module of the changed files, semantic router, CODEOWNERS

@psv666, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit ad2d01f161b9 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: psv666 <2693925048@qq.com>
@psv666 psv666 changed the title [Bugfix] Prevent evicted video frames from repopulating prewarm cache [Bugfix] Fix video prewarm cache retention and cancel-restart delay Sep 10, 2026
@psv666

psv666 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

@Gaohan123 @natureofnature PTAL

@natureofnature

Copy link
Copy Markdown
Collaborator

LGTM from static review of head a44fb5f. I found no blocking P1/P2 issues.

  • The cache fix addresses the reported race. Checking frame-buffer membership after decoding prevents an evicted frame from repopulating the cache. There is no await between the membership check and publication. Both successful decoding and failure-marker publication are guarded.

  • Removing the fixed 100 ms delay is reasonable. The handler still cancels and awaits the previous query task and awaits abort() before starting the next query. The actual AsyncOmni.generate() cancellation path also awaits Orchestrator abort acknowledgment. A fixed sleep provides no additional completion guarantee.

  • The scope is appropriate. The additional typing changes clarify the cache/sentinel contract. The regression tests check actual image collectability, repeated interruptions, and delayed abort acknowledgment.

One non-blocking suggestion: add a regression where a frame is evicted while decoding is pending and decoding subsequently fails. This would exercise the new early-return branch and verify that it neither retains a _BAD_FRAME entry nor emits a stale decode-failure event.

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 10, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123
Gaohan123 enabled auto-merge (squash) September 10, 2026 13:52
@Gaohan123
Gaohan123 merged commit c8bef00 into vllm-project:main Sep 10, 2026
6 of 9 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants