Skip to content

[Refactor] Declare model-local KV held outside the paged manager - #6171

Merged
hsliuustc0106 merged 15 commits into
vllm-project:mainfrom
linyueqian:feat/k3-model-local-kv-inventory
Sep 11, 2026
Merged

hsliuustc0106 merged 15 commits into
vllm-project:mainfrom
linyueqian:feat/k3-model-local-kv-inventory

Conversation

@linyueqian

@linyueqian linyueqian commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Part of #4855 (K3). Four models keep attention KV outside the engine's paged manager, as HuggingFace transformers cache objects. This has each of them declare what it holds and sums it at load, so the memory is visible rather than inferred.

Two rounds of external review shaped this. The first found the multiplicity model wrong and two of my claims unsupported; the second found that my replacement fixed those but introduced a regression of its own. Both sets of corrections are described below, including what I got wrong, because some of it was previously stated as fact in this description.

What each model declares

Model Cache Bounded by Per row Rows
Qwen3-TTS codec decoder (async-chunk only) sliding DynamicCache sliding_window - 1 = 71 4.44 MiB max_num_seqs, twice over for the working copy
Qwen3-TTS graph pool retained captures same 4.44 MiB fixed, from captured shapes
MiMo-Audio local transformer DynamicCache group_size + max(delay_pattern) = 11 704 KiB max_num_seqs
MiMo-Audio graph pool captured same 704 KiB fixed 261 (bucket sum)
MiniCPM-o Whisper encoder (duplex only) EncoderDecoderCache embed_positions rows = 1500 140.62 MiB duplex_max_sessions
ming_flash_omni talker StaticCache hardcoded max_cache_len = 2048 24.00 MiB fixed 1

No cache is bounded by max_model_len. Each has its own mechanism: a sliding window, a decode-loop trip count, an encoder frame limit, a constant. Sizing any of them by sequence length overstates it by orders of magnitude.

ming's geometry does not exist in this repo. self._llm_config is a Qwen2Config resolved from the checkpoint, so layers, kv-heads, head-dim and dtype are only knowable after load. That is what makes this a post-load query rather than a table, and it is the one row a static table could not have held.

Lifetime is not multiplicity

scope records when memory appears and goes. It does not decide how wide the allocation is. Width is declared by naming the driver: Fixed(n, because=...) for an extent the model controls, MaxNumSeqs() or DuplexMaxSessions() for one the engine controls. Fixed rejects a blank reason or a non-positive count, because an unexplained constant cannot be reviewed.

rows counts row-equivalents at peak and deliberately does not distinguish "one allocation N rows wide" from "N allocations of one row". They cost the same, and the object layout goes in allocation_note where it cannot be mistaken for arithmetic.

A declaration is inert when its path is. Qwen3-TTS holds KV only in async-chunk mode, so stateless decoding declares nothing. Where the model cannot see the deciding setting -- omni_config collapses "duplex off" and "duplex with one session" both to duplex_max_sessions=1 -- it states the condition in only_when rather than guessing.

Corrections, first round

The multiplier was wrong in two directions. Deriving it from scope made every non-MODEL cache scale with max_num_seqs. ming's talker is per-call but serialized -- forward takes runtime_additional_information[0] and runs the whole AR loop inline -- so it was over-reported by that factor. MiniCPM-o's encoder is per duplex session, capped by duplex_max_sessions, which is a different setting.

MiMo's bytes were right and its description was not. base_local_forward builds one DynamicCache whose batch dimension is the group size. One allocation of B rows costs what B allocations of one row cost, so the total was correct while the report said "N instances" of something there was one of.

Qwen3-TTS reported zero under enforce_eager. The declaration hung off the CUDA-graph wrapper, but the eager async-chunk path allocates its cache whether or not graphs were captured. The per-decode entry moved to the decoder; the wrapper declares only the graph-resident copies.

spec_from_hf_config defeated its own required field. The dataclass documented the batch extent as required; the helper every declarer uses defaulted it to 1.

I retract the claim that the collector had a bug I caught. The previous description said walking a CUDAGraphWrapper found nothing and that the fix was what made declarations appear. That is false: the wrapper's __getattr__ forwards to its runnable, so named_modules resolves and the original code worked. I inferred a bug from a stage that legitimately has no caches. The collector now uses the supported unwrap() accessor, which also removes a real double-counting path when a root-level model declares.

The accounting motivation was wrong for the largest figures. Graph capture happens during load_model, and determine_available_memory() then samples per-process memory through NVML, so the resident portion is already charged; subtracting the declared total would double-count it. What is genuinely unaccounted is the per-request half, allocated after profiling -- roughly 1.1 GiB for Qwen3-TTS at max_num_seqs=256. The consumer reports and does not subtract.

Corrections, second round

The fixes above were reviewed again, and the replacement had problems of its own.

I introduced a regression, and it is the reason this went back to draft. Having moved the per-decode entry to the decoder, I made it unconditional. Stateless decoding runs _forward_exact, which calls the transformer without use_cache and allocates no KV at all, so that declared roughly a gigabyte of nothing at max_num_seqs=256 -- worse than the revision before it, which reported the correct zero there. Declarations are now gated on async_chunk, which is the condition that actually decides. The eager path also deep-copies a working cache that is live alongside the retained one, so async-chunk mode declares two entries rather than one.

The object-count axis claimed a topology that was wrong in three of four declarers. MiniCPM-o keeps one cache object per session while asserting batch size 1; both graph pools keep one distinct object per captured bucket. All three folded their objects into a single wide allocation. allocations is gone; rows means row-equivalents, and layout is described in prose where it cannot be mistaken for arithmetic.

The "unresolved driver" mechanism was built on a false premise. OmniModelConfig.duplex_max_sessions exists and is read this way elsewhere in the tree. I concluded it did not because I grepped for it while my checkout had lost the file, and then designed around a gap that was not there. The machinery is deleted and the attribute read directly.

One bug fixed, one superseded

The TTS ratchet test asserted count == budget, re-imposing through pytest exactly what #6008 removed from the checker. A PR removing a branch then fails CI unless it also hand-edits the constant, which is what broke main for 5h19m in #5746. Budgets are a ceiling.

Superseded: MiniCPM-o's dead streaming cache. Measuring turned up that the vendored Whisper layer rebound its cache to None on every layer, so the encoder re-encoded its whole prefix each chunk. main fixed this independently while this PR was in draft, and did it better: it selects the kwarg name through self._past_key_values_kwarg rather than hardcoding the transformers v5 spelling, and guards the read-back with if len(attn_out) > 2 so a future 3-tuple still works, where this branch had deleted the branch outright. The merge takes main's version; nothing of that fix remains mine to claim.

Test Plan

pytest tests/model_executor/models/test_model_local_kv.py
pytest tests/tools/test_check_tts_adapter.py

Test Result

35 passed. The suite now calls the real declarers rather than comparing specs written in the test to caches built in the test. Verified by mutation twice: reintroducing the ming multiplier fails test_ming_does_not_scale_with_max_num_seqs, and removing the async-chunk gate fails test_qwen3_tts_declares_nothing_on_the_stateless_path. Both would have passed under the suite this replaced.

ming, against the real checkpoint config (Jonathan1909/Ming-flash-omni-2.0, talker/llm/config.json):

checkpoint geometry: L=24 kv=2 hidden=896 heads=14 dtype=torch.bfloat16
declared bytes_per_instance=25165824 (24.00 MiB)
real StaticCache bytes=25165824 (24.00 MiB)   MATCH

Qwen3-TTS and MiMo-Audio, real weights on an H200, under the previous revision's multiplicity model. Declared resident bytes equalled measured bytes exactly for Qwen3-TTS (5 x 4,653,056 = 23,265,280), and MiMo's 261 rows were read off the live model. Those measurements constrain geometry, which this revision does not change; the row-count rework is covered by the unit tests above and has not been re-run on GPU.

Not verified. ming and MiniCPM-o have not been booted. ming's checkpoint was not available; MiniCPM-o loads and declares but its pipeline needs cosyvoice2, which is absent from the machine I have.

Deliberately not fixed here. MiMo captures its CUDA graphs gated only on torch.cuda.is_available() (mimo_audio_llm.py:670), so 179.44 MiB stays resident even under enforce_eager: true. One line, but verifying it needs a model boot.

  • vLLM version: 0.27.0

Note on splitting

The ratchet fix is ~5 lines and unblocks anyone removing branches from serving_speech.py today. If the abstraction needs more debate, say so and I will pull that fix into its own PR rather than let it wait.

Four models keep their own attention KV state outside the engine's paged
manager (mimo_audio, ming_flash_omni, minicpmo_4_5, nemotron_voicechat).
They are HuggingFace transformers cache objects, not ad-hoc lists, and the
memory they hold is allocated after the profiling run that sized the KV
pool, so no per-stage footprint number accounts for it.

Writes down what exists, with line references, and states the two fixes
that are worth doing regardless of which RFC ends up owning the surface.
Proposes no design: unification is vllm-project#4855 K3 and should be scoped with the
owners of vllm-project#5244 rather than decided TTS-side.

Part of vllm-project#4855.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@linyueqian
linyueqian marked this pull request as draft August 13, 2026 17:40
@hsliuustc0106 hsliuustc0106 added the documentation Improvements or additions to documentation label Aug 14, 2026
MiniCPM-o's streaming audio cache was permanently dead. The vendored Whisper
layer read the cache back out of the attention return value:

    past_key_values = attn_out[2] if len(attn_out) > 2 else None

transformers 4.x returned the cache as a third element; v5 returns only
(attn_output, attn_weights), so attn_out[2] was unreachable and this rebound
the cache to None on every layer -- the encoder re-encoded its whole prefix on
every chunk. v5 Cache objects are mutated in place, so the object passed in is
already current and the read-back is both broken and unnecessary. Measured
against the real transformers 5.14.1 WhisperAttention with two 5-position
chunks: before, cache length None -> None; after, 5 -> 10.

The TTS ratchet test asserted count == budget, re-imposing through pytest what
vllm-project#6008 removed from the checker: with 27 branches and a budget of 27, removing
one branch gives 26 == 27 -> FAIL unless the constant is hand-edited too. That
is what broke main for 5h19m in vllm-project#5746. Budgets are a ceiling.

Adds model_local_kv.py: a post-load declaration protocol. Geometry is not
always static -- ming_flash_omni builds its cache from a checkpoint-side
Qwen2Config -- and no cache is bounded by max_model_len, so only the model
knows its own bound. scope x max_live_instances rather than a single
preallocated/grows flag, because MiMo's 704 KiB cache costs 179 MiB once
replicated per graph bucket. Describes only; allocation stays with
RFC vllm-project#5244 / PR vllm-project#6094.

Part of vllm-project#4855.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian changed the title [Doc] Inventory the model-local KV caches [BugFix][Refactor] Declare model-local KV caches; fix two silent failures Aug 15, 2026
@linyueqian
linyueqian marked this pull request as ready for review August 16, 2026 03:27
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@linyueqian linyueqian added the ready label to trigger buildkite CI label Aug 16, 2026


@runtime_checkable
class HasModelLocalKV(Protocol):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As shipped, nothing implements HasModelLocalKV and nothing in the engine calls total_declared_bytes — the protocol's only exercise is the test constructing specs by hand. All four declarations live as prose in the doc and literals in the test, so changing sliding_window in the Qwen3-TTS checkpoint tomorrow breaks nothing here, which is exactly the staleness the doc argues a table suffers from.

Consider landing one live edge in this PR — the Qwen3-TTS codec decoder is the natural first declarer (geometry fully known, numbers already derived) — or a minimal consumer that logs total_declared_bytes at startup, so the API can't rot unexercised until the K3 consumer lands.

One design question worth deciding before the first declarer freezes the API: for REQUEST scope, multiplicity is engine-side concurrency, not something the model knows — the test hardcodes max_live_instances=1 for a per-request cache. If the engine is expected to multiply by live requests, that contract should be explicit in ModelLocalKVSpec.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both of your suggestions landed since this comment, and you were right about the helper.

Live declarers now exist rather than prose: qwen3_tts/segmented_graph_wrapper.py, qwen3_tts/tokenizer_12hz/modeling_qwen3_tts_tokenizer_v2.py, mimo_audio/mimo_audio_llm.py, ming_flash_omni/talker_module.py with its forwarder in ming_flash_omni_talker.py, and minicpmo_4_5/minicpmo_4_5_omni_llm.py.

The engine consumer is OmniGPUModelRunner._log_model_local_kv (gpu_model_runner.py:215): it collects at startup and logs the total plus one line per declaration, wrapped so a reporting failure cannot break model load. So a sliding_window change in the Qwen3-TTS checkpoint now moves a logged number, and the test asserts the declaration matches the cache it describes.

On total_declared_bytes specifically, it stayed dead even after that consumer landed: the consumer needs the individual specs for its per-declaration lines and sums them inline, so the helper never acquired a production caller and the design doc never referenced it. Removed in 1f971f6. collect_model_local_kv_specs() is the single public entry point now.

On the REQUEST scope question, that contract is now explicit instead of a hardcoded test literal. Multiplicity moved onto ModelLocalKVSpec.rows as RowDriver: MAX_NUM_SEQS means one row per in-flight sequence with the value supplied by the engine, and FIXED carries rows_fixed plus a required rows_reason, validated in __post_init__. peak_bytes(max_num_seqs) and row_count(max_num_seqs) take the engine number as an argument. Scope stayed deliberately independent of size and diagnostic only, so a REQUEST-scoped cache the engine does not widen per sequence declares FIXED and has to say why.

@@ -0,0 +1,56 @@
# Model-local KV caches

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This page isn't linked from docs/design/index.md, so it stays out of the design-doc navigation even though RTD builds it. A one-line entry ("Runtime and stage execution" seems the closest fit) would fix it.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. docs/design/index.md links it under "Runtime and stage execution", which was the fit you suggested.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working refactor refactoring for better code scalability and quality labels Aug 16, 2026
@Gaohan123 Gaohan123 added this to the v0.28.0 milestone Aug 16, 2026
…ume them

The protocol shipped in the previous commit had no implementers and no
consumer, so total_declared_bytes() returned 0 for every model and the tests
asserted hand-typed constants that could not fail. This implements it.

All four cache owners now declare from the live config objects their caches
are built from: the Qwen3-TTS codec decoder, the MiMo-Audio local transformer,
the MiniCPM-o Whisper encoder and the ming_flash_omni talker.
collect_model_local_kv_specs() walks the module tree to find them, so no
intermediate class has to forward the call, and OmniGPUModelRunner.load_model
logs the total after load. GPUARModelRunner inherits it.

Separate model-side replication from engine concurrency. A model cannot know
how many requests are in flight, so max_live_instances now means model-side
copies only and scope decides whether the engine multiplies: MODEL-scoped
caches are captured once and shared, everything else scales with max_num_seqs
via peak_bytes(). Qwen3-TTS stage 1 declares 22.19 MiB resident plus 284 MiB
at max_num_seqs=64, which a single number could not express.

Running it against real weights corrected three things that were derived by
reading code:

- Qwen3-TTS retains five caches, not one. combined_states, icl_prefix_states
  and xvec_prefix_states each keep a DynamicCache alive for graph replay.
  Declared resident bytes now equal measured bytes exactly (23,265,280).
- StaticCache does not preallocate at construction. In transformers v5
  StaticLayer is lazy and a fresh cache measures 0 bytes; the first write to a
  layer then allocates all max_cache_len positions at once.
- collect_model_local_kv_specs() found nothing when the runner's model was a
  CUDAGraphWrapper, which is a plain callable holding .runnable rather than an
  nn.Module. That silently reported zero, the exact failure this is meant to
  surface. It unwraps first now, with a regression test.

Tests build the real transformers caches and compare allocated bytes against
the declaration, so a wrong spec fails.

Link the design doc from docs/design/index.md.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Resolve three conflicts:

- docs/design/index.md: keep both new entries.
- tests/tools/test_check_tts_adapter.py: main removed the legacy-detector
  budget and lowered MAX_MODEL_TYPE_BRANCHES 27 -> 20, so the branch's second
  assertion no longer has symbols to import. Keep main's single assertion but
  retain this branch's <= comparison; an equality assertion re-imposes through
  pytest exactly what vllm-project#6008 removed from the checker.
- vllm_omni/worker/gpu_model_runner.py: auto-merged; the model-local KV
  reporting hook is unchanged.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
"mis-size" trips the typos hook, which reads "mis" as a misspelling of
"miss". Only surfaced now because the branch was conflicting, so pre-commit
had not run on the previous head.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 17, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@FayeSpica PTAL NPU ci failures

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 19, 2026
@linyueqian linyueqian changed the title [BugFix][Refactor] Declare model-local KV caches; fix two silent failures [Refactor] Declare model-local KV held outside the paged manager Aug 19, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

NPU failure:

https://buildkite.com/vllm/vllm-omni-npu-ci/builds/5580/canvas?sid=01a01c6e-e49a-4f70-9d67-26e1d2922442&tab=output

=================================================================== short test summary info ====================================================================
--
  | FAILED tests/platforms/npu/test_minicpmo_code2wav_npugraph.py::test_minicpmo_cfm_estimator_npugraph_matches_eager - AttributeError: '_Flow' object has no attribute 'encoder'. Did you mean: 'decoder'?
  | ================================================================ 1 failed, 14 warnings in 9.26s ================================================================
  | 🚨 Error: The command exited with status 1

…pects

tests/platforms/npu/test_minicpmo_code2wav_npugraph.py fails on main, not just
on this branch: vllm-project#6274 made BatchedToken2Wav.__init__ undecorate
flow.encoder.forward_chunk, while the stub from vllm-project#5604 defines only a decoder,
so the adapter raises AttributeError before any graph-replay assertion runs.

The real CosyVoice2 flow has an encoder, and _undecorate_dynamo returns early
when the method is absent, so an empty namespace restores the test without
changing what it exercises.

Unrelated to the rest of this PR; it happens to be what turns NPU CI red here.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 20, 2026
@@ -0,0 +1,424 @@
# SPDX-License-Identifier: Apache-2.0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not move to the omni_config folder?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no vllm_omni/omni_config/ in the tree, so I read this as vllm_omni/config/. I would rather leave it under model_executor/models/, though I will move it if you still prefer that after this.

The module is a model-side declaration protocol rather than configuration. Models implement model_local_kv_specs() on themselves, and the helpers walk a live model object through named_modules(), unwrapping CUDA-graph wrappers and deduplicating a root declarer seen twice. Nothing in it is parsed from or serialized to a config file, and none of the values come from user input: they are read off the loaded checkpoint config and the model's own captured buckets.

docs/design/module/vllm_omni_config.md scopes that contract to vllm_omni/config/**, vllm_omni/deploy/** and pipeline.py, which is deploy and stage schema plus construction. Putting a Protocol implemented by model classes behind that boundary would separate it from the five declarers that implement it and give the config contract ownership of runtime model introspection.



@dataclass(frozen=True)
class EngineCapacity:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

does it only occupy two class args and one of them related to duplex? this abstraction looks over desgined

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and addressed in b55d434.

EngineCapacity is gone. Once the duplex driver was removed it held a single field, so peak_bytes() takes max_num_seqs directly instead of a wrapper whose only job was to carry it.

The four row classes (ModelLocalKVRows, Fixed, MaxNumSeqs, DuplexMaxSessions) collapsed into one RowDriver enum with two values, plus rows_fixed and rows_reason on the spec. Validation moved into __post_init__, so a FIXED count still cannot ship without a positive value and a stated reason. total_declared_bytes has since gone as well (1f971f6), leaving collect_model_local_kv_specs() as the only public entry point.

…chinery

Review feedback that the abstraction was over-designed, and that its
engine-capacity object was two fields with one of them duplex-specific. Both
land: the surface was carrying a driver, a config object and a conditional
annotation for a single declarer.

Four row classes become one enum. ModelLocalKVRows, Fixed, MaxNumSeqs and
DuplexMaxSessions are replaced by RowDriver with two values, plus rows_fixed
and rows_reason on the spec. Validation moves to __post_init__, so a FIXED
count still cannot ship without a positive value and a reason.

EngineCapacity is gone. With the duplex driver removed it held one field, so
peak_bytes() and total_declared_bytes() take max_num_seqs directly instead of
a wrapper whose only job was to carry it.

only_when is gone. It existed because duplex_max_sessions cannot distinguish
"duplex off" from "one session", which is a problem the protocol no longer has.

MiniCPM-o declares one session as its unit, with the session scaling noted in
prose. Its cache is one object per streaming session at batch size 1; the cap
is not visible from the model, and a cap-scaled figure would report memory a
non-duplex deployment never allocates. The multiplication is left to the
reader, who has the number.

Public names go from 11 to 7, and model_local_kv.py from 424 to 358 lines,
with no capability lost: every declarer expresses exactly what it did before.

Deliberately not justified by "duplex is experimental" -- vllm-project#6196 moves it out.
The argument is that one declarer does not earn a driver.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 22, 2026
Review feedback that nothing in the engine called it. The engine consumer
that landed since, OmniGPUModelRunner._log_model_local_kv, needs the
individual specs for its per-declaration lines and sums them inline, so the
helper had no production caller and was not referenced by the design doc.

collect_model_local_kv_specs stays the single public entry point. The tests
that used the helper as a byte-sum shorthand keep their intent through a
local _total_bytes().

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit d8f31eb79f3a produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 1, 2026
@Gaohan123 Gaohan123 modified the milestones: v0.28.0, v0.30.0 Sep 2, 2026
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 3, 2026
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 3, 2026
…av test

The main merge kept both encoder assignments in the _Flow stub: main's
nn.Identity() and this branch's older SimpleNamespace(). Assigning a
non-Module over a name already registered in _modules fails, so _Flow could
not be constructed and the NPU graph-replay assertions were never reached.

nn.Identity() satisfies the only requirement BatchedToken2Wav.__init__ has:
it reads flow.encoder and calls _undecorate_dynamo, which returns early when
forward_chunk is absent. The second assignment was redundant even before the
merge made it fatal.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 9, 2026
@hsliuustc0106
hsliuustc0106 merged commit 5c69fc2 into vllm-project:main Sep 11, 2026
8 of 9 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working documentation Improvements or additions to documentation high priority high priority issue, needs to be done asap ready label to trigger buildkite CI refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants