Repository navigation
[Core] Split Omni connector model runner mixin - #6903
hsliuustc0106 merged 2 commits into
Conversation
4a915f5 to
603cce7
Compare
|
This PR touches vllm_omni/distributed/, vllm_omni/worker/, tests/worker/, vllm_omni/utils/ (7 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be: Could one of you take a look when you get a chance? Thanks! |
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
603cce7 to
93a7221
Compare
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Codex Review: Didn't find any major issues. Delightful! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/omni_connector.md, docs/design/module/input_output_modality_contracts.md, docs/design/module/diffusion/parallelism.md. Module owners: @fake0fan @princepride @xuechendi Routing: @fake0fan via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @princepride via module of the changed files, module named in the PR description, semantic router; @xuechendi via module of the changed files, module named in the PR description, semantic router @natureofnature, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Reviewed One non-blocking design clarification: these are still coupled pieces of one mixin. At runtime.py:133–137, the runtime expects payload-defined hooks, including the I/O loops and full-payload flush. Initialization and cleanup call back into that subclass. Please document this host contract and the runner-adapter boundary in the class/module documentation, particularly since the move changes CODEOWNERS routing. Typed hooks would help future maintenance; a composition redesign is unnecessary for this extraction. Static comparison preserved all 86 method names (87 definitions counting the property getter/setter separately), with 86 definitions AST-identical. The sole method change is Before merging, please update the test evidence to identify the exact tested commit and commands/results: the PR body currently names |
Remove the dormant span-aware accumulation branch, both payload_span helper modules, and their dedicated tests. Built-in payload producers do not supply span bounds. Keep ordinary tensor/list accumulation, explicit overrides, endpoint schema fields, and metadata-only consumability rules. Add focused tests for retained accumulation behavior. Validation: 519 passed, 2 skipped in CPU regression tests; four real-weight Qwen3-Omni and Bagel E2Es passed. Independent reviews found no blockers. Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk> Signed-off-by: wenjie.yan <wenjyan@outlook.com>
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763) Signed-off-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> * [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253) Signed-off-by: liangmengh <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794) Signed-off-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> * [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240) Signed-off-by: Tianyao Wu <rayroy31@gmail.com> * [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Doc] Add AI usage policy for contributions (vllm-project#7305) Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> * [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539) Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> * [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730) Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> * [Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution (vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> * [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> * [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051) Signed-off-by: zjli2013 <leezhengjiang@126.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003) Signed-off-by: ZenAlexa <zimingwang945@gmail.com> * [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> * [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076) Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> * [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896) Signed-off-by: xutianle <xutianle@fudan.edu.cn> * [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Core] Split Omni connector model runner mixin (vllm-project#6903) Signed-off-by: natureofnature <wzliu@connect.hku.hk> * [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231) Signed-off-by: mglyn <1203789601@qq.com> * [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299) Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> * [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289) Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> * [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308) Signed-off-by: mglyn <1203789601@qq.com> * [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185) Signed-off-by: NumberWan <wantszkin2003@gmail.com> * [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685) Signed-off-by: zouyizhou <zouyizhou@huawei.com> * [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328) Signed-off-by: ZhengWG <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200) Signed-off-by: kunkunblueberry <1833921874@qq.com> * [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741) Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Nick Cao <ncao@redhat.com> * [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007) Signed-off-by: mershi <mershi@tencent.com> Co-authored-by: mershi <mershi@tencent.com> * [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155) Signed-off-by: hyw <yuweih205@gmail.com> * [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202) Signed-off-by: Sy03 <1370724210@qq.com> * [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301) Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> * [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292) Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336) Signed-off-by: andyluo7 <andy.luo@amd.com> * [NPU][CI] Add A5 and 310P CI support (vllm-project#6875) Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> * [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350) Signed-off-by: mglyn <1203789601@qq.com> * [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344) Signed-off-by: Sy03 <1370724210@qq.com> * [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230) Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453) Signed-off-by: herotai214 <herotai214@gmail.com> * [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356) Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597) Signed-off-by: wangyu <410167048@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: Samit <285365963@qq.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: Samit <285365963@qq.com> * [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI] Isolate layerwise offload memory measurements (vllm-project#6938) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560) Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> * [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362) Signed-off-by: tly <2200895168@qq.com> * [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363) Signed-off-by: psv666 <2693925048@qq.com> * Cosmos3 action policy improvements (vllm-project#6460) Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561) Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> * [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281) Signed-off-by: david6666666 <530634352@qq.com> * [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930) Signed-off-by: yancaocn <yancaochn@163.com> Co-authored-by: yancaocn <yancaochn@163.com> * [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005) Signed-off-by: samithuang <285365963@qq.com> * [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559) Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172) Signed-off-by: hyw <yuweih205@gmail.com> * [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171) Signed-off-by: hyw <yuweih205@gmail.com> * [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167) Signed-off-by: hyw <yuweih205@gmail.com> * [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841) Signed-off-by: wtz2333 <2955110911@qq.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> * [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853) Signed-off-by: XIN GAO <1037396230@qq.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430) Signed-off-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Codex <noreply@openai.com> * [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Co-authored-by: Sy03 <1370724210@qq.com> * [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441) Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427) Signed-off-by: Maciej Bala <mbala@nvidia.com> * [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094) Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390) Signed-off-by: Guangjian <hiro20833@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167) Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> * [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889) Signed-off-by: psv666 <2693925048@qq.com> * [NPU] upgrade to v0.29.0 (vllm-project#7433) Signed-off-by: Weiming Liao <liaowm5@gmail.com> * [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128) Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> * [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907) Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876) Signed-off-by: gerayking <399geray@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982) Signed-off-by: Qihan Kang <rollykanggg@gmail.com> * [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> --------- Signed-off-by: Asthenia <asthenia0412@gmail.com> Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: liangmengh <liangmengh@nvidia.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Signed-off-by: KrystalRay <keeleiray@gmail.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com> Signed-off-by: specture724 <specture724@gmail.com> Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: zjli2013 <leezhengjiang@126.com> Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Signed-off-by: ZenAlexa <zimingwang945@gmail.com> Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Signed-off-by: Allen Wu <allenwu2795@gmail.com> Signed-off-by: xutianle <xutianle@fudan.edu.cn> Signed-off-by: wangyu <410167048@qq.com> Signed-off-by: natureofnature <wzliu@connect.hku.hk> Signed-off-by: mglyn <1203789601@qq.com> Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Signed-off-by: NumberWan <wantszkin2003@gmail.com> Signed-off-by: zouyizhou <zouyizhou@huawei.com> Signed-off-by: ZhengWG <zwg0606@gmail.com> Signed-off-by: kunkunblueberry <1833921874@qq.com> Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Signed-off-by: mershi <mershi@tencent.com> Signed-off-by: hyw <yuweih205@gmail.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> Signed-off-by: Guangjian <hiro20833@gmail.com> Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Signed-off-by: herotai214 <herotai214@gmail.com> Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Signed-off-by: Samit <285365963@qq.com> Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Signed-off-by: tly <2200895168@qq.com> Signed-off-by: psv666 <2693925048@qq.com> Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Signed-off-by: Nick Cao <ncao@redhat.com> Signed-off-by: david6666666 <530634352@qq.com> Signed-off-by: yancaocn <yancaochn@163.com> Signed-off-by: samithuang <285365963@qq.com> Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Signed-off-by: wtz2333 <2955110911@qq.com> Signed-off-by: XIN GAO <1037396230@qq.com> Signed-off-by: QI JIA <qi.jia@shengshu.ai> Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Signed-off-by: gerayking <399geray@gmail.com> Signed-off-by: Qihan Kang <rollykanggg@gmail.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Signed-off-by: José Carlos <jose@valendra.tech> Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com> Co-authored-by: liangmenghuang <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Lei Ke <1141466880@qq.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com> Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com> Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com> Co-authored-by: boatman <1930807094@qq.com> Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com> Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com> Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com> Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> Co-authored-by: xutianle <24210290017@m.fudan.edu.cn> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> Co-authored-by: NATURE <wzliu@connect.hku.hk> Co-authored-by: Mu GuanLin <1203789601@qq.com> Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com> Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: NumberWan <wantszkin2003@gmail.com> Co-authored-by: zyz111222 <zouyizhou@huawei.com> Co-authored-by: Zheng Wengang <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com> Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com> Co-authored-by: Nick Cao <ncao@redhat.com> Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com> Co-authored-by: mershi <mershi@tencent.com> Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com> Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com> Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com> Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: Samit <285365963@qq.com> Co-authored-by: wkutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com> Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com> Co-authored-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com> Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com> Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com> Co-authored-by: yancaocn <yancaochn@163.com> Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: wtz2333 <2955110911@qq.com> Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com> Co-authored-by: Qi Jia <kuafou@gmail.com> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: DanaerLee <mrdanaer@gmail.com> Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com> Co-authored-by: junpengw67-max <junpengw67@gmail.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: geray <48796550+gerayking@users.noreply.github.com> Co-authored-by: KANG Qihan <3149604185@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Purpose
Move the 2,300-line
OmniConnectorModelRunnerMixinimplementation underdistributed/omni_connectors/model_runnerwhile keeping the existing workermodule as the public compatibility facade.
The implementation has two private layers with a single inheritance chain:
omni_connector_runtime.py: connector initialization/lifecycle, sharedrequest state, cleanup, rank routing, and KV transfer delegation
omni_connector_payload_transport.py: payload caching/building,full-payload and async-chunk flows, TP fanout, background connector I/O,
and scheduler feedback
These flows stay together because they share the same lock, per-request
queues, request-ID mapping, delivery state, and ordering contract.
The runner-facing Mixin import, class, methods, connector-selection helpers,
and mutable instance state remain available without changing model runners.
Connector configuration, request ordering, retries, cleanup, TP fanout, and
KV routing are unchanged.
Also remove the unused
payload_spanhelper modules, their dedicated tests,and the optional span-aware accumulation branch. Built-in producers do not
populate span bounds. Ordinary tensor/list accumulation, explicit overrides,
endpoint schema fields, and metadata-only payload handling remain intact.
This retires the helper import API and overlap-aware behavior for custom
payloads that explicitly supply spans.
Test Plan
vLLM Version: 0.28.0
vLLM-Omni Commit:
b1235938211333d02aa35e87b5f93117ffd0dbdeUpstream Base:
c704aeccae192d7f88e9138271682ddad1327a69Test Result
519 passed, 2 skipped: connector/runner, schema, scheduler, engine argument,Qwen3 streaming helper, and KV transfer tests. Includes new accumulator
concatenation/snapshot and override regression tests.
2 passed.1 passed.1 passed; returned text andaudio ASR transcription match exactly.
with identical normalized signatures; no new errors.
Real-weight E2E selectors, each run with
python3 -m pytest -s -v --run-level=advanced_model:This is representative regression coverage, not a full E2E suite or a
performance benchmark.
Earlier E2E run on
93a72218, before the span cleanup: