Skip to content

[Refactor] Stage Proc/Client for LLM and DiT - #5441

Closed
chickeyton wants to merge 42 commits into
vllm-project:mainfrom
chickeyton:clean_stage_proc_rebase
Closed

chickeyton wants to merge 42 commits into
vllm-project:mainfrom
chickeyton:clean_stage_proc_rebase

Conversation

@chickeyton

@chickeyton chickeyton commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

PLEASE FILL IN THE PR DESCRIPTION HERE.

Purpose

Refactor the following files

vllm_omni/diffusion/
inline_stage_diffusion_client.py
stage_diffusion_client.py
stage_diffusion_proc.py

vllm_omni/engine/
stage_client.py
stage_engine_core_client.py
stage_engine_core_proc_manager.py
stage_engine_core_proc.py
stage_pool.py

There will be subdirectories to contains the refactored code:

vllm_omni/diffusion/stage/
inline_stage_diffusion_client.py
stage_diffusion_client.py -> stage_diffusion_core_client.py
stage_diffusion_proc.py -> stage_diffusion_core_proc.py
stage_diffusion_core_proc_manager.py

vllm_omni/engine/stage/
stage_core_types.py
stage_client.py -> stage_core_client.py
stage_engine_core_client.py -> stage_llm_core_client.py
stage_engine_core_proc_manager.py -> stage_llm_core_proc_manager.py
stage_engine_core_proc.py -> stage_llm_core_proc.py
stage_pool -> stage_replica_pool.py
  • Better naming, accurate and informative
  • Remove unused, duplicated or redundant code
  • More elegant OOP structures

Inheritance Diagrams

flowchart BT
    ECP["EngineCoreProc"]
    SLCP["StageLLMCoreProc"]
    SDCP["StageDiffusionCoreProc"]

    CEPM["CoreEngineProcManager"]
    SLCPM["StageLLMCoreProcManager"]
    SDCPM["StageDiffusionCoreProcManager"]

    SLCP --> ECP
    SLCPM --> CEPM
    SLCPM -. "contains" .-o SLCP
    SDCPM -. "contains" .-o SDCP
Loading
flowchart BT
    AMPC["AsyncMPClient"]
    DAMPC["DPLBAsyncMPClient"]

    SCCB["StageCoreClientBase"]
    SLCCB["StageLLMCoreClientBase"]
    SLCC["StageLLMCoreClient"]
    DSLCC["DPLBStageLLMCoreClient"]
    SDCC["StageDiffusionCoreClient"]
    ISDC["InlineStageDiffusionClient"]

    SLCC  --> SCCB
    SLCCB --> SCCB
    DSLCC --> SLCCB 
    DSLCC --> DAMPC
    SLCC --> AMPC 
    SLCC --> SLCCB
    SDCC --> SCCB
    ISDC --> SCCB
Loading
flowchart BT
    ECR["EngineCoreRequest"]
    ECO["EngineCoreOutput"]
    ECOS["EngineCoreOutputs"]

    SCR["StageCoreRequest"]
    SCO["StageCoreOutput"]
    SCOS["StageCoreOutputs"]

    SLCR["StageLLMCoreRequest"]
    SLCO["StageLLMCoreOutput"]
    SLCOS["StageLLMCoreOutputs"]

    SDCR["StageDiffusionCoreRequest"]
    SDCO["StageDiffusionCoreOutput"]
    SDCOS["StageDiffusionCoreOutputs"]

    SLCR --> ECR
    SLCR --> SCR

    SLCO --> ECO
    SLCO --> SCO

    SLCOS --> ECOS
    SLCOS --> SCOS

    SDCR --> SCR
    SDCO --> SCO
    SDCOS --> SCOS
Loading

Changes

stage_core_types.py

  • Add StageCoreRequest, StageCoreOutput, StageCoreOutputs
  • Add StageLLMCoreRequest, StageLLMCoreOutput, StageLLMCoreOutputs
  • Add StageDiffusionCoreRequest, StageDiffusionCoreOutput, StageDiffusionCoreOutputs

stage_client.py -> stage_core_client.py

  • StagePoolClient, StagePoolLLMClient, StagePoolDiffusionClient not used, remove them
  • rename StageClientBase to StageCoreClientBase
  • update interface of StageCoreClientBase

stage_engine_core_client.py -> stage_llm_core_client.py

  • rename StageEngineCoreClientBase to StageLLMCoreClientBase (subclass from StageCoreClientBase)
  • rename StageEngineCoreClient to StageLLMCoreClient
  • rename DPLBStageEngineCoreClient to DPLBStageLLMCoreClient
  • make StageLLMCoreClientBase subclass from StageCoreClientBase
  • rename _default_process_engine_inputs to _default_process_core_inputs as a static function of StageLLMCoreClientBase
  • remove the monkey patch applied to vllm as which exists in patch.py already:
  • replace the references of EngineCoreRequest and EngineCoreOutputs with StageLLMCoreRequest and StageLLMCoreOutputs
  • rename _kv_sender_host to _core_host, _resolve_contact_host to _resolve_core_host
  • use concrete typing for arguments of engine_manager (renamed to proc_manager) and coodinator in make_async_mp_client:
proc_manager: StageLLMCoreProcManager | None
coordinator: DPCoordinator | None

engine/init.py

  • remove OmniEngineCoreOutput and OmniEngineCoreOutputs

patch.py

  • replace the references of OmniEngineCoreOutput, OmniEngineCoreOutputs and OmniEngineCoreRequest by StageLLMCoreOutput,
    StageLLMCoreOutputs and StageLLMCoreRequest

stage_engine_core_proc_manager.py -> stage_llm_core_proc_manager.py

  • rename StageEngineCoreProcManager to StageLLMCoreProcManager
  • in __init__, argument start_index (the configurated dp rank) and local_start_index are always set to 0 by callers, to be removed
  • rename local_engine_count to local_proc_count
  • remove local_dp_rank as not needed by StageLLMCoreProc

stage_engine_core_proc.py -> stage_llm_core_proc.py

  • rename StageEngineCoreProc to StageLLMCoreProc
  • Argument local_dp_rank in run_stage_core is not used, to be removed
  • remove the monkey patch for vllm, which is already exists in patch.py

stage_pool.py -> stage_replica_pool.py

  • rename StagePool to StageReplicaPool

inline_stage_diffusion_client.py

  • make InlineStageDiffusionClient subclass from StageClientBase

stage_diffusion_client.py -> stage_diffusion_core_client.py

  • rename StageDiffusionClient to StageDiffusionCoreClient and subclass from StageClientBase

stage_diffusion_proc.py -> stage_diffusion_core_proc.py

  • rename StageDiffusionProc to StageDiffusionCoreProc
  • move StageDiffusionProcManager to stage_diffusion_core_proc_manager.py

stage_diffusion_core_proc_manager.py

  • rename StageDiffusionProcManager to StageDiffusionCoreProcManager

Test Plan

vLLM Version: v0.25.0

Environment

Item Value
Host Test server 1
GPU 1× NVIDIA L20X (both stages colocated on one GPU)
vLLM 0.25.0

Model

  • Model: ByteDance-Seed/BAGEL-7B-MoT
  • Topology: two-stage (default)
    • Stage 0 — Thinker (AR / LLM): runs on the vLLM AR engine via the relocated StageLLMCoreProc.
    • Stage 1 — DiT (Diffusion): runs out-of-process via the relocated StageDiffusionCoreProc / StageDiffusionCoreClient / StageDiffusionCoreProcManager; the KV/latent handoff from Stage 0 crosses the ZMQ boundary through SharedMemoryConnector.
  • Task: TI2I / image-edit — an input image plus a text instruction produce an edited image.

Replica & parallel settings

Setting Stage 0 (Thinker) Stage 1 (DiT)
Replicas 1 1
devices "0" "0" (colocated with Stage 0)
Tensor parallel 1 1
Sequence / CFG / USP parallel — none (DiffusionParallelConfig() defaults)
Data parallel 1 (DP > 1 unsupported for the omni stage LLM engine) —
max_num_seqs 3 1
max_num_batched_tokens 32768 32768
gpu_memory_utilization 0.45 (default)
enforce_eager — true
enable_prefix_caching false false

No tensor / sequence / data parallelism was used — this is the single-GPU,
single-replica-per-stage default. The diffusion side uses DiffusionParallelConfig()
with its default (TP=1, no SP/CFG/USP parallel).

YAML config (vllm_omni/deploy/bagel.yaml)

async_chunk: false

stages:
  - stage_id: 0
    max_num_batched_tokens: 32768
    max_num_seqs: 3
    gpu_memory_utilization: 0.45
    trust_remote_code: true
    enable_prefix_caching: false
    devices: "0"
    default_sampling_params:
      temperature: 0.4
      top_p: 0.9
      top_k: 1
      max_tokens: 2048
      seed: 52
      detokenize: true
      repetition_penalty: 1.05

  - stage_id: 1
    max_num_batched_tokens: 32768
    max_num_seqs: 1
    enforce_eager: true
    trust_remote_code: true
    enable_prefix_caching: false
    devices: "0"
    input_connectors:
      from_stage_0: shared_memory_connector
    default_sampling_params:
      seed: 52

connectors:
  shared_memory_connector:
    name: SharedMemoryConnector

  rdma_connector:
    name: MooncakeTransferEngineConnector
    extra:
      host: "auto"
      zmq_port: 50051
      protocol: "rdma"
      device_name: ""
      memory_pool_size: 4294967296
      memory_pool_device: "cpu"

Launch / run command

# Server
vllm serve ByteDance-Seed/BAGEL-7B-MoT --omni --port 8091 \
    --deploy-config vllm_omni/deploy/bagel.yaml

# Request (img2img / TI2I)
IMAGE_BASE64=$(base64 -w 0 bagel_02_input.png)
curl http://localhost:8091/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "<|im_start|>Turn the main subject into a polished gold sculpture.<|im_end|>"},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,'"${IMAGE_BASE64}"'"}}
      ]
    }],
    "modalities": ["image"],
    "height": 512, "width": 512,
    "num_inference_steps": 30, "seed": 42
  }'

Example (1 of 10 jobs)

Input text (edit instruction):

Subject re-rendered as a polished gold sculpture; city-street setting preserved.

Input image — bagel_02_input.png (a green fire hydrant with a yellow cap and white leaf motif on a city sidewalk):

bagel_02_input

Output image — bagel_02.png, 512×512:

bagel_02

Test Result

Loaded in 175.4s  class=BagelPipeline  stages=2
[JOB  1] OK    21.9s  size=(512, 512)  -> bagel_01.png   (first job incl. warmup)
[JOB  2..10] OK ~8.6-8.8s each         -> bagel_02..10.png
BAGEL TI2I DONE: 10/10 succeeded
  • TI2I e2e: 10/10 images generated, all 512×512, subject preserved with the
    requested background/style edit applied.
  • Confirms the relocated out-of-process diffusion stage
    (StageDiffusionCoreProc/Client/ProcManager) and Thinker stage
    (StageLLMCoreProc) work end-to-end through the orchestrator after the move.
  • Also validated on this environment: 175/176 unit tests pass across the
    relocated stage modules, and single-stage Qwen-Image-Edit = 3/3.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@chickeyton

chickeyton commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Pre-check report

  • Mode: full
  • Type: general (refactor) — relocates stage proc/client into vllm_omni/engine/stage/ + vllm_omni/diffusion/stage/; the diffusion/ touches are module relocations, not a new diffusion model.
  • Base: origin/main @ 2bb61890
  • HEAD: 43acd350 · Diff: 49 files · 19 commits (all DCO-signed)
  • Re-run: 2026-07-27 (after fixing both code findings)

Dimension summary

Dimension Result Change since first run
PR title format ✗ unchanged (process, not code)
Code quality — kwargs ✓ —
Code quality — broad except ✓ — (relocated; metrics fix removed one)
Code quality — Any hints ✓ ⚠ → ✓ (2 msgspec fields now documented-intentional)
Code quality — hot-path copy ✓ —
Code quality — event-loop blocking ✓ —
Metrics-helper divergence from main ✓ ⚠ → ✓ (adopted centralized helpers)
Commit hygiene (squash) ⚠ unchanged (now 19 commits)
Dead code ✓ —
Test coverage ✓ — (176 unit still green after fix)
Rebase / CI gates ✓ 19 ahead / 0 behind; full pre-commit clean
Scope coherence ✓ —

Verdict: 1 blocking (PR title) · 1 warning (squash). Both code issues from the first run are fixed and validated. The only remaining items are process (title + squash), not code.


Fixed since first run

✅ Metrics-helper divergence → resolved (43acd350)

The relocated pool no longer carries the inline helpers that #5168 centralized.
stage_replica_pool.py now imports and uses count_audio_frames, count_image_pixels,
coerce_positive_int_scalar, iter_mm_outputs (from vllm_omni.metrics(.utils)) and
defs.resolve_audio_sample_rate[_or_none], and drops the duplicated inline
_count_audio_frames / _coerce_int_scalar / _count_image_value_pixels /
_iter_multimodal_outputs. _collect_audio_metrics / _infer_audio_sample_rate gained the
same use_default_sample_rate / use_default plumbing main adopted.

  • Net −58 lines (50 insertions / 108 deletions).
  • Verified: grep confirms 0 inline helper definitions remain; metrics-focused tests
    (-k "audio or sample_rate or image or metric") = 3 passed; full stage suite =
    176 passed, 0 failed on server 1.

✅ Loose Any msgspec fields → resolved (43acd350)

Investigation showed StageDiffusionCoreRequest.prompt and StageDiffusionCoreOutput.output
are msgspec wire fields — the client does encoder.encode(request) and the proc does
msgspec.convert(msg, StageDiffusionCoreRequest). They carry model-dependent / opaque payloads
(prompts with images, the full OmniRequestOutput) handled by the Omni msgpack enc/dec hooks.
Narrowing them would make msgspec.convert reject valid model-specific shapes on the wire,
so Any is the correct type — not laziness. Both fields now carry a comment documenting this,
turning an unexplained Any into a deliberate, reviewable one. (The other two Any in new
code — __init__(*args, **kwargs) vLLM-compat forwarding and collective_rpc_async -> Any —
were always justified.)


Remaining (process, not code)

✗ PR title format (blocking at PR-creation time)

No commit carries a conventional prefix. Open the PR with, e.g.:

[Core] Refactor stage proc/client into engine/stage & diffusion/stage subpackages

⚠ Commit hygiene — squash before merge

19 commits, several low-signal (enhancement, bugfix, fix arg name). Collapse into a small,
well-described set on squash-merge, preserving Signed-off-by.


Passing (✓) — evidence

  • Code quality: kwargs = only vLLM-compat super().__init__(*args, **kwargs) + test asserts;
    broad-except = relocated backlog (the metrics fix removed one inline except Exception);
    Any = all justified/documented; zero new .clone()/deepcopy/time.sleep/blocking-HTTP/lock-across-await;
    zero new SimpleNamespace test fakes.
  • Dead code: refactor deletes 7 legacy modules + the Protocol layer; ruff F401 clean; every
    new stage/ and diffusion/stage/ module is imported in production and exercised by tests.
  • Test coverage: 17 test files; validated on server 1 (L20X) — 176 unit tests pass,
    plus BAGEL-7B-MoT two-stage TI2I e2e = 10/10 and single-stage Qwen-Image-Edit = 3/3.
  • Rebase / CI gates: 19 ahead / 0 behind origin/main (2bb61890); full pre-commit suite
    passes (ruff, typos, whitespace, pytest-marks, pickle-imports); all 19 commits DCO-signed.
  • Scope: the non-stage touches are all OmniEngineCore*→StageLLMCore* rename propagation.

Recommended actions before opening the PR

  1. Set a [Core] PR title (blocking).
  2. (Optional) Squash the 19 commits into a clean set, preserving sign-offs.
  3. Write a PR body: motivation (dedupe stage runtime), what moved where, the
    OmniEngineCore*→StageLLMCore* rename, adoption of vllm_omni.metrics helpers, and the
    validation evidence (176 unit + BAGEL 10/10).

Report is advisory and local-only — nothing was posted to GitHub.

chickeyton and others added 13 commits July 27, 2026 15:59
Signed-off-by: chickeyton <ngton2014@gmail.com>
The proc's run_stage_core param is omni_coord_address, but the manager
passed it under key omni_coordinator_address, so it fell into **kwargs
and was forwarded into vLLM EngineCoreProc.__init__, raising TypeError.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>
The stage proc/client refactor added vllm_omni/engine/stage/* as new files,
so the rebase onto main applied cleanly but silently missed main's changes
to the modules those files were derived from. Carry them over:

stage_replica_pool.py (from engine/stage_pool.py):
- replica availability tracking (_unavailable_replicas,
  is_replica_available, available_replica_ids, mark_replica_unavailable,
  release_replica_bindings) and its use in select_replica_id / pollers
- add_client(..., replica_id=...) slot restore after unregister/re-register
- remove_client keeps the addr->replica_id entry
- native generation-token count preferred over output-derived count
- output_processor.add_request rollback via remove_request on submit failure

stage_diffusion_core_proc.py (from diffusion/stage_diffusion_proc.py):
- _is_executor_dead uses DiffusionExecutor.is_dead
- _process_request consumes step_streaming() (step() is deprecated)

Also follow two symbol moves made on main:
- FinalOutputModalityType -> vllm_omni.outputs.output_metadata
- MultimodalOutputProcessor -> vllm_omni.outputs.output_processor

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ffer test

main added test_build_add_request_message_preserves_model_intermediate_buffer
asserting isinstance(request, OmniEngineCoreRequest); this branch renamed that
type to StageLLMCoreRequest. The two changes were in separate hunks so the
rebase applied cleanly but left a dangling reference (NameError at runtime).

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…/client

These orchestrator tests exercised the pre-refactor stage APIs and were
already failing on clean_stage_proc (pre-rebase); they are not rebase
regressions. Align them with the refactored production surface:

- Import StageReplicaPool (as StagePool) instead of the legacy engine.stage_pool
  StagePool, so tests drive the pool the orchestrator now uses (its
  poll_diffusion_output returns a list, not a single output).
- FakeStageClient: rename process_engine_inputs -> process_core_inputs
  (orchestrator calls process_core_inputs).
- FakeStageClient: rename get_output_async -> get_outputs_async and add
  get_outputs_nowait, matching StageCoreClientBase (the new pool polls via
  get_outputs_async / get_outputs_nowait).

Not yet run to green end-to-end: the server deploy was interrupted before
the suite could be re-verified.

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The old top-level stage modules were replaced by the refactored
vllm_omni/engine/stage/ subpackage; production already imports the new
symbols. Delete the now-unused legacy modules and port their test-only
consumers to the new subpackage:

  stage_engine_core_proc{,_manager}.py  -> stage/stage_llm_core_proc{,_manager}.py
  stage_pool.py                         -> stage/stage_replica_pool.py (StageReplicaPool)
  stage_engine_core_client.py           -> stage/stage_llm_core_client.py (StageLLMCoreClient)
  diffusion/stage_diffusion_client.py   -> stage/stage_diffusion_core_client.py
  diffusion/stage_diffusion_proc.py     -> stage/stage_diffusion_core_proc.py

Test updates:
- Re-point imports, string mock.patch targets, and the stage-pool logger
  name to the new subpackage.
- Rename test_stage_engine_core_client.py -> test_stage_llm_core_client.py
  and test_stage_diffusion_proc.py -> test_stage_diffusion_core_proc.py.
- Adapt to real API changes in the new clients: replica_id now defaults to
  None (raises if unset), _kv_sender_host -> _core_host, and the diffusion
  client's typed request/batched-output wire protocol
  (add_request_async(StageDiffusionCoreRequest), get_outputs_nowait()).

stage_client.py is intentionally kept: it defines the shared StageClient/
StagePoolClient Protocols still used by the live runtime layer and is not
part of the old->new duplication.

Signed-off-by: chickeyton <ngton2014@gmail.com>
stage_client.py only ever provided static typing: the StageClient/
StagePoolClient Protocols (pure annotations) and StageClientBase (an empty
`pass` base). None of it carried runtime behavior, so it can be removed
outright rather than relocated:

- stage_runtime.py / async_omni_engine.py: StageClient/StagePoolClient
  annotations -> Any (the core and inline client families share no common
  ancestor but object, which is why the Protocol existed).
- async_omni_engine.py: drop the no-op cast(StageClient, ...); use the value
  directly. `cast` import retained (still used elsewhere).
- inline_stage_diffusion_client.py: InlineStageDiffusionClient no longer
  subclasses the empty StageClientBase (its __init__ never called super()).

Pure typing removal; no runtime behavior change. The unused
StagePoolLLMClient/StagePoolDiffusionClient Protocols go away with the file.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Migrate the in-process inline diffusion client onto the same typed contract as
the out-of-process StageDiffusionCoreClient so the pool drives both identically,
and remove the pool's isinstance special-casing.

Inline client:
- Subclasses StageCoreClientBase and implements the full contract with matching
  signatures: add_request_async(StageDiffusionCoreRequest), batched
  get_outputs_nowait/get_outputs_async returning StageDiffusionCoreOutputs,
  shutdown(timeout=None), and _engine_dead_reason. check_health keeps its active
  executor probe (overriding the base template).
- Requests carry sampling params as the plain-dict form (via
  sampling_params_to_dict); reconstructed in-process. Outputs are wrapped into
  StageDiffusionCoreOutput (success -> output, failure -> error), matching the
  out-of-process wire shape. This makes inline seed-deterministic like the
  out-of-process path (generator recreated from seed).

Pool (stage_replica_pool):
- _diffusion_add_request and poll_diffusion_output drop the
  isinstance(StageDiffusionCoreClient) branches; both client shapes use the
  typed request and batched output path.

Annotations:
- stage_runtime.py / async_omni_engine.py: restore StageCoreClientBase on the
  client-plumbing signatures (undoing the earlier Any degradation), now that
  every stage client is a StageCoreClientBase subclass.

Tests:
- test_inline_stage_diffusion_client.py: drive via StageDiffusionCoreRequest and
  assert on StageDiffusionCoreOutputs batches.
- test_orchestrator.py: FakeStageClient routes diffusion outputs through
  get_outputs_nowait as a StageDiffusionCoreOutputs batch.

NOTE: torch is unavailable locally, so this was verified by py_compile + static
cross-reference only. The diffusion/orchestrator suites and mypy must be run in
CI (or on a GPU box) before merge.

Signed-off-by: chickeyton <ngton2014@gmail.com>
The debranched pool forwards a single StageDiffusionCoreRequest to diffusion
clients (sampling params in plain-dict form) instead of unpacked positional
args, so assert on the typed request's fields.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Nothing imports symbols from the stage package top-level (all 75 consumers
import directly from submodules), so the lazy re-export machinery — the
_EXPORTS map, __getattr__/__dir__, __all__, and the TYPE_CHECKING mirror — was
dead weight. Reduce __init__ to a docstring; a trivial package init still
imports nothing heavy at module scope, preserving the lightweight-submodule
import property.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>

# Conflicts:
#	tests/engine/test_async_omni_engine_input.py
#	vllm_omni/engine/stage/stage_replica_pool.py
@Gaohan123 Gaohan123 modified the milestones: v0.26.0, v0.28.0 Aug 4, 2026
Signed-off-by: chickeyton <ngton2014@gmail.com>

# Conflicts:
#	vllm_omni/core/sched/omni_ar_scheduler.py
#	vllm_omni/core/sched/omni_generation_scheduler.py
…ean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>
@chickeyton
chickeyton force-pushed the clean_stage_proc_rebase branch from 4c1205c to 8e6af27 Compare August 5, 2026 07:09
@chickeyton
chickeyton requested a review from NickCao as a code owner August 7, 2026 02:50
… clean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>
…ean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>
… clean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

will the change the usage of diffusion DP? for example, we currently work heavily with distributed layerwise offload, it rereuiqs work with DP replica

… clean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>

# Conflicts:
#	docs/design/architecture_overview.md
#	vllm_omni/diffusion/stage/inline_stage_diffusion_client.py
… clean_stage_proc_rebase

Signed-off-by: chickeyton <ngton2014@gmail.com>

# Conflicts:
#	vllm_omni/engine/orchestrator.py
#	vllm_omni/engine/stage/stage_replica_pool.py
@chickeyton

Copy link
Copy Markdown
Contributor Author

will the change the usage of diffusion DP? for example, we currently work heavily with distributed layerwise offload, it rereuiqs work with DP replica

There is no functional nor behavioural change, just code relocation and renaming

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

please review the agent codes by yourself with comments first

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit fd880bfbd074 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@Gaohan123 Gaohan123 modified the milestones: v0.28.0, v0.30.0 Sep 2, 2026

async def get_outputs_async(self) -> StageEncodeCoreOutputs:
frame = await self._output.recv()
return self._decoder.decode(frame, StageEncodeCoreOutputs)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use msgspec.convert(self._decoder.decode(frame), StageEncodeCoreOutputs); decode doesn't accept a type argument.

from vllm_omni.engine.stage_client import StageClient, StagePoolClient
from vllm_omni.engine.stage_engine_core_client import StageEngineCoreClientBase
from vllm_omni.engine.stage.stage_core_client import StageCoreClientBase
from vllm_omni.engine.stage.stage_llm_core_client import StageLLMCoreClientBase

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update test_stage_runtime_passes_log_stats_to_llm_replica_launch to use StageLLMCoreClientBase; the stale name raises AttributeError.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@chickeyton this pull request has had no human commit, comment or review since 2026-09-16. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@chickeyton chickeyton closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants