Skip to content

[Feature] Transfer AR-to-diffusion payloads over NIXL - #6264

Merged
hsliuustc0106 merged 46 commits into
vllm-project:mainfrom
yuanwu2017:feat/ar2diffusion-stage-payload
Sep 23, 2026
Merged

hsliuustc0106 merged 46 commits into
vllm-project:mainfrom
yuanwu2017:feat/ar2diffusion-stage-payload

Conversation

@yuanwu2017

@yuanwu2017 yuanwu2017 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Transfer complete AR/encoder-to-diffusion payloads through an Omni connector rather than duplicating conditioning through engine-core IPC and the orchestrator. MiniMax-H3 payloads include text, visual/audio latents and layout metadata. Leader-only transfer and receiver TP broadcast support TP2 → TP4.

Updated dependency stack — 2026-09-14

Published and freshly tested: c9d7bf83c337db7cad5005af7946d0ed5244c3d7, tree 2d80115109128973bad63f4411af87666124bc3c.

New main streaming decode and image-response changes are retained. This is still a stacked integration, not #6264 independently applicable without #6939.

Fresh validation — PASS

Completed 2026-09-14 01:41:46 UTC, overall runner exit 0. Frozen source, index, refs and runner hashes remained unchanged.

  • Ruff lint/format: 37 changed Python files passed.
  • Expanded CPU: 870 passed, 1 skipped, 3 passing subtests.
  • H3 configuration: 21 passed.
  • Native NIXL/UCX: 8 passed in 237.34s, covering manager-created receivers, generic/T2VA/FL2VA/Ref2VA payloads, queried/direct metadata, exact mixed-device tensor semantics and source ownership/cleanup.
Model task HTTP Seconds MP4 bytes Equal NIXL put/get bytes
T2VA 200 76.963 5,242,391 359,024
FL2VA image 200 65.516 4,766,094 11,158,136
Ref2VA video + embedded AAC 200 201.757 6,094,847 64,963,792

All three requests passed exact producer/consumer key and positive-byte matching, H.264 1344×768, 107 frames/24fps, stereo 32-kHz audio, full ffmpeg decode and preview extraction. Keys:

  • video_sync-9e35132f9fa268e8-816c22ca_0_0
  • video_sync-92937f4a63654b16-a03e02c1_0_0
  • video_sync-b4a995398d3c805f-90ca89fd_0_0

These are fresh tests on the rebuilt tree, not reattributed historical results. The earlier September 13 0625 media-only successes used inline fallback after receiver initialization failed; they remain not NIXL receive evidence. The receiver fix is now inherited from merged main, and the current strict matcher is unchanged.

Runtime and scope

  • Immutable local image sha256:bce11ef0dfb05cd4c9e18695bccdd751b7c200514b14ee67fa14bf06b3f981e4: Python3.12.3, vLLM0.29.0, Torch2.13cu130, NIXL1.3.2.
  • Physical GPUs0,1,3,4,5,6 → logical0–5. Stage0 TP2/maxlen32768/util0.6; Stage1 TP4/Ulysses1/VAE4.
  • Stock vllm_omni/deploy/minimax_h3_disaggregated.yaml, CLI overrides {"0":{"max_model_len":32768},"1":{"tensor_parallel_size":4,"ulysses_degree":1}}.
  • Model snapshot 42ed227ee7df40d41602854ae760620d6eb651fe, FL2VA/Ref2VA partitions; 8 steps, seed42.
  • Source and cache mounted read-only; explicit proxies; all owned containers removed and six test GPUs released. GPU2 untouched.
  • Initial retry prerequisite: the image lacks Python Ruff; standalone host Ruff0.14.10 was used against frozen source without modifying the model image. The failed prerequisite attempt is preserved separately.

Local evidence: vllm-omni-pr6264-artifacts/restack-20260914-retry/ (local files, not public attachments). No cross-node, non-CUDA or Turbo-checkpoint model validation is claimed. Remote CI and approvals are separate.

DCO: inherited #6939 commit 957505518cfe583e70ed9971240a4d933ce825b2 still lacks its author's sign-off. This restack preserves author identity and does not fabricate a trailer or claim DCO is fixed.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@yuanwu2017

yuanwu2017 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor Author

A100 Docker E2E test results

Validated commit 315a3f4d9 with real MiniMaxAI/MiniMax-H3 FL2VA weights in vllm/vllm-openai:v0.27.0.

Focused regression suite

142 passed, 16 warnings, 3 subtests passed in 5.38s

Covered:

  • tests/diffusion/test_diffusion_stage_payload.py
  • tests/distributed/omni_connectors/test_tp_rank_aware.py
  • tests/distributed/omni_connectors/test_nixl_connector.py
  • tests/worker/test_omni_connector_mixin.py

Six-GPU H3 E2E

  • Hardware: 6x A100 80GB PCIe
  • Stage 0: text encoder TP2
  • Stage 1: diffusion TP4
  • Shared, read-only Hugging Face model cache
  • Request: 1344x768, 24 FPS, 4 seconds

Transport smoke test (num_inference_steps=2):

HTTP 200
Content-Type: video/mp4
Output size: 3,810,738 bytes
Total time: 48.86 s
NIXL puts: 1
NIXL gets: 1
Handshake timeouts: 0
Inline fallbacks: 0
Request errors: 0

Visual test (num_inference_steps=8):

HTTP 200
H.264 video: 1344x768 @ 24 FPS
AAC audio
Duration: 4.459 s
Output size: 4,575,064 bytes
Total time: 88.44 s
NIXL puts: 1
NIXL gets: 1
Inline fallbacks: 0
Request errors: 0

The preview confirms that the eight-step output is visually meaningful rather than the near-gray two-step smoke output:

MiniMax-H3 eight-step orange cat output

@yuanwu2017

Copy link
Copy Markdown
Contributor Author

Reproduction commands

The successful run mapped physical GPUs 0,1,3,4,5,6 to logical GPUs 0-5 inside the container. The deployment YAML then assigns logical GPUs 0,1 to the text encoder (TP2) and 2,3,4,5 to diffusion (TP4).

Start the six-GPU server

Run from the repository root and adjust HF_CACHE if needed:

export REPO_ROOT="$PWD"
export HF_CACHE=/mnt/disk5/HF_CACHE
export H3_SNAPSHOT="$HF_CACHE/hub/models--MiniMaxAI--MiniMax-H3/snapshots/42ed227ee7df40d41602854ae760620d6eb651fe"

docker run --rm   --name h3-nixl-e2e-v027   --gpus '"device=0,1,3,4,5,6"'   --network host   --ipc host   -v "$REPO_ROOT:/workspace/vllm-omni"   -v "$HF_CACHE:/root/.cache/huggingface:ro"   -w /workspace/vllm-omni   -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5   -e HF_HOME=/root/.cache/huggingface   -e HF_HUB_OFFLINE=1   -e TRANSFORMERS_OFFLINE=1   -e HF_MODULES_CACHE=/tmp/hf_modules   -e VLLM_LOGGING_LEVEL=DEBUG   -e VLLM_WORKER_MULTIPROC_METHOD=spawn   -e VLLM_OMNI_VIDEO_SYNC_TIMEOUT=14400   -e SETUPTOOLS_SCM_PRETEND_VERSION=0.26.1.dev0   --entrypoint /bin/bash   vllm/vllm-openai:v0.27.0   -lc 'mkdir -p /tmp/hf_modules && exec vllm serve     /root/.cache/huggingface/hub/models--MiniMaxAI--MiniMax-H3/snapshots/42ed227ee7df40d41602854ae760620d6eb651fe     --omni     --host 0.0.0.0     --port 8091     --trust-remote-code     --task-type fl2va     --deploy-config /workspace/vllm-omni/vllm_omni/deploy/minimax_h3_disaggregated.yaml'

Wait until /health returns 200:

curl -fsS http://127.0.0.1:8091/health

Submit the eight-step visual request

curl -sS --max-time 14400   -D /tmp/h3-cat-headers.txt   -o /tmp/h3-cat-output.mp4   -w 'http_code=%{http_code}\ncontent_type=%{content_type}\nsize_download=%{size_download}\ntime_total=%{time_total}\n'   http://127.0.0.1:8091/v1/videos/sync   -F 'prompt=A fluffy orange tabby cat with bright green eyes walks across a sunlit wooden kitchen floor, pauses beside a blue ceramic bowl, looks directly into the camera, then playfully bats a small red ball. Warm natural daylight, detailed orange fur, realistic cinematic video, smooth camera movement, vivid colors, clear subject.'   -F 'width=1344'   -F 'height=768'   -F 'fps=24'   -F 'num_inference_steps=8'   -F 'seed=123'   -F 'flow_shift=12.0'   -F 'extra_params={"task":"t2va","duration":4.0,"aspect_ratio":"16:9","audio_flow_shift":3.0}'

Expected response characteristics from this run:

HTTP 200
Content-Type: video/mp4
H.264 1344x768 @ 24 FPS + AAC audio
Duration: 4.459 s
Output size: 4,575,064 bytes

Run the focused regression suite

docker run --rm   --entrypoint /usr/bin/python3   -v "$REPO_ROOT:/workspace/vllm-omni:ro"   -w /workspace/vllm-omni   -e PYTHONPATH=/workspace/vllm-omni   vllm-omni-h3-nixl:test   -m pytest -q   tests/diffusion/test_diffusion_stage_payload.py   tests/distributed/omni_connectors/test_tp_rank_aware.py   tests/distributed/omni_connectors/test_nixl_connector.py   tests/worker/test_omni_connector_mixin.py

Result: 142 passed, 3 subtests passed.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: MinimaxH3.

Model owners: @david6666666

@yuanwu2017, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@yuanwu2017
yuanwu2017 marked this pull request as draft August 17, 2026 07:10
@yuanwu2017

Copy link
Copy Markdown
Contributor Author

Retest after 001514f97

Retested the latest head (001514f97) in the matching vLLM 0.27 Docker environment.

Focused payload/H3 tests:

41 passed, 15 warnings in 0.43s

Six-GPU MiniMax-H3 E2E (text encoder TP2 -> diffusion TP4), using a 4-second T2VA request with 2 inference steps:

HTTP 200
Content-Type: video/mp4
Request time: 38.56 s
Output size: 2,381,284 bytes
H.264 1344x768 @ 24 FPS + AAC
Duration: 4.459 s

Request-scoped connector evidence:

NIXL put:       1 (645,719 bytes)
NIXL get:       1 (645,719 bytes, same key)
NIXL failures:  0
Stage fallbacks: 0
Tracebacks:     0
Errors:         0

Scope note: H3 exercises the AR text-encoder producer -> diffusion consumer path. The new 001514f97 behavior (a diffusion producer dropping transferred inline keys after a successful put) is covered by the focused producer tests, but H3's two-stage topology does not have a downstream stage after diffusion and therefore does not exercise that exact branch in E2E.

@yuanwu2017

Copy link
Copy Markdown
Contributor Author

Regression results (completed September 22, 2026)

@hsliuustc0106 @xuechendi Please see the completed regression results below.

Revision scope: The full matrix tested PR commit b9844ca2106f64443b564d822969911584ab69bf, paired against main 86fcfe95f12010061592e0d1dfe939d9cd81bee2. It completed at 2026-09-22T02:00:54Z after waiting for GPU resources. The PR has since advanced: its head was 9ba6e90c99540b8a664f2d61162d4e51c2ddac9b when preparing this report. These full-matrix results do not certify the current head.

Results on b9844ca

Coverage Result
Unit regression: configuration, H3 contracts, workers, orchestration, payload/KV coexistence, NIXL and shared memory 1,555 passed, 6 skipped; 3 additional subtests passed
Real four-GPU payload probes: HSDP shard/replicate 4x1, 2x2, 1x4, and TP2 x SP2 All passed: 16 rank records and 48 rank/scenario checks across delivered payload, inline fallback and missing payload; real shared memory and distributed groups
Native NIXL structured-payload matrix 8 passed in 241.81s; original published 60-second process joins retained, no local 180-second timeout patch
Six-GPU disaggregated H3: T2VA, FL2VA, Ref2VA All three passed: HTTP 200, media validation and request-matched NIXL put/get checks
Paired main/PR aggregated H3 and dense FastH3 30 normal generation requests, including warmups, and 12 recovery requests returned HTTP 200; 12 invalid-request checks failed as detailed below

Six-GPU request measurements (individual runs, not a performance benchmark):

Task Request time Output bytes
T2VA 78.684s 5,084,597
FL2VA 67.275s 4,731,648
Ref2VA 204.478s 6,108,800

The paired aggregated tests covered T2VA (including a 50-step request), FL2VA, reference-video Ref2VA and mixed image/audio Ref2VA. Dense FastH3 generation coverage was T2VA only.

Retained failures

The overall queue exited 1, not all green. Both pinned main and PR returned HTTP 500 instead of the expected 4xx for:

  • Aggregated Ref2VA: invalid short audio, once per revision.
  • FastH3: invalid steps, task, request LoRA, video flow shift and audio flow shift, once each per revision.

That is 6 failing checks per revision, 12 total. All 12 subsequent valid recovery requests returned HTTP 200. These failures reproduce on the paired main baseline; they have not been waived or relabeled as passes.

Environment and verification boundaries

  • Same-host NVIDIA A100 80GB testing; four-GPU probes/paired generation and six-GPU disaggregated generation.
  • Pinned runtime image: sha256:bce11ef0dfb05cd4c9e18695bccdd751b7c200514b14ee67fa14bf06b3f981e4. Source snapshots were mounted read-only: this is source-mounted validation, not a rebuilt release-image test.
  • MiniMax-H3 snapshot: 42ed227ee7df40d41602854ae760620d6eb651fe; dense FastH3 adapter SHA256: 4ce198c83132251b7fd0de2503823aa49c53983f068318f66cb19eaefb7fcc12.
  • Six-GPU startup harness budgets: --init-timeout 1500 --stage-init-timeout 900. Native test assertions and process timeouts were unchanged.
  • Recorded source/input integrity checks passed and owned test containers were removed.
  • No cross-node, non-CUDA, separate Turbo-checkpoint, or manual media-quality/numerical-parity certification is claimed.

Subsequent merge validation (separate evidence)

Merge commit f0cec9b6db40a0529d21ccb20d6202599d2a317e, incorporating main 0b5f8832b91e10537f1c96f33fbbd2bedb7bca51, received focused validation: 1,273 passed, 3 skipped, plus 3 subtests passed, with Ruff and formatting checks passing. This combined the H3 contract module (220 passed, 2 skipped) and adjacent configuration/worker/orchestrator/connector coverage (1,053 passed, 1 skipped, plus 3 subtests).

The initial CPU-only H3 run had five failures because that runtime's CPU platform did not implement empty_cache; the unchanged module passed with CUDA platform selection. Full model E2E was not rerun on f0cec9b6d, and no validation of the newer 9ba6e90c9 head is claimed here.

@xuechendi xuechendi left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codes and Design Doc is clear to me, approved

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 22, 2026
@hsliuustc0106 hsliuustc0106 added cuda-test Used to trigger vllm-omni cuda CI separately. and removed ready label to trigger buildkite CI diffusion codes related to diffusion models labels Sep 22, 2026
@hsliuustc0106 hsliuustc0106 added the ready label to trigger buildkite CI label Sep 23, 2026
@hsliuustc0106
hsliuustc0106 merged commit cb5f508 into vllm-project:main Sep 23, 2026
5 of 6 checks passed
fusinjay added a commit to fusinjay/vllm-omni that referenced this pull request Sep 24, 2026
main already gained receive-timeout safety for
test_unknown_key_is_queried_once through its force_timeout
parametrization (vllm-project#6264), so keep the upstream test body and fold this
branch's remaining guarantee into it: the request socket is retired
exactly on the timeout path and stays reusable when the producer
serves a reply.

Signed-off-by: fusinjay <1747683542@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner cuda-test Used to trigger vllm-omni cuda CI separately. enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants