Skip to content

fix(e2e): re-point HunyuanImage3 offline test to text_to_image.py - #7525

Merged
Gaohan123 merged 6 commits into
vllm-project:mainfrom
chethanuk:fix/issue-7495-bug-nightly-ci-failed-testse2eaccuracytesthunyuanimage3pixelaccuracypytesthunyuanimage3pixelaccuracyoffline-cant
Sep 19, 2026
Merged

Gaohan123 merged 6 commits into
vllm-project:mainfrom
chethanuk:fix/issue-7495-bug-nightly-ci-failed-testse2eaccuracytesthunyuanimage3pixelaccuracypytesthunyuanimage3pixelaccuracyoffline-cant

Conversation

@chethanuk

Copy link
Copy Markdown
Contributor

Purpose

Fixes nightly/CI failure tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py::test_hunyuan_image3_pixel_accuracy_offline — can't open file '.../hunyuan_image3/end2end.py': [Errno 2].

Root cause: commit a7ab91fb8 (#5559) deleted examples/offline_inference/hunyuan_image3/end2end.py and migrated HunyuanImage-3.0 to the shared examples/offline_inference/text_to_image/text_to_image.py, but tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:76 still invoked the deleted path, exiting with status 2 (/usr/bin/python3: can't open file ...).

Fix (one file, tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py):

  • _OFFLINE_SCRIPT re-pointed to examples/offline_inference/text_to_image/text_to_image.py.
  • _run_vllm_omni_hunyuan_image3_offline flag remap: --modality dropped, --prompts → --prompt, --steps → --num-inference-steps, --bot-task none + --sys-type en_unified → --extra-body '{"bot_task": null}' + --use-system-prompt en_unified + --trust-remote-code, and --output now a direct file path (no output_*.png glob/re-save). See text_to_image.py:106,128,140,323,364,383.

Fixes #7495

Test Plan

vLLM Version: n/a

vLLM-Omni Commit: 0a857ad (this branch, 1 commit vs main)

Local (non-GPU):

  • python3 -m py_compile tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py — OK
  • git grep -n "hunyuan_image3.*end2end" — only tests/buildkite/test_vllm_omni_package_discovery.py:59 (end2end_probe.py fixture) remains
  • Flag existence check: --prompt :106, --output :128, --num-inference-steps :140, --extra-body :323, --use-system-prompt :364 (en_unified in choices), --trust-remote-code :383 in text_to_image.py

GPU (exact CI commands, per plan):

  • CUDA 4×H100 (.buildkite/cuda/test-nightly.yml:286): pytest -s -v tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py -m full_model --run-level full_model (add -k offline for local)
  • NPU 4×A3 (.buildkite/npu/test-npu-nightly.yml:179): -m "full_model and A3" --run-level full_model

Note: old script let pipeline_hunyuan_image3.py:1235 build the generator, new script seeds via torch.Generator(device).manual_seed(seed) (text_to_image.py:489); seeding triage: if offline SSIM/PSNR regresses while ..._online passes, check noise seeding first before touching thresholds.

Test Result

  • py_compile OK, worktree/primary clean, single-file diff (10+, 14-).
  • No runtime on 4×H100/A3 in this env; full pytest -m full_model to be confirmed on CI.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.

Commit a7ab91f removed examples/offline_inference/hunyuan_image3/end2end.py
and moved HunyuanImage-3.0 to the shared text_to_image/text_to_image.py;
tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py still invoked
the deleted path, causing test_hunyuan_image3_pixel_accuracy_offline to
fail with No such file / exit 2.

Re-point _OFFLINE_SCRIPT and remap CLI flags per the shared script:
--modality dropped, --prompts->--prompt, --steps->--num-inference-steps,
--bot-task/--sys-type->--extra-body '{"bot_task": null}' /
--use-system-prompt en_unified plus --trust-remote-code, and write
directly to the file path instead of globbing output_*.png.

Fixes vllm-project#7495

Signed-off-by: ChethanUK <chethanuk@outlook.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: HunyuanImage.

Model owners: @Bounty-hunter @NickCao @yenuo26

Routing: @Bounty-hunter via semantic router, model owner; @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS

@chethanuk, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@NickCao

NickCao commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator

@vllm-omni-review-bot

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict under experiment vllm-omni-strict-5050-20260829.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Scan:

Category Result
Tests / verification 5 finding(s) below
Security no finding reported
Docs / comments no finding reported
Behavior / compatibility no finding reported
Correctness no finding reported

Validated:

  • [resolved] FileNotFoundError for hunyuan_image3/end2end.py: _OFFLINE_SCRIPT now points at text_to_image.py:76; residual = accuracy/seed/AR parity not covered by path fix.
  • [resolved] #7495 Errno 2: path remount closes missing-file exit 2; residual is semantic parity of remounted launch.
  • [resolved] Missing end2end.py Errno 2 (issue #7495): fixed by re-point; residual: no CPU-level is_file assert, so path rot still needs full_model subprocess to surface.
  • [claim-verified] hunyuan_image3.*end2end leftovers: only tests/buildkite/test_vllm_omni_package_discovery.py:59 end2end_probe.py fixture.
  • [claim-verified] CI guard :full_moon: HunyuanImage3-DIT · Accuracy Test (.buildkite/cuda/test-nightly.yml:280-286) matches PR Test Plan selector; NPU twin test-npu-nightly.yml:179.
  • [claim-verified] CI lane HunyuanImage3-DIT Accuracy Test selects this file via diffusion_hunyuan_image3_accuracy in ci_source_file_dependencies.yml:366-368.

Keep two majors on the remount’s real contract risks—offline CLI omitting --enable-expert-parallel so runtime EP=False overrides deploy YAML while online keeps EP=True, and text_to_image.py forcing HunyuanImage3 through AR prefill/bot_task null vs the online DiT-only bot_task:"none" path—plus one minor asking for the exact nightly full_model GPU SSIM/PSNR result given the new generator seeding, and one nit for an optional _OFFLINE_SCRIPT.is_file() guard. Drop resolved/no_issue/excluded acknowledgments and flag-parity checks; fold the overlapping CI/seeding confirms into the seeding comment.

Verdict: REQUEST CHANGES

Findings

  • **[P1] This remount's offline subprocess ends at --enforce-eager and never passes -…** — tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.pyThis remount's offline subprocess ends at--enforce-eagerand never passes--enable-expert-parallel, while _DEPLOY_CONFIGstill setsenable_expert_parallel: True(test file L98). Unchanged by this diff, present in the PR-time tree:text_to_image.pydefines--enable-expert-parallelasstore_trueand always sets"enable_expert_parallel": args.enable_expert_parallelinto Omni kwargs (default False), andstage_config.pywrites non-None runtime overrides into nestedparallel_config, so CLI False overrides the written deploy YAML. The online twin in the same file only passes --deploy-config(and serve rejects CLI EP under--omni), so it keeps EP=True. Add --enable-expert-parallel` to the offline argv so offline/online MoE topology match the deploy contract.

Evidence: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:98 "enable_expert_parallel": True,; tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:258 "--enforce-eager", (offline argv ends here; no --enable-expert-parallel); unchanged by this diff, present in the PR-time tree: examples/offline_inference/text_to_image/text_to_image.py:265-267 "--enable-expert-parallel", / action="store_true",; examples/offline_inference/text_to_image/text_to_image.py:559 "enable_expert_parallel": args.enable_expert_parallel,; vllm_omni/config/stage_config.py:144 if value is None or key not in parallel_fields: then L150 parallel_config_dict[key] = runtime_overrides.pop(key) (False is applied over YAML); tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:190-192 online server_args = [ / "--deploy-config", / deploy_config, (no EP CLI flag).

Suggestion: "--enforce-eager",
"--enable-expert-parallel",

  • **[P1] This remount points offline at text_to_image.py and passes --extra-body '{"b…** — tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.pyThis remount points offline attext_to_image.pyand passes--extra-body '{"bot_task": null}'(L253–254), while the online twin still postsbot_task: "none"(L212) under DiT-only deploy. HunyuanImage3 always registersar_input_builder, so offline always runs get_ar_input_builderand injects ARprompt_token_ids (text_to_image.pyL725–731; registry L252–256) even for DiT-only. Online single-stage/v1/images stays on the string-prompt path (api_server.pyL1763 / L1831–1842) wherenormalize_hunyuan_single_stage_bot_task("none")→"auto" (request_layout.pyL105–119).apply_declared_extra_args drops None (param_utils.pyL24), so null never reaches DiT extras. Do not "fix" by sending string"none" on the AR path (_BOT_TASK_PRESETShas no"none"`). Keep offline on the DiT-only string-prompt contract matching online, or prove SSIM/PSNR vs baseline on nightly GPU and refresh the baseline before treating the remount as closed.

Evidence: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:212 "bot_task": "none", / L253-254 --extra-body / '{"bot_task": null}' / L76 _OFFLINE_SCRIPT = ... text_to_image.py / L78-80 DiT-only "pipeline": "hunyuan_image3_dit". Unchanged by this diff, present in the PR-time tree: examples/offline_inference/text_to_image/text_to_image.py:725 ar_input_builder = get_ar_input_builder(model_class_name) and L468-469 if ar_inputs.prompt_token_ids is not None: prompt_dict["prompt_token_ids"] = ar_inputs.prompt_token_ids. Unchanged: vllm_omni/model_extras/registry.py:247-256 comment says DiT-only deploys key on HunyuanImage3ForCausalMM with "ar_input_builder": build_hunyuan_image3_ar_stage_inputs. Unchanged: vllm_omni/entrypoints/openai/api_server.py:1763 if len(stage_configs) > 1: else L1832 prompt: OmniTextPrompt = {"prompt": request.prompt, "modalities": ["image"]} and L1841-1842 if request.bot_task is not None: extra_args["bot_task"] = request.bot_task. Unchanged: vllm_omni/diffusion/models/hunyuan_image3/request_layout.py:105-119 if isinstance(bot_task, str) and bot_task.lower() == "none": bot_task = None … return tokenizer_bot_task or "auto". Unchanged: vllm_omni/diffusion/utils/param_utils.py:24 declared = {key: user_kwargs[key] for key in declared_params if user_kwargs.get(key) is not None}. Unchanged: vllm_omni/diffusion/models/hunyuan_image3/prompt_utils.py:109-115 _BOT_TASK_PRESETS keys are only None, think, recaption, think_recaption, vanilla (no "none").

  • [P2] #7495 Errno 2 closed by remount: test_hunyuan_image3_pixel_accuracy.py:76 point… — ``
    #7495 Errno 2 closed by remount: test_hunyuan_image3_pixel_accuracy.py:76 points _OFFLINE_SCRIPT at text_to_image/text_to_image.py (hunyuan_image3/end2end.py absent). Residual (minor): semantic parity of remounted launch args vs the old specialized example.

Evidence: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:76 _OFFLINE_SCRIPT = _REPO_ROOT / "examples" / "offline_inference" / "text_to_image" / "text_to_image.py" — remount to existing shared example. Launch remaps Hunyuan knobs at :251-255 --use-system-prompt / en_unified / --extra-body / '{"bot_task": null}'. Unchanged by needing the deleted tree: PR-time examples/offline_inference/hunyuan_image3/ has 0 files (old end2end.py path gone).

  • **[P2] Issue #7495 missing end2end.py is fixed by re-pointing _OFFLINE_SCRIPT to te…** — `` Issue #7495 missing end2end.py is fixed by re-pointing _OFFLINE_SCRIPTtotext_to_image.py(L76). Residual: no CPU-level_OFFLINE_SCRIPT.is_file()/exists()assert beforesubprocess.run` (L231–234), so future script path rot still only surfaces under full_model subprocess (Errno 2).

Evidence: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:76 _OFFLINE_SCRIPT = _REPO_ROOT / "examples" / "offline_inference" / "text_to_image" / "text_to_image.py" — re-point fix; tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:231-234 subprocess.run([sys.executable, str(_OFFLINE_SCRIPT), ...]) — no intervening _OFFLINE_SCRIPT.is_file()/exists() assert; only post-run assert output_path.exists() at L262. PR-time tree: examples/offline_inference/text_to_image/text_to_image.py exists; examples/offline_inference/hunyuan_image3/ absent.

  • [P3] This remount fixes #7495 by pointing _OFFLINE_SCRIPT at text_to_image.py (L… — tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py
    This remount fixes #7495 by pointing _OFFLINE_SCRIPT at text_to_image.py (L76), but the covering tests remain @pytest.mark.full_model + 4-card hardware_test (L297–298, L312–313), so another path rot would again surface only on GPU nightly. Optional: add assert _OFFLINE_SCRIPT.is_file() beside the constant so missing-script failures fail at import/collection without 4×H100.

Evidence: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:76 _OFFLINE_SCRIPT = _REPO_ROOT / "examples" / "offline_inference" / "text_to_image" / "text_to_image.py" — no is_file() guard on adjacent lines; unchanged by this diff's marker placement, present in the PR-time tree: tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py:312-313 @pytest.mark.full_model / @hardware_test(res={"cuda": "H100", "npu": "A3"}, num_cards=4) (same pattern at L297–298)

Suggestion: _OFFLINE_SCRIPT = _REPO_ROOT / "examples" / "offline_inference" / "text_to_image" / "text_to_image.py"
assert _OFFLINE_SCRIPT.is_file(), f"Offline script missing: {_OFFLINE_SCRIPT}"

@hsliuustc0106 hsliuustc0106 added the CI/CD codes related to changes to CI/CD label Sep 15, 2026
@yenuo26

yenuo26 commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

@chethanuk Please check review

@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 19, 2026
…eaccuracytesthunyuanimage3pixelaccuracypytesthunyuanimage3pixelaccuracyoffline-cant
Only construct AR inputs when the deployment contains an AR stage. Keep
the DiT-only accuracy request on the same string-prompt/bot_task=none path
as online serving, and explicitly enable expert parallelism in the shared
offline CLI so its default cannot override the deploy configuration.

Cover script existence, real CLI parsing and both DiT-only/AR+DiT request
construction. Patch only tokenizer loading in existing tests so model
imports can still access the rest of transformers.

Validation: 80 related unit tests and all applicable pre-commit hooks pass.
The new regressions fail twice against the original PR source as expected.
Both full accuracy cases pass on 4 L20X with a local Triton MoE override:
SSIM 0.987250, PSNR 35.653664 dB, with identical online/offline PNG hashes.
The exact nightly FlashInfer CUTLASS configuration cannot initialize on
L20X; H100/B200/A3 validation remains with CI. No thresholds were changed.

AI assistance: Codex inspected the review, implemented the code and tests,
and ran local validation.

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
The latest main rejects unknown diffusion configuration fields. The shared
offline example still supplied mode="text-to-image", which has no consumer
and makes the HunyuanImage3 offline accuracy launch fail before loading.
Remove it and exercise the real ingress validator in the request-building
regression test instead of allowing every mocked Omni keyword silently.

Validation on main 698f716: 458 CPU tests passed, 2 skipped, 4 deselected;
all applicable pre-commit hooks passed. Four-card L20X/Triton offline
accuracy passed with SSIM 0.987251 and PSNR 35.653656 dB, matching the online
case and its PNG hash. Original accuracy thresholds remain unchanged.

AI assistance: Codex implemented and validated this compatibility fix.

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
@Gaohan123

Copy link
Copy Markdown
Collaborator

Following up on the review findings and the request to address them. Fixes are pushed in 357101838 and a21b939c7, with main 698f71601 merged.

  1. P1 — expert-parallel topology: added --enable-expert-parallel to the offline argv. The shared CLI now preserves the deploy config's EP=True behavior. The CPU regression parses the actual launch arguments and checks this flag.

  2. P1 — DiT-only vs AR input contract: the shared runner now invokes the AR input builder only when the deployment's sampling-parameter list contains a non-diffusion stage. DiT-only requests retain the plain string prompt and pass bot_task="none", matching online serving; they do not send "none" into the AR builder. Regression coverage exercises both DiT-only and AR+DiT paths, including tokenizer invocation, prompt token IDs and AR stop tokens.

  3. P2 — launch semantics / seed and accuracy parity: real four-card online and offline generation passed against the existing baseline on 4 NVIDIA L20X, using a local-only moe_backend=triton override. Final results for both paths: SSIM 0.987251, PSNR 35.653656 dB, exceeding the original 0.97 / 30 dB thresholds. Pixel-difference checks also passed, and online/offline PNG SHA-256 hashes are identical. The model was HunyuanImage-3.0-Instruct revision 2ec2c78bee7d4b94157341fba86c4c2c7b1858b2, seed 42, 50 steps, 1024×1024, guidance 2.5. No baseline or threshold was changed.

    Hardware validation limit: the exact nightly command below was attempted, but both cases failed initialization because its configured unquantized FlashInfer CUTLASS MoE kernel does not support L20X. The passing Triton runs do not establish results for the exact H100/B200 FlashInfer or A3/NPU lanes; those remain to be confirmed on supported CI hardware.

    gpu run --gpus 4 -- python -m pytest -s -v tests/e2e/accuracy/test_hunyuan_image3_pixel_accuracy.py -m full_model --run-level full_model
  4. P2/P3 — catch future script-path regressions without full-model inference: test_hunyuan_accuracy_offline_cli_contract in tests/model_extras/test_shared_script_ar_integration.py asserts that the actual subprocess script path is a file, exercises the real CLI parser, and checks the direct output-file contract. This runs under core_model and cpu in regular CI, rather than requiring four-card inference to detect a missing script. The check is in the regression test rather than an import-time assertion.

The original PR source failed the new EP and DiT-only contract regressions as expected. Final related CPU validation: 458 passed, 2 skipped, 4 deselected; all applicable pre-commit hooks passed. After merging main's stricter configuration validation, the offline run also exposed an obsolete mode="text-to-image" kwarg. It was removed, and the regression now exercises the real diffusion ingress validator; the final offline accuracy rerun passed.

Validation used the repository uv environment, HF_HOME=/workspace/models, exclusive gpu run reservations, Python 3.12.13, vLLM 0.29.0, torch 2.13.0+cu129, transformers 5.14.1 and diffusers 0.40.0. The focused CPU checks can be reproduced with:

gpu run --gpus 1 -- python -m pytest tests/model_extras/test_shared_script_ar_integration.py tests/model_extras/test_model_extras.py tests/examples/offline_inference/test_image_task_prompts.py tests/config/test_omni_config.py tests/engine/test_stage_engine_args.py tests/entrypoints/test_async_omni_diffusion_config.py -m 'core_model and cpu' --run-level core_model -q

AI assistance: Codex inspected the review, implemented the fixes and tests, ran validation, and drafted this response.

@Gaohan123 Gaohan123 added nightly-test label to trigger buildkite nightly test CI and removed nightly-test label to trigger buildkite nightly test CI labels Sep 19, 2026
…eaccuracytesthunyuanimage3pixelaccuracypytesthunyuanimage3pixelaccuracyoffline-cant

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 19, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 19, 2026 17:05
@Gaohan123
Gaohan123 merged commit 573ec4c into vllm-project:main Sep 19, 2026
9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…lm-project#7525)

Signed-off-by: ChethanUK <chethanuk@outlook.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…lm-project#7525)

Signed-off-by: ChethanUK <chethanuk@outlook.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD ready label to trigger buildkite CI

Projects

None yet

6 participants