Skip to content

[Model][MiniMax-H3] Support FastH3 model-level CPU offload - #6949

Open
EchoHayate wants to merge 10 commits into
vllm-project:mainfrom
EchoHayate:fix/fasth3-model-cpu-offload
Open

EchoHayate wants to merge 10 commits into
vllm-project:mainfrom
EchoHayate:fix/fasth3-model-cpu-offload

Conversation

@EchoHayate

@EchoHayate EchoHayate commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Allow Dense/Data-Free FastH3 with the existing model-level CPU-offload path.
  • Keep VSA plus model-level offload, ordinary layerwise offload, and distributed layerwise offload fail-closed.
  • Add a standalone single-A100 recipe with pinned downloads, server command, and T2VA request.

Model-level offload follows ordinary checkpoint loading: FastH3 deltas are fused and validated before the offloader is installed. The loader contract tests verify that this path does not create a HostWeightPlan.

Validation

  • At bf5b0aaa: 59 focused tests passed; changed-file pre-commit hooks and git diff --check passed.
  • The subsequent 1086cf79 change adds only the standalone A100 recipe; its applicable pre-commit hooks and whitespace checks passed.
  • Current head: 1086cf79018a959631167a17ad754baaf7ab95f7. GitHub build, pre-commit, DCO, and documentation checks pass.

Single-A100 E2E evidence

The previously published controlled run used one NVIDIA A100 80GB PCIe, model-level CPU offload, concurrency one, and one warmup followed by three measured requests per arm. Model revision, prompt, seed, output dimensions/length, and attention backend were fixed.

Workload Base, 50 steps FastH3, 4 steps E2E speedup
672x384, 107 frames 124.734 s 25.362 s 4.918x
1024x576, 124 frames 389.992 s 53.429 s 7.299x

Outputs contained decodable video and non-silent audio; repeated warmed outputs were byte-identical within each arm. At 7dce3afb, a separate 672x384 rerun measured 120.352 s versus 26.391 s, with output hashes matching the earlier run.

These are historical GPU results, not a GPU rerun of 1086cf79. The speedup comes from FastH3's four-step inference, not a new offload optimization. This does not establish Base/FastH3 perceptual-quality parity or production throughput.

The measured recipe uses legacy --enable-cpu-offload, including VAE staging. The newer dit/text_encoder component selector keeps VAEs resident and does not inherit those memory measurements.

Detailed evidence and revision boundaries: #6949 (comment)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models enhancement New feature or request labels Sep 3, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches tests/diffusion/, recipes/MiniMaxAI/, vllm_omni/diffusion/ (4 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@NickCao @yenuo26 @Bounty-hunter

Could one of you take a look when you get a chance? Thanks!

…offload

Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

EchoHayate commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor Author

A100 evidence update

I completed the planned controlled Base H3 versus FastH3 run for this PR on one
NVIDIA A100 80GB PCIe GPU with model-level CPU offload.

The comparison kept the model revision, hardware, prompt, seed, FPS, output
shape/length, attention backend, and concurrency fixed. Each arm used one
warmup request followed by three measured requests. Base used 50 steps and
FastH3 Dense/Data-Free used four steps.

Workload Base median E2E FastH3 median E2E E2E speedup Base median denoise FastH3 median denoise Denoise speedup
672x384, 107 frames 124.734 s 25.362 s 4.918x 108.677 s 14.094 s 7.711x
1024x576, 124 frames 389.992 s 53.429 s 7.299x 372.467 s 35.860 s 10.387x

Peak HBM was 65961/65989 MiB for Base/FastH3 at 672x384 and 68456/68527
MiB at 1024x576. The change therefore enables the existing model-level
offload path; it does not claim an additional HBM reduction.

All 16 final requests returned HTTP 200. The outputs contain decodable H.264
video at the requested resolution and 24 FPS plus non-silent 32 kHz stereo AAC
audio. All three warmed outputs are byte-identical within each arm and
workload. Base and FastH3 output digests differ, while logs confirm 50 versus
four inference steps and one FastH3 fusion/activation per server lifetime.
Automated checks found no sustained black segment, one-second freeze, fatal
error, leaked benchmark container, or retained GPU allocation after shutdown.

The source-level scope remains intentionally narrow:

  • allow Dense/Data-Free FastH3 with model-level CPU offload;
  • keep VSA, ordinary layerwise offload, and DLO combinations fail-closed;
  • do not change the schedule, fusion math, request-level LoRA, HWR, or direct
    mmap paths.

After merging the latest upstream main, the two complete MiniMax-H3
FastH3/offload test files plus the PR-added loader contract pass:
57 passed, 15 warnings in 7.84s. All changed-file pre-commit hooks and
git diff --check also pass.

I also reran 672x384 end to end at the current PR head
7dce3afb... (one warmup plus one measured request per arm). The measured
result was 120.352 s Base versus 26.391 s FastH3 (4.560x E2E), with
108.180 s versus 14.147 s denoise time (7.647x). Both measured MP4 digests
exactly match the corresponding outputs from the earlier formal run, and the
same media, error, shutdown, and GPU-cleanup checks passed.

The complete raw text evidence, exact artifact revisions, hashes, commands,
resource samples, and media metadata are retained separately. This is a
single-A100 functional/E2E result, not an official H100/B300 gate or a
production-throughput claim.

@EchoHayate
EchoHayate marked this pull request as ready for review September 3, 2026 05:10
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: MinimaxH3.

Model owners: @david6666666

@EchoHayate, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@EchoHayate

Copy link
Copy Markdown
Contributor Author

Author self-review

I reviewed the production diff, tests, recipe documentation, and the exercised runtime path at the current PR head.

  • The capability change is intentionally limited to Dense/Data-Free FastH3 with the existing model-level CPU-offload path.
  • VSA plus CPU offload, ordinary layerwise offload, DLO, HWR, Ref2VA, and request-level LoRA remain outside this PR and continue to fail closed.
  • I checked that model-level CPU offload stays on the normal loader path and does not create a HostWeightPlan; the tests also cover the capability matrix, loader/fusion ordering, and release metadata.
  • Focused affected tests pass: 57 passed, 15 warnings. Changed-file pre-commit hooks and git diff --check pass.
  • I ran controlled Base 50-step versus FastH3 four-step E2E validation on one A100 80GB at two workloads. Both paths started, generated structurally valid video plus non-silent audio, shut down cleanly, and released GPU allocations. The measured speedup is attributed to FastH3 four-step inference, not to a new offload optimization.
  • I found no P0-P2 issue in the focused diff. The evidence establishes single-A100 functional compatibility and E2E behavior, but does not claim perceptual-quality parity, production throughput, or an official H100/B300 gate.

Current head reviewed: 7dce3afb5498ad95535856a8fa3269482dc0533d.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 3, 2026
…offload

Co-authored-by: TRAE CLI <traecli@bytedance.com>

# Conflicts:
#	tests/diffusion/model_loader/test_diffusers_loader.py
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Current-head refresh / author self-review

  • Refreshed onto upstream main at 1cd92553cf36. The only textual conflict was in tests/diffusion/model_loader/test_diffusers_loader.py; the resolution preserves both the upstream component-selective offload tests and this PR's contract that model-level CPU offload stays on the ordinary loader path without producing a HostWeightPlan.
  • Rechecked the production diff against current main: the capability remains limited to Dense/Data-Free FastH3 plus model-level CPU offload. VSA plus CPU offload, ordinary layerwise offload, DLO, HWR, Ref2VA, and request-level LoRA remain fail-closed.
  • Fresh local verification at this head: the complete MiniMax-H3 FastH3 and offload test files plus the loader contract pass (59 passed, 15 warnings); the five loader tests around the resolved conflict pass separately (5 passed, 15 warnings). Targeted Ruff 0.14.10 check/format, Python byte compilation, and git diff --check pass.
  • Evidence boundary: the earlier controlled single-A100 E2E results were produced at 7dce3afb. This refresh does not claim those measurements as current-head GPU evidence; a current-head smoke remains separate from the CPU/static verification above.

Current head reviewed: e9bacc785a8f66356062c9da9d5f6af90e02adfa.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

update docs since #5929 merged

@EchoHayate

EchoHayate commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Updated for #5929 in bf5b0aaa.

  • The FastH3 recipe now recommends --diffusion-offload-config '{"mode":"module","components":["dit","text_encoder"]}' for new configurations.
  • It documents the MiniMax-H3 compatibility boundary explicitly: the compact selector keeps the VAEs resident, while legacy --enable-cpu-offload also stages the VAEs. The existing single-A100 measurements used the legacy full-topology path, so the recipe does not attribute those peak-HBM results to the compact selector.
  • The loader contract test now exercises the compact module configuration and verifies that it still uses ordinary checkpoint loading without a HostWeightPlan, preserving FastH3 fusion before offloader installation.

Fresh focused verification: 59 passed, 14 warnings; all changed-file pre-commit hooks and git diff --check pass.

Current head reviewed: bf5b0aaa29055e4a2a562ce30befd44fe12443ee.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate
EchoHayate force-pushed the fix/fasth3-model-cpu-offload branch from 4977e7a to bf5b0aa Compare September 6, 2026 12:51
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

please add an indepdent A100 recipe pelase

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Independent A100 recipe update / author self-review

Added the requested standalone single-A100 recipe in 1086cf79.

  • It includes pinned MiniMax-H3 FL2VA and Dense/Data-Free FastH3 download commands.
  • It provides the complete one-GPU server command and the complete 672x384, 4.4-second, four-step T2VA curl request.
  • The server uses the legacy --enable-cpu-offload compatibility path that was actually exercised on the A100 and stages the full MiniMax-H3 topology, including the VAEs.
  • The recipe keeps the [diffusion][feature] Add component-selective offload policies #5929 compact module selector separate and does not attribute the legacy path peak-HBM evidence to that selector.
  • The documented latency, denoise time, frame count, and sampled peak HBM were checked against the retained raw A100 run artifacts. No new GPU or production-throughput claim was added.
  • Fresh checks for the changed recipe: all applicable pre-commit hooks passed, including markdownlint-cli2 and typos; git diff --check passed.

Current head reviewed: 1086cf79018a959631167a17ad754baaf7ab95f7.

@EchoHayate

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106, the requested standalone A100 recipe is included in 1086cf79, with pinned downloads and complete server/request commands. I have also updated the PR description to reflect the already-published A100 evidence and its revision/offload boundaries; no code or new GPU results were added today.

The PR is ready for review and current CI is green. Could you take another look, or help route the remaining review to the appropriate owner?

@EchoHayate

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106, a gentle follow-up on my September 14 message. The requested standalone A100 recipe and the #5929 documentation update are included in 1086cf79; current CI is green, and the published A100 evidence retains its revision/offload boundaries.

Could you confirm any remaining blockers and help route this to an available reviewer? An expected review timeframe would also help. Thanks!

@@ -607,17 +607,22 @@ def check_serving_contract(
raise ValueError("FastH3 preview v1 distills T2VA only, so it cannot serve a Ref2VA partition")
offloads = [

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please check:
check_serving_contract() still inspects only the legacy enable_*_offload booleans. Compact diffusion_offload_config users keep those booleans false, so {"mode":"layer",...}, distributed layer options, and VSA plus {"mode":"module",...} bypass the intended fail-closed matrix (the VSA check at line 622 has the same problem). Please resolve the strategy with resolve_offload(od_config) and reject LAYER_WISE/DISTRIBUTED_LAYER_WISE; for VSA, reject MODEL_LEVEL as well.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d85f0d7. check_serving_contract() now uses resolve_offload(od_config).strategy rather than the legacy booleans. It rejects LAYER_WISE and DISTRIBUTED_LAYER_WISE, and also rejects MODEL_LEVEL for VSA; Dense/Data-Free module offload remains supported.

The regression cases cover compact module allow/reject behavior, compact layer mode, distributed AllGather, and rank-local transfer with positive resident_layers, with legacy booleans left false. Replaying the published-head guard in memory against the nine offload cases reproduced four missing-rejection failures; the repaired focused suite passes 64 tests. All applicable changed-file pre-commit hooks passed, including mypy, CI marks, and SPDX.

These are local macOS CPU checks, not new GPU/E2E or performance evidence. Could you recheck the guard change?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Bounty-hunter Could you recheck this thread when convenient? The fix in the reply above is preserved on refreshed head 7b3aacbf6cfdf6e726185cd3c2ca71cc2cfda887.

check_serving_contract() now calls resolve_offload_strategy(od_config) (the current helper returns resolve_offload(config).strategy). Both layerwise strategies remain rejected, as does model-level offload for VSA; Dense/Data-Free module offload remains allowed. The current-head author self-review records 76 focused tests passing, including compact-policy guard coverage.

GitHub marks this thread outdated but still unresolved. If the current change addresses your concern, could you mark it resolved, or point out anything still missing? I have left the resolution to you. The local test result is not a new GPU/E2E qualification.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Author self-review — compact-offload guard follow-up

Current head reviewed: d85f0d732fb43610b2099e031a52ae4b6237c1f1.

  • Scope: two files only. The FastH3 startup guard now consumes the canonical resolved offload strategy; regression tests cover both legacy and compact configuration entry points. No offloader implementation, kernel, recipe, schedule, or supported feature was added.
  • Invariants: Dense/Data-Free module offload remains allowed. Both layer-wise strategies remain rejected, including compact AllGather and rank-local resident_layers > 0; VSA with module offload remains rejected.
  • Regression evidence: replaying only the published 1086cf79 guard in memory against the nine offload cases produced 4 failed, 5 passed, all failures being missing expected rejection. The repaired full focused selection below reports 64 passed, 15 warnings.
  • Static review and checks: reviewed the production/test diff, direct pipeline caller, and shared strategy resolver. No new P0–P2 findings in this incremental review. All applicable changed-file pre-commit hooks passed, including mypy, test marks, and SPDX; git diff --check passed.
  • Environment/boundaries: local macOS CPU, Python 3.12.13, PyTorch 2.11.0, and vLLM source 98dff2a81d747d1dba01a47f939f48c3526d4206. No new GPU/E2E results, performance claim, or full-loader-suite acceptance. Earlier A100 evidence remains tied to its original revision and offload topology.

Focused reproduction, with this checkout and the compatible vLLM source on PYTHONPATH:

HF_HUB_OFFLINE=1 python -m pytest \
  -p no:cacheprovider \
  -o addopts='--strict-markers --strict-config' \
  tests/diffusion/models/minimax_h3/test_minimax_h3_fasth3.py \
  tests/diffusion/models/minimax_h3/test_minimax_h3_offload.py \
  tests/diffusion/model_loader/test_diffusers_loader.py::test_compact_model_cpu_offload_uses_ordinary_loader_without_host_weight_plan \
  -q --disable-warnings

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both asks from 09-06 are addressed and the core fix is the right shape: check_serving_contract now resolves the offload strategy through the component-selective API instead of blanket-rejecting every offload flag — model-level CPU offload is correctly allowed (it installs after ordinary loading, so the fusion completeness check still runs) while layerwise paths stay rejected, and the VSA variant is explicitly fenced to Dense/Data-Free with a clear error. The recipe's new single-A100 section with pinned revisions and the component-selective docs update close out the documentation asks, and the fasth3 contract tests cover both the allowed and rejected paths. Checks green.

Merge upstream 7ab582a, preserve the reviewed offload capability matrix through the canonical strategy accessor, and clarify recipe evidence and non-offloaded VSA instructions.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Author self-review — upstream conflict refresh

Current head reviewed: 28037909b993956aa3be2f7c608597dd49d4ab8e.

Refreshed against upstream 7ab582ae2545476284bef8903bec1c0ddc44a082.
This resolves the conflict with the resolved-offload-policy refactor without
expanding the previously reviewed FastH3 support matrix.

  • Scope/invariants checked: the guard now uses the upstream
    resolve_offload_strategy accessor. Dense/Data-Free model-level CPU offload
    remains supported through both legacy and compact inputs; ordinary checkpoint
    loading still precedes offloader installation and preserves adapter fusion.
  • Fail-closed combinations: both layer-wise strategies remain rejected,
    including compact AllGather and rank-local transfer with positive resident
    layers. VSA plus model-level offload and the preview-v1 Ref2VA partition remain
    rejected. Upstream adapter-target and test-fixture updates were preserved.
  • Independent static review: identified and corrected an ambiguous VSA
    recipe instruction that could inherit CPU-offload flags from the preceding
    example. VSA instructions now explicitly require a non-offloaded command.
    The reviewer confirmed that the finding is resolved.
  • Checks: all applicable pre-commit hooks passed for the four PR-delta
    files, including mypy, Ruff, typos, Markdown, CI marks, and SPDX. The scoped
    diff whitespace check also passed.
  • Focused CPU tests: the FastH3 contract suite, MiniMax-H3 offload suite,
    and compact ordinary-loader regression passed: 76 passed, 15 warnings,
    with no failures, errors, or skips.
  • Broader test boundary: the expanded loader/static-gate selection reported
    125 passed, 6 failed. All six failures reproduced on an unmodified archive
    of the pinned upstream revision: two Linux mount-info requirements, one local
    Torch MPS allocator assertion, two unsupported-platform Helios cases, and an
    existing Qwen Image legacy-flag reader. They were not patched into this PR,
    and this is not a claim that the broader suite is green.
  • Runtime/evidence boundary: macOS CPU validation with Python 3.12,
    PyTorch 2.11, and the compatible vLLM source overlay. No new GPU/E2E or
    performance run. The recipe now explicitly states that its earlier A100
    measurements predate this refresh and do not validate the refreshed code's
    latency or peak HBM.

Focused reproduction (with compatible vLLM source available on PYTHONPATH):

HF_HUB_OFFLINE=1 python -m pytest \
  -p no:cacheprovider \
  -o addopts='--strict-markers --strict-config' \
  tests/diffusion/models/minimax_h3/test_minimax_h3_fasth3.py \
  tests/diffusion/models/minimax_h3/test_minimax_h3_offload.py \
  tests/diffusion/model_loader/test_diffusers_loader.py::test_compact_model_cpu_offload_uses_ordinary_loader_without_host_weight_plan \
  -q --disable-warnings

@hsliuustc0106 hsliuustc0106 added the ready label to trigger buildkite CI label Sep 28, 2026
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 28037909b993 refers to an earlier head; the pull request now points at 64a36a1b0cc2.

@EchoHayate

Copy link
Copy Markdown
Contributor Author

CI failure triage — September 28

@hsliuustc0106, I inspected the Buildkite failures on
28037909b993956aa3be2f7c608597dd49d4ab8e. Could you help route the remaining
CI decisions to the appropriate owners?

Confirmed inherited failure

Main build #16231
reports 7101 passed, 1 failed, 48 skipped in the diffusion CPU selection.
The sole failure is
test_no_runtime_module_reads_the_legacy_offload_flags, identifying
diffusion/models/qwen_image/pipeline_qwen_image.py:344.
Both that source file and the gate test are unchanged versus incoming upstream
7ab582ae2545476284bef8903bec1c0ddc44a082. The identical failure was reproduced
on an unmodified archive of that revision during the September 27 checks and
disclosed in the preceding self-review. A blind retry will not address that
deterministic source check.

AMD failures still needing attribution

AMD build #12868
has eight failed tests across the four diffusion shards:

  • the same Qwen legacy-reader gate;
  • three test_hunyuan_image3_peft_qkv_forward cases attempting
    vllm::rocm_unquantized_gemm on CPU;
  • three Wan decoder fast-path cases;
  • one Ulysses UAA mask-layout case.

Those test files are also unchanged by this PR, but I have not reproduced
the other AMD failures on the base under the same ROCm environment, so I am
not calling them all inherited or transient. The separate R2-01 GPU job
aborted with exit 134, but is marked soft_failed; it should not be confused
with the hard-failing diffusion shards.

Could you point me to an existing upstream fix for the Qwen gate and an AMD
CI owner or same-environment base run for comparison? That would let us choose
between a targeted upstream refresh and retries of genuinely transient jobs,
without folding unrelated runtime changes into this FastH3 compatibility PR.

No new code, tolerance change, or CI retry has been made in this triage.
Wheel builds, pre-commit, DCO, docs, Intel and NPU statuses currently pass;
the main and AMD Buildkite failures remain open.

@EchoHayate

Copy link
Copy Markdown
Contributor Author

September 30 CI follow-up

This FastH3 compatibility fix is still planned. I located concrete upstream
changes relevant to the September 28 CI triage:

These are upstream fixes or reported evidence, not a fresh base reproduction
by me, and not proof that every failure on this PR is now cleared. The checks
on head 28037909b993 still report the old main/AMD failures.

@hsliuustc0106, the concrete next step is an upstream refresh with the existing
FastH3 support matrix unchanged, followed by scoped regression checks and
fresh CI. Is waiting for #8165 before that refresh the preferred sequencing?
I will not cherry-pick unrelated test-policy changes or blindly retry the old
head. No code, CI retry, or new GPU/performance result is part of this comment.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@EchoHayate this PR is labeled ready + high priority, but CI is failing on the latest commit:

Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks!

EchoHayate and others added 2 commits October 4, 2026 15:13
Record the verified October 1 upstream merge at 423f343 without changing the reviewed FastH3 support matrix.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Merge upstream 3ab726c, including consolidated AMD CI repairs, while preserving the existing four-file FastH3 offload delta and historical evidence boundaries.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Author self-review — October 4 CI refresh

@hsliuustc0106 I refreshed this branch onto upstream
3ab726c126d760f6052c815b9bdced135cd86154, including the consolidated
CI repairs in #8342. #8165 was closed as superseded by that merged PR;
there is no separate routing cherry-pick in this update.

  • Scope: the four-file FastH3 PR delta is unchanged in content
    (stable patch ID aaf22c8eacca09bcf37e21ff96959509606e286c).
    The refresh preserves upstream loader/HSDP additions and test-policy
    changes without modifying them.

  • Support boundaries: Dense/Data-Free model-level CPU offload remains
    supported via legacy and compact configuration. Both layerwise
    strategies, VSA with model-level offload, and preview-v1 Ref2VA remain
    rejected. Ordinary loading still precedes offloader installation.

  • Fresh focused checks: 84 passed across the FastH3/offload suites,
    compact ordinary-loader regression, legacy-reader gate and AMD pipeline
    tests. Separately, all seven selected Wan CPU-plumbing and Hunyuan LoRA
    regression cases passed. All applicable four-file pre-commit hooks and
    the scoped whitespace check passed.

  • Broader comparison: candidate 285 passed / 28 failed / 2 skipped;
    untouched pinned upstream 284 passed / the same 28 failed / 2 skipped.
    All 314 shared test outcomes and failure messages match; the
    candidate-only loader regression passes. These failures are retained
    local platform/offline limitations, not waived or fixed by this PR.

  • Evidence boundary: these are macOS CPU/source-overlay checks, not
    an AMD runtime qualification or a new GPU benchmark. Earlier A100
    results retain their original revision and offload-topology boundaries.

  • Independent pre-publication review: a separate agent reviewed the
    exact four-file delta, upstream loader compatibility, and validation
    artifacts; no concrete P0–P2 findings. This is not a maintainer approval.
    The tests do not provide a single end-to-end runner/fusion/offload-install
    integration check; those contracts were checked separately.

Current head reviewed: 64a36a1b0cc239069be96c398630b536b3c53071.

Please evaluate the new exact-head CI rather than retrying the old
28037909 builds. I will follow up on any remaining failures with
head-specific evidence.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Author self-review — upstream conflict refresh

Refreshed against upstream 874b8b3ad. Resolved the MiniMax-H3 recipe conflicts while preserving upstream FlashInfer VSA, selective AdaLN offload, sampler, and VDN-H3 documentation. The FastH3 production guard is unchanged.

Dense/Data-Free module offload still uses ordinary checkpoint loading and completes fusion/validation before offloader installation. Layerwise and distributed-layerwise offload, VSA plus module offload, and preview-v1 Ref2VA remain fail-closed. The recipe now clarifies the separate AdaLN offload boundary and removes module/layer offload flags from the VSA example.

Validation:

  • 76 focused tests passed; all applicable four-file pre-commit hooks and diff checks passed.
  • The additional upstream-interaction selection returned 15 passed / 5 failed on both this candidate and unmodified upstream in the same local environment. All five failures require the unavailable compiled silu_and_mul operator. No fixture or operator bypass was used.
  • Source overlay: vLLM v0.31.0 (db9527a4). This local macOS/CPU environment is not the full upstream CI image.
  • Direct and independent review found no concrete P0–P2 issue in the four-file PR delta. The committed delta matches the tested patch.

This refresh does not renew the historical A100 benchmarks or establish GPU/E2E or production-performance acceptance. GitHub reports the refreshed head as mergeable; new-head basic CI was still pending when checked.

Current head reviewed: 7b3aacbf6cfdf6e726185cd3c2ca71cc2cfda887.

@EchoHayate

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 Following up on your October 4 CI request: the refreshed head is 7b3aacbf6cfdf6e726185cd3c2ca71cc2cfda887, based on upstream 874b8b3ad.

All five basic checks on this head are now green: Python 3.11/3.12 wheel builds, pre-commit, DCO and Read the Docs. The scope, local verification and evidence limitations are documented in the author self-review above.

I cannot see a buildkite/vllm-omni or buildkite/vllm-omni-amd-ci result attached to this head. Could you confirm whether the required current-head builds need maintainer authorization/triggering, and help initiate them if appropriate? Please also let me know if there is any other outstanding merge gate or author-side work. I am not treating the basic checks as GPU/AMD CI acceptance or reusing the old-head Buildkite results.

The existing approval remains recorded; this is a request to verify the remaining gates, not to bypass them.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants