Repository navigation
[Core] Expose HWR payload validation in diffusion startup - #7144
hsliuustc0106 wants to merge 1 commit into
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/host_weight_runtime.md. Module owners: @hsliuustc0106 @hsliuustc0106, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: no human activity for 16 days@hsliuustc0106 this pull request has had no human commit, comment or review since 2026-09-06. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 3c608d57-9012-433c-b404-fcb63b109bbb) — check |
1 similar comment
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 3c608d57-9012-433c-b404-fcb63b109bbb) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: cf5e007d-dba3-4a70-816a-725c025699e6) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 4e243d98-90fe-477e-8694-d3d8700aa0f2) — check |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
0 actionable finding(s).
CI at
d989aa6fe2af(2026-10-10T05:29:57.446014+00:00): verification incomplete; required-check status is unknown. Observed Buildkite:buildkite/omni-release(passed).
Note: The assigned review arm
strict/zcode/GLM-5.3-Flashcould not complete this review, so it was produced by the fallback armdirect/cursor/auto. It is excluded from the routing experiment.
Full review analysis
PR description
Diffusion startup can now choose how strictly Host Weight Runtime checks a warm artifact. --host-weight-runtime-validation and host_weight_runtime_validation travel through the serve CLI, orchestrator args, the default diffusion stage, stage deploy config, and OmniDiffusionConfig into the loader’s IntegrityPolicy.local_lookup. The default stays manifest_and_metadata, which checks identity and tensor metadata. full_checksum hashes payload bytes on warm acquisition: preferred mode reloads canonical weights and can publish a replacement, and required mode fails startup. Artifact identity is unchanged, and the check runs at load time.
Change flow
flowchart LR
A["[NEW] CLI flag host_weight_runtime_validation"]:::new --> B["[CHANGED] Stage and OmniDiffusionConfig"]:::changed
B --> C["[CHANGED] Loader IntegrityPolicy local_lookup"]:::changed
C --> D["[EXISTING] Warm HWR acquire"]:::existing
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
Purpose
Fixes #7141. Part of #7107.
The diffusion loader always selected metadata-only HWR validation, so serving users could not opt into the existing full payload checksums. Add
--host-weight-runtime-validation/host_weight_runtime_validationthrough CLI, offline and stage configuration to the loader'sIntegrityPolicy.Keep
manifest_and_metadataas the default;full_checksumreads and hashes payloads on warm acquisition. Existing preferred mode reloads canonical weights on corruption and can publish a replacement; required mode fails startup before restoring corrupt weights. Policy selection does not change artifact identity or domain metadata.Document startup cost and the limits of validation. Add the mixin's existing
od_config: OmniDiffusionConfigcontract so the new access is typed.Test Plan
Real CPU loader/store tests populate canonical BF16 weights, mutate only payload bytes while preserving header/size, then verify both validation levels in preferred/required mode. Preferred full-checksum recovery produces correct tensors and a subsequent warm hit. Config tests cover forwarding, defaults and invalid values.
CUDA_VISIBLE_DEVICES='' VLLM_TARGET_DEVICE=cpu \ /tmp/codex-pr6607-vllm028/bin/python -m pytest -n 0 -q \ tests/host_weight_runtime \ tests/diffusion/model_loader/test_diffusers_loader.py \ tests/config/test_omni_config.py \ tests/entrypoints/test_async_omni_diffusion_config.pyvLLM Version: 0.28.0; Python 3.12.13; torch 2.13.0+cu129; safetensors 0.8.0.
vLLM-Omni Commit: based on
039808e0d97d7969a2cb102d074c1e45b8ceef0b.Test Result
test_get_all_weights,test_load_model) that attempt CUDA with GPUs hidden and raiseNo CUDA GPUs are available. All HWR and new policy tests passed.