Skip to content

[Core] Expose HWR payload validation in diffusion startup - #7144

Open
hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-integrity-policy
Open

hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-integrity-policy

Conversation

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Purpose

Fixes #7141. Part of #7107.

The diffusion loader always selected metadata-only HWR validation, so serving users could not opt into the existing full payload checksums. Add --host-weight-runtime-validation / host_weight_runtime_validation through CLI, offline and stage configuration to the loader's IntegrityPolicy.

Keep manifest_and_metadata as the default; full_checksum reads and hashes payloads on warm acquisition. Existing preferred mode reloads canonical weights on corruption and can publish a replacement; required mode fails startup before restoring corrupt weights. Policy selection does not change artifact identity or domain metadata.

Document startup cost and the limits of validation. Add the mixin's existing od_config: OmniDiffusionConfig contract so the new access is typed.

Test Plan

Real CPU loader/store tests populate canonical BF16 weights, mutate only payload bytes while preserving header/size, then verify both validation levels in preferred/required mode. Preferred full-checksum recovery produces correct tensors and a subsequent warm hit. Config tests cover forwarding, defaults and invalid values.

CUDA_VISIBLE_DEVICES='' VLLM_TARGET_DEVICE=cpu \
  /tmp/codex-pr6607-vllm028/bin/python -m pytest -n 0 -q \
  tests/host_weight_runtime \
  tests/diffusion/model_loader/test_diffusers_loader.py \
  tests/config/test_omni_config.py \
  tests/entrypoints/test_async_omni_diffusion_config.py

vLLM Version: 0.28.0; Python 3.12.13; torch 2.13.0+cu129; safetensors 0.8.0.

vLLM-Omni Commit: based on 039808e0d97d7969a2cb102d074c1e45b8ceef0b.

Test Result

  • Combined suites: 366 passed, 1 skipped, 2 failed. Both failures are existing Helios model integration tests (test_get_all_weights, test_load_model) that attempt CUDA with GPUs hidden and raise No CUDA GPUs are available. All HWR and new policy tests passed.
  • Final focused HWR loader/config/CLI selection: 18 passed.
  • All applicable local hooks pass except mypy: 32 diagnostics, all present among the 38 baseline diagnostics after normalizing line numbers. No new errors; declaring the existing mixin config contract removes six. The comparison uses unchanged base production files.
  • Exact repository hooks/pins use an external config selecting installed Node v22.23.1 for markdownlint; repository configuration is unchanged. Full precheck and diff check completed.
  • No GPU workload completed, no model-quality or startup-performance claim, and no automatic running-service recovery claim.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/host_weight_runtime.md.

Module owners: @hsliuustc0106

@hsliuustc0106, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit d989aa6fe2af produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106
hsliuustc0106 removed the request for review from Gaohan123 September 6, 2026 09:10
@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request labels Sep 6, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 16 days

@hsliuustc0106 this pull request has had no human commit, comment or review since 2026-09-06. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 3c608d57-9012-433c-b404-fcb63b109bbb) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

1 similar comment
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 3c608d57-9012-433c-b404-fcb63b109bbb) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: cf5e007d-dba3-4a70-816a-725c025699e6) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 4e243d98-90fe-477e-8694-d3d8700aa0f2) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

0 actionable finding(s).

CI at d989aa6fe2af (2026-10-10T05:29:57.446014+00:00): verification incomplete; required-check status is unknown. Observed Buildkite: buildkite/omni-release (passed).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

Diffusion startup can now choose how strictly Host Weight Runtime checks a warm artifact. --host-weight-runtime-validation and host_weight_runtime_validation travel through the serve CLI, orchestrator args, the default diffusion stage, stage deploy config, and OmniDiffusionConfig into the loader’s IntegrityPolicy.local_lookup. The default stays manifest_and_metadata, which checks identity and tensor metadata. full_checksum hashes payload bytes on warm acquisition: preferred mode reloads canonical weights and can publish a replacement, and required mode fails startup. Artifact identity is unchanged, and the check runs at load time.

Change flow

flowchart LR
    A["[NEW] CLI flag host_weight_runtime_validation"]:::new --> B["[CHANGED] Stage and OmniDiffusionConfig"]:::changed
    B --> C["[CHANGED] Loader IntegrityPolicy local_lookup"]:::changed
    C --> D["[EXISTING] Warm HWR acquire"]:::existing
    classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
    classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
    classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
    classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[RFC] Expose HWR payload validation policy in diffusion startup

2 participants