Skip to content

[Core] Add explicit HWR registration bypass and progress logs - #7147

Open
hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-registration-policy
Open

hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-registration-policy

Conversation

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Purpose

Fixes #7146. Part of #7107; related to #6726, which remains unresolved.

Add --dlo-host-registration-mode disabled / dlo_host_registration_mode="disabled" to select bounded host staging without registering HWR mappings. This avoids using a guessed undersized byte budget as a disable switch. Preserve the auto default, zero-budget unlimited semantics, and staging's existing pin-memory policy.

Carry the option through CLI, offline/stage configuration and OffloadConfig. Disabled mode exits before platform registration. Auto mode logs registration entry with device/budget and each CUDA region's entry/completion, making the last entered boundary visible if a synchronous call stalls. Existing rollback and lease retention are unchanged.

This is explicit avoidance and diagnostics, not safe cancellation of a CUDA call or a fix for the driver's underlying hang. No automatic timeout fallback or cross-rank coordination is introduced.

Test Plan

Existing CPU offloader fixtures exercise disabled/auto selection with zero and positive budgets, no registration call in disabled mode, preserved staging pinning policy, bounded slots, and lease cleanup. A fake CUDA runtime verifies that region-entry logs exist inside each call, completion logs only follow success, and failures still roll back. Config/CLI tests cover forwarding and invalid/ineligible settings.

CUDA_VISIBLE_DEVICES='' VLLM_TARGET_DEVICE=cpu \
  /tmp/codex-pr6607-vllm028/bin/python -m pytest -n 0 -q \
  tests/diffusion/offloader/test_distributed_layerwise_backend.py \
  tests/diffusion/offloader/test_cuda_host_registration.py \
  tests/config/test_omni_config.py \
  tests/entrypoints/test_async_omni_diffusion_config.py

vLLM Version: 0.28.0; Python 3.12.13; torch 2.13.0+cu129; safetensors 0.8.0.

vLLM-Omni Commit: based on 039808e0d97d7969a2cb102d074c1e45b8ceef0b.

Test Result

  • Focused selection/diagnostic/config checks: 14 passed.
  • Complete four-suite CPU run: 310 passed, 1 skipped. Three existing CPU tests now request the existing mocked runtime fixture; previously they attempted to construct real CUDA streams before their assertions. No assertion was removed.
  • All applicable local hooks pass except five mypy diagnostics proven identical on base and branch after normalizing line numbers. Final changed-test hooks pass. Full precheck completed without new blockers.
  • Exact repository hook pins use an external config selecting installed Node v22.23.1 for markdownlint; repository configuration unchanged.
  • No accelerator workload completed. Fake-runtime tests do not establish real pinning, CUDA latency, driver recovery or inference correctness. The reported RTX 4090 incident still needs native-stack/device investigation under [Bug]: HWR mmap registration (cudaHostRegister) never completes on 4x RTX 4090, TP2xUSP2 no-AllGather DLO, MiniMax-H3 #6726.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/offloader.md.

Module owners: @yuanheng-zhao @david6666666 @lishunyang12

@hsliuustc0106, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 729ccd7f11df produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request labels Sep 6, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 16 days

@hsliuustc0106 this pull request has had no human commit, comment or review since 2026-09-06. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: a45022ce-c0d0-43f7-87a9-21305ba3dd4a) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: cdb41500-2595-4aac-b708-d1a33c65f27b) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 44f98e70-7b11-4264-91bf-99e204196ed7) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

0 actionable finding(s).

CI at 729ccd7f11df (2026-10-10T05:35:34.294343+00:00): verification incomplete; required-check status is unknown. Observed Buildkite: buildkite/omni-release (passed).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

This change adds an explicit HWR registration policy, dlo_host_registration_mode, with auto as the default and disabled as an opt-in. disabled is accepted only for no-AllGather distributed layerwise offload with Host Weight Runtime enabled, and the offloader returns before cudaHostRegister so bounded host staging is used while the existing pin-memory policy stays in place. auto keeps zero-budget unlimited registration and now logs the device, budget, and each CUDA region's entry and successful completion so a stalled synchronous register call leaves the last entered boundary visible. Rollback, lease retention, and the underlying driver hang are unchanged.

Change flow

flowchart LR
  cli["[CHANGED] CLI, deploy, and diffusion config"]:::changed
  transport["[CHANGED] OffloadConfig validation"]:::changed
  bypass["[NEW] Disabled mode skips CUDA registration"]:::new
  logs["[CHANGED] Auto mode region entry and completion logs"]:::changed
  staging["[EXISTING] Two-slot bounded host staging"]:::existing
  cli --> transport
  transport --> bypass
  transport --> logs
  bypass --> staging
  logs --> staging
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[RFC] Add explicit HWR registration bypass and progress diagnostics

2 participants