Skip to content

[Core] Add opt-in serialized HWR build admission - #7140

Open
hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-build-admission
Open

hsliuustc0106 wants to merge 1 commit into
mainfrom
codex-hwr-build-admission

Conversation

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Purpose

Fixes #7135. Part of #7107.

Different HWR identities can build concurrently and race the same capacity checks. Add opt-in CapacityPolicy(build_admission="serialized", max_store_bytes=...) to admit one producer/publication operation per domain. Warm hits bypass admission and same-key callers still join published artifacts. The existing coordination deadline bounds admission waiting; timeout and acquisition errors use existing typed failure routing.

Reuse the existing filesystem lock and release it after every outcome or process death. Concurrent domains retain their exact schema-1 policy; serialized domains use schema 2, preventing mixed workers from bypassing admission. Opting in requires a new domain root. Default behavior is unchanged.

This coordinates cooperative producers, not a filesystem byte quota: metadata and outside writers still affect usage. It does not add parallel reservations, automatic reclamation, a CLI option, or cancellation of a hung producer.

Test Plan

Real CPU processes exercise same/different-key coordination in both modes, capacity rejection after another build commits, crash recovery and tensor correctness. Additional tests cover warm-hit bypass, timeout ownership and preferred/required routing, producer failure, release errors, and policy compatibility.

CUDA_VISIBLE_DEVICES='' VLLM_TARGET_DEVICE=cpu \
  /tmp/codex-pr6607-vllm028/bin/python -m pytest -n 0 -q tests/host_weight_runtime

vLLM Version: 0.28.0; Python 3.12.13; torch 2.13.0+cu129; safetensors 0.8.0.

vLLM-Omni Commit: based on 039808e0d97d7969a2cb102d074c1e45b8ceef0b.

Test Result

  • Complete HWR CPU suite: 139 passed.
  • All applicable local pre-commit gates passed. Exact repository hooks/pins use an external config selecting installed Node v22.23.1 for markdownlint; repository configuration is unchanged.
  • Full precheck and diff-scoped simplification review completed without blockers. No new dependency, model-specific Python example, or accelerator call.
  • No GPU or performance claim; this policy trades concurrent cold production for serialized admission.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/host_weight_runtime.md.

Module owners: @hsliuustc0106

@hsliuustc0106, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 90815cc86632 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request labels Sep 6, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 16 days

@hsliuustc0106 this pull request has had no human commit, comment or review since 2026-09-06. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Serialize HWR builds for bounded-capacity domains

2 participants