Repository navigation
[Core] Add opt-in serialized HWR build admission - #7140
hsliuustc0106 wants to merge 1 commit into
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/host_weight_runtime.md. Module owners: @hsliuustc0106 @hsliuustc0106, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: no human activity for 16 days@hsliuustc0106 this pull request has had no human commit, comment or review since 2026-09-06. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Purpose
Fixes #7135. Part of #7107.
Different HWR identities can build concurrently and race the same capacity checks. Add opt-in
CapacityPolicy(build_admission="serialized", max_store_bytes=...)to admit one producer/publication operation per domain. Warm hits bypass admission and same-key callers still join published artifacts. The existing coordination deadline bounds admission waiting; timeout and acquisition errors use existing typed failure routing.Reuse the existing filesystem lock and release it after every outcome or process death. Concurrent domains retain their exact schema-1 policy; serialized domains use schema 2, preventing mixed workers from bypassing admission. Opting in requires a new domain root. Default behavior is unchanged.
This coordinates cooperative producers, not a filesystem byte quota: metadata and outside writers still affect usage. It does not add parallel reservations, automatic reclamation, a CLI option, or cancellation of a hung producer.
Test Plan
Real CPU processes exercise same/different-key coordination in both modes, capacity rejection after another build commits, crash recovery and tensor correctness. Additional tests cover warm-hit bypass, timeout ownership and preferred/required routing, producer failure, release errors, and policy compatibility.
CUDA_VISIBLE_DEVICES='' VLLM_TARGET_DEVICE=cpu \ /tmp/codex-pr6607-vllm028/bin/python -m pytest -n 0 -q tests/host_weight_runtimevLLM Version: 0.28.0; Python 3.12.13; torch 2.13.0+cu129; safetensors 0.8.0.
vLLM-Omni Commit: based on
039808e0d97d7969a2cb102d074c1e45b8ceef0b.Test Result