Repository navigation
[Bugfix] Give SHM connector lock files a segment-derived lifecycle (#7636 Issue 24) - #7840
yanbao1217 wants to merge 2 commits into
Conversation
|
This PR appears to belong to: docs/design/module/omni_connector.md, docs/design/module/input_output_modality_contracts.md, docs/design/module/diffusion/parallelism.md. Module owners: @princepride @fake0fan @natureofnature Routing: @princepride via module of the changed files, module named in the PR description, CODEOWNERS; @fake0fan via module of the changed files, module named in the PR description; @natureofnature via module of the changed files, module named in the PR description @yanbao1217, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
4d0ac49 to
efcccd4
Compare
…ssue 24 of vllm-project#7636) Lock files /dev/shm/shm_{key}_lockfile.lock were only removed on the successful receiving read or by cleanup()/close() for keys still in the sender's in-memory _pending_keys, so put() failures, failed reads whose segment was already unlinked, and abnormally terminated processes (whose segments resource_tracker reaps) left zero-byte lock files behind forever. Contract: a lock file has no lifecycle of its own - it exists exactly while its segment does. Changes: - put(): remove the lock file when shm_write_bytes raises (the segment never came to life), re-raising afterwards - _get_data_with_lock(): remove the lock file on failure when the segment is already gone; keep it while the segment lives (retryable) - once per process, the first constructed connector sweeps orphan lock files (segment absent + mtime past a 60 s grace + non-blocking flock acquirable + inode/uid recheck under the lock) - class docstring spells out the contract; _pending_keys is now purely cleanup()/close() bookkeeping Adds TestLockFileLifecycle covering the matrix requested in review pullrequestreview-5197739857: roundtrip, put failure, failed reads with dead/live segments, sweep preconditions (stale, live segment, fresh, held, unrelated files), once-per-process, and a crashed subprocess. Signed-off-by: yanbao1217 <yanliuwebsite493@gmail.com>
The forbidden-imports hook rejects stdlib re: the lock-file name match now uses prefix/suffix slicing (equivalent, dependency-free). mypy-3.10 flags using a list.append result in a boolean context: the once-per-process test stub is a plain function now. Signed-off-by: yanbao1217 <yanliuwebsite493@gmail.com>
efcccd4 to
fc17ff9
Compare
|
Self-review:
|
|
Out of scope for this PR: we are implementing file lock manually using fcntl in multiple places, these should be replaced with the filelock library. |
|
Thanks for taking a look! Happy to track the migration in a separate issue if that helps. |
Omni ReviewBot: no human activity for 7 days@yanbao1217 this pull request has had no human commit, comment or review since 2026-09-24. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on cursor (cursor-grok-4.6-high) under experiment |
Follow-up of the Unified Full-duplex Framework (#7413): addresses issue 24 of #7636 (SHM connector lock-file lifecycle contract), from @Sy0307's review summary (pullrequestreview-5197739857).
Purpose
SharedMemoryConnectorcreates a zero-byte lock file/dev/shm/shm_{key}_lockfile.lockper transfer as anflockcarrier. Lock files were only removed on the successful receiving read, or bycleanup()/close()for keys still in the sender's in-process_pending_keysset:What broke:
put()— the lock file is created beforeshm_write_bytes()runs (shm_connector.py:49onmain), the key is tracked only after success (:56) — leaked the lock file even on a gracefulclose();shm_read_bytes()had already unlinked left the lock file behind forever (:81-86removed it only whendeserialized): read succeeds,deserialize_obj()fails, segment gone, lock orphaned, no removal path left at all;close()) stranded one lock file per unconsumed transfer:resource_trackerreaps the segments at owner death — the warnings seen after the duplex process exits — while the lock files are plain files in tmpfs with no backstop at all; they accumulate until the machine reboots.Root cause: the lock file had a lifecycle of its own, tracked only in process memory. Fix — the contract: a lock file has no lifecycle of its own; it exists exactly while its segment does.
put(): whenshm_write_bytes()raises, remove the lock file while still holding it, then re-raise._get_data_with_lock(): on failure, remove the lock file when the segment is already gone; keep it while the segment lives (the transfer can be retried).open() -> flock()window) and segment absence;O_RDONLYopen + non-blockingflock(fails while a critical section holds the lock); recheck under the lock — segment still absent, mtime still past grace, path still resolves to the locked inode, owned by this uid — then remove while holding the lock._pending_keysbecomes purelycleanup()/close()bookkeeping.cleanup()/close()behavior is unchanged. No threads, no timers, no lock-file content format.Residual risks accepted: a µs-scale interleaving (consumer releases the lock → same-key
put()creates a new inode → the sweep'sunlinklands on the new file) can produce "segment alive, lock gone" — the receivingget()then returnsNoneand the sender'sclose()reclaims the segment; fixing that would need a lock-file content protocol, deliberately avoided here. Ashm_write_bytes()failure after segment creation but before returning can still leak a half-written segment (pre-existing). A forked child inherits the once-per-process flag; the next fresh interpreter sweeps.Note for maintainers:
SharedMemoryConnector.cleanup()currently has no production callers (only unit tests; the chunk adapter'scleanup()is a method on the adapter, not the connector). Lock reclamation in practice rides on the receiving read,close()at shutdown, and now the startup sweep — wiringcleanup()into request teardown might be worth a separate discussion. This file is byte-identical to pre-#7413main; the leak predates the framework.Test Plan
New
TestLockFileLifecycleintests/distributed/omni_connectors/test_shm_connector.py(13 cases,core_model and cpu): roundtrip; failedput()plus same-key retry; failed reads with dead/live segments (the deserialize-after-read case is the leak that had no removal path at all); sweep preconditions (stale orphan, live segment, fresh file, heldflock, unrelated files); once-per-process; and a crashed-subprocess orphan swept afterwards.core_model and cpu./dev/shmholds zero*lockfile*residue after the full directory run.put()failure, receiver deserialize failure after read, sender SIGKILL, sender clean exit withoutclose(), partial multi-chunk consumption with both processes killed — on unpatchedmainand on this branch.ruff check/ruff format,mypy==1.11.1(--follow-imports silent --ignore-missing-imports),tools/pre_commit/check_forbidden_imports.pyon both changed files.test_bagel_shared_memory_connector.py,advanced_model) runs in CI on merge; it was not run locally.Quick repro of the bug (on
mainthe second command fails withImportErrorand the lock file stays forever; on this branch it printsstill there: False):pytest tests/distributed/omni_connectors/test_shm_connector.py -m "core_model and cpu"vLLM Version: 0.28.0 (local cpu suites; the connector path does not touch vLLM APIs) — CI lane runs the repo's 0.29 line
vLLM-Omni Commit: f90c267 (tested base), rebased on 23d8c83
Test Result
CPU-marked suites on a dev box (an RTX 5090 is available but not exercised by these suites), Python 3.12, editable checkout of this branch:
pytest tests/distributed/omni_connectors/test_shm_connector.py -m "core_model and cpu": 29 passed (16 pre-existing + 13 new).pytest tests/distributed/omni_connectors -m "core_model and cpu": 379 passed, 1 skipped, 30 deselected in 28 s. The one skip istest_omni_connector_configs.py:66("No config files found or directory missing") — environment-conditional and pre-existing; the deselected are non-CPU marks (mooncake/NIXL guards,advanced_modelGPU e2e)./dev/shmcontains zero*lockfile*residue after the full directory run.mainevery path leaves lock files behind — theresource_trackerwarnings and the reaped-but-locked state from the issue reproduce one-to-one; on this branch all paths are clean and the baselines are unchanged.refor the name match — replaced with prefix/suffix slicing in 4d0ac49); SPDX headers untouched; the new tests carry the module'score_model and cpumarks.put()failure cleanup swallowOSErroraround best-effortos.removecalls, matching the existingcleanup()/close()idiom in the same file; nothing on a fail-fast path swallows.🤖 Generated with Claude Code