Repository navigation
[Bugfix][MoRIIO] Keep discovery heartbeats running while workers hold the GIL - #59441
Conversation
Extract the process-based discovery sender from vllm-project#58968 onto main. Preserve the discovery payload and rank-zero startup guard, stop the helper on explicit shutdown or parent pipe EOF, and retain the GIL-starvation and parent-exit lifecycle regressions. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: whx-sjtu <xiaowang990929@gmail.com>
944b575 to
512a877
Compare
| failures = 0 | ||
| with zmq.Context() as context, context.socket(zmq.DEALER) as sock: | ||
| sock.setsockopt(zmq.LINGER, 0) | ||
| sock.setsockopt(zmq.SNDTIMEO, 1000) |
There was a problem hiding this comment.
Could we preserve recovery after a temporary discovery/router outage here?
This adds SNDTIMEO=1000, and lines 75–79 terminate the helper after max_retries consecutive ZMQErrors. The parent never monitors or restarts _process, so once the helper exits this worker permanently stops refreshing its discovery registration even if the router/network later recovers.
This is slightly different from the current _ping path on main: the socket uses ZeroMQ's default infinite send timeout, so once its outbound queue is full the sender blocks and can resume when connectivity comes back rather than terminating the heartbeat owner.
Could we either:
- keep retrying recoverable send errors for the lifetime of the worker, or
- have the parent detect/restart a failed heartbeat subprocess?
It would also be useful to add a regression test that makes discovery unavailable and then restores it, verifying that the same live worker resumes registration without requiring a vLLM restart.
There was a problem hiding this comment.
Agreed. Working on it.
Retry temporary ZMQ send timeouts for the worker lifetime. Surface unexpected heartbeat child termination, including exit code zero, from worker completion polling. Cover real router outage recovery and child exits. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: whx-sjtu <xiaowang990929@gmail.com>
|
✅ @whx-sjtu, CI is now available for this PR.
|
|
/amd-ci run |
|
✅ Triggered Buildkite AMD CI #14007 for commit |
|
/ci run |
|
❌ This PR is 9 commits behind upstream |
|
/ci run |
|
✅ Triggered Buildkite CI #92574 for commit |
|
/ci run |
|
✅ Triggered Buildkite CI #92760 for commit |
CI selector (shadow): 2 test steps (5 jobs) instead of 8 (13 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (2)
Would skip (today's rules run them) (6)
Would add (today's rules do not run them) (0)none AMD mirrors: would skip (16)
AMD mirrors: would add (0)none 4 changed files · base |
Overview
Keep MoRIIO discovery registration alive during worker GIL holds and temporary
router outages. Report unexpected heartbeat subprocess exit through worker
completion polling. This independent fix was extracted from #58968, which no
longer contains the implementation; it has no dependency on the K3 data path.
Claims
Validation
At pre-review HEAD
512a877511, the new router-outage test fails to resumeregistration; killing the child or ending its input pipe leaves worker
completion polling silent (3 failed).
MoRIIO TP-ACK suites. Tests use real subprocesses and loopback ZeroMQ for
GIL independence, initial registration, router disappearance until the send
queue times out, registration after the router returns, parent death,
repeated shutdown, and propagation of both signal and zero-code child exits.
VLLM_TARGET_DEVICE=cpu PYTHONPATH="$PWD" .venv/bin/python -m pytest -q \ tests/v1/kv_connector/unit/test_moriio_proxy_routing.py \ tests/v1/kv_connector/unit/test_moriio_tp_ack.py .venv/bin/pre-commit run --files \ vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_heartbeat.py \ vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py \ tests/v1/kv_connector/unit/test_moriio_proxy_routing.py \ tests/v1/kv_connector/unit/test_moriio_tp_ack.py .venv/bin/pre-commit run mypy-3.12 --hook-stage manual --files \ vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_heartbeat.py \ vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py \ tests/v1/kv_connector/unit/test_moriio_proxy_routing.py \ tests/v1/kv_connector/unit/test_moriio_tp_ack.pyValidation covers process and worker lifecycle on CPU. No model-output,
accuracy, throughput, or end-to-end model-serving result is claimed for this
isolated control-plane change.
Details
The finite retry limit previously treated ZeroMQ send timeouts during temporary
disconnection as terminal. Those timeouts now retry for the worker lifetime;
other repeated ZMQ errors still terminate the child. The parent previously
never inspected child status. It now checks on each worker completion poll,
including unexpected exit code zero. An idle worker observes a fatal child
exit when it next polls; this does not introduce a background restart manager.
Door: Two-way. Reverting restores the threaded sender without changing the
discovery wire format. Blast radius: MoRIIO workers with a discovery proxy.
AI assistance was used for implementation, tests, extraction and PR preparation.
Pull Request Checklist
/pr-checklistskill.