Repository navigation
Conversation
On the no-worker-PID branch, probe the local claim owner. Only a dead owner authorizes release; hold a live or unprovable launch through manual and TTL reclaim to prevent duplicate workers. Preserve foreign-host policy and worker-fingerprint handling.
|
🤖 Hermes Agent automated review: PR reviewed. Diff analyzed (189 lines changed). Check CI status and manual review recommended. |
SummaryReady to merge. Defends against the pid-less-claim race condition. Positives
Suggestions
Questions
|
Popen precedes the pid stamp, so a dead claimer can leave a detached, heartbeating worker. Current-run heartbeat/spawned evidence now keeps the claim held on the manual, TTL and stale paths; prior-run evidence is ignored. Real-subprocess orphan regression is RED on 739682d.
|
Round 2 (addresses review B1: dead claimer != no worker spawned). The no-pid branch now checks for heartbeat/spawned events bound to the CURRENT run before it probes the claimer. Any such evidence keeps the claim liveness_unprovable, so it stays held and no second worker is spawned. Evidence from a prior run is ignored, so a never-spawned retry still releases. Class sweep over every caller of _terminate_reclaimed_worker / _worker_survived_termination:
Tests: the real-subprocess orphan arm is RED on the previous head (3 failed) and GREEN on this head. The incident, alive-claimer and foreign-host arms are still green. Adjacent suites are green. |
… context A dead claimer does not prove no worker exists (Popen precedes the pid stamp). With no current-run heartbeat/spawned event, hold the pid-less claim until DEAD_CLAIMER_LAUNCH_BOUND_SECONDS (900 s) after the claim; with such evidence, release once it is older than DEFAULT_CLAIM_HEARTBEAT_MAX_STALE_SECONDS so a worker that heartbeated and died cannot hold the card forever. Omitting conn/task_id now fails closed.
|
Update (09bdbd5): a dead claimer is no longer treated as proof that no worker exists, because Popen runs before the pid stamp. A pid-less claim is now held for 900 s after the claim when the current run has no heartbeat or spawned event. When it does have one, the claim is held until the newest event is older than DEFAULT_CLAIM_HEARTBEAT_MAX_STALE_SECONDS. Omitting conn/task_id now fails closed. Note, pre-existing and not changed here: the dashboard _set_status_direct upstream terminates the worker after commit, and nothing gates the release on survival. |
Problem
A dispatcher can claim a card and die before stamping its worker PID. A missing PID by itself is ambiguous: the claimer may still be launching a worker. Releasing its claim risks a duplicate; repeatedly deferring an already-dead claimer wedges the card.
Change
Probe the host-local claimer PID in
claim_lockwhen no worker PID was stamped.psutil.Process(pid)is the Windows-safe equivalent ofos.kill(pid, 0);NoSuchProcessproves the claimer is gone and lets the normal reclaim recordclaimer_pid_dead. A live, inaccessible, or malformed claimer holds its claim. This applies to both TTL and manual reclaim; foreign-host release policy and existing worker PID fingerprint semantics are unchanged. The upstream tree split the dispatcher intokanban_db_dispatch.py, so this ports the fork-first fix into that shape.Fork-first PR: ANG-Ventures#983
Tests
New manual/TTL dead- and live-claimer regressions failed against upstream main before implementation (4 failures, one foreign-host control passed). Targeted upstream run: 60 passed, 1 skipped, 1 deselected. The deselected
test_infrastructure_spawn_refusal_never_charges_the_cardfails identically on clean upstream/main on this macOS host (verified independently); it is unrelated to this diff. Ruff andgit diff --checkpass.