fix(kanban): claim guard sees open prior runs and every spawn of a run - #946
Conversation
db31dd0 to
e4212cc
Compare
|
🤖 merged-by: apollo · lane: owner · gate: BYPASS: FR paused by Ace 2026-09-22 · why: t_3a06ba8f argus r6 PASS-WITH-CAVEATS @e4212cc: claim guard sees open prior runs and every spawn of a run (reclaim prev_pid=null class) |
e4212cc to
caaf1af
Compare
Round-4 review (t_3a06ba8f) found two prior-owner shapes the second-claim guard could not see: - a ready/review card with a leaked OPEN current_run_id: the scan filtered ended_at IS NOT NULL, then claim_task's invariant recovery closed that run and minted a second worker beside the still-live owner; - a run stamped twice: only the newest spawned event was probed, so a dead later PID vouched for a live earlier one. _prior_worker_still_alive now scans every prior run (open or ended) and _spawned_owner_alive probes every spawned PID in the run's claim interval. _set_worker_pid is the only worker_pid writer and appends `spawned` in the same txn, so the event record covers an open run's row pid. Tests (dispatch boundary, proven-dead controls): leaked open run with a real /bin/sleep owner, review-door variant, double-stamp older/newer alive, and a TTL-expiry NULL-PID hold with a dead-PID requeue control. Verified: new tests RED on 250b521 (3 failed); mutants RED: ended-only scan (2), newest-spawn-only (1), TTL survivor predicate replaced with `termination_attempted and not terminated` (1). Focused kanban suites 249 passed, 1 skipped.
Round-6 of t_3a06ba8f (Argus r5 B-1, B-2). B-2: _spawned_owner_alive treated any live process at a recorded PID as the prior owner, so a recycled PID stranded a ready/review card forever (live via #921; regression vs pre-card base). A live PID now counts as the owner only if its psutil create_time lies in the run's causal window [claimed_at - 1 s, spawned_at + 2 s] (claimed_at falls back to the run's started_at). The window is one-sided around the spawn on purpose: the spawned event is written after Popen and can lag the real worker by up to the DB busy timeout, so a symmetric +/-N window would call genuine workers dead. Unreadable create_time fails CLOSED ("unverified"). Refusals carry owner_identity. Surfacing: new diagnostics rule claim_refused_live_owner fires for READY and REVIEW cards refused for prior_worker_still_alive past a 5 min grace (error, critical at 30 min), clears on the next claimed event. Visible via `hermes kanban diagnostics` and the dashboard diagnostics surface. B-1: test_model_reread_on_retry_spawn returned PID 2 as its "crashed" worker; on Linux that is kthreadd (always alive), so the open-run guard refused the retry and CI slice 14/16 went red. The stub now returns the PID of an exited, reaped child. Verified: second_claim_class + model_override 54 passed; diagnostics rule 7 passed. Old fixture FAILS under Linux PID-2 emulation, new passes. Mutants RED: identity check removed (both recycled arms), symmetric +/-3 s window (lock-lagged arm), fail-open on unreadable ctime, rule unregistered, rule ready-only (review arms). Broad kanban glob on main df43599 + this branch: 1860 passed, 5 failed, 3 skipped; the same 5 (4 review_surfaces, 1 notify artifact) fail on unmodified main 03de88f.
caaf1af to
e00d74e
Compare
|
Rebased onto main 9f857c4 (head e00d74e). Not a pure mechanical rebase, so please read this before merging. While this PR was DIRTY, main landed #956 ("reject recycled PIDs as prior worker owners"). #956 fixes the same B-2 class as this PR's r6 commit, but independently. It adds How the two were merged into one guard:
Verification on e00d74e:
|
|
🤖 merged-by: apollo · lane: kanban-merge-pass · gate: BYPASS: FleetReview paused by Ace 2026-09-22 (state/fleetreview-pause marker present) · why: t_d5113765: reclaim with prev_pid=null, rebased over #956, CI 33 pass; Argus removed from card review (Ace 13:08), CI green, mergeable |
…ecycled PID (t_0ae83825) Rebased onto main after #946/#956/t_21dfa673. The owner window now carries the spawned event's start_token, so termination and liveness paths use the same clock-step-immune identity as the claim guard; legacy rows keep the causal window. enforce_max_runtime passes conn/task_id per main's run-context class guard.
FleetReviewReview: post-merge · head
Post-merge review ( Reviewed with 1 of 2 model families — xai unavailable. profile: light (rule: default light: lines 691<1000000, files 5<1000000, hunks 11<1000000, no hot path) · round 0 · members: B-assert-ctx, B-state, F, G · families: openai Confidence: 1/5 Findings
FleetReview provenance · models: B=gpt-6-sol, D=grok-4.6, F=gpt-6-sol · cost: $1.27 (estimated) · duration: 6m 35s · rounds: 1 · files examined: 5 |
FleetReviewFleetReview's daily member-call budget is spent (60/600 for 2026-09-28 UTC); review skipped. FleetReview · reviewKind: skipped-budget |
Follow-up to #921 (merged at 250b521 while Argus's round-4 review of t_3a06ba8f was still FAIL/HOLD). This carries the round-4 fix, which #921 does not include.
Defects (Argus round 4, t_3a06ba8f)
_prior_worker_still_aliveonly scanned runs withended_at IS NOT NULL. A ready or review card with a leaked opencurrent_run_idpassed the guard.claim_task's invariant recovery then closed that run and started a second worker next to a live owner. Repro:dispatch_oncespawned a second worker while the original/bin/sleepPID was still running.spawnedevent of a run was probed, so a dead newer PID vouched for a live older one._worker_survived_termination(...)withtermination_attempted and not terminatedstayed green.Change
_set_worker_pidis the onlyworker_pidwriter and always appendsspawnedin the same transaction, so an open run's row PID is already covered by the event record. A separate row-PID branch was tried; its mutant survived because the branch was redundant, so it was dropped.Tests
New tests in
test_kanban_second_claim_class.py. They go through the realdispatch_oncespawn boundary and each negative case has a proven-dead control:/bin/sleepowner (no spawn, run stays open), plus a dead-owner control (recovered and dispatched)Verification
probe_gaps.py: ENDED_CONTROL, ACTIVE_UNENDED and TWO_SPAWNS all showsecond_spawned=false.test_kanban_notifyartifact delivery and 4×test_kanban_review_surfaces. They fail the same way on 250b521, so they are not caused by this change.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.