fix(bin): stop false wedge alarms for deliberately stood-down workers - #3281
AgardnerAU wants to merge 49 commits into
Conversation
|
|
Speaking as Kun's firstmate: Reviewed HEAD Contract-class: new-default. Unconfigured fleet now has first-class VISION (per rule):
Attestation HTML binds |
…pair and rollback
…e before wiring retire
… refuse an absent CLI
…ed run projections
d402739 to
562d760
Compare
Resolve the run-attribution overlap in favour of main's landed selection. main's fm_nm_select_run and fm_nm_runs_status_for_worktree now own which no-mistakes run is attributed to a worktree, including the newest-same-branch rule this branch had added its own corroboration pass for, so fm-crew-state reads runs through main's path and this branch keeps only the worker-state layer on top of it. fm_nm_branch_run_verdict survives as the branch-run question fm-control stand-down asks.
|
Speaking as Kun's firstmate: re-triage after newer author activity (HEAD moved since 2026-08-29 stamp). Outcome: waiting-ci (attestation now MATCH; first-time fork CI re-approved this pass for new HEAD and is pending) contract-class: new-default — unconfigured fleet gains first-class VISION.md
Attestation: MATCH |
|
CI run 35379109555, job 105717538312, failed on one assertion: I could not reproduce it on either side. Against this branch head On the mechanism: One thing you should weigh yourselves: this branch adds about 100 lines of tests to Would you be willing to re-run that job? A re-run is the cheapest way to tell a timing-dependent failure from a real regression. |
|
Follow-up to my previous comment: I can now reproduce that failure deterministically, which turns the earlier "could not reproduce" into a causal explanation. The fixture A missing file rather than wrong contents, and the same assertion text. That is a pre-existing race in the fixture, independent of this change. To be straight about what does not support that conclusion: I also ran the suite at 24-way concurrency, and all 24 runs failed — but on different assertions ( Could you re-run job 105717538312? I cannot do it myself — Everything else is green: 18 of 19 checks, including the body-compliance run. |
|
Speaking as Kun's firstmate: re-triage after newer author activity (HEAD still HEAD Attestation: MATCH. CI run Contract-class: new-default (prior class stands). Unconfigured fleet gains first-class VISION.md (per rule)
Outcome: waiting-author (CI red; author owns the contributions budget-hang fix). Firstmate flag: no — not otherwise ready while CI-red. Security FYI: none (fail-closed; live agents stay in wedge detection). |
…r-false-wedge-alarms
…r-false-wedge-alarms # Conflicts: # bin/fm-control.sh # docs/agent-control.md # docs/scripts.md
test_budget_exhaustion_keeps_prior_record's `hang` mode faulted the first forge call of the observation, so the one line proving the observation had started was written by the very process the budget deadline was about to kill. With a one-second budget, a contended runner could spend the whole bound spawning the fixture, leaving $FORGE/calls absent and failing "budget exhaustion (hang) never started the observation". Move the hang onto a later call so the calls that prove the observation started are ordinary completed reads, and raise the budget to five seconds, which is still `-le 5` and so still classifies a killed read as budget exhaustion rather than an unavailable forge. The assertions are unchanged. Record preservation remains the proof that exhaustion actually happened, so the case cannot pass vacuously: disabling the hang fault still fails the test.
|
Rebased onto current Three conflicts, additive on both sides: One change outside the original scope, flagged because it is a reviewer-visible claim rather than a merge decision. CI failed at "budget exhaustion (hang) never started the observation" in The hang now targets a later call, so that evidence comes from completed reads rather than the read being killed, and the budget is five seconds - still I believe the pattern is wider and have left the rest alone: the |
|
Updated with current upstream main. The PR is mergeable and ready for CI once the fork workflow runs are approved. |
|
Refreshed onto current upstream main at 9c15ae1; the two conflicts (operational-home-layout state list, teardown rm list) were additive and both sides are kept. The budget-hang fixture fix from the earlier refresh is carried unchanged (upstream main has not changed that test); tests/fm-contributions.test.sh passed 3/3 runs idle and 5/5 under heavy CPU load locally. Mergeable and ready for CI - the CI and Require no-mistakes runs are awaiting workflow approval. |
Intent
On 2026-09-30 the captain ordered a refresh of #3281 as part of the next wave of work, including the fix for its red contributions test: GitHub reports it as conflicting with upstream main again, so it cannot merge.
Context: the upstream maintainer's latest verdict (2026-09-19, on head 2804abb) was waiting-author: "Attestation: MATCH. CI run 35379109555 FAILURE (Behavior portable serial 7 / tests/fm-contributions.test.sh budget-hang assertion; author reproduced and explained fixture race) ... Prior workflow approval stands; not re-approving a red tip." The maintainer classes the change new-default, so it will not auto-merge even when green. A refresh on 2026-09-27 moved the head to 882ce51; main has since moved again. Upstream closes a PR after 14 days of author silence, and a stale close is irreversible.
The PR is on branch
fm/fm-stooddown-worker-false-wedge-alarmsfrom the forkAgardnerAU/firstmateintokunchenguid:main. Reuse that branch and that PR; never replace them.PR 3281 itself stops false wedge alarms for deliberately stood-down workers: firstmate can record that a ship or scout worker was deliberately stood down (fm-control stand-down / repair-worker-state, a per-task worker-state record owned by bin/fm-worker-state-lib.sh), and the watcher and crew-state stop raising wedge alarms for that task only while its endpoint remains provably worker-free; relaunch and teardown retire the record.
What Changed
fm-controlverbs,stand-downandrepair-worker-state, on tmux only. They are backed by a new per-taskstate/<id>.worker-staterecord thatbin/fm-worker-state-lib.showns.stand-downstops a held ship or scout worker and writes the record only after it proves that the agent is gone. It refuses when an unacknowledged inbox instruction is pending, or when the branch's no-mistakes run cannot be proved quiet. That check uses the newfm_nm_branch_run_verdictinbin/fm-nm-run-lib.sh.repair-worker-statekeeps a valid record only while its endpoint is proved dead. In every other case, it clears the record and the task goes back to ordinary supervision.fm-watchskips stale-pane and wedge detection for a task only while it has a stood-down record and its endpoint still has no agent.fm-crew-statereports such a task asparked · source: worker-state, or asunknownif the endpoint is gone. A live run still decides whatfm-crew-statereports.fm-sendrefuses new input to a proved worker-free held task.fm-spawn --relaunchrefuses when the record is invalid. Otherwise it clears the record before it arms the replacement, and puts the record back if the relaunch aborts before publication.fm-teardowndeletes the record. The docs (agent-control.md,architecture.md,configuration.md,scripts.md,AGENTS.md) now describe these verbs. The tests cover stand-down, repair, relaunch, crew-state, send and watch behaviour. The contribution budget test no longer depends on when its fixture starts.🤖 Generated with Claude Code
Risk Assessment
✅ Low: This round's only new content is a merge of upstream main (f470a01). It has one textual conflict, in tests/fm-contributions.test.sh, which keeps the branch's hang-fault fix. That fix is now consistent with upstream's atomic clock_bump and with the other callers of the
hangfault, and fm_run_timed still group-kills the slow fixture call. The other upstream changes merged here are busy-deferral escalation, the fire-and-forget retry ring, the FM_TASK_INBOX export and the teardown dirt report. Each was traced against the stood-down exemption in fm-watch, fm-send, fm-spawn and fm-teardown and found compatible: a stood-down pane is classified dead, so it takes the unavailable-endpoint or retry-clear arms and never reaches busy deferral. The feature logic was already reviewed in earlier rounds and is unchanged.Testing
I built a disposable lab home with real Claude Code workers on a private tmux socket. I drove the real fm-control, fm-crew-state, fm-watch, fm-send, fm-spawn relaunch and fm-teardown scripts through each live scenario. The watcher stayed quiet for 45s on a proven stood-down task, which is the main promise of the change. The same watcher reported a stale wake within seconds when a live worker sat behind the record. The adversarial guards held: - A crash is not relabelled as a hold. - Repair never infers a hold from an absent worker. - A vanished endpoint is reported as unknown, not as parked. - fm-send refuses input while the task is held. - A failed relaunch restores the record. The colocated suites for control, relaunch, crew-state and contributions passed, and contributions passed three times in a row. The new watch-triage tests also passed. All nine live scenarios passed. Two scenarios are untested live: the contributions budget-hang test and the in-flight-run stand-down refusal were covered only by colocated test runs. Cleanup was partial: the permission policy refused removal of the lab temp directory, but its tmux server is stopped.bash tests/fm-contributions.test.shthree times, both budget-exhaustion modes passing each time), not a scenar…bash tests/fm-control.test.shagainst a stubbed axi CLI. Driving it live needs a real no-mistakes run in flight on the lab b…Evidence: Stand-down of a live Claude worker, repeat run, crew-state before and after
Source: Stand-down of a live Claude worker, repeat run, crew-state before and after
Evidence: Watcher raises no stale or wedge alarm for a stood-down task (45s)
Source: Watcher raises no stale or wedge alarm for a stood-down task (45s)
Evidence: Live worker behind a stale record is still reported as stale
Source: Live worker behind a stale record is still reported as stale
Evidence: repair-worker-state clears a record that a live agent contradicts
Source: repair-worker-state clears a record that a live agent contradicts
Evidence: fm-send refuses input to a stood-down task
Source: fm-send refuses input to a stood-down task
Evidence: Failed relaunch restores the stood-down record
Source: Failed relaunch restores the stood-down record
Evidence: Successful relaunch removes the record and starts a live worker
Source: Successful relaunch removes the record and starts a live worker
Evidence: Refused teardown (no treehouse) keeps the record
Source: Refused teardown (no treehouse) keeps the record
Evidence: Teardown removes the record
Source: Teardown removes the record
Evidence: A crashed worker is not relabelled as a hold; a declared pause then allows stand-down
Source: A crashed worker is not relabelled as a hold; a declared pause then allows stand-down
Evidence: Vanished endpoint reports unknown, not a healthy park
Source: Vanished endpoint reports unknown, not a healthy park
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
rm -rfof the disposable lab home/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I. The lab's tmux server was stopped and its private socket directory was removed byfm-lab-home.sh teardown, so nothing is still running. The leftover directory is only a temp-dir fixture and is outside the worktree. Remove it by hand if you want it gone. The lab's Claude workers also added a folder-trust entry for the lab worktree paths to Claude's own config.bash tests/fm-contributions.test.shthree times, both budget-exhaustion modes passing each time), not a scenar…bash tests/fm-control.test.shagainst a stubbed axi CLI. Driving it live needs a real no-mistakes run in flight on the lab b…Lab setup:bin/fm-lab-home.sh createplus a private tmux socket frombin/fm-lab-home.sh tmux-dir. Each task had a realclaudeworker in a tmux window, a real git worktree and astate/<id>.meta.fm-control.sh held stand-downon a live Claude worker, then run again; record contents and mode checked;fm-crew-state.sh heldrun before and afterfm-watch.sh(FM_POLL=2, FM_STALE_ESCALATE_SECS=6) observed for 45s with only the stood-down taskLiveclauderestarted by hand behind the stale record, thenfm-crew-state.shandfm-watch.shrunfm-control.sh held repair-worker-statewith a live agent, then run againfm-send.sh held "please continue"while the task was stood downfm-control.sh held relaunch --note ...: a failed attempt (invalid brief) and then a successful onefm-teardown.sh heldon a stood-down task, using a lab-onlytreehouse returnshimWorker killed with SIGKILL, thenfm-control.sh crash stand-down(no declared hold), thenrepair-worker-state, then apaused:status line, thenstand-downWindow of a stood-down task killed, thenfm-crew-state.sh crashbash tests/fm-contributions.test.shrun three timesbash tests/fm-control.test.shbash tests/fm-control-relaunch.test.shbash tests/fm-crew-state.test.shbash tests/fm-watch-triage.test.sh: the two new stood-down tests passed and there were no failures in 94 results✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.