Skip to content

fix(bin): stop false wedge alarms for deliberately stood-down workers - #3281

Open
AgardnerAU wants to merge 49 commits into
kunchenguid:mainfrom
AgardnerAU:fm/fm-stooddown-worker-false-wedge-alarms
Open

AgardnerAU wants to merge 49 commits into
kunchenguid:mainfrom
AgardnerAU:fm/fm-stooddown-worker-false-wedge-alarms

Conversation

@AgardnerAU

@AgardnerAU AgardnerAU commented Aug 29, 2026 •

Copy link
Copy Markdown

Intent

On 2026-09-30 the captain ordered a refresh of #3281 as part of the next wave of work, including the fix for its red contributions test: GitHub reports it as conflicting with upstream main again, so it cannot merge.

Context: the upstream maintainer's latest verdict (2026-09-19, on head 2804abb) was waiting-author: "Attestation: MATCH. CI run 35379109555 FAILURE (Behavior portable serial 7 / tests/fm-contributions.test.sh budget-hang assertion; author reproduced and explained fixture race) ... Prior workflow approval stands; not re-approving a red tip." The maintainer classes the change new-default, so it will not auto-merge even when green. A refresh on 2026-09-27 moved the head to 882ce51; main has since moved again. Upstream closes a PR after 14 days of author silence, and a stale close is irreversible.

The PR is on branch fm/fm-stooddown-worker-false-wedge-alarms from the fork AgardnerAU/firstmate into kunchenguid:main. Reuse that branch and that PR; never replace them.

PR 3281 itself stops false wedge alarms for deliberately stood-down workers: firstmate can record that a ship or scout worker was deliberately stood down (fm-control stand-down / repair-worker-state, a per-task worker-state record owned by bin/fm-worker-state-lib.sh), and the watcher and crew-state stop raising wedge alarms for that task only while its endpoint remains provably worker-free; relaunch and teardown retire the record.

What Changed

  • Added two fm-control verbs, stand-down and repair-worker-state, on tmux only. They are backed by a new per-task state/<id>.worker-state record that bin/fm-worker-state-lib.sh owns. stand-down stops a held ship or scout worker and writes the record only after it proves that the agent is gone. It refuses when an unacknowledged inbox instruction is pending, or when the branch's no-mistakes run cannot be proved quiet. That check uses the new fm_nm_branch_run_verdict in bin/fm-nm-run-lib.sh. repair-worker-state keeps a valid record only while its endpoint is proved dead. In every other case, it clears the record and the task goes back to ordinary supervision.
  • fm-watch skips stale-pane and wedge detection for a task only while it has a stood-down record and its endpoint still has no agent. fm-crew-state reports such a task as parked · source: worker-state, or as unknown if the endpoint is gone. A live run still decides what fm-crew-state reports. fm-send refuses new input to a proved worker-free held task.
  • fm-spawn --relaunch refuses when the record is invalid. Otherwise it clears the record before it arms the replacement, and puts the record back if the relaunch aborts before publication. fm-teardown deletes the record. The docs (agent-control.md, architecture.md, configuration.md, scripts.md, AGENTS.md) now describe these verbs. The tests cover stand-down, repair, relaunch, crew-state, send and watch behaviour. The contribution budget test no longer depends on when its fixture starts.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: This round's only new content is a merge of upstream main (f470a01). It has one textual conflict, in tests/fm-contributions.test.sh, which keeps the branch's hang-fault fix. That fix is now consistent with upstream's atomic clock_bump and with the other callers of the hang fault, and fm_run_timed still group-kills the slow fixture call. The other upstream changes merged here are busy-deferral escalation, the fire-and-forget retry ring, the FM_TASK_INBOX export and the teardown dirt report. Each was traced against the stood-down exemption in fm-watch, fm-send, fm-spawn and fm-teardown and found compatible: a stood-down pane is classified dead, so it takes the unavailable-endpoint or retry-clear arms and never reaches busy deferral. The feature logic was already reviewed in earlier rounds and is unchanged.

Testing

I built a disposable lab home with real Claude Code workers on a private tmux socket. I drove the real fm-control, fm-crew-state, fm-watch, fm-send, fm-spawn relaunch and fm-teardown scripts through each live scenario. The watcher stayed quiet for 45s on a proven stood-down task, which is the main promise of the change. The same watcher reported a stale wake within seconds when a live worker sat behind the record. The adversarial guards held: - A crash is not relabelled as a hold. - Repair never infers a hold from an absent worker. - A vanished endpoint is reported as unknown, not as parked. - fm-send refuses input while the task is held. - A failed relaunch restores the record. The colocated suites for control, relaunch, crew-state and contributions passed, and contributions passed three times in a row. The new watch-triage tests also passed. All nine live scenarios passed. Two scenarios are untested live: the contributions budget-hang test and the in-flight-run stand-down refusal were covered only by colocated test runs. Cleanup was partial: the permission policy refused removal of the lab temp directory, but its tmux server is stopped.

  • Live validation: ✅ go - 9 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Operator runs fm-control stand-down on a live Claude worker: the agent exits to the shell, a 0600 stood-down record bound to the task and endpoint is written, a repeat run is idempotent, and crew-stat… ✅ pass live 01-stand-down-live-claude.txt
Watcher polling a stood-down task whose endpoint stays worker-free emits no stale or wedge wake and creates no pane-hash tracking for 45s ✅ pass live 02-watcher-ignores-stood-down.txt
Adversarial: a live Claude worker restarted behind a stale stood-down record is reported as a contradiction by crew-state and is surfaced as stale by the watcher ✅ pass live 03-live-worker-behind-stale-record.txt
repair-worker-state clears a record that a live agent contradicts, warns, and a repeat run is a no-op ✅ pass live 04-repair-clears-contradicted-record.txt
fm-send refuses an instruction to a stood-down task and creates no inbox entry ✅ pass live 05-send-refused-and-relaunch-retires-record.txt
fm-control relaunch removes the stood-down record and a live worker resumes; a relaunch that aborts before launch restores the record ✅ pass live 06-relaunch-retires-record.txt, 06a-failed-relaunch-restores-record.txt
fm-teardown removes the worker-state record together with the task's other records ✅ pass live 07-teardown-retires-record.txt (the plain git worktree needed a lab-only treehouse return shim; fm-teardown itself was real); 07a shows a refused teardown keeps the record
Adversarial: stand-down on a worker that crashed with no declared hold refuses and writes no record, and repair does not infer a hold; after a paused: status line, stand-down succeeds ✅ pass live 08-crash-is-not-a-stand-down.txt
Adversarial: a stood-down task whose tmux window has vanished is reported as unknown with the lost endpoint named, not as a healthy park ✅ pass live 09-vanished-endpoint-not-healthy-park.txt
The contributions budget-hang test that failed upstream CI passes repeatedly ⏸️ untested no The prior payload did not establish a live result. It records only colocated test runs (bash tests/fm-contributions.test.sh three times, both budget-exhaustion modes passing each time), not a scenar…
Stand-down refuses while the branch owns an in-flight no-mistakes run ⏸️ untested no The prior payload did not establish a live result. It was covered only by bash tests/fm-control.test.sh against a stubbed axi CLI. Driving it live needs a real no-mistakes run in flight on the lab b…
Evidence: Stand-down of a live Claude worker, repeat run, crew-state before and after

Source: Stand-down of a live Claude worker, repeat run, crew-state before and after

$ fm-crew-state.sh held   # live claude worker, no record
state: unknown · source: pane · harness state unavailable (unknown missing)
rc=0
$ fm-control.sh held stand-down
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
stood-down held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
$ cat state/held.worker-state
schema=1
task_id=held
endpoint=fmses:fm-held
state=stood-down
-rw-------@
$ pane after stand-down:
claude
AG-MBP-96% claude
AG-MBP-96%
pane_current_command=zsh
$ fm-crew-state.sh held   # after stand-down
state: parked · source: worker-state · worker deliberately stood down
rc=0
$ fm-control.sh held stand-down   # repeat
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
already-stood-down held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
Evidence: Watcher raises no stale or wedge alarm for a stood-down task (45s)

Source: Watcher raises no stale or wedge alarm for a stood-down task (45s)

# watcher run A (initial status signal drained): only the stood-down task 'held', real dead claude endpoint fmses:fm-held, FM_POLL=2 FM_STALE_ESCALATE_SECS=6, observed 45s
watcher still blocking after 45s: yes
--- stdout (wake lines):
--- wake queue:
--- stale/wedge tracking files for the held pane:
(none)
Evidence: Live worker behind a stale record is still reported as stale

Source: Live worker behind a stale record is still reported as stale

# a live claude started by hand in the held pane; the stood-down record is left in place
pane_current_command=2.1.289
$ cat state/held.worker-state
schema=1
task_id=held
endpoint=fmses:fm-held
state=stood-down
$ fm-crew-state.sh held
state: unknown · source: worker-state · record says stood down but endpoint has a live worker
$ fm-watch.sh (FM_POLL=2 FM_STALE_ESCALATE_SECS=6), up to 90s
watcher still blocking at end: no (exited after ~4s)
--- stdout (wake lines):
check: rearm-resurface
--- wake queue:
--- stale tracking files:
--- second watcher run (first exited on an unrelated check: rearm-resurface after the previous kill; acknowledged):
watcher still blocking at end: no (loop iterations: 4 x 2s)
--- stdout (wake lines):
stale: fmses:fm-held
--- wake queue:
1791184043	3	stale	fmses:fm-held	stale: fmses:fm-held
--- stale tracking files:
.count-fmses_fm-held
.hash-fmses_fm-held
Evidence: repair-worker-state clears a record that a live agent contradicts

Source: repair-worker-state clears a record that a live agent contradicts

# live claude behind the stale record: repair-worker-state
$ fm-control.sh held repair-worker-state
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
warning: task held declared no worker but its endpoint has a live agent; the declaration was cleared and the task is back under ordinary supervision
cleared-live-worker agent-state=alive held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
$ ls state/held.worker-state
ls: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/held.worker-state: No such file or directory
$ fm-control.sh held repair-worker-state  # repeat
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
no-record agent-state=alive held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
Evidence: fm-send refuses input to a stood-down task

Source: fm-send refuses input to a stood-down task

$ fm-control.sh held stand-down
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
stood-down held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
$ fm-send.sh held "please continue"   # task deliberately has no worker
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
error: task held deliberately has no worker; relaunch it before sending an instruction
rc=1
ls: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/held.inbox: No such file or directory
$ fm-control.sh held relaunch
env: timeout: No such file or directory
rc=127
$ ls state/held.worker-state
/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/held.worker-state
pane_current_command=zsh
$ fm-crew-state.sh held
state: parked · source: worker-state · worker deliberately stood down
Evidence: Failed relaunch restores the stood-down record

Source: Failed relaunch restores the stood-down record

$ fm-control.sh held relaunch --note "resume after a deliberate hold"
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
error: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/data/held/brief.md must contain nonempty ## Captain's intent and ## Firstmate spec subsections (or a nonempty legacy # Task body) before spawn
error: the replacement agent for held could not be launched on claude
error: held's agent was stopped but the replacement did not launch; no agent is running, and its work plus the recorded progress note are preserved at /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=1
$ ls state/held.worker-state
/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/held.worker-state
pane_current_command=zsh
$ fm-crew-state.sh held
state: parked · source: worker-state · worker deliberately stood down
Evidence: Successful relaunch removes the record and starts a live worker

Source: Successful relaunch removes the record and starts a live worker

$ fm-control.sh held relaunch --note "resume after a deliberate hold"
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
warning: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/data/held/launch-brief.md records no ship branch; defaulting to legacy branch fm/held
warning: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/data/held/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
relaunched held harness=claude from=claude model=default effort=default backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
$ ls state/held.worker-state
ls: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/held.worker-state: No such file or directory
pane_current_command=2.1.289
$ fm-crew-state.sh held
state: working · source: pane · harness busy (claude-hook)
Evidence: Refused teardown (no treehouse) keeps the record

Source: Refused teardown (no treehouse) keeps the record

$ fm-control.sh held stand-down
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
stood-down held harness=claude backend=tmux endpoint=fmses:fm-held worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held
rc=0
$ ls state/held.*
held.control-relaunch
held.control-relaunch.brief-prior
held.control-relaunch.meta-prior
held.control-relaunch.note
held.git-hooks
held.inbox
held.meta
held.status
held.turn-ended
held.worker-state
$ fm-teardown.sh held
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
teardown: reaping leaked worktree process(es) for held: 59348 72562
teardown: force-killing leaked worktree process(es) for held: 72562
A new version of treehouse is available: v2.3.0 → v3.1.2
Run "treehouse update" to update

worktree /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I/wt-held is not managed by treehouse
error: treehouse return failed for worktree /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held; teardown aborted
rc=1
$ ls state/held.*  (after teardown)
held.control-relaunch
held.control-relaunch.brief-prior
held.control-relaunch.meta-prior
held.control-relaunch.note
held.git-hooks
held.inbox
held.meta
held.status
held.turn-ended
held.worker-state
Evidence: Teardown removes the record

Source: Teardown removes the record

# lab-only: plain git worktree, so a shim `treehouse return` (git worktree remove) stands in for the pool manager; fm-teardown itself is real
$ cat state/held.worker-state  (before)
schema=1
task_id=held
endpoint=fmses:fm-held
state=stood-down
$ fm-teardown.sh held
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I/projects/demo: skipped: no origin remote
teardown held complete (window fmses:fm-held, worktree /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-held)
Backlog: held just finished (this home keeps no markdown backlog at /private/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I/data/backlog.md). Update /private/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I/data/backlog.md - move held to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
rc=0
$ ls state/held.*  (after teardown)
(no held.* records remain)
Evidence: A crashed worker is not relabelled as a hold; a declared pause then allows stand-down

Source: A crashed worker is not relabelled as a hold; a declared pause then allows stand-down

# task 'crash': its real claude worker was SIGKILLed (simulated crash); no stand-down was ever run
pane_current_command=zsh
$ fm-crew-state.sh crash
state: unknown · source: pane · harness state unavailable (unknown missing)
$ fm-control.sh crash stand-down   # must NOT relabel a crash as a deliberate hold
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
error: task crash's agent is already dead without a stand-down record; refusing to relabel a possible worker failure as intentional. Declare the hold first by appending a 'paused: <reason>' line to /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/crash.status, then re-run stand-down
rc=1
$ ls state/crash.worker-state
ls: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/crash.worker-state: No such file or directory
$ fm-control.sh crash repair-worker-state   # must not infer a hold from absence
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
no-record agent-state=dead crash harness=claude backend=tmux endpoint=fmses:fm-crash worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-crash
rc=0
ls: /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/state/crash.worker-state: No such file or directory
--- declare the hold the documented way, then stand-down:
$ echo "paused: waiting on upstream review" >> state/crash.status
$ fm-control.sh crash stand-down
fm-gate-refuse: gate agent lifecycle permitted only against lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I
stood-down crash harness=claude backend=tmux endpoint=fmses:fm-crash worktree=/var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T//fm-lab.lPJs4I/wt-crash
rc=0
schema=1
task_id=crash
endpoint=fmses:fm-crash
state=stood-down
$ fm-crew-state.sh crash
state: parked · source: worker-state · worker deliberately stood down
Evidence: Vanished endpoint reports unknown, not a healthy park

Source: Vanished endpoint reports unknown, not a healthy park

# the stood-down task's tmux window is destroyed (session survives); record still on disk
schema=1
task_id=crash
endpoint=fmses:fm-crash
state=stood-down
$ fm-crew-state.sh crash
state: unknown · source: worker-state · stood down but its endpoint is gone: fmses:fm-crash (relaunch cannot restore it in place)
- Outcome: ⚠️ 1 info across 1 run (22m39s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 info
  • ℹ️ Lab cleanup was only partial. The permission policy refused rm -rf of the disposable lab home /var/folders/41/64hrmnwx11q5lw3l9d7zmfgr0000gn/T/fm-lab.lPJs4I. The lab's tmux server was stopped and its private socket directory was removed by fm-lab-home.sh teardown, so nothing is still running. The leftover directory is only a temp-dir fixture and is outside the worktree. Remove it by hand if you want it gone. The lab's Claude workers also added a folder-trust entry for the lab worktree paths to Claude's own config.
  • Live validation: ✅ go - 9 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Operator runs fm-control stand-down on a live Claude worker: the agent exits to the shell, a 0600 stood-down record bound to the task and endpoint is written, a repeat run is idempotent, and crew-stat… ✅ pass live 01-stand-down-live-claude.txt
Watcher polling a stood-down task whose endpoint stays worker-free emits no stale or wedge wake and creates no pane-hash tracking for 45s ✅ pass live 02-watcher-ignores-stood-down.txt
Adversarial: a live Claude worker restarted behind a stale stood-down record is reported as a contradiction by crew-state and is surfaced as stale by the watcher ✅ pass live 03-live-worker-behind-stale-record.txt
repair-worker-state clears a record that a live agent contradicts, warns, and a repeat run is a no-op ✅ pass live 04-repair-clears-contradicted-record.txt
fm-send refuses an instruction to a stood-down task and creates no inbox entry ✅ pass live 05-send-refused-and-relaunch-retires-record.txt
fm-control relaunch removes the stood-down record and a live worker resumes; a relaunch that aborts before launch restores the record ✅ pass live 06-relaunch-retires-record.txt, 06a-failed-relaunch-restores-record.txt
fm-teardown removes the worker-state record together with the task's other records ✅ pass live 07-teardown-retires-record.txt (the plain git worktree needed a lab-only treehouse return shim; fm-teardown itself was real); 07a shows a refused teardown keeps the record
Adversarial: stand-down on a worker that crashed with no declared hold refuses and writes no record, and repair does not infer a hold; after a paused: status line, stand-down succeeds ✅ pass live 08-crash-is-not-a-stand-down.txt
Adversarial: a stood-down task whose tmux window has vanished is reported as unknown with the lost endpoint named, not as a healthy park ✅ pass live 09-vanished-endpoint-not-healthy-park.txt
The contributions budget-hang test that failed upstream CI passes repeatedly ⏸️ untested no The prior payload did not establish a live result. It records only colocated test runs (bash tests/fm-contributions.test.sh three times, both budget-exhaustion modes passing each time), not a scenar…
Stand-down refuses while the branch owns an in-flight no-mistakes run ⏸️ untested no The prior payload did not establish a live result. It was covered only by bash tests/fm-control.test.sh against a stubbed axi CLI. Driving it live needs a real no-mistakes run in flight on the lab b…
  • Lab setup: bin/fm-lab-home.sh create plus a private tmux socket from bin/fm-lab-home.sh tmux-dir. Each task had a real claude worker in a tmux window, a real git worktree and a state/&lt;id&gt;.meta.
  • fm-control.sh held stand-down on a live Claude worker, then run again; record contents and mode checked; fm-crew-state.sh held run before and after
  • fm-watch.sh (FM_POLL=2, FM_STALE_ESCALATE_SECS=6) observed for 45s with only the stood-down task
  • Live claude restarted by hand behind the stale record, then fm-crew-state.sh and fm-watch.sh run
  • fm-control.sh held repair-worker-state with a live agent, then run again
  • fm-send.sh held &#34;please continue&#34; while the task was stood down
  • fm-control.sh held relaunch --note ...: a failed attempt (invalid brief) and then a successful one
  • fm-teardown.sh held on a stood-down task, using a lab-only treehouse return shim
  • Worker killed with SIGKILL, then fm-control.sh crash stand-down (no declared hold), then repair-worker-state, then a paused: status line, then stand-down
  • Window of a stood-down task killed, then fm-crew-state.sh crash
  • bash tests/fm-contributions.test.sh run three times
  • bash tests/fm-control.test.sh
  • bash tests/fm-control-relaunch.test.sh
  • bash tests/fm-crew-state.test.sh
  • bash tests/fm-watch-triage.test.sh: the two new stood-down tests passed and there were no failures in 94 results
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Adds worker lifecycle control and stand-down state tracking.

The PR does not appear safe to merge until the waiting-worker inbox retry can proceed when an older ordinary instruction remains escalated.

Reviews (10) · Last reviewed commit: "Merge upstream main into stood-down work..."

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Reviewed HEAD d402739060cecca80622526f652da54954ecfc1f vs main. Whole thread read (body, Greptile 5/5, no human reviews). Files: new bin/fm-worker-state-lib.sh, bin/fm-control.sh (stand-down / repair-worker-state), crew-state / nm-run-lib / send / spawn / teardown / watch, docs, tests. No .github/workflows/*. AgardnerAU is not blocked. Not disguised security. Fail-closed: unproven records clear toward supervision; live agents stay in wedge detection; Herdr relaunch remains; stand-down is tmux-only.

Contract-class: new-default. Unconfigured fleet now has first-class stand-down / repair-worker-state verbs, a durable worker-state record, watcher exemption for proven holds, and fm-send refusal to a proven worker-free task. That is a new default lifecycle path, not a restore of a promised existing hold.

VISION (per rule):

  • One captain, one interface — aligns (stops false dead-worker wedges on intentional holds).
  • Authority is explicit — tension (new capability ships as default verbs/docs, not an opt-in flag).
  • Scripts own the mechanics — aligns (deterministic record + classifiers; no agent judgment mixed in).
  • A restart is a non-event — aligns (durable endpoint-bound record; interrupted standing-down stays supervised).
  • Delegation with a spine — aligns (refuses unproven holds; does not cancel in-flight runs).
  • The fleet outlives any vendor — aligns (tmux-only worker-state; Herdr ordinary relaunch kept).
  • Scope — aligns (lifecycle supervision, not merge policy).

Attestation HTML binds head_sha b8a8bd077c4aa9401ccfc56420b2b8c9e071ce5c, not this HEAD. no-mistakes is blocking. Fork CI approved this pass: 33250636532 (CI), 33250636533 (Require no-mistakes). Waiting on the author to rebind attestation. Not with the captain (not otherwise ready). No auto-merge. I am not opening a competing PR.

@AgardnerAU
AgardnerAU force-pushed the fm/fm-stooddown-worker-false-wedge-alarms branch from d402739 to 562d760 Compare August 30, 2026 02:39
@AgardnerAU AgardnerAU changed the title fix(bin): distinguish intentional stand-downs from dead workers feat(bin): represent deliberately stood-down workers Aug 30, 2026
Resolve the run-attribution overlap in favour of main's landed selection.
main's fm_nm_select_run and fm_nm_runs_status_for_worktree now own which
no-mistakes run is attributed to a worktree, including the newest-same-branch
rule this branch had added its own corroboration pass for, so fm-crew-state
reads runs through main's path and this branch keeps only the worker-state
layer on top of it. fm_nm_branch_run_verdict survives as the branch-run
question fm-control stand-down asks.
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: re-triage after newer author activity (HEAD moved since 2026-08-29 stamp).

Outcome: waiting-ci (attestation now MATCH; first-time fork CI re-approved this pass for new HEAD and is pending)

contract-class: new-default — unconfigured fleet gains first-class stand-down / repair-worker-state verbs, durable worker-state records, watcher exemption for proven holds, and fm-send refusal to a proven worker-free task. Still not a restore of a previously specified hold path (prior class stands).

VISION.md

  • One captain, one interface: pass — stops false dead-worker wedges on intentional holds.
  • Authority explicit never inferred: tension — new default lifecycle verbs/docs, not an opt-in flag.
  • Scripts own mechanics: pass — deterministic worker-state record + classifiers.
  • Restart non-event: pass — durable endpoint-bound record; interrupted standing-down stays supervised.
  • Delegation with spine: pass — refuses unproven holds; does not cancel in-flight NM runs.
  • Fleet outlives vendor: pass — tmux-only worker-state; Herdr ordinary relaunch kept.
  • Scope: pass — lifecycle supervision, not merge policy.

Attestation: MATCH 2804abbd59f9f066d10c1e323b0acbe768c6cbd1 (prior stamp had MISMATCH) · CI/NM: pending after approve (35379109555, 35379109799) · Mergeable: MERGEABLE/UNSTABLE · workflow-zero: yes · Firstmate-flag: no (not otherwise ready — defer default-behavior card) · security-fyi: no · 14d-stale: no (author push 2026-09-18)

@AgardnerAU

Copy link
Copy Markdown
Author

CI run 35379109555, job 105717538312, failed on one assertion: budget exhaustion (hang) never started the observation in tests/fm-contributions.test.sh, preceded by grep: .../budget-hang/forge/calls: No such file or directory. The other 17 jobs and the body-compliance run passed.

I could not reproduce it on either side. Against this branch head 2804abbd and against unmodified main at 9bc051ff, in separate clones: 5 sequential runs plus 8 concurrent runs each, so 13 runs per tree and 26 in total, all passing.

On the mechanism: test_budget_exhaustion_keeps_prior_record hang runs the poll under FM_CONTRIBUTIONS_BUDGET=1 against a fixture gh that sleeps 4 seconds. That fixture appends its arguments to $FORGE/calls before sleeping, and the assertion greps that file, so it holds only if the subprocess starts and completes its first write inside the one-second budget. If the bounded call ends it first, the file is never created at all — which matches the symptom here, a missing file rather than wrong contents. The sibling exhaust case, which has no sleep, passed.

One thing you should weigh yourselves: this branch adds about 100 lines of tests to tests/fm-control-relaunch.test.sh, which runs in the same shard (portable serial 7) and took 64 seconds there, in a shard totalling 512 seconds. Added test time cannot introduce the race, since it is inherent to the fixture above, but on a contended shard it can plausibly change the timing that decides it.

Would you be willing to re-run that job? A re-run is the cheapest way to tell a timing-dependent failure from a real regression.

@AgardnerAU

Copy link
Copy Markdown
Author

Follow-up to my previous comment: I can now reproduce that failure deterministically, which turns the earlier "could not reproduce" into a causal explanation.

The fixture gh in test_budget_exhaustion_keeps_prior_record hang appends its arguments to $FORGE/calls on its third line and then sleeps 4 seconds, while the poll runs under FM_CONTRIBUTIONS_BUDGET=1. The assertion therefore holds only if that subprocess reaches its third line inside one second. Injecting a 2-second startup delay ahead of that append — what a contended runner does to process startup — and changing no assertion, produces the reported symptom bit-for-bit:

grep: /.../budget-hang/forge/calls: No such file or directory
not ok - budget exhaustion (hang) never started the observation

A missing file rather than wrong contents, and the same assertion text. That is a pre-existing race in the fixture, independent of this change.

To be straight about what does not support that conclusion: I also ran the suite at 24-way concurrency, and all 24 runs failed — but on different assertions (1 contribution regressions, and watcher did not surface the new contribution signal). None of them produced this job's symptom. That run suggests the suite is timing-sensitive under load generally, but it is not evidence for this specific failure and I am not offering it as such.

Could you re-run job 105717538312? I cannot do it myself — gh run rerun reports that it needs admin rights on this repository. I would rather not push a commit to force a fresh run either: when this branch was pushed at 18:15 both workflow runs went to action_required and waited for a maintainer anyway, and a new push would also move the head that the attestation in the PR body is bound to, which is the mismatch this update set out to fix.

Everything else is green: 18 of 19 checks, including the body-compliance run.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: re-triage after newer author activity (HEAD still 2804abbd…; author diagnosed CI failure).

HEAD 2804abbd59f9f066d10c1e323b0acbe768c6cbd1. MERGEABLE/UNSTABLE vs main 1bb72cc5f880. Author AgardnerAU is not blocked.

Attestation: MATCH. CI run 35379109555 FAILURE (Behavior portable serial 7 / tests/fm-contributions.test.sh budget-hang assertion; author reproduced and explained fixture race). NM body-compliance 35379109799 SUCCESS. Prior workflow approval stands; not re-approving a red tip.

Contract-class: new-default (prior class stands). Unconfigured fleet gains first-class stand-down / repair-worker-state, durable *.worker-state, watcher exemption for proven holds, and fm-send refusal to a proven worker-free task. That is a new default lifecycle surface, not restore of a previously specified hold path. No auto-merge.

VISION.md (per rule)

  • One captain, one interface: aligns — stops false dead-worker wedges on intentional holds.
  • Authority explicit never inferred: tension — new default lifecycle verbs/docs without an opt-in gate (hence new-default, not restore).
  • Scripts own mechanics: aligns — deterministic worker-state lib + control verbs.
  • A restart is a non-event: aligns — durable stand-down record survives session death.
  • Delegation with a spine: aligns — explicit hold vs dead-worker distinction.
  • Fleet outlives vendors: aligns — tmux-only lifecycle; Herdr relaunch preserved.
  • Scope: aligns — command-layer supervision, not workshop.

Outcome: waiting-author (CI red; author owns the contributions budget-hang fix). Firstmate flag: no — not otherwise ready while CI-red. Security FYI: none (fail-closed; live agents stay in wedge detection).

…r-false-wedge-alarms

# Conflicts:
#	bin/fm-control.sh
#	docs/agent-control.md
#	docs/scripts.md
test_budget_exhaustion_keeps_prior_record's `hang` mode faulted the first
forge call of the observation, so the one line proving the observation had
started was written by the very process the budget deadline was about to
kill. With a one-second budget, a contended runner could spend the whole
bound spawning the fixture, leaving $FORGE/calls absent and failing
"budget exhaustion (hang) never started the observation".

Move the hang onto a later call so the calls that prove the observation
started are ordinary completed reads, and raise the budget to five seconds,
which is still `-le 5` and so still classifies a killed read as budget
exhaustion rather than an unavailable forge.

The assertions are unchanged. Record preservation remains the proof that
exhaustion actually happened, so the case cannot pass vacuously: disabling
the hang fault still fails the test.
@AgardnerAU

Copy link
Copy Markdown
Author

Rebased onto current main and ready for CI.

Three conflicts, additive on both sides: main added the endpoint-absence proof shared by exit and relaunch, this branch adds the stand-down and repair-worker-state verbs. Both survive in bin/fm-control.sh, docs/agent-control.md and docs/scripts.md. Nothing this PR promises changed.

One change outside the original scope, flagged because it is a reviewer-visible claim rather than a merge decision. CI failed at "budget exhaustion (hang) never started the observation" in tests/fm-contributions.test.sh, preceded by grep: .../budget-hang/forge/calls: No such file or directory. The fault is in the fixture, not this branch: hang targeted the first forge call, so the line proving the observation had started was written by the very process the one-second budget was about to kill, and a slow process start on a contended runner consumed the whole bound.

The hang now targets a later call, so that evidence comes from completed reads rather than the read being killed, and the budget is five seconds - still -le 5, so a killed read is still budget exhaustion rather than an unavailable forge. No assertion was removed or relaxed; record preservation still proves exhaustion occurred, so the case cannot pass vacuously. Verified by reproducing the failure with an injected fixture startup delay, confirming the fixed test survives a 3-second delay on every fixture call, and confirming it still fails with the hang fault disabled.

I believe the pattern is wider and have left the rest alone: the FORGE_LATENCY=3 case has roughly two seconds of headroom against a five-second bound, and "eight 3-second PR reads complete fresh within one 20-second poll cycle" is timing-sensitive by construction. An earlier run at 24-way concurrency failed on different assertions. Those are for separate follow-up.

@AgardnerAU AgardnerAU changed the title feat(bin): represent deliberately stood-down workers fix(bin): stop false wedge alarms for deliberately stood-down workers Sep 27, 2026
@AgardnerAU

Copy link
Copy Markdown
Author

Updated with current upstream main. The PR is mergeable and ready for CI once the fork workflow runs are approved.

@AgardnerAU

Copy link
Copy Markdown
Author

Refreshed onto current upstream main at 9c15ae1; the two conflicts (operational-home-layout state list, teardown rm list) were additive and both sides are kept. The budget-hang fixture fix from the earlier refresh is carried unchanged (upstream main has not changed that test); tests/fm-contributions.test.sh passed 3/3 runs idle and 5/5 under heavy CPU load locally. Mergeable and ready for CI - the CI and Require no-mistakes runs are awaiting workflow approval.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants