Conversation
…dicting refusals A task whose recorded endpoint was destroyed rather than stopped could only be recovered by hand-editing its record: fm-spawn --relaunch sent the operator to fm-control exit, and exit refused a tmux `missing` endpoint with no command that could act on it. Herdr already had a reclaim path; tmux was left deadlocked because a task record carries no tmux socket identity. - The control plane's one absence proof (fm_control_endpoint_absence_verdict) now proves a tmux endpoint gone by what a duplicate launch would collide with: no process of this user that classifies as a verified harness is working in the recorded worktree (new fm_agent_process_worktree_scan, /proc on Linux, lsof and ps elsewhere). An agent found there, or an unreadable process table, still refuses, and ambiguous or unreadable endpoint states refuse as before. - fm-spawn --relaunch rebinds a proven-gone tmux task to one fresh window in the same worktree under the same task identity, closing that window again if the record is not republished. - Refusals now name a command that works for the state they read: `alive` points at fm-control relaunch, `ambiguous`/`unreadable` at fm-peek and a retry, an unproven absence at the relaunch that reclaims it once that changes, and the relaunch rollback messages say to rerun relaunch. - A no-mistakes ship's relaunch progress note tells the replacement to reattach to a validation run already under way rather than start or abort one; the control plane itself never touches the run. Evidence (tests/fm-control-relaunch.test.sh, same selected tests on the base commit 30ef650 vs this branch): before: not ok - exit must report a proven-gone tmux endpoint rather than refuse it before: not ok - the refusal should name the one command that replaces a live agent (missing: 'fm-control.sh rl15 relaunch') before: not ok - the refusal must not send the operator to a verb that refuses this state too (unexpected: 'fm-control.sh rl63 exit') after: ok - tmux: a destroyed window is reclaimed in the same worktree under the same task after: ok - fm-spawn --relaunch: refuses a live endpoint and names the command that replaces its agent after: ok - reclaim: an unclassifiable endpoint is still refused, so two agents cannot share one
…0 strip; drop relaunch note
…free tmux husk window
|
| WID=$(fm_backend_tmux_create_task "$SES" "$W" "$WT") || exit 1 | ||
| T="$SES:$W" | ||
| WT_TARGET=$WID | ||
| TMUX_REBIND_ABORT_TARGET=$T |
There was a problem hiding this comment.
Abort targets a shared name If two homes reclaim the same task ID on one tmux server at the same time, both can pass the window check and create
fm-<id> windows. This line saves the shared name for abort cleanup instead of the stable window ID returned by creation. If one reclaim then fails before publishing its record, cleanup can close the other home's window or leave its own behind, and later commands cannot reliably distinguish the two. Bind cleanup to the window this attempt created.
Knowledge Base Used: Backend adapters
| || fail "the reclaim must open exactly one window ($what)" | ||
| window=$(meta_field "$dir" "$id" window) | ||
| case "$window" in | ||
| *":fm-$id") ;; |
There was a problem hiding this comment.
Rebound session goes untested The test checks only the
:fm-<id> suffix of window=, although reclaim is meant to bind the task to the session this seat selects. These test commands unset TMUX, so that session should be firstmate. A regression that kept the old session would pass this assertion, even when that session no longer exists and reclaim cannot complete. Assert the full session:window value.
| if [ -f "$dir/fake/nm-calls" ] && grep -Eq '(^|[[:space:]])(run|respond|abort|sync|cancel|rerun)([[:space:]]|$)' "$dir/fake/nm-calls"; then | ||
| fail "the reclaim must never abort, answer, or start the task's validation run ($what): $(cat "$dir/fake/nm-calls")" | ||
| fi |
There was a problem hiding this comment.
Validation tripwire misses actions The new recorder shadows only
no-mistakes, and this check rejects only six named verbs. A future reclaim call through the axi entrypoint, or an action outside that list, would still pass the test's claim that reclaim never acts on a parked validation run. Capture the entrypoints and actions the assertion intends to protect, or narrow its stated guarantee.
|
Superseded by #6007, which carries the same change rebased onto current main at b070249. This branch's head 642034d is the pre-rebase lineage and could not be updated without a force-push, so the work was republished on a new branch instead. Closing this in favour of #6007; the branch here is left intact and nothing is discarded. |
Intent
Firstmate's control plane deadlocks when a task's recorded endpoint has been destroyed rather than merely stopped, and it cost manual record surgery tonight on a live task.
Observed on task fm-teardown-shared-slot-retire, a claude ship lane whose herdr pane was destroyed while its no-mistakes run sat parked at a review gate. Its work was safe: the branch was committed in the recorded worktree and the parked run was still held. Recovery was nonetheless impossible through the supported commands:
bin/fm-control.sh <id> relaunchlaunched and then could not confirm a running agent, reporting the work preserved.bin/fm-spawn.sh <id> --relaunchrefused: "endpoint reads 'missing'; a relaunch requires a positively agent-free endpoint (stop the agent first with bin/fm-control.sh exit)".bin/fm-control.sh <id> exitrefused: "recorded endpoint is gone, so there is no agent to stop; reconcile the task before any further control action" (bin/fm-control.sh:461).So relaunch sends the operator to exit, exit sends the operator to reconcile, and no command performs that reconcile. The operator's only route was to create a pane by hand, launch the harness by hand, and rewrite the window and herdr fields in state/.meta with sed - which is exactly the unguarded record editing the control plane exists to prevent.
A related symptom in the same area, useful as evidence but NOT this task's scope: when herdr reports a pane's agent binding as lost while the agent is in fact alive and advancing, bin/fm-send.sh declines to ring the doorbell and reports that the agent has exited. That is the opposite error - a live agent read as gone - and the two together suggest the agent-state classification is being trusted for decisions it cannot support.
What Changed
fm_agent_process_worktree_scan(bin/fm-agent-process-lib.sh), a/proc-based scan for a verified harness process working in a task's recorded worktree, and wired it intofm_control_endpoint_absence_verdictso a tmuxmissingendpoint can now resolve togone(no agent in the worktree) instead of always returningunproven; an agent found there, an unreadable process table, or a host without/procstill refuses.bin/fm-spawn.sh --relaunchnow rebinds a proven-gone tmux task: it ensures the container session, creates one fresh window in the recorded worktree on the server this seat addresses, republisheswindow=, and closes that window from the abort trap if anything refuses before the record names it.bin/backends/tmux.sh'screate_taskreads the window inventory three-way and refuses both an existing same-named window and an unreadable inventory rather than inferring absence from an empty stream.relaunch→exit→ "reconcile" with a sharedfm_control_unclassified_next_stephelper, soexit,relaunch, and the launch owner each name a runnable next command for the verb actually invoked; relaunch rollback messages now name the retry command. Tests cover the new scan, the tmux reclaim path and the refusals (tests/lib.shgainsfm_proc_scan_availableandfm_wait_for_agent_argv0), anddocs/agent-control.md,docs/scripts.md, and the stuck-crewmate-recovery skill drop the "no tmux reclaim" claim.Risk Assessment
Testing
I drove the whole endpoint-gone reclaim against the real product in a disposable lab home with a real private tmux server, a real treehouse worktree and a real Claude harness, never a stub. The deadlock the intent describes reproduced live against the base commit - relaunch, fm-spawn --relaunch and exit each refused and pointed at one another - and on this branch the same state gives exit an endpoint-gone verdict and relaunch a working reclaim that opens one fresh window, rebinds the record, and brings a real agent back up in the same worktree while the commit, the never-committed file and the parked-gate status line all survive. The adversarial probes all held: a real Claude still working in the recorded worktree blocked both verbs with the refusal naming its PID and left it alone, a foreign window already holding the task's name was refused and stayed open, and every refusal named a command that actually ran when I ran it verbatim. The repository's targeted suites, including the real-tmux backend smoke test, pass as the baseline. This product surface is a terminal control plane rather than a graphical one, so the reviewer-visible artifact is the captured live tmux pane showing the reclaimed agent plus the CLI transcripts; there is no rendered web or GUI surface to screenshot.
endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 ..., rc=0bin/fm-control.sh labscout1 relaunch --note "<why>"; running it succeeded (rc=0) and reused the surviving window @2 rather th…spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 worktree=...rc=0, new window @3, recorded worktree keptbash tests/fm-backend-tmux-smoke.test.shagainst a real tmux server on a private socket: 'refuses an unreadable window inventory rather than assuming the name is free' and 'refuses a duplicate witho…Evidence: Live reclaim transcript (base-commit deadlock, reclaim, adversarial probes)
Source: Live reclaim transcript (base-commit deadlock, reclaim, adversarial probes)
Evidence: Reclaimed tmux window pane - the real Claude agent running in the reclaimed endpoint
Source: Reclaimed tmux window pane - the real Claude agent running in the reclaimed endpoint
Evidence: Final lab state: rebound record, unchanged HEAD, surviving uncommitted file and parked status line
Source: Final lab state: rebound record, unchanged HEAD, surviving uncommitted file and parked status line
Evidence: Base-commit deadlock, reproduced live on the same lab task
Evidence: Same state on this branch: exit reports it, relaunch reclaims it
Evidence: Adversarial: a real Claude still working in the worktree blocks both verbs
Evidence: Adversarial: a foreign same-named window is refused, never closed
Pipeline
Updates from git push no-mistakes
... (17 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:
tests/fm-control-relaunch.test.sh:1861-make_nm_recorder(tests/fm-control-relaunch.test.sh:1859-1868) stages a fakeno-mistakeson PATH andassert_tmux_reclaimsthen grepsnm-callsforrun|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI -fm-control.sh,fm-spawn.sh,fm-control-lib.sh, andbin/backends/tmux.shcontain no call tono-mistakesor anaxirun verb - sonm-callsis never created and the guard cannot fail for the right reason. It also shadows only theno-mistakesname, not theaxientrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gatesurvives), theHEADcomparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add theaxiname, or drop it, since the surrounding assertions carry the requirement.bin/fm-control.sh:1063-interruptis the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint readsambiguousmakesbin/fm-control.sh <id> interruptdie with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, whileexit(bin/fm-control.sh:605) andrelaunch(bin/fm-spawn.sh:1756) through the same state now end withfm_control_unclassified_next_steppointing atbin/fm-peek.sh <id>and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it becauseinterruptreads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.bin/fm-agent-process-lib.sh:147- The branch's whole deliverable - the tmux reclaim - is unavailable on macOS, so on the default backend the exact dead-end the intent reports survives unchanged. Concrete sequence on a macOS seat: a tmux task's window is destroyed;fm_backend_tmux_agent_statereadsmissing;fm_control_endpoint_absence_verdict(bin/fm-control-lib.sh:353) callsfm_agent_process_worktree_scan, whose second guard[ ! -r /proc/self/cmdline ] || [ ! -L /proc/self/cwd ](bin/fm-agent-process-lib.sh:146-150) fires because macOS has no /proc; the verdict isunproven;bin/fm-control.sh <id> exitdies at :602 andbin/fm-spawn.sh <id> --relaunchdies at :1755 - and the refusal text itself ends "once that server is gone there is no such seat, and no route". That is verbatim the state the intent calls unacceptable: "So relaunch sends the operator to exit, exit sends the operator to reconcile, and no command performs that reconcile. The operator's only route was to create a pane by hand, launch the harness by hand, and rewrite the window and herdr fields in state/<id>.meta with sed - which is exactly the unguarded record editing the control plane exists to prevent." macOS is a first-class platform in this repo, not a hypothetical: bin/fm-backend.sh:175 gates cmux fallbacks onuname = Darwin, bin/fm-install-treehouse.sh:39 ships Darwin-arm64/Darwin-x86_64 builds, and bin/fm-agent-process-lib.sh:81 documents macOScommsemantics for this very classifier. tmux is the default/reference backend (bin/backends/tmux.sh:4, andbackend=absent means tmux), so this is not a niche path. The change also carries no automated evidence there: every new case skips viafm_proc_scan_available(tests/lib.sh:504, tests/fm-control-relaunch.test.sh:1902, :1911, :1920, :1929, :1944, :2035, tests/fm-control.test.sh:783), so on macOS the reclaim is simultaneously unavailable and unexercised. This needs the author's decision, not a repair: the remedy is a second, non-/proc absence proof, which EXTENDS the change, and a prior fix round in this same run (6f88055 "drop lsof scan fallback, refuse absence proof without /proc", after 8e58903 and 50f7b2c spent two rounds on lsof parsing bugs) deliberately removed exactly that. Deciding Linux-only is defensible - Herdr reclaims everywhere and the refusal is honest rather than misleading - but the intent marks the reconcile route as required and does not authorize the gap, so say explicitly whether tmux-on-macOS is out of scope, and if so scope docs/agent-control.md:110-115 and the capability matrix at :186-192 to say the tmux reclaim is Linux-only.docs/agent-control.md:113- Three surfaces describe the tmuxgoneverdict as proving more than the code establishes; this is the same family as the placement overclaim the author corrected in round 8 (998afee), left behind at these sibling sites. (1) docs/agent-control.md:113 - "An agent found there refuses and names its process id, and so does any process whose working directory could not be read" overstates the guard: bin/fm-agent-process-lib.sh:175 runs[ "$(fm_agent_process_classify ...)" = agent ] || continueBEFORE the unreadable-cwd refusal at :176-179, so a non-agent process with an unreadable cwd is silently skipped, never refused. The correct sentence is "any agent process whose working directory could not be read". (2) bin/fm-control-lib.sh:317 - "gone - absence is PROVEN: no agent, and no endpoint at the address the record names" contradicts its own next clause two lines down ("A window may still survive elsewhere - on a tmux server this seat cannot address"): a record's tmux address is a baresession:windowwith no socket component, so the identical address can name a live window on another server, and on tmux the scan proves only the agent absent, never the endpoint. (3) The same overclaim rides along at bin/fm-control.sh:582 ("nothing survived at the recorded address for this verb to preserve") and its user-facing copy in theexitrow at docs/agent-control.md:36. All four are wording; no behavior is wrong. Scope each to what tmux actually proves - no agent process working in the recorded worktree - and leave the herdr phrasing, which does prove the endpoint absent, as it is.bin/backends/tmux.sh:105- The change promotesfm_backend_tmux_create_task's duplicate-window refusal to the reclaim's safety boundary - the new comment at bin/backends/tmux.sh:83-92 says closing an unrecorded window "would destroy another home's deliberately preserved endpoint", and docs/agent-control.md:118 states the reclaim refuses rather than replacing such a window - but the guard itself fails open on a read it could not make.tmux list-windows -t "$ses" -F '#{window_name}' | grep -qx "$wname"(bin/backends/tmux.sh:105) cannot distinguish "the session has no such window" from "the read failed": grep sees an empty stream either way and returns 1, so the create proceeds. Sequence: a reclaim runs from a seat whose ambient session already holds another home'sfm-<id>window;tmux list-windowsfails transiently;tmux new-windowat :109 succeeds and tmux permits the duplicate name; bin/fm-spawn.sh:3471 persists the NAME formT="$SES:$W"as the record'swindow=, and every later name-targeted read and write - the env exports at bin/fm-spawn.sh:5112-5132, the launch itself at :5227, andspawn_send_key "$T" Enterat :5234 - resolves to whichever window tmux matches first, which can be the foreign one the refusal exists to protect. The sibling in this same file already does this correctly:fm_backend_tmux_killcallsfm_backend_tmux_window_inventoryand branches on its three-way status (bin/backends/tmux.sh:198-206), refusing on status 1 precisely so an unreadable read is never mistaken for absence. Classifying this as ask-user rather than auto-fix because the remedy changes a shared helper's contract for the ordinary spawn path at bin/fm-spawn.sh:3557 too, which is wider than this change's scope; the pre-existing helper is unchanged here, only the new reliance on it is.🔧 Fix applied.
4 issues (1 warning, 3 infos) still open:
tests/fm-control-relaunch.test.sh:1861-make_nm_recorder(tests/fm-control-relaunch.test.sh:1859-1868) stages a fakeno-mistakeson PATH andassert_tmux_reclaimsthen grepsnm-callsforrun|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI -fm-control.sh,fm-spawn.sh,fm-control-lib.sh, andbin/backends/tmux.shcontain no call tono-mistakesor anaxirun verb - sonm-callsis never created and the guard cannot fail for the right reason. It also shadows only theno-mistakesname, not theaxientrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gatesurvives), theHEADcomparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add theaxiname, or drop it, since the surrounding assertions carry the requirement.bin/fm-control.sh:1063-interruptis the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint readsambiguousmakesbin/fm-control.sh <id> interruptdie with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, whileexit(bin/fm-control.sh:605) andrelaunch(bin/fm-spawn.sh:1756) through the same state now end withfm_control_unclassified_next_steppointing atbin/fm-peek.sh <id>and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it becauseinterruptreads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.bin/fm-spawn.sh:1752- Every refusal and rollback message this change added namesbin/fm-control.sh <id> relaunchas the next step, but that verb dies immediately for a ship or scout task unless--note/--note-fileis also passed (bin/fm-control.sh:979:[ "$NOTE_SET" = 1 ] && [ -n "$NOTE" ] || die "relaunch of a $KIND task requires --note..."). An operator who copies the printed command verbatim gets a second refusal, which is the same cross-pointing this branch set out to end (commit 9350468 "stop refusals promising routes the state cannot offer"). The change's own test admits it: tests/fm-control-relaunch.test.sh:1712 comments "The named command must actually work for this state, rather than send the operator on to another refusal", then at :1713 runsrun_control "$dir" rl15 relaunch --note "replace the live agent"- adding an argument the refusal never mentions - on a task created byadd_ship_task(:1704). So the test provesrelaunch --noteworks, not that the command as printed works. docs/agent-control.md:180 states the stronger claim outright: "Every such refusal names a command that acts on the state it read: analiveendpoint points atbin/fm-control.sh <id> relaunch, which stops that agent first". Every site in the changed code where the same invariant must hold: bin/fm-spawn.sh:1752 (newalivebranch, primary); bin/fm-spawn.sh:1743 ("before bin/fm-control.sh $ID relaunch can reclaim the task" on an unprovable absence); bin/fm-spawn.sh:1756 via bin/fm-control-lib.sh:402 and :405 ("retry bin/fm-control.sh %s %s" with verb=relaunch); bin/fm-control.sh:604 ("before bin/fm-control.sh $ID $VERB can act", where $VERB is relaunch inside the relaunch transaction); and the three rollback lines this change added at bin/fm-control.sh:746, :775, :778 ("rerun bin/fm-control.sh $ID relaunch") - the rollback case is the most reachable, since a failed replacement launch is exactly when an operator reruns the printed command. Classified ask-user because the remedy is a product-copy decision, not a mechanical repair: fm-control-lib.sh's shared helper has no access to KIND, so either the printed command grows a--note "<why>"placeholder unconditionally, or the two callers that do know KIND (bin/fm-spawn.sh, bin/fm-control.sh) append it and the helper gains a parameter. Pick one; do not harden further..agents/skills/stuck-crewmate-recovery/SKILL.md:43- Round 9's accepted instruction scoped the tmux reclaim's Linux-only caveat into docs/agent-control.md, and that landed correctly (docs/agent-control.md:113 "the tmux reclaim is Linux-only", :114 "On macOS, or any other host without/proc, a tmux task whose endpoint was destroyed cannot be reclaimed at all", plus the capability-matrix paragraph at :195-197). This sibling agent-facing surface was left behind: the rewritten sentence now tells every agent reading it that "An endpoint that is not merely idle but destroyed - a Herdr pane or workspace removed in churn, a tmux window or server that no longer exists - is recovered by that samebin/fm-control.sh <id> relaunch... nothing special is needed", with no platform qualifier. On a macOS seat that is false for the tmux half: bin/fm-agent-process-lib.sh:146-150 refuses withunreadablebecause there is no /proc, so both verbs refuse. The following sentence ("Do not work around a refusal by respawning or by editing the task record") keeps the agent from doing damage, so this is guidance drift rather than a safety hole. Wording only - the smallest honest remedy is one clause naming that the tmux reclaim needs /proc and the Herdr reclaim does not, matching what docs/agent-control.md already says; add no component.🔧 Fix applied.
4 issues (1 warning, 3 infos) still open:
tests/fm-control-relaunch.test.sh:1861-make_nm_recorder(tests/fm-control-relaunch.test.sh:1859-1868) stages a fakeno-mistakeson PATH andassert_tmux_reclaimsthen grepsnm-callsforrun|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI -fm-control.sh,fm-spawn.sh,fm-control-lib.sh, andbin/backends/tmux.shcontain no call tono-mistakesor anaxirun verb - sonm-callsis never created and the guard cannot fail for the right reason. It also shadows only theno-mistakesname, not theaxientrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gatesurvives), theHEADcomparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add theaxiname, or drop it, since the surrounding assertions carry the requirement.bin/fm-control.sh:1063-interruptis the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint readsambiguousmakesbin/fm-control.sh <id> interruptdie with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, whileexit(bin/fm-control.sh:605) andrelaunch(bin/fm-spawn.sh:1756) through the same state now end withfm_control_unclassified_next_steppointing atbin/fm-peek.sh <id>and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it becauseinterruptreads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.bin/fm-spawn.sh:1752- Every refusal and rollback message this change added namesbin/fm-control.sh <id> relaunchas the next step, but that verb dies immediately for a ship or scout task unless--note/--note-fileis also passed (bin/fm-control.sh:979:[ "$NOTE_SET" = 1 ] && [ -n "$NOTE" ] || die "relaunch of a $KIND task requires --note..."). An operator who copies the printed command verbatim gets a second refusal, which is the same cross-pointing this branch set out to end (commit 9350468 "stop refusals promising routes the state cannot offer"). The change's own test admits it: tests/fm-control-relaunch.test.sh:1712 comments "The named command must actually work for this state, rather than send the operator on to another refusal", then at :1713 runsrun_control "$dir" rl15 relaunch --note "replace the live agent"- adding an argument the refusal never mentions - on a task created byadd_ship_task(:1704). So the test provesrelaunch --noteworks, not that the command as printed works. docs/agent-control.md:180 states the stronger claim outright: "Every such refusal names a command that acts on the state it read: analiveendpoint points atbin/fm-control.sh <id> relaunch, which stops that agent first". Every site in the changed code where the same invariant must hold: bin/fm-spawn.sh:1752 (newalivebranch, primary); bin/fm-spawn.sh:1743 ("before bin/fm-control.sh $ID relaunch can reclaim the task" on an unprovable absence); bin/fm-spawn.sh:1756 via bin/fm-control-lib.sh:402 and :405 ("retry bin/fm-control.sh %s %s" with verb=relaunch); bin/fm-control.sh:604 ("before bin/fm-control.sh $ID $VERB can act", where $VERB is relaunch inside the relaunch transaction); and the three rollback lines this change added at bin/fm-control.sh:746, :775, :778 ("rerun bin/fm-control.sh $ID relaunch") - the rollback case is the most reachable, since a failed replacement launch is exactly when an operator reruns the printed command. Classified ask-user because the remedy is a product-copy decision, not a mechanical repair: fm-control-lib.sh's shared helper has no access to KIND, so either the printed command grows a--note "<why>"placeholder unconditionally, or the two callers that do know KIND (bin/fm-spawn.sh, bin/fm-control.sh) append it and the helper gains a parameter. Pick one; do not harden further.bin/fm-spawn.sh:1743- On a host without /proc the tmux absence proof returnsunreadablewith a reason that ends "...once that server is gone there is no such seat, and no route" (bin/fm-agent-process-lib.sh:147-149). The parenthetical this branch appended then wraps that reason with "that reason is what would have to change before bin/fm-control.sh <id> relaunch --note &fix(watcher): make check wakes lossless via watcher-side suppression #34;<why>&fix(watcher): make check wakes lossless via watcher-side suppression #34; can reclaim the task", so one message tells a macOS operator both "no route" and "relaunch reclaims it once the reason changes" - and the reason is the absent /proc, which cannot change on that seat. docs/agent-control.md:113-114 names macOS as a real seat, and docs/agent-control.md:178 states the branch's own invariant ("Every such refusal names a command that acts on the state it read, spelled as the operator can run it"), which this one does not satisfy there: relaunch on that host refuses for the identical reason, forever. Concrete sequence: on macOS, a tmux ship task whose window and server were destroyed -bin/fm-spawn.sh <id> --relaunchprints this message; the operator runs the namedbin/fm-control.sh <id> relaunch --note "<why>"and gets the same /proc refusal through bin/fm-control.sh:607. Sibling sites where the same wrapper is printed over the same reason: bin/fm-control.sh:607 (exit, via $VERB_RETRY). Both sentences were added by this branch's fix rounds (9350468, re-worded in 3d3ac1f), not by the author's original commit. Mitigating: the reason text itself does name the only real route (a seat that still addresses the recorded tmux server), so the operator is not left with nothing - which is why this is informational rather than a blocker. Remedy is message copy only, so it needs the author's call rather than a mechanical fix: either scope the "what would have to change" clause to the proofs that can change (an agent still in the worktree, an unreadable process table) and omit it when the reason is a missing /proc, or drop the clause and let the reason stand alone. Add no component and no conditional message machinery beyond that choice.🔧 Fix applied.
3 issues (1 warning, 2 infos) still open:
tests/fm-control-relaunch.test.sh:1861-make_nm_recorder(tests/fm-control-relaunch.test.sh:1859-1868) stages a fakeno-mistakeson PATH andassert_tmux_reclaimsthen grepsnm-callsforrun|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI -fm-control.sh,fm-spawn.sh,fm-control-lib.sh, andbin/backends/tmux.shcontain no call tono-mistakesor anaxirun verb - sonm-callsis never created and the guard cannot fail for the right reason. It also shadows only theno-mistakesname, not theaxientrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gatesurvives), theHEADcomparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add theaxiname, or drop it, since the surrounding assertions carry the requirement.bin/fm-control.sh:1063-interruptis the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint readsambiguousmakesbin/fm-control.sh <id> interruptdie with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, whileexit(bin/fm-control.sh:605) andrelaunch(bin/fm-spawn.sh:1756) through the same state now end withfm_control_unclassified_next_steppointing atbin/fm-peek.sh <id>and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it becauseinterruptreads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.docs/agent-control.md:180- The invariant sentence claims more than the code now does. Line 179 names four refusal classes -alive,ambiguous,unreadable, "and so does any endpoint whose absence is not provable" - and line 180 then asserts "Every such refusal names a command that acts on the state it read, spelled as the operator can run it", enumerating only the first three after the colon. The fourth class names no such command. Concrete sequence: a Linux tmux ship task whose window readsmissingwhile an agent process still works in the recorded worktree -bin/fm-spawn.sh <id> --relaunchprints bin/fm-spawn.sh:1743, which ends "(bin/fm-control.sh <id> exit and relaunch refuse it for the same reason)", i.e. it names two commands that explicitly do NOT act on the state; the reason text supplied by bin/fm-control-lib.sh:363 gives prose actions ("stop that process where it runs, or address its tmux server from this seat, then retry") but no spelled command. The same holds for the other three reasons the proof can return: the no-/proc reason (bin/fm-agent-process-lib.sh:147) ends "there is no such seat, and no route"; the unresolvable-worktree reason (bin/fm-agent-process-lib.sh:143) and the backend-load failure (bin/fm-control-lib.sh:358) carry no next step at all. Sibling site where the same class is printed: bin/fm-control.sh:607, whose exit refusal ends "and relaunch refuses for the same reason". This drift was created inside this run's fix rounds: 9350468/3d3ac1f added the line-180 claim while the unprovable-absence refusals still carried a route clause, and 19f6f78 (the round the human explicitly ordered) deleted that clause from bin/fm-spawn.sh:1743 and bin/fm-control.sh:607 without narrowing the doc. Remedy corrects what the change already does rather than extending it: scope line 180 to the three positively-read verdicts and state plainly that an unprovable absence names the reason and whatever route the proof found, up to and including none. Do not re-add a route clause to either message and add no conditional message machinery.🔧 Fix applied.
3 infos still open:
tests/fm-control-relaunch.test.sh:1861-make_nm_recorder(tests/fm-control-relaunch.test.sh:1859-1868) stages a fakeno-mistakeson PATH andassert_tmux_reclaimsthen grepsnm-callsforrun|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI -fm-control.sh,fm-spawn.sh,fm-control-lib.sh, andbin/backends/tmux.shcontain no call tono-mistakesor anaxirun verb - sonm-callsis never created and the guard cannot fail for the right reason. It also shadows only theno-mistakesname, not theaxientrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gatesurvives), theHEADcomparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add theaxiname, or drop it, since the surrounding assertions carry the requirement.bin/fm-control.sh:1063-interruptis the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint readsambiguousmakesbin/fm-control.sh <id> interruptdie with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, whileexit(bin/fm-control.sh:605) andrelaunch(bin/fm-spawn.sh:1756) through the same state now end withfm_control_unclassified_next_steppointing atbin/fm-peek.sh <id>and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it becauseinterruptreads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.tests/fm-control-relaunch.test.sh:1916- The reclaim's most operator-visible new behaviour is unverified: the record's session part moves from the recorded session to the reclaiming seat's.assert_tmux_reclaimsmatches the rebound record withcase "$window" in *":fm-$id")(tests/fm-control-relaunch.test.sh:1916), which passes identically whether the field readsfirstmate:fm-rl60(what the code does) orfmses:fm-rl60(the recorded session). Concrete trace:add_ship_taskrecordswindow=fmses:fm-rl60(tests/fm-control-relaunch.test.sh:204,215),run_controlnow unsets TMUX (:242), so bin/fm-spawn.sh:3470'sfm_backend_tmux_container_ensureresolvesfirstmateand bin/fm-spawn.sh:4768 writeswindow=firstmate:fm-rl60- a value no assertion reads. The change states this as a contract in three places (docs/agent-control.md:116-117 "a tmux reclaim can move the task into the reclaiming seat's session", :130 "on tmux that binding includes the session", bin/fm-control.sh:66-69 "tmux: on the server and in the session THIS seat addresses"), and it is not cosmetic: a regression that pinned the recorded session would maketmux new-window -t "fmses:"fail outright for the two cases the suite covers where that session no longer exists (tests/fm-control-relaunch.test.sh:1962 session-missing, :1971 server-dead), re-creating the deadlock the change exists to remove. The stub cannot catch it either - itsnew-windowbranch records only the-nname intocreated-windows(tests/fm-control-relaunch.test.sh:158-174), never the-ttarget - so the record field is the only observable. Sibling site where the same invariant is unpinned: tests/fm-control-relaunch.test.sh:1986, wheretest_spawn_relaunch_alone_reclaims_a_gone_tmux_endpointchecks the worktree field but never readswindowat all. Remedy is mechanical and test-only: assert the reboundwindowequalsfirstmate:fm-$idinassert_tmux_reclaims, and read the same field in the launch-owner case.✅ **Test** - passed
✅ No issues found.
endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 ..., rc=0bin/fm-control.sh labscout1 relaunch --note "<why>"; running it succeeded (rc=0) and reused the surviving window @2 rather th…spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 worktree=...rc=0, new window @3, recorded worktree keptbash tests/fm-backend-tmux-smoke.test.shagainst a real tmux server on a private socket: 'refuses an unreadable window inventory rather than assuming the name is free' and 'refuses a duplicate witho…bin/fm-lab-home.sh create $LAB+tmux -L fm-labprivate socket - disposable lab home and isolated tmux serverbin/fm-brief.sh labscout1 labproj --scoutthenbin/fm-spawn.sh labscout1 $LAB/projects/labproj --scout --backend tmux --harness claude- real live spawn, real treehouse worktree, real Claude agenttmux -L fm-lab rename-windowthenbin/fm-spawn.sh labscout1 --relaunch --harness claudeandbin/fm-control.sh labscout1 exit- adversarial: endpoint reads missing while a real Claude still works in the worktreegit archive 30ef650 | tar -xthen$LAB/base/bin/fm-control.sh labscout1 relaunch --note,$LAB/base/bin/fm-spawn.sh labscout1 --relaunch,$LAB/base/bin/fm-control.sh labscout1 exit- base-commit deadlock reproductionbin/fm-control.sh labscout1 exiton the destroyed endpoint - expects endpoint-gonebin/fm-control.sh labscout1 relaunch --note "the tmux window was destroyed; pick the work back up"- the reclaimtmux -L fm-lab list-windows -a,cat $LAB/state/labscout1.meta,git -C $WT rev-parse HEAD,git -C $WT status --porcelain,cat $LAB/state/labscout1.status- post-reclaim state verificationbin/fm-spawn.sh labscout1 --relaunch --harness claudeagainst the live reclaimed endpoint, then thebin/fm-control.sh labscout1 relaunch --note "<why>"it named, run verbatimbin/fm-spawn.sh labscout1 --relaunch --harness claudealone after a second window destroy - launch owner reclaims on its ownreclaim driven from a pane in sessionotherthat already held a decoyfm-labscout1window - duplicate-name refusal, then rerun of the command the rollback namedbash tests/fm-backend-tmux-smoke.test.sh- real-tmux create/duplicate/unreadable-inventory guardsbash tests/fm-control-relaunch.test.shbash tests/fm-control.test.shbin/backends/tmux.sh:112- Judgment call, deliberately left as code comments rather than prose. The change makes fm_backend_tmux_create_task refuse when a session's window inventory cannot be read (previously a failed read silently passed for "no duplicate"), and that helper serves ordinary tmux spawns as well as the new reclaim. No prose surface owns spawn-time window-create refusals today: docs/tmux-backend.md covers setup, liveness, and composer behavior only, and docs/agent-control.md:119 documents the duplicate-name refusal solely in the reclaim context. The new refusal's rationale and text live in the function's own header comment, which is the correct owner for a local safety invariant, so I added no prose and created no new section (the placement policy forbids inventing a surface to close a perceived gap). Flagging it in case a future operator-facing note about tmux spawn refusals is wanted; nothing is currently wrong or contradicted.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.