Skip to content

fix(bin): reclaim a tmux task whose endpoint is gone - #5980

Closed
keenvc wants to merge 14 commits into
kunchenguid:mainfrom
keenvc:fm/fm-endpoint-gone-reconcile
Closed

keenvc wants to merge 14 commits into
kunchenguid:mainfrom
keenvc:fm/fm-endpoint-gone-reconcile

Conversation

@keenvc

@keenvc keenvc commented Sep 28, 2026

Copy link
Copy Markdown

Intent

Firstmate's control plane deadlocks when a task's recorded endpoint has been destroyed rather than merely stopped, and it cost manual record surgery tonight on a live task.

Observed on task fm-teardown-shared-slot-retire, a claude ship lane whose herdr pane was destroyed while its no-mistakes run sat parked at a review gate. Its work was safe: the branch was committed in the recorded worktree and the parked run was still held. Recovery was nonetheless impossible through the supported commands:

  • bin/fm-control.sh <id> relaunch launched and then could not confirm a running agent, reporting the work preserved.
  • bin/fm-spawn.sh <id> --relaunch refused: "endpoint reads 'missing'; a relaunch requires a positively agent-free endpoint (stop the agent first with bin/fm-control.sh exit)".
  • bin/fm-control.sh <id> exit refused: "recorded endpoint is gone, so there is no agent to stop; reconcile the task before any further control action" (bin/fm-control.sh:461).
    So relaunch sends the operator to exit, exit sends the operator to reconcile, and no command performs that reconcile. The operator's only route was to create a pane by hand, launch the harness by hand, and rewrite the window and herdr fields in state/.meta with sed - which is exactly the unguarded record editing the control plane exists to prevent.

A related symptom in the same area, useful as evidence but NOT this task's scope: when herdr reports a pane's agent binding as lost while the agent is in fact alive and advancing, bin/fm-send.sh declines to ring the doorbell and reports that the agent has exited. That is the opposite error - a live agent read as gone - and the two together suggest the agent-state classification is being trusted for decisions it cannot support.

What Changed

  • Added fm_agent_process_worktree_scan (bin/fm-agent-process-lib.sh), a /proc-based scan for a verified harness process working in a task's recorded worktree, and wired it into fm_control_endpoint_absence_verdict so a tmux missing endpoint can now resolve to gone (no agent in the worktree) instead of always returning unproven; an agent found there, an unreadable process table, or a host without /proc still refuses.
  • bin/fm-spawn.sh --relaunch now rebinds a proven-gone tmux task: it ensures the container session, creates one fresh window in the recorded worktree on the server this seat addresses, republishes window=, and closes that window from the abort trap if anything refuses before the record names it. bin/backends/tmux.sh's create_task reads the window inventory three-way and refuses both an existing same-named window and an unreadable inventory rather than inferring absence from an empty stream.
  • Replaced the refusal texts that pointed relaunch → exit → "reconcile" with a shared fm_control_unclassified_next_step helper, so exit, relaunch, and the launch owner each name a runnable next command for the verb actually invoked; relaunch rollback messages now name the retry command. Tests cover the new scan, the tmux reclaim path and the refusals (tests/lib.sh gains fm_proc_scan_available and fm_wait_for_agent_argv0), and docs/agent-control.md, docs/scripts.md, and the stuck-crewmate-recovery skill drop the "no tmux reclaim" claim.

Risk Assessment

⚠️ Medium: The core absence proof is sound - the one sequence that must never pass (an agent alive in the worktree on an unaddressable tmux server) is caught by the kernel cwd read, and every unreadable or agent-found path still refuses - but a pipeline fix round added an unrequested tmux husk-replacement that converts a safety refusal on the ordinary spawn path into a destructive kill-window reachable across Firstmate homes, which needs human sign-off before merge.

Testing

I drove the whole endpoint-gone reclaim against the real product in a disposable lab home with a real private tmux server, a real treehouse worktree and a real Claude harness, never a stub. The deadlock the intent describes reproduced live against the base commit - relaunch, fm-spawn --relaunch and exit each refused and pointed at one another - and on this branch the same state gives exit an endpoint-gone verdict and relaunch a working reclaim that opens one fresh window, rebinds the record, and brings a real agent back up in the same worktree while the commit, the never-committed file and the parked-gate status line all survive. The adversarial probes all held: a real Claude still working in the recorded worktree blocked both verbs with the refusal naming its PID and left it alone, a foreign window already holding the task's name was refused and stayed open, and every refusal named a command that actually ran when I ran it verbatim. The repository's targeted suites, including the real-tmux backend smoke test, pass as the baseline. This product surface is a terminal control plane rather than a graphical one, so the reviewer-visible artifact is the captured live tmux pane showing the reclaimed agent plus the CLI transcripts; there is no rendered web or GUI surface to screenshot.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
An operator whose tmux task window was destroyed reclaims the task with bin/fm-control.sh <id> relaunch --note, and a real agent comes back up in the same worktree ✅ pass live live-reclaim-transcript.md section 3 and reclaimed-window-pane.txt: relaunch rc=0, a new tmux window id (@2 replacing the destroyed @1), record rebound, and a real Claude running in it
The reported deadlock reproduces before the fix: on the base commit, relaunch, fm-spawn --relaunch and exit all refuse the same destroyed endpoint and leave no route ✅ pass live live-reclaim-transcript.md section 2 - base-commit scripts extracted with git archive 30ef650 and driven against the same live lab task; all three refuse
bin/fm-control.sh <id> exit on a destroyed tmux endpoint reports endpoint-gone instead of the 'reconcile the task first' dead end ✅ pass live live-reclaim-transcript.md section 3: endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 ..., rc=0
The reclaim preserves the work the incident said was safe: committed branch state, an uncommitted file, and a status log parked at a validation gate ✅ pass live final-lab-state.txt: HEAD ccf2864 before == after, uncommitted.txt still present and untracked, the 'parked at a no-mistakes review gate' status line still in state/labscout1.status after four reclaim…
Adversarial: an agent still alive in the recorded worktree stops the reclaim - neither verb launches a second agent, and the refusal names the live process ✅ pass live live-reclaim-transcript.md section 1: both verbs refuse naming 'agent process 1918852 claude', no window created, PID 1918852 still alive afterwards
Adversarial: a foreign window already holding fm-<id> in this seat's session is refused rather than closed or duplicated, and the rollback's named rerun works once it is cleared ✅ pass live live-reclaim-transcript.md section 6: refusal names bin/fm-peek.sh, decoy window @5 left open, no second window; after closing it the same command succeeded and rebound to other:fm-labscout1 (@6)
Every refusal names a command the operator can actually run: the alive-endpoint refusal names relaunch --note "<why>" and that command, run verbatim, replaces the agent ✅ pass live live-reclaim-transcript.md section 4: fm-spawn --relaunch refused naming bin/fm-control.sh labscout1 relaunch --note &#34;&lt;why&gt;&#34;; running it succeeded (rc=0) and reused the surviving window @2 rather th…
bin/fm-spawn.sh <id> --relaunch on its own reaches the same verdict and reclaims a destroyed endpoint, so the two commands cannot disagree ✅ pass live live-reclaim-transcript.md section 5: spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 worktree=... rc=0, new window @3, recorded worktree kept
The tmux create guard refuses when the window inventory cannot be read, rather than assuming the name is free and opening a duplicate ✅ pass live bash tests/fm-backend-tmux-smoke.test.sh against a real tmux server on a private socket: 'refuses an unreadable window inventory rather than assuming the name is free' and 'refuses a duplicate witho…
Evidence: Live reclaim transcript (base-commit deadlock, reclaim, adversarial probes)

Source: Live reclaim transcript (base-commit deadlock, reclaim, adversarial probes)

# Live drive: reclaiming a tmux task whose endpoint was destroyed

Product driven as an operator runs it: a disposable lab firstmate home
(`bin/fm-lab-home.sh create`), a real private tmux server (`tmux -L fm-lab`,
`TMUX_TMPDIR=$LAB/tmux`), a real throwaway git project, a real `treehouse get`
worktree, and a real `claude` harness. No fakes, no stubs.

Task `labscout1` was created by the real launch path:

    $ bin/fm-spawn.sh labscout1 $LAB/projects/labproj --scout --backend tmux --harness claude
    spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 \
      worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj

A real Claude agent came up in `firstmate:fm-labscout1` (window id @1) and began
working. Work was then staged in the worktree the way the incident describes it:
one commit (HEAD ccf2864), one never-committed file, and a status line
`working: parked at a no-mistakes review gate`.

--------------------------------------------------------------------------------
## 1. Adversarial: endpoint unreadable while the agent is still alive

The recorded window was renamed out from under the record (`tmux rename-window`),
so `firstmate:fm-labscout1` reads `missing` while the real Claude keeps running in
the worktree. Both verbs must refuse rather than launch a second agent.

    $ bin/fm-spawn.sh labscout1 --relaunch --harness claude
    error: task labscout1's recorded endpoint firstmate:fm-labscout1 reads 'missing',
    but agent process 1918852 claude is still working in the recorded worktree
    /tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj, so its window may
    be on a tmux server this seat cannot address; stop that process where it runs, or
    address its tmux server from this seat, then retry. An endpoint that cannot be
    proven absent may still hold a live agent on this task's worktree; refusing rather
    than launching a second agent into it (bin/fm-control.sh labscout1 exit and
    relaunch refuse it for the same reason)
    rc=1

    $ bin/fm-control.sh labscout1 exit
    error: task labscout1's endpoint firstmate:fm-labscout1 reads 'missing', but agent
    process 1918852 claude is still working in the recorded worktree ...; exit will not
    claim an agent stopped at an address it cannot trust, nor send lifecycle input to
    one, and relaunch refuses for the same reason
    rc=1

    $ tmux -L fm-lab list-windows -a -F '#{session_name}:#{window_name}'
    firstmate:bash
    firstmate:stranded-elsewhere        <- no second window created

PID 1918852 was still alive after both refusals. The refusal names the exact live
process it found.

--------------------------------------------------------------------------------
## 2. The reported deadlock, reproduced at the base commit (30ef650)

The window was then destroyed (`tmux kill-window`), taking the agent with it; zero
processes remained in the worktree. The base-commit scripts were extracted with
`git archive 30ef650` and driven against the same live lab home:

    $ $LAB/base/bin/fm-control.sh labscout1 relaunch --note "..."
    error: task labscout1's endpoint firstmate:fm-labscout1 reads 'missing', but tmux
    absence cannot be proven from a task record: ...
    error: relaunch of labscout1 failed while stopping the old agent and its state is
    'missing', so it was not proven stopped; its original instructions were restored
    and the durable record was retained for recovery

    $ $LAB/base/bin/fm-spawn.sh labscout1 --relaunch --harness claude
    error: task labscout1's recorded endpoint firstmate:fm-labscout1 reads 'missing',
    but tmux absence cannot be proven from a task record: ... refusing rather than
    launching a second agent into it

    $ $LAB/base/bin/fm-control.sh labscout1 exit
    error: task labscout1's endpoint firstmate:fm-labscout1 reads 'missing', but tmux
    absence cannot be proven ...; exit will not claim an agent stopped at an address it
    cannot trust, nor send lifecycle input to one

All three supported commands refuse; there is no route. This is the deadlock the
intent reports, reproduced live before the fix.

--------------------------------------------------------------------------------
## 3. The same state on this branch: exit reports it, relaunch reclaims it

    $ bin/fm-control.sh labscout1 exit
    endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 \
      worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
    rc=0

    $ bin/fm-control.sh labscout1 relaunch --note "the tmux window was destroyed; pick the work back up"
    relaunched labscout1 harness=claude from=claude model=default effort=default \
      backend=tmux endpoint=firstmate:fm-labscout1 \
      worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
    rc=0

    $ tmux -L fm-lab list-windows -a -F '#{session_name}:#{window_name} id=#{window_id}'
    firstmate:bash id=@0
    firstmate:fm-labscout1 id=@2       <- a NEW window (the destroyed one was @1)

A real Claude came up in the new window, in the same worktree, and picked the work
back up (pane capture: reclaimed-window-pane.txt). Nothing else moved:

  * worktree HEAD ccf2864 before == ccf2864 after
  * the never-committed file is still there, untouched
  * the `parked at a no-mistakes review gate` status line survived
  * the progress note reached the replacement's brief
  * no record was edited by hand at any point

--------------------------------------------------------------------------------
## 4. Refusals name a command that actually runs

With the reclaimed agent alive, the launch owner refuses and names its route:

    $ bin/fm-spawn.sh labscout1 --relaunch --harness claude
    error: task labscout1's endpoint reads 'alive': an agent still runs there, and a
    relaunch requires a positively agent-free endpoint, so replace it with
    bin/fm-control.sh labscout1 relaunch --note "<why>", which stops that agent before
    launching this one
    rc=1

Running exactly the command it named (with a real why) succeeded and reused the
surviving window @2 rather than opening a second one.

--------------------------------------------------------------------------------
## 5. The launch owner reclaims on its own

Window destroyed again; `bin/fm-spawn.sh --relaunch` alone reaches the same verdict:

    $ bin/fm-spawn.sh labscout1 --relaunch --harness claude
    spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 \
      worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
    rc=0
    firstmate:fm-labscout1 id=@3       <- again a new window, same worktree

--------------------------------------------------------------------------------
## 6. Adversarial: a foreign window already holds the name

A second session `other` was given a decoy window named `fm-labscout1` (standing in
for another home's deliberately preserved endpoint), and the reclaim was driven from
inside that session, so the new window would land there:

    $ bin/fm-control.sh labscout1 relaunch --note "..."      # run from a pane in session `other`
    error: window other:fm-labscout1 already exists; read it with
    bin/fm-peek.sh other:fm-labscout1 to see what holds it. No control-plane command
    closes a window this home's records do not name, so a leftover or another home's
    window has to be closed where it runs before this task can open its own
    error: the replacement agent for labscout1 could not be launched on claude
    error: labscout1's agent was stopped but the replacement did not launch; no agent is
    running, and its work plus the recorded progress note are preserved at ...; rerun
    bin/fm-control.sh labscout1 relaunch --note "<why>" once the launch failure above is
    resolved
    rc=1

    other:fm-labscout1 id=@5           <- the foreign window is still open, untouched

Closing the decoy and rerunning the command the rollback named succeeded:

    relaunched labscout1 ... endpoint=other:fm-labscout1 ...
    rc=0
    other:fm-labscout1 id=@6           <- fresh window, record rebound to this seat's session

The lab home, its tmux server, its worktree pool and its project were removed in the
same turn.
Evidence: Reclaimed tmux window pane - the real Claude agent running in the reclaimed endpoint

Source: Reclaimed tmux window pane - the real Claude agent running in the reclaimed endpoint


 ▐▛███▛█   Claude Code v2.1.283
▝▜██████▀  Opus 5 with high effort · Claude Max
 ▝▝   ▝▝   /tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj


❯ : Firstmate operational input waiting: read '/tmp/fm-lab.Emxt7W/state/operational-inbox/1790577269-e08a5de7bd81e971.msg' and handle its contents as Firstmate operational input.

● I'll read the operational input.

● Getting git HEAD and read existing report
  ⎿  $ pwd && echo "---REVPARSE---" && git rev-parse HEAD && echo "---STATUSBR---" && git status --short --branch && echo "---LOG---" && git log --oneline -3 && echo "---EXISTING REPORT---" && cat
     /tmp/fm-lab.Emxt7W/data/labscout1/report.md





























✢ Topsy-turvying… (23s · ↓ 1.0k tokens · thinking with high effort)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ?2
  [OMC#5.5.0L] | Model: Opus 5 | 5h:[--------]5%(4h45m) wk:[###-----]38%(4d16h) fable:[--------]0%(4d16h) | thinking | session:0m | ctx:[#---------]6% | 🔧2
  ⏵⏵ bypass permissions on (shift+tab to cycle) · ← 2 agents
Evidence: Final lab state: rebound record, unchanged HEAD, surviving uncommitted file and parked status line

Source: Final lab state: rebound record, unchanged HEAD, surviving uncommitted file and parked status line

# Final lab state after four reclaims of one task (task labscout1, backend tmux)

## tmux windows on the lab socket
firstmate:bash window_id=@0
other:bash window_id=@4
other:fm-labscout1 window_id=@6

## state/labscout1.meta (record after the last reclaim)
window=other:fm-labscout1
endpoint_task_id=labscout1
worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
project=/tmp/fm-lab.Emxt7W/projects/labproj
harness=claude
kind=scout
tasktmp=/tmp/fm-labscout1
model=default
effort=default
busy_gen=g1790577268.2445190.28331
spawn_gen=s1790577269.2433259.28594
decisions_reviewed=1
decision_keys=
control_relaunch_tx=2426236.20260928T063425Z.9845

## worktree HEAD (recorded before the endpoint was destroyed, then now)
ccf2864fdc0a3c2b5cb0ce16c254497d60b9dd1f
ccf2864fdc0a3c2b5cb0ce16c254497d60b9dd1f

## uncommitted file staged before the destroy
never committed
?? .omc/
?? uncommitted.txt

## status log (the parked-gate line survived every reclaim)
working [at=1790577082]: parked at a no-mistakes review gate
done [at=1790577091]: worktree usable; HEAD 886fd3be0d375eca2e6461aeb43dbceb00864ad5 recorded in data/labscout1/report.md
Evidence: Base-commit deadlock, reproduced live on the same lab task
$ $LAB/base/bin/fm-control.sh labscout1 relaunch --note "..."
error: task labscout1's endpoint firstmate:fm-labscout1 reads 'missing', but tmux absence cannot be proven from a task record: ...
error: relaunch of labscout1 failed while stopping the old agent and its state is 'missing', so it was not proven stopped

$ $LAB/base/bin/fm-spawn.sh labscout1 --relaunch --harness claude
error: ... refusing rather than launching a second agent into it

$ $LAB/base/bin/fm-control.sh labscout1 exit
error: ... exit will not claim an agent stopped at an address it cannot trust, nor send lifecycle input to one
Evidence: Same state on this branch: exit reports it, relaunch reclaims it
$ bin/fm-control.sh labscout1 exit
endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
rc=0

$ bin/fm-control.sh labscout1 relaunch --note "the tmux window was destroyed; pick the work back up"
relaunched labscout1 harness=claude from=claude model=default effort=default backend=tmux endpoint=firstmate:fm-labscout1 worktree=/tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj
rc=0

$ tmux -L fm-lab list-windows -a -F '#{session_name}:#{window_name} id=#{window_id}'
firstmate:bash id=@0
firstmate:fm-labscout1 id=@2 # the destroyed window was @1
Evidence: Adversarial: a real Claude still working in the worktree blocks both verbs
$ bin/fm-spawn.sh labscout1 --relaunch --harness claude
error: task labscout1's recorded endpoint firstmate:fm-labscout1 reads 'missing', but agent process 1918852 claude is still working in the recorded worktree /tmp/fm-lab.Emxt7W/treehouse/.treehouse/labproj-45721d/1/labproj, so its window may be on a tmux server this seat cannot address; stop that process where it runs, or address its tmux server from this seat, then retry. ... refusing rather than launching a second agent into it
rc=1

$ tmux -L fm-lab list-windows -a -F '#{session_name}:#{window_name}'
firstmate:bash
firstmate:stranded-elsewhere # no second window, PID 1918852 still alive
Evidence: Adversarial: a foreign same-named window is refused, never closed
$ bin/fm-control.sh labscout1 relaunch --note "..." # driven from a pane in session `other`
error: window other:fm-labscout1 already exists; read it with bin/fm-peek.sh other:fm-labscout1 to see what holds it. No control-plane command closes a window this home's records do not name ...
error: labscout1's agent was stopped but the replacement did not launch; ... rerun bin/fm-control.sh labscout1 relaunch --note "<why>" once the launch failure above is resolved
rc=1

other:fm-labscout1 id=@5 # the foreign window is still open, untouched

Pipeline

Updates from git push no-mistakes

... (17 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 3 infos

🔧 Fix applied.
5 issues (1 warning, 4 infos) still open:

  • ℹ️ tests/fm-control-relaunch.test.sh:1861 - make_nm_recorder (tests/fm-control-relaunch.test.sh:1859-1868) stages a fake no-mistakes on PATH and assert_tmux_reclaims then greps nm-calls for run|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI - fm-control.sh, fm-spawn.sh, fm-control-lib.sh, and bin/backends/tmux.sh contain no call to no-mistakes or an axi run verb - so nm-calls is never created and the guard cannot fail for the right reason. It also shadows only the no-mistakes name, not the axi entrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gate survives), the HEAD comparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add the axi name, or drop it, since the surrounding assertions carry the requirement.
  • ℹ️ bin/fm-control.sh:1063 - interrupt is the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint reads ambiguous makes bin/fm-control.sh &lt;id&gt; interrupt die with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, while exit (bin/fm-control.sh:605) and relaunch (bin/fm-spawn.sh:1756) through the same state now end with fm_control_unclassified_next_step pointing at bin/fm-peek.sh &lt;id&gt; and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it because interrupt reads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.
  • ⚠️ bin/fm-agent-process-lib.sh:147 - The branch's whole deliverable - the tmux reclaim - is unavailable on macOS, so on the default backend the exact dead-end the intent reports survives unchanged. Concrete sequence on a macOS seat: a tmux task's window is destroyed; fm_backend_tmux_agent_state reads missing; fm_control_endpoint_absence_verdict (bin/fm-control-lib.sh:353) calls fm_agent_process_worktree_scan, whose second guard [ ! -r /proc/self/cmdline ] || [ ! -L /proc/self/cwd ] (bin/fm-agent-process-lib.sh:146-150) fires because macOS has no /proc; the verdict is unproven; bin/fm-control.sh &lt;id&gt; exit dies at :602 and bin/fm-spawn.sh &lt;id&gt; --relaunch dies at :1755 - and the refusal text itself ends "once that server is gone there is no such seat, and no route". That is verbatim the state the intent calls unacceptable: "So relaunch sends the operator to exit, exit sends the operator to reconcile, and no command performs that reconcile. The operator's only route was to create a pane by hand, launch the harness by hand, and rewrite the window and herdr fields in state/<id>.meta with sed - which is exactly the unguarded record editing the control plane exists to prevent." macOS is a first-class platform in this repo, not a hypothetical: bin/fm-backend.sh:175 gates cmux fallbacks on uname = Darwin, bin/fm-install-treehouse.sh:39 ships Darwin-arm64/Darwin-x86_64 builds, and bin/fm-agent-process-lib.sh:81 documents macOS comm semantics for this very classifier. tmux is the default/reference backend (bin/backends/tmux.sh:4, and backend= absent means tmux), so this is not a niche path. The change also carries no automated evidence there: every new case skips via fm_proc_scan_available (tests/lib.sh:504, tests/fm-control-relaunch.test.sh:1902, :1911, :1920, :1929, :1944, :2035, tests/fm-control.test.sh:783), so on macOS the reclaim is simultaneously unavailable and unexercised. This needs the author's decision, not a repair: the remedy is a second, non-/proc absence proof, which EXTENDS the change, and a prior fix round in this same run (6f88055 "drop lsof scan fallback, refuse absence proof without /proc", after 8e58903 and 50f7b2c spent two rounds on lsof parsing bugs) deliberately removed exactly that. Deciding Linux-only is defensible - Herdr reclaims everywhere and the refusal is honest rather than misleading - but the intent marks the reconcile route as required and does not authorize the gap, so say explicitly whether tmux-on-macOS is out of scope, and if so scope docs/agent-control.md:110-115 and the capability matrix at :186-192 to say the tmux reclaim is Linux-only.
  • ℹ️ docs/agent-control.md:113 - Three surfaces describe the tmux gone verdict as proving more than the code establishes; this is the same family as the placement overclaim the author corrected in round 8 (998afee), left behind at these sibling sites. (1) docs/agent-control.md:113 - "An agent found there refuses and names its process id, and so does any process whose working directory could not be read" overstates the guard: bin/fm-agent-process-lib.sh:175 runs [ &#34;$(fm_agent_process_classify ...)&#34; = agent ] || continue BEFORE the unreadable-cwd refusal at :176-179, so a non-agent process with an unreadable cwd is silently skipped, never refused. The correct sentence is "any agent process whose working directory could not be read". (2) bin/fm-control-lib.sh:317 - "gone - absence is PROVEN: no agent, and no endpoint at the address the record names" contradicts its own next clause two lines down ("A window may still survive elsewhere - on a tmux server this seat cannot address"): a record's tmux address is a bare session:window with no socket component, so the identical address can name a live window on another server, and on tmux the scan proves only the agent absent, never the endpoint. (3) The same overclaim rides along at bin/fm-control.sh:582 ("nothing survived at the recorded address for this verb to preserve") and its user-facing copy in the exit row at docs/agent-control.md:36. All four are wording; no behavior is wrong. Scope each to what tmux actually proves - no agent process working in the recorded worktree - and leave the herdr phrasing, which does prove the endpoint absent, as it is.
  • ℹ️ bin/backends/tmux.sh:105 - The change promotes fm_backend_tmux_create_task's duplicate-window refusal to the reclaim's safety boundary - the new comment at bin/backends/tmux.sh:83-92 says closing an unrecorded window "would destroy another home's deliberately preserved endpoint", and docs/agent-control.md:118 states the reclaim refuses rather than replacing such a window - but the guard itself fails open on a read it could not make. tmux list-windows -t &#34;$ses&#34; -F &#39;#{window_name}&#39; | grep -qx &#34;$wname&#34; (bin/backends/tmux.sh:105) cannot distinguish "the session has no such window" from "the read failed": grep sees an empty stream either way and returns 1, so the create proceeds. Sequence: a reclaim runs from a seat whose ambient session already holds another home's fm-&lt;id&gt; window; tmux list-windows fails transiently; tmux new-window at :109 succeeds and tmux permits the duplicate name; bin/fm-spawn.sh:3471 persists the NAME form T=&#34;$SES:$W&#34; as the record's window=, and every later name-targeted read and write - the env exports at bin/fm-spawn.sh:5112-5132, the launch itself at :5227, and spawn_send_key &#34;$T&#34; Enter at :5234 - resolves to whichever window tmux matches first, which can be the foreign one the refusal exists to protect. The sibling in this same file already does this correctly: fm_backend_tmux_kill calls fm_backend_tmux_window_inventory and branches on its three-way status (bin/backends/tmux.sh:198-206), refusing on status 1 precisely so an unreadable read is never mistaken for absence. Classifying this as ask-user rather than auto-fix because the remedy changes a shared helper's contract for the ordinary spawn path at bin/fm-spawn.sh:3557 too, which is wider than this change's scope; the pre-existing helper is unchanged here, only the new reliance on it is.

🔧 Fix applied.
4 issues (1 warning, 3 infos) still open:

  • ℹ️ tests/fm-control-relaunch.test.sh:1861 - make_nm_recorder (tests/fm-control-relaunch.test.sh:1859-1868) stages a fake no-mistakes on PATH and assert_tmux_reclaims then greps nm-calls for run|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI - fm-control.sh, fm-spawn.sh, fm-control-lib.sh, and bin/backends/tmux.sh contain no call to no-mistakes or an axi run verb - so nm-calls is never created and the guard cannot fail for the right reason. It also shadows only the no-mistakes name, not the axi entrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gate survives), the HEAD comparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add the axi name, or drop it, since the surrounding assertions carry the requirement.
  • ℹ️ bin/fm-control.sh:1063 - interrupt is the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint reads ambiguous makes bin/fm-control.sh &lt;id&gt; interrupt die with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, while exit (bin/fm-control.sh:605) and relaunch (bin/fm-spawn.sh:1756) through the same state now end with fm_control_unclassified_next_step pointing at bin/fm-peek.sh &lt;id&gt; and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it because interrupt reads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.
  • ⚠️ bin/fm-spawn.sh:1752 - Every refusal and rollback message this change added names bin/fm-control.sh &lt;id&gt; relaunch as the next step, but that verb dies immediately for a ship or scout task unless --note/--note-file is also passed (bin/fm-control.sh:979: [ &#34;$NOTE_SET&#34; = 1 ] &amp;&amp; [ -n &#34;$NOTE&#34; ] || die &#34;relaunch of a $KIND task requires --note...&#34;). An operator who copies the printed command verbatim gets a second refusal, which is the same cross-pointing this branch set out to end (commit 9350468 "stop refusals promising routes the state cannot offer"). The change's own test admits it: tests/fm-control-relaunch.test.sh:1712 comments "The named command must actually work for this state, rather than send the operator on to another refusal", then at :1713 runs run_control &#34;$dir&#34; rl15 relaunch --note &#34;replace the live agent&#34; - adding an argument the refusal never mentions - on a task created by add_ship_task (:1704). So the test proves relaunch --note works, not that the command as printed works. docs/agent-control.md:180 states the stronger claim outright: "Every such refusal names a command that acts on the state it read: an alive endpoint points at bin/fm-control.sh &lt;id&gt; relaunch, which stops that agent first". Every site in the changed code where the same invariant must hold: bin/fm-spawn.sh:1752 (new alive branch, primary); bin/fm-spawn.sh:1743 ("before bin/fm-control.sh $ID relaunch can reclaim the task" on an unprovable absence); bin/fm-spawn.sh:1756 via bin/fm-control-lib.sh:402 and :405 ("retry bin/fm-control.sh %s %s" with verb=relaunch); bin/fm-control.sh:604 ("before bin/fm-control.sh $ID $VERB can act", where $VERB is relaunch inside the relaunch transaction); and the three rollback lines this change added at bin/fm-control.sh:746, :775, :778 ("rerun bin/fm-control.sh $ID relaunch") - the rollback case is the most reachable, since a failed replacement launch is exactly when an operator reruns the printed command. Classified ask-user because the remedy is a product-copy decision, not a mechanical repair: fm-control-lib.sh's shared helper has no access to KIND, so either the printed command grows a --note &#34;&lt;why&gt;&#34; placeholder unconditionally, or the two callers that do know KIND (bin/fm-spawn.sh, bin/fm-control.sh) append it and the helper gains a parameter. Pick one; do not harden further.
  • ℹ️ .agents/skills/stuck-crewmate-recovery/SKILL.md:43 - Round 9's accepted instruction scoped the tmux reclaim's Linux-only caveat into docs/agent-control.md, and that landed correctly (docs/agent-control.md:113 "the tmux reclaim is Linux-only", :114 "On macOS, or any other host without /proc, a tmux task whose endpoint was destroyed cannot be reclaimed at all", plus the capability-matrix paragraph at :195-197). This sibling agent-facing surface was left behind: the rewritten sentence now tells every agent reading it that "An endpoint that is not merely idle but destroyed - a Herdr pane or workspace removed in churn, a tmux window or server that no longer exists - is recovered by that same bin/fm-control.sh &lt;id&gt; relaunch ... nothing special is needed", with no platform qualifier. On a macOS seat that is false for the tmux half: bin/fm-agent-process-lib.sh:146-150 refuses with unreadable because there is no /proc, so both verbs refuse. The following sentence ("Do not work around a refusal by respawning or by editing the task record") keeps the agent from doing damage, so this is guidance drift rather than a safety hole. Wording only - the smallest honest remedy is one clause naming that the tmux reclaim needs /proc and the Herdr reclaim does not, matching what docs/agent-control.md already says; add no component.

🔧 Fix applied.
4 issues (1 warning, 3 infos) still open:

  • ℹ️ tests/fm-control-relaunch.test.sh:1861 - make_nm_recorder (tests/fm-control-relaunch.test.sh:1859-1868) stages a fake no-mistakes on PATH and assert_tmux_reclaims then greps nm-calls for run|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI - fm-control.sh, fm-spawn.sh, fm-control-lib.sh, and bin/backends/tmux.sh contain no call to no-mistakes or an axi run verb - so nm-calls is never created and the guard cannot fail for the right reason. It also shadows only the no-mistakes name, not the axi entrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gate survives), the HEAD comparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add the axi name, or drop it, since the surrounding assertions carry the requirement.
  • ℹ️ bin/fm-control.sh:1063 - interrupt is the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint reads ambiguous makes bin/fm-control.sh &lt;id&gt; interrupt die with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, while exit (bin/fm-control.sh:605) and relaunch (bin/fm-spawn.sh:1756) through the same state now end with fm_control_unclassified_next_step pointing at bin/fm-peek.sh &lt;id&gt; and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it because interrupt reads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.
  • ⚠️ bin/fm-spawn.sh:1752 - Every refusal and rollback message this change added names bin/fm-control.sh &lt;id&gt; relaunch as the next step, but that verb dies immediately for a ship or scout task unless --note/--note-file is also passed (bin/fm-control.sh:979: [ &#34;$NOTE_SET&#34; = 1 ] &amp;&amp; [ -n &#34;$NOTE&#34; ] || die &#34;relaunch of a $KIND task requires --note...&#34;). An operator who copies the printed command verbatim gets a second refusal, which is the same cross-pointing this branch set out to end (commit 9350468 "stop refusals promising routes the state cannot offer"). The change's own test admits it: tests/fm-control-relaunch.test.sh:1712 comments "The named command must actually work for this state, rather than send the operator on to another refusal", then at :1713 runs run_control &#34;$dir&#34; rl15 relaunch --note &#34;replace the live agent&#34; - adding an argument the refusal never mentions - on a task created by add_ship_task (:1704). So the test proves relaunch --note works, not that the command as printed works. docs/agent-control.md:180 states the stronger claim outright: "Every such refusal names a command that acts on the state it read: an alive endpoint points at bin/fm-control.sh &lt;id&gt; relaunch, which stops that agent first". Every site in the changed code where the same invariant must hold: bin/fm-spawn.sh:1752 (new alive branch, primary); bin/fm-spawn.sh:1743 ("before bin/fm-control.sh $ID relaunch can reclaim the task" on an unprovable absence); bin/fm-spawn.sh:1756 via bin/fm-control-lib.sh:402 and :405 ("retry bin/fm-control.sh %s %s" with verb=relaunch); bin/fm-control.sh:604 ("before bin/fm-control.sh $ID $VERB can act", where $VERB is relaunch inside the relaunch transaction); and the three rollback lines this change added at bin/fm-control.sh:746, :775, :778 ("rerun bin/fm-control.sh $ID relaunch") - the rollback case is the most reachable, since a failed replacement launch is exactly when an operator reruns the printed command. Classified ask-user because the remedy is a product-copy decision, not a mechanical repair: fm-control-lib.sh's shared helper has no access to KIND, so either the printed command grows a --note &#34;&lt;why&gt;&#34; placeholder unconditionally, or the two callers that do know KIND (bin/fm-spawn.sh, bin/fm-control.sh) append it and the helper gains a parameter. Pick one; do not harden further.
  • ℹ️ bin/fm-spawn.sh:1743 - On a host without /proc the tmux absence proof returns unreadable with a reason that ends "...once that server is gone there is no such seat, and no route" (bin/fm-agent-process-lib.sh:147-149). The parenthetical this branch appended then wraps that reason with "that reason is what would have to change before bin/fm-control.sh <id> relaunch --note &fix(watcher): make check wakes lossless via watcher-side suppression #34;<why>&fix(watcher): make check wakes lossless via watcher-side suppression #34; can reclaim the task", so one message tells a macOS operator both "no route" and "relaunch reclaims it once the reason changes" - and the reason is the absent /proc, which cannot change on that seat. docs/agent-control.md:113-114 names macOS as a real seat, and docs/agent-control.md:178 states the branch's own invariant ("Every such refusal names a command that acts on the state it read, spelled as the operator can run it"), which this one does not satisfy there: relaunch on that host refuses for the identical reason, forever. Concrete sequence: on macOS, a tmux ship task whose window and server were destroyed - bin/fm-spawn.sh &lt;id&gt; --relaunch prints this message; the operator runs the named bin/fm-control.sh &lt;id&gt; relaunch --note &#34;&lt;why&gt;&#34; and gets the same /proc refusal through bin/fm-control.sh:607. Sibling sites where the same wrapper is printed over the same reason: bin/fm-control.sh:607 (exit, via $VERB_RETRY). Both sentences were added by this branch's fix rounds (9350468, re-worded in 3d3ac1f), not by the author's original commit. Mitigating: the reason text itself does name the only real route (a seat that still addresses the recorded tmux server), so the operator is not left with nothing - which is why this is informational rather than a blocker. Remedy is message copy only, so it needs the author's call rather than a mechanical fix: either scope the "what would have to change" clause to the proofs that can change (an agent still in the worktree, an unreadable process table) and omit it when the reason is a missing /proc, or drop the clause and let the reason stand alone. Add no component and no conditional message machinery beyond that choice.

🔧 Fix applied.
3 issues (1 warning, 2 infos) still open:

  • ℹ️ tests/fm-control-relaunch.test.sh:1861 - make_nm_recorder (tests/fm-control-relaunch.test.sh:1859-1868) stages a fake no-mistakes on PATH and assert_tmux_reclaims then greps nm-calls for run|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI - fm-control.sh, fm-spawn.sh, fm-control-lib.sh, and bin/backends/tmux.sh contain no call to no-mistakes or an axi run verb - so nm-calls is never created and the guard cannot fail for the right reason. It also shadows only the no-mistakes name, not the axi entrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gate survives), the HEAD comparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add the axi name, or drop it, since the surrounding assertions carry the requirement.
  • ℹ️ bin/fm-control.sh:1063 - interrupt is the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint reads ambiguous makes bin/fm-control.sh &lt;id&gt; interrupt die with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, while exit (bin/fm-control.sh:605) and relaunch (bin/fm-spawn.sh:1756) through the same state now end with fm_control_unclassified_next_step pointing at bin/fm-peek.sh &lt;id&gt; and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it because interrupt reads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.
  • ⚠️ docs/agent-control.md:180 - The invariant sentence claims more than the code now does. Line 179 names four refusal classes - alive, ambiguous, unreadable, "and so does any endpoint whose absence is not provable" - and line 180 then asserts "Every such refusal names a command that acts on the state it read, spelled as the operator can run it", enumerating only the first three after the colon. The fourth class names no such command. Concrete sequence: a Linux tmux ship task whose window reads missing while an agent process still works in the recorded worktree - bin/fm-spawn.sh &lt;id&gt; --relaunch prints bin/fm-spawn.sh:1743, which ends "(bin/fm-control.sh <id> exit and relaunch refuse it for the same reason)", i.e. it names two commands that explicitly do NOT act on the state; the reason text supplied by bin/fm-control-lib.sh:363 gives prose actions ("stop that process where it runs, or address its tmux server from this seat, then retry") but no spelled command. The same holds for the other three reasons the proof can return: the no-/proc reason (bin/fm-agent-process-lib.sh:147) ends "there is no such seat, and no route"; the unresolvable-worktree reason (bin/fm-agent-process-lib.sh:143) and the backend-load failure (bin/fm-control-lib.sh:358) carry no next step at all. Sibling site where the same class is printed: bin/fm-control.sh:607, whose exit refusal ends "and relaunch refuses for the same reason". This drift was created inside this run's fix rounds: 9350468/3d3ac1f added the line-180 claim while the unprovable-absence refusals still carried a route clause, and 19f6f78 (the round the human explicitly ordered) deleted that clause from bin/fm-spawn.sh:1743 and bin/fm-control.sh:607 without narrowing the doc. Remedy corrects what the change already does rather than extending it: scope line 180 to the three positively-read verdicts and state plainly that an unprovable absence names the reason and whatever route the proof found, up to and including none. Do not re-add a route clause to either message and add no conditional message machinery.

🔧 Fix applied.
3 infos still open:

  • ℹ️ tests/fm-control-relaunch.test.sh:1861 - make_nm_recorder (tests/fm-control-relaunch.test.sh:1859-1868) stages a fake no-mistakes on PATH and assert_tmux_reclaims then greps nm-calls for run|respond|abort|sync|cancel|rerun, claiming it proves "the reclaim must never abort, answer, or start the task's validation run". Nothing in the reclaim path invokes any validation CLI - fm-control.sh, fm-spawn.sh, fm-control-lib.sh, and bin/backends/tmux.sh contain no call to no-mistakes or an axi run verb - so nm-calls is never created and the guard cannot fail for the right reason. It also shadows only the no-mistakes name, not the axi entrypoint the same pipeline is driven through, so a future call through that name would slip past it. The real evidence for the intent's "a parked validation run is left exactly as it was" is already in the same helper and does hold: the status-log assertion (parked at a validation gate survives), the HEAD comparison, and the uncommitted-file check. Noting the weak guard rather than prescribing work: either keep it as a cheap tripwire and add the axi name, or drop it, since the surrounding assertions carry the requirement.
  • ℹ️ bin/fm-control.sh:1063 - interrupt is the one lifecycle refusal of an unattributed endpoint left without the new next-step sentence. A task whose endpoint reads ambiguous makes bin/fm-control.sh &lt;id&gt; interrupt die with "refusing to send a lifecycle key into an unattributed endpoint" and nothing else, while exit (bin/fm-control.sh:605) and relaunch (bin/fm-spawn.sh:1756) through the same state now end with fm_control_unclassified_next_step pointing at bin/fm-peek.sh &lt;id&gt; and a retry. The helper's own contract at bin/fm-control-lib.sh:380-387 deliberately scopes itself to exit, relaunch, and the launch owner, so this is a stated boundary rather than a missed site - noting it because interrupt reads the same state vocabulary in the same file and is the only refusal left in this area with nowhere for the operator to go. No behaviour is wrong and the change is unaffected either way.
  • ℹ️ tests/fm-control-relaunch.test.sh:1916 - The reclaim's most operator-visible new behaviour is unverified: the record's session part moves from the recorded session to the reclaiming seat's. assert_tmux_reclaims matches the rebound record with case &#34;$window&#34; in *&#34;:fm-$id&#34;) (tests/fm-control-relaunch.test.sh:1916), which passes identically whether the field reads firstmate:fm-rl60 (what the code does) or fmses:fm-rl60 (the recorded session). Concrete trace: add_ship_task records window=fmses:fm-rl60 (tests/fm-control-relaunch.test.sh:204,215), run_control now unsets TMUX (:242), so bin/fm-spawn.sh:3470's fm_backend_tmux_container_ensure resolves firstmate and bin/fm-spawn.sh:4768 writes window=firstmate:fm-rl60 - a value no assertion reads. The change states this as a contract in three places (docs/agent-control.md:116-117 "a tmux reclaim can move the task into the reclaiming seat's session", :130 "on tmux that binding includes the session", bin/fm-control.sh:66-69 "tmux: on the server and in the session THIS seat addresses"), and it is not cosmetic: a regression that pinned the recorded session would make tmux new-window -t &#34;fmses:&#34; fail outright for the two cases the suite covers where that session no longer exists (tests/fm-control-relaunch.test.sh:1962 session-missing, :1971 server-dead), re-creating the deadlock the change exists to remove. The stub cannot catch it either - its new-window branch records only the -n name into created-windows (tests/fm-control-relaunch.test.sh:158-174), never the -t target - so the record field is the only observable. Sibling site where the same invariant is unpinned: tests/fm-control-relaunch.test.sh:1986, where test_spawn_relaunch_alone_reclaims_a_gone_tmux_endpoint checks the worktree field but never reads window at all. Remedy is mechanical and test-only: assert the rebound window equals firstmate:fm-$id in assert_tmux_reclaims, and read the same field in the launch-owner case.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
An operator whose tmux task window was destroyed reclaims the task with bin/fm-control.sh <id> relaunch --note, and a real agent comes back up in the same worktree ✅ pass live live-reclaim-transcript.md section 3 and reclaimed-window-pane.txt: relaunch rc=0, a new tmux window id (@2 replacing the destroyed @1), record rebound, and a real Claude running in it
The reported deadlock reproduces before the fix: on the base commit, relaunch, fm-spawn --relaunch and exit all refuse the same destroyed endpoint and leave no route ✅ pass live live-reclaim-transcript.md section 2 - base-commit scripts extracted with git archive 30ef650 and driven against the same live lab task; all three refuse
bin/fm-control.sh <id> exit on a destroyed tmux endpoint reports endpoint-gone instead of the 'reconcile the task first' dead end ✅ pass live live-reclaim-transcript.md section 3: endpoint-gone labscout1 harness=claude backend=tmux endpoint=firstmate:fm-labscout1 ..., rc=0
The reclaim preserves the work the incident said was safe: committed branch state, an uncommitted file, and a status log parked at a validation gate ✅ pass live final-lab-state.txt: HEAD ccf2864 before == after, uncommitted.txt still present and untracked, the 'parked at a no-mistakes review gate' status line still in state/labscout1.status after four reclaim…
Adversarial: an agent still alive in the recorded worktree stops the reclaim - neither verb launches a second agent, and the refusal names the live process ✅ pass live live-reclaim-transcript.md section 1: both verbs refuse naming 'agent process 1918852 claude', no window created, PID 1918852 still alive afterwards
Adversarial: a foreign window already holding fm-<id> in this seat's session is refused rather than closed or duplicated, and the rollback's named rerun works once it is cleared ✅ pass live live-reclaim-transcript.md section 6: refusal names bin/fm-peek.sh, decoy window @5 left open, no second window; after closing it the same command succeeded and rebound to other:fm-labscout1 (@6)
Every refusal names a command the operator can actually run: the alive-endpoint refusal names relaunch --note "<why>" and that command, run verbatim, replaces the agent ✅ pass live live-reclaim-transcript.md section 4: fm-spawn --relaunch refused naming bin/fm-control.sh labscout1 relaunch --note &#34;&lt;why&gt;&#34;; running it succeeded (rc=0) and reused the surviving window @2 rather th…
bin/fm-spawn.sh <id> --relaunch on its own reaches the same verdict and reclaims a destroyed endpoint, so the two commands cannot disagree ✅ pass live live-reclaim-transcript.md section 5: spawned labscout1 harness=claude kind=scout window=firstmate:fm-labscout1 worktree=... rc=0, new window @3, recorded worktree kept
The tmux create guard refuses when the window inventory cannot be read, rather than assuming the name is free and opening a duplicate ✅ pass live bash tests/fm-backend-tmux-smoke.test.sh against a real tmux server on a private socket: 'refuses an unreadable window inventory rather than assuming the name is free' and 'refuses a duplicate witho…
  • bin/fm-lab-home.sh create $LAB + tmux -L fm-lab private socket - disposable lab home and isolated tmux server
  • bin/fm-brief.sh labscout1 labproj --scout then bin/fm-spawn.sh labscout1 $LAB/projects/labproj --scout --backend tmux --harness claude - real live spawn, real treehouse worktree, real Claude agent
  • tmux -L fm-lab rename-window then bin/fm-spawn.sh labscout1 --relaunch --harness claude and bin/fm-control.sh labscout1 exit - adversarial: endpoint reads missing while a real Claude still works in the worktree
  • git archive 30ef650 | tar -x then $LAB/base/bin/fm-control.sh labscout1 relaunch --note, $LAB/base/bin/fm-spawn.sh labscout1 --relaunch, $LAB/base/bin/fm-control.sh labscout1 exit - base-commit deadlock reproduction
  • bin/fm-control.sh labscout1 exit on the destroyed endpoint - expects endpoint-gone
  • bin/fm-control.sh labscout1 relaunch --note &#34;the tmux window was destroyed; pick the work back up&#34; - the reclaim
  • tmux -L fm-lab list-windows -a, cat $LAB/state/labscout1.meta, git -C $WT rev-parse HEAD, git -C $WT status --porcelain, cat $LAB/state/labscout1.status - post-reclaim state verification
  • bin/fm-spawn.sh labscout1 --relaunch --harness claude against the live reclaimed endpoint, then the bin/fm-control.sh labscout1 relaunch --note &#34;&lt;why&gt;&#34; it named, run verbatim
  • bin/fm-spawn.sh labscout1 --relaunch --harness claude alone after a second window destroy - launch owner reclaims on its own
  • reclaim driven from a pane in session other that already held a decoy fm-labscout1 window - duplicate-name refusal, then rerun of the command the rollback named
  • bash tests/fm-backend-tmux-smoke.test.sh - real-tmux create/duplicate/unreadable-inventory guards
  • bash tests/fm-control-relaunch.test.sh
  • bash tests/fm-control.test.sh
⚠️ **Document** - 1 info
  • ℹ️ bin/backends/tmux.sh:112 - Judgment call, deliberately left as code comments rather than prose. The change makes fm_backend_tmux_create_task refuse when a session's window inventory cannot be read (previously a failed read silently passed for "no duplicate"), and that helper serves ordinary tmux spawns as well as the new reclaim. No prose surface owns spawn-time window-create refusals today: docs/tmux-backend.md covers setup, liveness, and composer behavior only, and docs/agent-control.md:119 documents the duplicate-name refusal solely in the reclaim context. The new refusal's rationale and text live in the function's own header comment, which is the correct owner for a local safety invariant, so I added no prose and created no new section (the placement policy forbids inventing a surface to close a perceived gap). Flagging it in case a future operator-facing note about tmux spawn refusals is wanted; nothing is currently wrong or contradicted.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…dicting refusals

A task whose recorded endpoint was destroyed rather than stopped could only be
recovered by hand-editing its record: fm-spawn --relaunch sent the operator to
fm-control exit, and exit refused a tmux `missing` endpoint with no command that
could act on it. Herdr already had a reclaim path; tmux was left deadlocked
because a task record carries no tmux socket identity.

- The control plane's one absence proof (fm_control_endpoint_absence_verdict)
  now proves a tmux endpoint gone by what a duplicate launch would collide with:
  no process of this user that classifies as a verified harness is working in
  the recorded worktree (new fm_agent_process_worktree_scan, /proc on Linux,
  lsof and ps elsewhere). An agent found there, or an unreadable process table,
  still refuses, and ambiguous or unreadable endpoint states refuse as before.
- fm-spawn --relaunch rebinds a proven-gone tmux task to one fresh window in the
  same worktree under the same task identity, closing that window again if the
  record is not republished.
- Refusals now name a command that works for the state they read: `alive`
  points at fm-control relaunch, `ambiguous`/`unreadable` at fm-peek and a
  retry, an unproven absence at the relaunch that reclaims it once that changes,
  and the relaunch rollback messages say to rerun relaunch.
- A no-mistakes ship's relaunch progress note tells the replacement to reattach
  to a validation run already under way rather than start or abort one; the
  control plane itself never touches the run.

Evidence (tests/fm-control-relaunch.test.sh, same selected tests on the base
commit 30ef650 vs this branch):
  before: not ok - exit must report a proven-gone tmux endpoint rather than refuse it
  before: not ok - the refusal should name the one command that replaces a live agent (missing: 'fm-control.sh rl15 relaunch')
  before: not ok - the refusal must not send the operator to a verb that refuses this state too (unexpected: 'fm-control.sh rl63 exit')
  after:  ok - tmux: a destroyed window is reclaimed in the same worktree under the same task
  after:  ok - fm-spawn --relaunch: refuses a live endpoint and names the command that replaces its agent
  after:  ok - reclaim: an unclassifiable endpoint is still refused, so two agents cannot share one
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[High risk] Adds process-table scanning to task recovery logic.

The PR should not merge until tmux reclaim abort cleanup is bound to the window it created rather than a shared name.

Reviews (1) · Last reviewed commit: "no-mistakes(document): scope Herdr rebin..."

Comment thread bin/fm-spawn.sh
WID=$(fm_backend_tmux_create_task "$SES" "$W" "$WT") || exit 1
T="$SES:$W"
WT_TARGET=$WID
TMUX_REBIND_ABORT_TARGET=$T

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Abort targets a shared name If two homes reclaim the same task ID on one tmux server at the same time, both can pass the window check and create fm-<id> windows. This line saves the shared name for abort cleanup instead of the stable window ID returned by creation. If one reclaim then fails before publishing its record, cleanup can close the other home's window or leave its own behind, and later commands cannot reliably distinguish the two. Bind cleanup to the window this attempt created.

Knowledge Base Used: Backend adapters

|| fail "the reclaim must open exactly one window ($what)"
window=$(meta_field "$dir" "$id" window)
case "$window" in
*":fm-$id") ;;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Rebound session goes untested The test checks only the :fm-<id> suffix of window=, although reclaim is meant to bind the task to the session this seat selects. These test commands unset TMUX, so that session should be firstmate. A regression that kept the old session would pass this assertion, even when that session no longer exists and reclaim cannot complete. Assert the full session:window value.

Comment on lines +1934 to +1936
if [ -f "$dir/fake/nm-calls" ] && grep -Eq '(^|[[:space:]])(run|respond|abort|sync|cancel|rerun)([[:space:]]|$)' "$dir/fake/nm-calls"; then
fail "the reclaim must never abort, answer, or start the task's validation run ($what): $(cat "$dir/fake/nm-calls")"
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Validation tripwire misses actions The new recorder shadows only no-mistakes, and this check rejects only six named verbs. A future reclaim call through the axi entrypoint, or an action outside that list, would still pass the test's claim that reclaim never acts on a parked validation run. Capture the entrypoints and actions the assertion intends to protect, or narrow its stated guarantee.

@keenvc

keenvc commented Sep 28, 2026

Copy link
Copy Markdown
Author

Superseded by #6007, which carries the same change rebased onto current main at b070249. This branch's head 642034d is the pre-rebase lineage and could not be updated without a force-push, so the work was republished on a new branch instead. Closing this in favour of #6007; the branch here is left intact and nothing is discarded.

@keenvc keenvc closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant