Skip to content

fix(bin): refuse teardown of a worktree reassigned to another task - #2

Merged
LeonidShamis merged 5 commits into
mainfrom
fm/firstmate-teardown-reassigned-worktree
Sep 5, 2026
Merged

LeonidShamis merged 5 commits into
mainfrom
fm/firstmate-teardown-reassigned-worktree

Conversation

@LeonidShamis

Copy link
Copy Markdown
Owner

Intent

Fix: bin/fm-teardown.sh can kill another task's agent when it is re-run after a partial failure, because it terminates processes in, and returns, the worktree path recorded in the task's meta without verifying the worktree still belongs to that task.

Observed 2026-09-05 on the Herdr backend, but the defect is backend-independent:

  1. fm-teardown.sh A returned A's treehouse worktree to the pool ("Worktree returned to pool."), then failed the Herdr pane close ("herdr presentation cleanup target is the captain's active tab; refusing a close that cannot preserve focus") and exited non-zero with "retaining every durable task record", so state/A.meta (with worktree=) stayed on disk.
  2. fm-spawn.sh B ran; treehouse handed task B the same now-free slot-1 worktree.
  3. fm-teardown.sh A was re-run. It terminated every lingering process in the recorded slot-1 path ("Terminated lingering processes: bash, claude, ..."), which were now B's agent, and returned the worktree again. B died with no status line and its branch was deleted by the pool return, with no warning or refusal.

Required behavior:

  1. Before terminating processes in or returning a worktree, teardown must prove the worktree is still bound to the task being torn down. Use the strongest signal available: a task-owned marker inside the worktree written at spawn (there are already per-task worktree files such as the hook token pointers; read bin/fm-spawn.sh and bin/fm-teardown.sh to pick or add one that every backend and harness writes), the worktree's current branch matching the task's fm/ branch or the recorded head, and/or treehouse's lease state. If the worktree is not provably this task's, teardown must refuse the process kill and the pool return loudly, naming the task it appears to belong to when that can be read, and must still be safe to re-run once the records are reconciled.
  2. Make the partial-failure path converge: when the pool return has already succeeded but a later step (pane close, presentation cleanup) fails, teardown must record that the worktree is no longer this task's (for example by clearing or annotating worktree= in the meta, per the meta contract in the script headers) so a re-run skips the worktree steps instead of repeating them. Read the header of bin/fm-teardown.sh (it documents the treehouse-return retry contract) and state/.meta field ownership in the producing script headers before choosing the mechanism; keep each contract stated once, in its owner.
  3. Never discard unlanded work: the new refusal must not force, stash, or delete anything.
  4. Regression tests: extend tests/fm-teardown.test.sh and/or tests/fm-teardown-endpoint-safety.test.sh (read them fully first and follow their fixture pattern) with: (a) a re-run after a simulated post-return failure does not kill processes in or return a worktree that a second task has since acquired, and refuses loudly; (b) an ordinary teardown still terminates and returns its own worktree; (c) a re-run after the meta has been reconciled completes cleanly. Existing tests must keep passing.
  5. Update bin/fm-teardown.sh's header for the new ownership check and the converged re-run contract, and any --help text that describes the return step (fm-teardown.sh has no --help handler, so its header Usage block is the owner).

Firstmate repo rules that apply to this change (it touches firstmate's own shared tracked material in bin/): follow .agents/skills/firstmate-coding-guidelines/SKILL.md - one sentence per line in Markdown, plain dashes never em dashes, bin/*.sh must pass bin/fm-lint.sh (pinned shellcheck 0.11.0), colocate regression tests in tests/ named .test.sh following tests/lib.sh conventions and extend an existing test file where one already covers the subject, tests must exercise behavior through the executable and never assert on implementation source bytes, never add an agent co-author to commits, read the header comment of every script touched first because each header is the single owner of that script's contract and must be updated when the contract changes, keep the change minimal and focused on this defect and do not refactor surrounding code, and do not touch data/, state/, config/, or projects/ of any firstmate home.

Implementation decisions made while doing the work, which a reviewer reading only the diff would not know:

  • The chosen ownership marker is a new gitignored .fm-task-owner file written by bin/fm-spawn.sh into every non-secondmate task worktree, containing "task=" and "state=". The existing per-task worktree files (.fm-grok-turnend, .fm-kimi-turnend, .claude/settings.local.json, .opencode plugins) were all rejected as the signal because each is written by only one harness; the task explicitly asked for one that every backend and harness writes, so a new one was added. It is registered through the existing exclude_path helper so it stays out of git's view, and it was also added to validate_worktree_teardown_safety's untracked-file filter as belt-and-braces in case exclude_path silently no-ops.
  • Ownership evidence is deliberately graded rather than a single check. The marker is consulted FIRST and is decisive, because every spawn (including a relaunch) rewrites it, so it names whichever task took the worktree last. An earlier iteration checked the cross-task meta claim first; that was proven wrong by the real-Herdr end-to-end suite, where a live task record legitimately carries a stale worktree= line pointing at a slot another task has since taken. Marker-first is the corrected ordering.
  • A marker only counts while the task it names still has a record where the marker says that task's records live (marker_owner_still_recorded). This staleness rule exists so a leftover marker from an already finished task can never permanently wedge an unrelated teardown; it was also required to make the real-Herdr presentation e2e pass.
  • The marker is removed alongside the task's other per-task worktree files, immediately BEFORE the pool return, not after. Removing it after the return would race a spawn that had already re-leased the slot and could delete that new task's marker. Removing it before is safe because a failed return leaves the slot still leased to this task, so no other task can hold it and the lost marker cannot mislead a re-run.
  • With no live marker (a task spawned before the marker existed), teardown falls back to: another live record in this home claiming the same canonical path, then the worktree sitting on another live task's fm/ branch, then accepting its own fm/ branch. When no signal is available either way (for example a hand-made or non-git worktree), it proceeds on the recorded path with a warning rather than refusing. That deliberate permissiveness preserves compatibility for tasks already running without a marker and for existing test fixtures; it is documented in the script header as the compatibility boundary. Refusal is reserved for evidence that the worktree belongs to someone else.
  • The ownership check applies under --force too. --force is authority to discard THIS task's work, never another task's.
  • Re-run convergence uses a new worktree_returned=1 line in state/.meta rather than clearing worktree=. The recorded path is kept for diagnostics and for the completion message. bin/fm-teardown.sh is now the owner of that field and its header documents it; bin/fm-spawn.sh's preserve_relaunch_meta owned-key list gained worktree_returned so a relaunch drops it when it rebinds a worktree. A successful teardown removes the metadata entirely, so the field is only ever visible between a partial failure and its re-run.
  • record_worktree_returned rewrites the meta through mktemp + cp -p (to preserve the original permissions) + awk + mv -f, following the existing atomic meta-rewrite pattern in bin/fm-pr-check.sh. awk is used rather than grep so a no-match cannot be mistaken for a failure.
  • When the return succeeded but the convergence record could not be written, teardown fails loudly instead of leaving the two records disagreeing silently.
  • On a converged re-run the leaked-process reap is narrowed to the task's own tasktmp root only, since the worktree is no longer this task's. The Orca branch was restructured minimally so the terminal kill still happens on a converged re-run while the worktree-specific steps are skipped.

Test evidence and known environment limitation:

  • Four regression tests were added: three in tests/fm-teardown.test.sh (ordinary teardown still reaps and returns its own worktree; a re-run after a genuine post-return failure refuses loudly and neither kills the successor task's process nor returns or branch-deletes its worktree, covering both the cross-record and the marker signals; a re-run after the records are reconciled completes without a second return) and one in tests/fm-spawn-worktree-settle.test.sh (spawn marks the SETTLED worktree, not the transient stale path, and the marker is invisible to git). The post-return failure is driven for real, not simulated in the test's own logic: an unreadable grok turn-end token makes remove_grok_turnend_auth fail after the pool return, which is the same shape as the observed Herdr pane-close failure.
  • A new add_lsof_cwd_map fixture stub reports live processes in lsof -Fpn form from a pid/cwd map and drops pids that have already exited, so the reap's identity rechecks behave as they do against real lsof.
  • bin/fm-lint.sh passes (shellcheck 0.11.0, actionlint 1.7.12). All 59 test suites that reference fm-teardown.sh or fm-spawn.sh were run.
  • One test, tests/fm-teardown.test.sh's leaked-process-reap case, fails on this development host because lsof is not installed there. It was verified to fail identically on unmodified main from a clean temporary worktree, so it is pre-existing and unrelated to this change.
  • tests/fm-backend-herdr-presentation-e2e.test.sh (real Herdr, ~5 minutes) initially failed with the first iteration of the ownership check and was used to drive the marker-first ordering and the marker staleness rule; it passes on unmodified main and passes with the final change.

What Changed

  • bin/fm-spawn.sh now writes a gitignored .fm-task-owner marker (task=<id>, state=<state dir>) into every non-secondmate task worktree on every spawn, and bin/fm-teardown.sh checks that marker (falling back to another live task record claiming the same path, then the worktree's fm/<id> branch) before killing processes in, resetting, or returning the recorded worktree; when the evidence binds the worktree to a different task, teardown refuses loudly, names that task, and forces, stashes, or discards nothing, including under --force.
  • After a successful treehouse pool return or Orca worktree removal, teardown records worktree_returned=1 in state/<id>.meta so a rerun after a later failure (pane close, presentation cleanup) skips every worktree step and reaps only the task's own tasktmp; a failed Orca removal now aborts with the record unmarked, and bin/fm-spawn.sh --relaunch and bin/fm-control.sh's safe checkpoint refuse to relaunch a task whose worktree was already returned.
  • Regression tests in tests/fm-teardown.test.sh, tests/fm-spawn-worktree-settle.test.sh, and tests/fm-control-relaunch.test.sh cover an ordinary teardown still reaping and returning its own worktree, a rerun after a real post-return failure refusing to touch a successor task's worktree, a rerun after reconciliation completing cleanly, spawn marking the settled worktree, and the relaunch refusal; script headers and docs/agent-control.md / docs/architecture.md document the ownership check and rerun-convergence contract.

Risk Assessment

✅ Low: The re-review confirms each fix-round change against the code: the marker-signal test can now only pass through the marker path, the Orca removal failure fails closed with the record retained as the header states, and relaunch refuses a returned-marked record in both fm-spawn and fm-control before touching the agent, leaving no reachable path to the original cross-task teardown that I could substantiate.

Testing

Ran the three touched suites plus the endpoint-safety and Orca suites; all new and existing cases pass except five real-lsof reap cases in tests/fm-teardown.test.sh that fail identically on the base commit because lsof is not installed on this host. Captured end-user CLI transcripts of bin/fm-teardown.sh for the incident scenario on both target and base: on target the rerun refuses loudly, names the task that now owns the worktree from both the cross-record and marker signals, and returns the slot once, while on base the rerun reaps the successor's process and returns the slot twice. Also captured the persisted meta showing worktree_returned=1 after a partial failure and a reconciled rerun completing without a second return.

Evidence: Evidence index (what each file shows)

Source: Evidence index (what each file shows)

# Test evidence: teardown never tears down a worktree reassigned to another task

Branch fm/firstmate-teardown-reassigned-worktree, target 7467a90 against base c03bfbe.
Host note: lsof is not installed on this host, which is why the five real-lsof reap cases fail identically on base and target.

## Product-level transcripts (bin/fm-teardown.sh stderr/stdout, pool return log, persisted meta)

- transcripts/reassigned-worktree-refusal/ (target): first.stderr is the partial run (return succeeded, later step failed), second.stderr is the rerun refusing because task-x2's record claims the worktree, third.stderr is the rerun refusing because the worktree's .fm-task-owner marker names task-x2, treehouse.log shows exactly one pool return across all three runs.
- transcripts-base-c03bfbe/reassigned-worktree-refusal/ (base, same scenario): second.stderr shows the rerun reaping the successor task's process and treehouse.log shows the slot returned twice, which is the reported defect.
- transcripts/reconciled-rerun/ (target): second.stderr shows the rerun skipping every worktree step, second.stdout shows it completing, treehouse.log shows one return.
- transcripts-base-c03bfbe/reconciled-rerun/ (base): treehouse.log shows the rerun returning the already-returned slot a second time.
- transcripts/own-worktree-return/ (target): an ordinary teardown still reaps its own leaked process and returns its own worktree.
- converged-meta-transcript.txt: state/task-x1.meta before and after a real partial failure, showing worktree_returned=1 recorded, and an unreconciled rerun skipping the worktree steps with still one pool return.

## Suite logs

- fm-teardown.test.log: full tests/fm-teardown.test.sh run on target (stops at the pre-existing lsof failure after the new cases pass).
- fm-teardown-ownership-cases.log and fm-teardown-ownership-cases-base-c03bfbe.log: the three new regression cases on target (pass) and on base scripts (the two defect cases fail, the ordinary case passes).
- fm-teardown-remaining-cases-target.log and fm-teardown-remaining-cases-base-c03bfbe.log: the ten reap cases that follow the lsof failure, run one at a time on target and base, with identical results.
- fm-control-relaunch.test.log, fm-spawn-worktree-settle.test.log, fm-teardown-endpoint-safety.test.log, fm-backend-orca.test.log: full runs of the other touched or adjacent suites, all passing.

## Helpers used to produce the evidence

- run-selected-teardown-tests.sh and run-one-teardown-test.sh: run named test functions of tests/fm-teardown.test.sh individually and copy each case sandbox's transcripts out before cleanup.
- show-converged-meta.sh: drives one real partial-failure teardown with the suite's fixtures and prints the persisted meta.
Evidence: Target: rerun refusal transcript, cross-record signal (bin/fm-teardown.sh stderr)

Source: Target: rerun refusal transcript, cross-record signal (bin/fm-teardown.sh stderr)

REFUSED: worktree /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt is no longer task task-x1's: task task-x2's own record claims the same worktree. Killing its processes or returning it would destroy another task's work. Reconcile the task records first (clear the stale worktree binding on task-x1), then rerun teardown; nothing was killed, returned, or discarded.

REFUSED: worktree /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt is no longer task task-x1's: task task-x2's own record claims the same worktree.
Killing its processes or returning it would destroy another task's work.
Reconcile the task records first (clear the stale worktree binding on task-x1), then rerun teardown; nothing was killed, returned, or discarded.
Evidence: Target: rerun refusal transcript, .fm-task-owner marker signal (bin/fm-teardown.sh stderr)

Source: Target: rerun refusal transcript, .fm-task-owner marker signal (bin/fm-teardown.sh stderr)

REFUSED: worktree /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt is no longer task task-x1's: it is marked as task task-x2's in /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/state. Killing its processes or returning it would destroy another task's work. Reconcile the task records first (clear the stale worktree binding on task-x1), then rerun teardown; nothing was killed, returned, or discarded.

REFUSED: worktree /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt is no longer task task-x1's: it is marked as task task-x2's in /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/state.
Killing its processes or returning it would destroy another task's work.
Reconcile the task records first (clear the stale worktree binding on task-x1), then rerun teardown; nothing was killed, returned, or discarded.
Evidence: Target: pool return log across the partial run and both refused reruns (exactly one return)

Source: Target: pool return log across the partial run and both refused reruns (exactly one return)

return --force /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt

return --force /tmp/fm-teardown-tests.Rx7PUW/reassigned-worktree-refusal/wt
Evidence: Base c03bfbe: same scenario reproduces the defect (rerun reaps the successor's process, slot returned twice)

Source: Base c03bfbe: same scenario reproduces the defect (rerun reaps the successor's process, slot returned twice)

second.stderr: teardown: reaping leaked worktree process(es) for task-x1: 423829 treehouse.log: return --force /tmp/fm-teardown-tests.E8MSjo/reassigned-worktree-refusal/wt return --force /tmp/fm-teardown-tests.E8MSjo/reassigned-worktree-refusal/wt

return --force /tmp/fm-teardown-tests.E8MSjo/reassigned-worktree-refusal/wt
return --force /tmp/fm-teardown-tests.E8MSjo/reassigned-worktree-refusal/wt
Evidence: Persisted state: meta after a real partial failure and the converged rerun

Source: Persisted state: meta after a real partial failure and the converged rerun

## state/task-x1.meta after the partial failure (record retained and converged) window=firstmate:fm-task-x1 endpoint_task_id=task-x1 worktree=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt project=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/project kind=ship mode=no-mistakes worktree_returned=1 ## rerun without reconciling: exit and stderr exit=1 teardown: worktree /tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt was already returned for task-x1 by an earlier run; skipping every worktree step ## treehouse calls after the unreconciled rerun (must still be one) return --force /tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt

## state/task-x1.meta before teardown
window=firstmate:fm-task-x1
endpoint_task_id=task-x1
worktree=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt
project=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/project
kind=ship
mode=no-mistakes
## worktree marker before teardown (.fm-task-owner)
(none: fixture worktree was not spawned by fm-spawn.sh)
## first teardown exit=1 (pool return succeeded, later step failed)
## treehouse calls so far
return --force /tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt
## state/task-x1.meta after the partial failure (record retained and converged)
window=firstmate:fm-task-x1
endpoint_task_id=task-x1
worktree=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt
project=/tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/project
kind=ship
mode=no-mistakes
worktree_returned=1
## rerun without reconciling: exit and stderr
exit=1
teardown: worktree /tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt was already returned for task-x1 by an earlier run; skipping every worktree step
## treehouse calls after the unreconciled rerun (must still be one)
return --force /tmp/fm-teardown-tests.O3KPz8/converged-meta-demo/wt
Evidence: Target: reconciled rerun completes and skips the worktree steps

Source: Target: reconciled rerun completes and skips the worktree steps

stderr: teardown: worktree .../reconciled-rerun/wt was already returned for task-x1 by an earlier run; skipping every worktree step stdout: teardown task-x1 complete (window firstmate:fm-task-x1, worktree .../reconciled-rerun/wt) treehouse.log: one return only

/tmp/fm-teardown-tests.jic79R/reconciled-rerun/project: already current
teardown task-x1 complete (window firstmate:fm-task-x1, worktree /tmp/fm-teardown-tests.jic79R/reconciled-rerun/wt)
Backlog: task-x1 just finished. Run tasks-axi done task-x1 --pr PR_URL, then run tasks-axi ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due.
Evidence: Target: ordinary teardown still reaps its own process and returns its own worktree

Source: Target: ordinary teardown still reaps its own process and returns its own worktree

teardown: reaping leaked worktree process(es) for task-x1: 414987 teardown task-x1 complete (window firstmate:fm-task-x1, worktree .../own-worktree-return/wt) treehouse.log: return --force .../own-worktree-return/wt

teardown: reaping leaked worktree process(es) for task-x1: 414987
Evidence: New regression cases: pass on target, fail on base scripts

Source: New regression cases: pass on target, fail on base scripts

target: ok - an ordinary teardown still reaps and returns the worktree that is its own ok - a rerun after a post-return failure refuses loudly instead of tearing down a reassigned worktree ok - a rerun after the records are reconciled completes without repeating the worktree steps base c03bfbe: ok - an ordinary teardown still reaps and returns the worktree that is its own not ok - reassigned-worktree-refusal: the rerun killed the successor task's process not ok - reconciled-rerun: the rerun returned the already-returned worktree again

ok - an ordinary teardown still reaps and returns the worktree that is its own
[test_ordinary_teardown_reaps_and_returns_its_own_worktree] exit=0
not ok - reassigned-worktree-refusal: the rerun killed the successor task's process
[test_rerun_after_partial_failure_refuses_a_reassigned_worktree] exit=1
not ok - reconciled-rerun: the rerun returned the already-returned worktree again
[test_rerun_after_reconciled_records_completes_cleanly] exit=1
Evidence: Full tests/fm-teardown.test.sh run on target

Source: Full tests/fm-teardown.test.sh run on target

ok - local-only worktree with HEAD on a fork remote is torn down and the home summary is refreshed
ok - an ordinary teardown still reaps and returns the worktree that is its own
ok - a rerun after a post-return failure refuses loudly instead of tearing down a reassigned worktree
ok - a rerun after the records are reconciled completes without repeating the worktree steps
warning: lsof is unavailable; cannot resolve the tmux pane process group for task-x1
ok - teardown prompts tasks-axi backlog refresh when compatible
warning: lsof is unavailable; cannot resolve the tmux pane process group for task-x1
ok - teardown honors config/backlog-backend=manual even when tasks-axi is compatible
ok - local-only worktree with truly unpushed work is refused (safety preserved)
ok - local-only worktree with work merged into local main is torn down (no regression)
ok - no-mistakes worktree with HEAD on origin is torn down (no regression)
ok - no-mistakes worktree with genuinely unlanded work is refused (safety preserved)
ok - local-only worktree with unpushed work is torn down under --force (escape hatch)
ok - teardown completes when an exact busy-state sidecar is already absent
ok - herdr teardown removes pane-owned escalation dedupe state
ok - herdr flat teardown refuses before returning the isolated copy under lock contention and the retry completes cleanly
ok - herdr flat teardown never erases records when pane presence is unparseable
ok - herdr flat teardown preflight refuses before every destructive change
ok - forced secondmate teardown preflights every Herdr child before cleanup mutation
ok - forced secondmate teardown holds every descendant lifecycle and metadata lock
ok - forced secondmate teardown retains Herdr child identity until exact pane disappearance
ok - forced teardown retains a nested secondmate home and its grandchild's Herdr identity when the grandchild close is unconfirmed
ok - herdr projection teardown retires its journal only after confirming the exact recorded pane is gone
ok - herdr projection teardown retains every record when post-close presence is unknown
ok - herdr projection teardown surfaces failed focus restoration without turning confirmed cleanup into a hard failure
ok - squash-merged + deleted-branch worktree (PR merged) is torn down (the fix)
ok - squash-merged PR accepts a local HEAD that is an ancestor of the final PR head
ok - teardown discovers a merged PR by branch name and tears down when no pr= was ever recorded
ok - squash-merged PR accepts replayed unpushed local patches contained in the PR head
ok - merged PR does not allow teardown after a later local commit
ok - fm-pr-check does not refresh PR head after HEAD moves
ok - fm-pr-check records the remote PR head when the local worktree lags
ok - worktree whose content already landed in the default branch is torn down (content fallback)
ok - content fallback refreshes origin default before comparing trees
ok - dirty worktree is refused even when its committed work has landed (dirty always wins)
ok - gh lookup error with content not in default refuses (fail-safe)
ok - provably-stale worktree index.lock (old, no live holder) is cleared and teardown succeeds
ok - live-held worktree index.lock is never removed and teardown refuses
ok - lsof errors leave worktree index.lock in place and refuse teardown
ok - stale lock cleanup rechecks and refuses dirty worktree before return
ok - normal repo index.lock is resolved from the worktree and cleared when stale
ok - lock mtime read failures leave worktree index.lock in place and refuse teardown
ok - transient index.lock cleared after first failed return is retried successfully without force-remove
ok - persistent index.lock exhausts retries and refuses without force-removing the lock
ok - empty retry wait overrides use the default without aborting teardown
ok - fractional legacy retry wait remains supported without arithmetic
ok - a task's own parked no-mistakes run is aborted, not orphaned, before the worker is removed
ok - teardown refuses before reap or removal when a task-owned run remains parked
ok - a different run cannot confirm the targeted abort
ok - empty post-abort status is not accepted as confirmation
ok - the CLI's exact run-not-found signal confirms completion
ok - a parked run on another branch is never aborted by this task's teardown (ownership is precise)
ok - a task-owned autonomous running step is left alone rather than aborted
not ok - leaked-process-reap: leaked worktree process survived teardown

real	0m55.176s
user	0m17.460s
sys	0m28.999s
exit=1
Evidence: Reap cases after the lsof failure: target vs base (identical, lsof missing on host)

Source: Reap cases after the lsof failure: target vs base (identical, lsof missing on host)

not ok - leaked-process-reap: leaked worktree process survived teardown
[test_leaked_worktree_process_is_reaped] exit=1
not ok - leaked-tasktmp-reap: leaked tasktmp process survived teardown
[test_leaked_tasktmp_process_is_reaped] exit=1
ok - missing lsof falls back to reaping the tmux pane process group
[test_lsof_absent_reaps_tmux_process_group] exit=0
ok - an erroring lsof scan refuses teardown and preserves the task
[test_lsof_error_refuses_before_removal] exit=0
ok - a reused pid with a different start time is never force-killed
[test_reused_pid_identity_is_not_force_killed] exit=0
not ok - exec-changed-process: teardown should succeed: expected exit 0, got 1
[test_exec_changed_process_is_still_reaped] exit=1
not ok - grace-spawn-convergence: TERM handler did not spawn a child
[test_process_spawned_during_grace_is_reaped_on_later_pass] exit=1
ok - persistent leaked processes refuse teardown after bounded retries
[test_persistent_scan_refuses_after_bounded_retries] exit=0
ok - a process exiting during identity lookup does not block teardown
[test_process_exit_during_identity_lookup_does_not_refuse] exit=0
not ok - abort-then-reap-then-remove-order: the leaked process was not yet reaped when the worktree return ran
[test_run_abort_precedes_process_reap_precedes_worktree_removal] exit=1
Evidence: tests/fm-control-relaunch.test.sh run (relaunch refuses a returned worktree)

Source: tests/fm-control-relaunch.test.sh run (relaunch refuses a returned worktree)

ok - fm-control relaunch: a same-harness relaunch replaces the agent in the same endpoint and worktree
ok - fm-control relaunch: durable task metadata survives replacement launch publication
ok - fm-control relaunch: trace and concurrent task metadata publications serialize
ok - fm-control relaunch: disabling tracing clears metadata and pane context
ok - fm-control relaunch: the progress note lands in the instructions the replacement reads
ok - fm-control relaunch: a ship task refuses without the progress note its replacement needs
ok - fm-control relaunch: switching harness is one ordinary relaunch, and the old wiring goes with the old agent
ok - fm-control relaunch: a harness switch resets model and effort unless they are named too
ok - fm-control relaunch: a prefixed recorded harness can switch adapters transactionally
ok - fm-control relaunch: a prefixed command requires an explicit replacement harness
ok - fm-control relaunch: a same-harness relaunch keeps the profile axes it was running with
ok - fm-control relaunch: explicit model and effort win over the recorded ones
ok - fm-control relaunch: refuses to relaunch onto an adapter with no verified mechanics
ok - fm-control relaunch: the retired incarnation's global turn-end token is revoked
ok - fm-control relaunch: wiring cleanup failure refuses replacement arming
ok - fm-control-lib: one owner resolves each harness's turn-end registry entry, and refuses a malformed token
ok - fm-control relaunch: a secondmate relaunch re-resolves its durable configured harness pin
ok - fm-control relaunch: invalid configured effort is ignored before stop
ok - fm-control relaunch: an adapter unverified for this task kind refuses before the agent is stopped
ok - fm-control relaunch: explicit secondmate harness resets unnamed profile axes
ok - fm-control relaunch: a ship task keeps its recorded harness instead of re-reading crew config
ok - fm-spawn --relaunch: with no explicit harness it reuses the task's recorded one, never the crew default
ok - fm-spawn --relaunch: wiring armed under a prefixed harness name is still retired
ok - fm-spawn --relaunch: switching away from muse retires its session binding
ok - fm-spawn --relaunch: switching away from cursor retires its session binding
ok - fm-control relaunch: an unaccountable local copy refuses before the agent is touched
ok - fm-control relaunch: a worker with nothing to work from is never launched
ok - fm-control relaunch: a refusal before the agent is stopped leaves the durable record untouched
ok - fm-control relaunch: checkpoint inspection failures refuse before stopping
ok - fm-control relaunch: a launch failure after the stop keeps the prior record and reports the real state
ok - fm-control relaunch: unpublished rollback keeps concurrent durable metadata
ok - fm-control relaunch: post-publication failure keeps the new durable record
ok - fm-control relaunch: partial stop reconciles actual agent state
ok - fm-control relaunch: failed journal replacement preserves durable phase
ok - fm-spawn relaunch: prepublication abort removes replacement state
ok - fm-control relaunch: the checkpoint records the exact unlanded work it preserved
ok - fm-control relaunch: a secondmate's child work is accounted for and its charter is left alone
ok - fm-control relaunch: a secondmate home that is not this secondmate's is refused
ok - fm-control relaunch: unreadable and untraversable child state fails checkpoint
ok - fm-control relaunch: two control actions on one task serialize instead of interleaving
ok - fm-spawn relaunch: direct entry participates in lifecycle serialization
ok - fm-promote: promotion participates in lifecycle serialization
ok - fm-spawn --relaunch: refuses to launch a second agent into a live endpoint
ok - fm-spawn --relaunch: every identity axis comes from the record, and a contradicting flag refuses
ok - fm-spawn --relaunch: an unrecorded task is refused
ok - fm-spawn --relaunch: refuses to start a replacement outside the copy holding the work
ok - fm-spawn --relaunch: refuses to adopt a worktree an earlier teardown already returned
ok - fm-control relaunch: a returned worktree refuses before the agent is touched
Evidence: tests/fm-spawn-worktree-settle.test.sh run (spawn writes the owner marker)

Source: tests/fm-spawn-worktree-settle.test.sh run (spawn writes the owner marker)

ok - a single transient stale pane_current_path read is not accepted as the worktree
ok - an already-settled pane confirms via the existing inter-poll sleep, not an extra full cycle
ok - spawn marks the settled worktree as this task's, out of git's view
# all fm-spawn-worktree-settle tests passed
Evidence: tests/fm-backend-orca.test.sh run (Orca removal failure keeps the record)

Source: tests/fm-backend-orca.test.sh run (Orca removal failure keeps the record)

ok - fm_backend_orca_capture: parses result.terminal.tail and calls terminal read
ok - fm_backend_orca_capture: falls back to result text fields
ok - fm_backend_orca_capture: fails closed on Orca read error JSON
ok - fm_backend_orca_runtime_check: accepts reachable ready runtime
ok - fm_backend_orca_runtime_check: fails closed when runtime is not ready
ok - fm_backend_orca_send_text_submit: verifies empty composer after Enter with one bounded read
ok - fm_backend_orca_send_text_submit: a borderless claude composer confirms delivery (the missing #2029 shape)
ok - fm_backend_orca_composer_state: a stale startup banner cannot outrank the live composer row
ok - fm_backend_orca_send_text_submit: retries Enter while composer remains pending
ok - fm_backend_orca_composer_state: a slash-command popup's argument-hint placeholder still reads pending
ok - fm_backend_orca_composer_state: a bare dead-shell prompt reads unknown (unsafe-for-injection), never empty
ok - fm_backend_orca_send_text_submit: a slash-command popup's placeholder fill on Enter #1 does not short-circuit as submitted; Enter #2 is retried and lands it
ok - fm_backend_orca_send_literal: sends text without submitting
ok - fm_backend_orca_send_text_submit: reports send-failed when Orca send fails
ok - Orca send helpers: fail closed on ok:false JSON
ok - fm_backend_orca_send_key: Enter maps to empty enter, C-c maps to interrupt
ok - fm_backend_orca_send_key: refuses unsupported keys loudly
ok - fm_backend_orca_send_key: refuses Escape instead of mapping it to interrupt
ok - fm_backend_orca_kill: calls terminal close and stays best-effort
ok - fm_backend_orca_remove_worktree: refuses empty worktree ids
ok - fm_backend_orca_remove_worktree: fails closed on ok:false JSON
ok - fm_backend_orca_worktree_path: resolves an Orca worktree id to its path
ok - fm-backend dispatcher: accepts orca and routes capture through bin/backends/orca.sh
ok - fm_backend_orca_json_get: ignores undocumented terminal id shapes
ok - Orca lifecycle helpers: register repo, create worktree, create terminal, parse stable ids
ok - fm_backend_orca_worktree_create: removes created worktree when path is missing
ok - fm-spawn.sh --backend orca: preserves metadata when pathless cleanup fails
ok - fm-spawn.sh --backend orca: reuses implicit terminal, records metadata, launches harness
ok - fm-spawn.sh --backend orca --secondmate: refuses before secondmate-home mutation
ok - fm-spawn.sh --backend orca: refuses before mutation when Orca runtime is not ready
ok - fm-spawn.sh --backend orca: refuses non-isolated worktrees and closes implicit terminals
ok - fm-spawn.sh --backend orca: removes worktree when terminal creation fails
ok - fm-spawn.sh --backend orca: preserves metadata when abort cleanup fails
ok - fm-spawn.sh --backend orca: releases terminal and worktree on later aborts
ok - fm-peek/fm-send/fm-crew-state route through backend=orca metadata and its durable inbox
ok - fm-peek/fm-crew-state: Orca read error JSON fails closed
ok - fm_backend_target_exists: Orca ok:false read JSON is not live
ok - fm-teardown.sh backend=orca: scout report gate then helper-backed worktree removal
ok - fm-teardown.sh backend=orca: scout teardown refuses id/path mismatches
ok - fm-teardown.sh backend=orca: releases terminal/worktree when path is absent
ok - fm-teardown.sh backend=orca: preserves metadata on remove ok:false JSON
ok - fm-teardown.sh backend=orca: scout report gate precedes pathless helper cleanup
ok - fm-teardown.sh backend=orca: ship teardown fails closed when worktree path is missing
ok - fm-teardown.sh backend=orca: ship teardown requires a matching Orca id path
ok - fm-teardown.sh backend=orca: ship teardown fails closed when id resolution fails
ok - fm-teardown.sh backend=orca: ship teardown refuses id/path mismatches
ok - fm-teardown.sh backend=orca: refuses missing worktree ids before cleanup
ok - fm-teardown.sh backend=orca: refuses incomplete worktree-only endpoint metadata before runtime dispatch
ok - fm-teardown.sh --force: removes Orca secondmate children through Orca
ok - fm-teardown.sh --force: refuses Orca child id/path mismatches
ok - fm-teardown.sh --force: refuses partial Orca secondmate children before runtime dispatch
Evidence: tests/fm-teardown-endpoint-safety.test.sh run

Source: tests/fm-teardown-endpoint-safety.test.sh run

ok - fm-teardown: missing, empty, malformed, ambiguous, and task-mismatched endpoints refuse before every mutation or runtime call
ok - fm-teardown: a concurrent lifecycle action refuses before mutation
ok - fm-teardown: destructive cleanup serializes with metadata writers
ok - cleanup identity: valid tmux, Herdr, Zellij, Orca, and cmux records validate while every empty backend target refuses
ok - tmux backend: direct empty target returns nonzero without invoking tmux
ok - process cleanup: creation-time PID identity removes only the exact child and preserves the control child
ok - fm-teardown: exact tmux cleanup preserves invalid and prefix-matched neighbors while removing only the recorded target

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (3) ✅
  • ⚠️ bin/fm-spawn.sh:2721 - The relaunch path can still bind task A to a worktree another task now holds, and the new marker then makes A's teardown treat that worktree as its own. Sequence: fm-teardown.sh A returns slot-1 and records worktree_returned=1, then fails at the pane close; treehouse leases slot-1 to B; A's endpoint now reads 'dead' (the return killed its agent), so bin/fm-control.sh relaunch A passes the agent-free check and fm-spawn.sh --relaunch adopts the recorded worktree unconditionally (RELAUNCH_WT at bin/fm-spawn.sh:1047-1051 only checks the directory exists), overwrites B's .fm-task-owner with task=A at bin/fm-spawn.sh:2369, and preserve_relaunch_meta drops worktree_returned. A's next teardown now sees a marker naming itself, passes assert_worktree_owned_by_task, reaps B's agent and returns B's worktree - the original defect via a sibling path. The intent's note that 'a relaunch drops it when it rebinds a worktree' is not what the code does: relaunch never rebinds, it reuses the recorded path. Recommended boundary: in the --relaunch adoption block (next to the missing-worktree refusal at bin/fm-spawn.sh:1048), refuse when the prior meta carries worktree_returned=1, stating the worktree was already returned and teardown must be rerun; keep dropping the key only for the non-relaunch spawn that genuinely acquires a new worktree.
  • ⚠️ bin/fm-teardown.sh:2861 - In the Orca branch record_worktree_returned runs unconditionally after fm_backend_remove_worktree, whose exit status is discarded. If orca worktree rm fails (tool missing, id not found, transient error), the meta still gains worktree_returned=1, so a rerun after a later failure skips every worktree step and the final rm of the meta leaves an Orca worktree that was never removed, with the record having claimed otherwise. The treehouse branch gates the record on a successful return; the Orca branch should do the same, e.g. if fm_backend_remove_worktree &#34;$BACKEND&#34; &#34;$ORCA_WORKTREE_ID&#34;; then record_worktree_returned || {...}; fi (or fail loudly on removal failure, matching the header's 'never leave the two records disagreeing silently' contract).

🔧 Fix: refuse relaunch into returned worktree; gate Orca return record
1 error still open:

  • 🚨 bin/fm-teardown.sh:2868 - The fix round turned an Orca removal failure from a fail-closed abort into a warn-and-continue path that then deletes the task record. Under the script's set -eu the base's bare fm_backend_remove_worktree call exited teardown on failure and kept state/<id>.meta; now the else branch only prints a warning and execution falls through to the unconditional rm -f ... &#34;$STATE/$ID.meta&#34; near line 3002. Trace: backend=orca, orca worktree rm fails (tool missing, id not found, transient error) -> warning says 'a rerun retries the removal' -> teardown exits 0 and removes the meta -> no record remains to rerun from and the Orca worktree is orphaned. This regresses base behavior and makes the new header sentence at line 50 ('leaves the record unmarked, so a rerun retries the removal') false. Fix: fail loudly in the else branch, e.g. echo &#34;error: Orca worktree $WT (${ORCA_WORKTREE_ID:-no id}) could not be removed for $ID; retaining the task record so a rerun retries the removal&#34; &gt;&amp;2; exit 1, matching the treehouse return-failure path at line ~2895 and the header contract.

🔧 Fix: fail closed when Orca worktree removal fails
2 issues (1 warning, 1 info) still open:

  • ⚠️ tests/fm-teardown.test.sh:2764 - The marker-signal sub-case cannot prove the marker check works. Before the third run the worktree is still on branch fm/task-x2 (checked out earlier in the test) and task-x2.meta still exists, so assert_worktree_owned_by_task's branch fallback ('it is on task task-x2's branch') refuses with 'task-x2' in stderr even if the marker path were broken (e.g. marker_owner_still_recorded always returning 1, or worktree_owner_marker_value returning nothing). The assertions only check exit != 0, 'task-x2' in stderr, and the return count, all of which the branch fallback also satisfies. Fix: before the third run detach the worktree (git -C &#34;$case_dir/wt&#34; checkout -q --detach) or switch it to a non-fm/ branch so only the marker remains, and assert the marker-specific evidence text ('is marked as task task-x2') rather than just the task id.
  • ℹ️ bin/fm-spawn.sh:1055 - Relaunch refuses only on the meta's worktree_returned=1 and never reads the worktree's own .fm-task-owner marker. If teardown's pool return succeeds but record_worktree_returned fails (teardown exits 1 with the record unmarked and the marker already removed), a later task can lease the slot and write its marker, and a subsequent relaunch of the retired task would still adopt the path and overwrite that marker plus clear the other task's harness wiring. The path is narrow (a filesystem error in the state dir, followed by an operator relaunching instead of reconciling as the error instructs, and the relaunch's shell-in-worktree check surviving the return), so this is noted as a residual gap rather than a blocker; reading the marker in the --relaunch adoption block and refusing when it names another task would close it cheaply.

🔧 Fix: isolate marker-signal refusal test from branch fallback
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-teardown.test.sh (full suite on target; all cases pass up to and including the three new ownership cases, then stops at the pre-existing lsof-dependent leaked-process-reap failure)
  • bash tests/fm-control-relaunch.test.sh (pass, including the two new returned-worktree refusal cases)
  • bash tests/fm-spawn-worktree-settle.test.sh (pass, including the new task-owner marker case)
  • bash tests/fm-teardown-endpoint-safety.test.sh (pass)
  • bash tests/fm-backend-orca.test.sh (pass, including fm-teardown.sh backend=orca: preserves metadata on remove ok:false JSON, which covers the fail-closed Orca removal path)
  • run-selected-teardown-tests.sh tests test_ordinary_teardown_reaps_and_returns_its_own_worktree test_rerun_after_partial_failure_refuses_a_reassigned_worktree test_rerun_after_reconciled_records_completes_cleanly on target: all three pass, transcripts captured
  • Same three test functions run against the base commit c03bfbe scripts (extracted with git archive to a temp dir): reassigned-worktree case fails with the rerun killed the successor task&#39;s process, reconciled-rerun case fails with the rerun returned the already-returned worktree again, ordinary case passes
  • The ten reap cases after the lsof failure (test_leaked_worktree_process_is_reaped through test_run_abort_precedes_process_reap_precedes_worktree_removal) run one at a time on target and on base: identical results, five pre-existing failures caused by missing lsof on this host, five pass
  • Manual check show-converged-meta.sh: one real partial-failure teardown, then printed state/task-x1.meta showing worktree_returned=1, then an unreconciled rerun that skips the worktree steps with the pool return count still one
  • git status --porcelain --ignored in the worktree: clean, no transient artifacts left
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

fm-teardown.sh terminated processes in, and returned, the worktree path
recorded in a task's metadata without checking the worktree still belonged
to that task. After a partial failure - the pool return succeeded but a
later step (a Herdr pane close) failed, so the task record stayed on disk -
treehouse could hand the freed slot to the next task before teardown was
rerun. The rerun then killed the new task's agent and deleted its branch
with no warning or refusal.

Ownership proof: every non-secondmate spawn now writes a gitignored
.fm-task-owner marker into the task worktree naming the task and the state
directory holding its records. Teardown reads it before killing processes
in, resetting, or returning the worktree, and refuses loudly - naming the
other task - when it belongs to someone else. The marker is rewritten by
every spawn, so it names whichever task took the worktree last; it speaks
only for a task that still has a record where it says that task's records
live, so a leftover never wedges an unrelated teardown. A worktree with no
live marker falls back to another live record in this home claiming the same
path, then to the fm/<task-id> branch. The refusal forces, stashes, and
discards nothing, and rerunning after the records are reconciled completes.

Rerun convergence: a successful pool return or Orca worktree removal is
recorded as worktree_returned=1 in the task record, so a rerun after a later
failure skips every worktree step instead of repeating it, reaping only the
task's own temp root. A return that succeeded but could not be recorded
fails loudly rather than leaving the two records disagreeing silently.

Both script headers carry the new contracts. Regression tests cover an
ordinary teardown still reaping and returning its own worktree, a rerun
refusing loudly on a worktree a second task has since acquired or marked,
a rerun completing cleanly once the records are reconciled, and spawn
marking the settled worktree out of git's view.
@LeonidShamis
LeonidShamis merged commit df933e8 into main Sep 5, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant