Skip to content

fix(bin): relaunch an in-flight worker when its session ends - #4

Merged
MrGTV-love merged 11 commits into
mainfrom
fm/fm-idle-exit-relaunch-ship
Sep 29, 2026
Merged

MrGTV-love merged 11 commits into
mainfrom
fm/fm-idle-exit-relaunch-ship

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain's standing ask (2026-09-27): "we should not have any stoppages due to session limits." On the captain board (2026-09-28, card fm-idle-exit-relaunch) he chose option A: "build auto-restart plus one debug worker". Option A as he approved it: a small firstmate change where the existing session-end event triggers the existing relaunch path while the task is open, and Claude's debug log is enabled on one idle worker so the next stop names the signal. Result he approved: "A stopped lane restarts in minutes with nothing lost; the next stop tells us who sends the signal." Idle Claude Code worker sessions (2.1.283) take Claude's graceful-shutdown path after an unknown signal (SIGTERM, SIGHUP or orphan detection); the pane prints "Resume this session with:" and returns to the shell. One 01:23:03Z event ended several sessions at once. The proxy, usage limits, jetsam and the turn-end guard were ruled out. The per-worktree SessionEnd hook written by fm-spawn already records event=session-end in the task's busy-state file, but nothing acts on it, so the lane stays stopped until someone notices. bin/fm-control.sh relaunch already handles an already-stopped agent (exit_result=already-stopped).

What Changed

  • The watcher now relaunches an in-flight worker after its session ends. A new library, bin/fm-session-end-relaunch-lib.sh, runs from each bin/fm-watch.sh cycle. It looks for an open ship or scout whose busy record says event=session-end and whose pane shows no live agent. It then runs bin/fm-control.sh <id> relaunch, which keeps the recorded worktree and all uncommitted work.
    • It skips a task if the task is done or failed, paused, held for the captain, already inside a control relaunch, or stopped on purpose with the exit verb. fm-control.sh exit now writes a state/<id>.control-exit marker for the current busy generation, so the watcher can tell a deliberate stop.
    • It tries at most once per 30 minutes and 3 times per day for each task, and it runs at most one relaunch per watcher cycle. The relaunch time limit is the watcher stale grace minus 60 seconds, and the watcher beacon is touched first, so a blocked watcher does not look dead.
    • It sends a check: <id> auto-relaunched after session-end wake, a failed wake, or a paused wake. Teardown removes the new marker, ledger, and handled files.
  • New opt-in flag --claude-debug for bin/fm-control.sh <id> relaunch and bin/fm-spawn.sh --relaunch. It works only with a claude harness. It adds Claude's --debug to the launch and sets CLAUDE_CODE_DIAGNOSTICS_FILE to state/<id>.claude-diagnostics.jsonl. Claude writes the shutdown_signal event to that file, so the next stop shows which signal ended the session.
  • Other changes: fm_exec_timed in bin/fm-timeout-lib.sh now works when BASHPID is not set. AGENTS.md and the docs (architecture, agent-control, configuration, turnend-guard, operational-home-layout skill) describe the new wakes, files, and grace limit. The new tests/fm-session-end-relaunch.test.sh covers this change, and the control and teardown tests are extended.

Risk Assessment

✅ Low: The fix round follows the user's three instructions. Only dead is eligible now, and the test checks that a missing endpoint gets no control call, no wake, and no ledger row. --claude-debug is refused before batch dispatch on every fresh spawn form. The bound is the grace minus 60s, and the beacon is touched before the blocking call. The new tests run the real code and would fail without the fixes. The only remaining note is a non-default configuration edge.

Testing

The focused session-end unit test file passed as a baseline. Then I stood up a disposable lab home with a private tmux server, spawned a real Claude scout worker, and stopped it with real signals. The real watcher auto-relaunched the lane with nothing lost. Every guard (cap, deliberate exit, done, missing window, failed-attempt no-loop) and every --claude-debug rule held, and the debug diagnostics file named the SIGHUP that stopped the worker. The worker's Claude account was at its weekly usage limit, so the relaunched agent could not do model work. The runtime checks (agent alive, same worktree, file kept, note appended) prove the restart itself. The auto-relaunch passes ran the watcher from a plain clone of the target commit, because the gate check refuses them from a gate worktree. All evidence is CLI/TUI transcripts, since there is no graphical UI surface. The lab was fully removed afterward.

  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
An idle Claude worker stopped by SIGTERM records session-end: the pane prints 'Resume this session with:', returns to the shell, the busy record has event=session-end, and the agent state is dead ✅ pass live s1-signal-transcript.txt
The real watcher auto-relaunches that open lane within one cycle. The new Claude session is alive in the same worktree, the uncommitted notes.txt is kept, the session-end note is appended, the journal… ✅ pass live s1-auto-relaunch.txt (watcher run from a plain clone of 469b432 inside the lab)
A second session-end within 30 minutes is not relaunched. The watcher wakes once with 'auto-relaunch paused after 1 attempt in 1800s' and does not repeat on the next pass ✅ pass live s2-cap-holds.txt
fm-control relaunch --claude-debug launches claude --debug with CLAUDE_CODE_DIAGNOSTICS_FILE. The next stop (SIGHUP) is named in state/<id>.claude-diagnostics.jsonl as shutdown_signal SIGHUP ✅ pass live s3-claude-debug-names-signal.txt
A deliberate fm-control exit is not auto-relaunched: the busy record is retired, the control-exit marker is written, the watcher makes no attempt and the agent stays dead ✅ pass live s4-deliberate-exit-skipped.txt
A lane whose latest status is done is not auto-relaunched after a session-end ✅ pass live s5-done-lane-skipped.txt
A failed auto-relaunch is ledgered, wakes once with 'auto-relaunch failed after session-end: <detail>', and is not retried on the next watcher pass ✅ pass live s6-failed-relaunch-no-loop.txt (failure came from the gate check in the gate-worktree run)
A lane whose tmux window is gone (missing endpoint, session-end recorded) is not auto-relaunched ✅ pass live s7-missing-endpoint-skipped.txt
--claude-debug is refused with a clear error (rc=1, no task record created) on fresh scout, ship, secondmate and batch spawns, and on fm-control exit ✅ pass live s8-claude-debug-refused-outside-relaunch.txt
--claude-debug on a relaunch to a non-Claude harness (codex) is refused before the checkpoint, and the task record is unchanged ✅ pass live s9-claude-debug-refused-non-claude.txt
Teardown of the lane removes the session-end ledger, handled and control-exit files, and keeps the diagnostics file as evidence ✅ pass live s10-teardown-clears-session-end-files.txt
Evidence: Idle Claude worker stopped by SIGTERM: Resume line, session-end record, dead classifier

Source: Idle Claude worker stopped by SIGTERM: Resume line, session-end record, dead classifier

=== before signal: busy record
v1 gen=g1790658879.69348.25374 seq=3 state=idle source=claude-hook event=stop-failure ts=1790658887
=== kill -TERM 75921 at 2026-09-29T05:17:14Z
=== after signal: pane tail
 and matches only commands containing a literal * at that position. Use Bash(rm*/.tac/governance/*) for wildcard matching.
Permission ask rule (../../../../../../../.claude/settings.json): Bash(mv*/.tac/governance/*:*) mixes * with the trailing :* prefix syntax, so it is matched as a literal prefix (the * is not expanded)
 and matches only commands containing a literal * at that position. Use Bash(mv*/.tac/governance/*) for wildcard matching.
Permission ask rule (../../../../../../../.claude/settings.json): Bash(*>*/.tac/governance/*:*) mixes * with the trailing :* prefix syntax, so it is matched as a literal prefix (the * is not expanded)
 and matches only commands containing a literal * at that position. Use Bash(*>*/.tac/governance/*) for wildcard matching.
Resume this session with:
claude --resume 164ba3bc-3ac0-4223-9559-b4162c4c6ba2
charlesabrooker@GTT-CB demo %
=== after signal: busy record
v1 gen=g1790658879.69348.25374 seq=4 state=idle source=claude-hook event=session-end ts=1790659035
=== agent state via firstmate backend classifier
dead
Evidence: Watcher auto-relaunch: same worktree, uncommitted file kept, note appended, exit_result=already-stopped

Source: Watcher auto-relaunch: same worktree, uncommitted file kept, note appended, exit_result=already-stopped

=== fixture reset: removed the gate-refused attempt ledger (lab artefact)
=== before: busy record
v1 gen=g1790658879.69348.25374 seq=4 state=idle source=claude-hook event=session-end ts=1790659035
=== before: agent state
dead
=== before: worktree status
?? notes.txt
=== before: meta
window=primary:fm-se-scout
endpoint_task_id=se-scout
worktree=~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
project=~/tmp/fm-lab.LdoTZs/projects/demo
harness=claude
kind=scout
tasktmp=/tmp/fm-se-scout
model=default
effort=low
busy_gen=g1790658879.69348.25374
spawn_gen=s1790658880.60581.20148
=== watcher run from plain checkout; returned after 16s; stdout:
check: se-scout auto-relaunched after session-end
rc=0
=== stderr tail
=== after: agent state
alive
=== after: meta
window=primary:fm-se-scout
endpoint_task_id=se-scout
worktree=~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
project=~/tmp/fm-lab.LdoTZs/projects/demo
harness=claude
kind=scout
tasktmp=/tmp/fm-se-scout
model=default
effort=low
busy_gen=g1790659228.62371.28887
spawn_gen=s1790659229.57071.24946
control_relaunch_tx=56127.20260929T052024Z.7521
=== after: worktree status (uncommitted file kept)
?? notes.txt
notes.txt: alpha
=== after: busy record
v1 gen=g1790659228.62371.28887 seq=3 state=idle source=claude-hook event=stop-failure ts=1790659237
=== after: ledger
1790659222	attempt
1790659233	relaunched
=== after: handled
g1790658879.69348.25374	4	relaunched
=== after: control-relaunch journal
v1
task=se-scout
phase=complete
ts=2026-09-29T05:20:33Z
backend=tmux
endpoint=primary:fm-se-scout
worktree=~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
kind=scout
from_harness=claude
from_model=default
from_effort=low
to_harness=claude
to_model=default
to_effort=low
worktree_head=3d2a3f013582487fcc14ff66dadf86fcd45e01a8
worktree_dirty=yes
note_file=~/tmp/fm-lab.LdoTZs/state/se-scout.control-relaunch.note
exit_result=already-stopped
=== after: wake queue rows
1790659050	1	check	session-end-relaunch-failed-se-scout-g1790658879.69348.25374-4	check: se-scout auto-relaunch failed after session-end: error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/31c777dfc236.git)
1790659233	2	check	session-end-relaunch-se-scout-g1790658879.69348.25374-4	check: se-scout auto-relaunched after session-end
=== after: progress note appended to brief (tail)

First, check your instruction inbox: list ~/tmp/fm-lab.LdoTZs/state/se-scout.inbox/*.msg, act on
each message in numeric order, then mv each handled file into
~/tmp/fm-lab.LdoTZs/state/se-scout.inbox/handled/. A steer sent before the relaunch survives there.

The previous worker session ended while this task was still open. The local copy and every uncommitted change were left as the previous worker left them. Continue from the instructions and the instruction inbox.
=== after: pane (non-empty lines, tail)
charlesabrooker@GTT-CB demo % export GOTMPDIR=/tmp/fm-se-scout/gotmp
charlesabrooker@GTT-CB demo % export COMPACT_ADVISER_DISABLE=1
charlesabrooker@GTT-CB demo % export FM_TASK_ID=se-scout
charlesabrooker@GTT-CB demo % . '/tmp/fm-se-scout+703d7b0e70401f023d1bf2fd58e6cd4e3edc5e60bd668223c869e2559e8e5964/launch.s1790659229.57071.24946.sh'
 ▐▛███▛█   Claude Code v2.1.284
▝▜██████▀  Opus 5.5 (1M context) with low effort · Claude Max
 ▝▝   ▝▝   ~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
⚠ Your login expires in 3 days · run /login to renew
❯ : Firstmate operational input waiting: read '~/tmp/fm-lab.LdoTZs/state/operational-inbox/1790659231-b7d011752c49e389.msg' and handle its contents as Firstmate operational
  input.
  ⎿  You've hit your weekly limit · resets 6am (America/Chicago)
     Use your limit reset to reset it now: clau.de/reset
⏺ Usage limit reached · continuing automatically at 6am · esc to cancel
✻ Cogitated for 1s · done 12:20 AM
                                                                                                     tmux detected · scroll with PgUp/PgDn · or add 'set -g mouse on' to ~/.tmux.conf for wheel scroll
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⚠ Usage limit reached · limit resets 6am
    Continuing automatically at 6am · esc to cancel · /usage-credits to continue now
  Opus 5.5 (1M context) ░░░░░░░░░░ 0% 0/0 demo:643c33a7
  ⏵⏵ bypass permissions on (shift+tab to cycle) · ← 1 agent
Evidence: 30-minute cap: one paused wake, no relaunch, no repeat wake

Source: 30-minute cap: one paused wake, no relaunch, no repeat wake

=== kill -TERM 66639 (relaunched worker) at 2026-09-29T05:21:03Z
=== busy record
v1 gen=g1790659228.62371.28887 seq=4 state=idle source=claude-hook event=session-end ts=1790659263
=== agent state
dead
=== watcher pass A stdout:
check: se-scout auto-relaunch paused after 1 attempt in 1800s; session-end still recorded
rc=0
=== watcher pass B stdout:
check: rearm-resurface
rc=0
=== agent state after both passes (still dead = not relaunched)
dead
=== ledger
1790659222	attempt
1790659233	relaunched
=== handled
g1790659228.62371.28887	4	capped-min
Evidence: --claude-debug relaunch; diagnostics file names SIGHUP

Source: --claude-debug relaunch; diagnostics file names SIGHUP

=== $ bin/fm-control.sh se-scout relaunch --claude-debug --note '...'
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/tmp/fm-lab.LdoTZs
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/tmp/fm-lab.LdoTZs
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no live watcher process holds this home lock (last beat: 14s ago).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  After draining queued wakes, repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
relaunched se-scout harness=claude from=claude model=default effort=low backend=tmux endpoint=primary:fm-se-scout worktree=~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
rc=0
=== launch script (/tmp/fm-se-scout+703d7b0e70401f023d1bf2fd58e6cd4e3edc5e60bd668223c869e2559e8e5964/launch.s1790659291.35724.3367.sh) claude line
CLAUDE_CODE_DIAGNOSTICS_FILE='~/tmp/fm-lab.LdoTZs/state/se-scout.claude-diagnostics.jsonl' 
claude --debug --dangerously-skip-permissions --add-di
=== agent state
alive
=== claude process args
42552 claude --debug --dangerously-skip-permissions --add-dir ~/tmp/fm-lab.LdoTZs/state/operational
=== kill -HUP 42552 at 2026-09-29T05:21:46Z
=== pane tail
 and matches only commands containing a literal * at that position. Use Bash(mv*/.tac/governance/*) for wildcard matching.
Permission ask rule (../../../../../../../.claude/settings.json): Bash(*>*/.tac/governance/*:*) mixes * with the trailing :* prefix syntax, so it is matched as a literal prefix (the * is not expanded)
 and matches only commands containing a literal * at that position. Use Bash(*>*/.tac/governance/*) for wildcard matching.
charlesabrooker@GTT-CB demo %
=== busy record
v1 gen=g1790659291.38296.31827 seq=2 state=idle source=claude-hook event=session-end ts=1790659307
=== diagnostics file: ~/tmp/fm-lab.LdoTZs/state/se-scout.claude-diagnostics.jsonl
-rw-r--r--@ 1 charlesabrooker  staff  13458 Sep 29 00:21 ~/tmp/fm-lab.LdoTZs/state/se-scout.claude-diagnostics.jsonl
=== shutdown events in diagnostics file
{"timestamp":"2026-09-29T05:21:46.927Z","level":"info","event":"shutdown_signal","data":{"signal":"SIGHUP"}}
{"timestamp":"2026-09-29T05:21:46.933Z","level":"info","event":"session_activity_at_shutdown","data":{"owner_key":"cli","refcount":1,"active":{"hook_exec":1},"oldest_activity_ms":8534}}
Evidence: Deliberate fm-control exit is not auto-relaunched

Source: Deliberate fm-control exit is not auto-relaunched

=== fixture: live worker, cap ledger cleared
=== agent state
alive
=== $ bin/fm-control.sh se-scout exit
stopped se-scout harness=claude backend=tmux endpoint=primary:fm-se-scout worktree=~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo
rc=
=== busy record after exit
(retired)
=== control-exit marker
gen=g1790659332.83244.7000
=== agent state
dead
=== watcher (plain checkout) pass 1:
check: rearm-resurface
rc=0
=== watcher pass 2:
check: rearm-resurface
rc=0
=== agent state (still dead = not relaunched)
dead
=== session-end ledger
(none - no attempt made)
Evidence: Done lane is not auto-relaunched

Source: Done lane is not auto-relaunched

=== fixture: live worker relaunched, status line 'done' appended, cap ledger cleared, then kill -TERM 77055
=== last status line
done [at=1790659415]: lab fixture marks this scout done
=== busy record
v1 gen=g1790659404.70706.12481 seq=4 state=idle source=claude-hook event=session-end ts=1790659416
=== watcher (plain checkout) stdout
check: rearm-resurface
rc=0
=== agent state (dead = not relaunched)
dead
=== ledger
(none - no attempt made)
Evidence: Failed relaunch (gate refusal) wakes once and does not retry

Source: Failed relaunch (gate refusal) wakes once and does not retry

=== watcher pass 1 (gate worktree code) stdout:
check: se-scout auto-relaunch failed after session-end: error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/31c777dfc236.git)
rc=0
=== ledger after pass 1
1790659049	attempt
1790659050	failed
=== handled after pass 1
g1790658879.69348.25374	4	failed
=== watcher pass 2 stdout (45s):
check: rearm-resurface
rc=0
=== ledger after pass 2
1790659049	attempt
1790659050	failed
Evidence: Missing tmux window is not auto-relaunched

Source: Missing tmux window is not auto-relaunched

=== fixture: live worker relaunched, no status file, cap ledger cleared, then tmux kill-window primary:fm-se-scout
=== busy record
v1 gen=g1790659451.35640.11991 seq=4 state=idle source=claude-hook event=session-end ts=1790659460
=== agent state
missing
=== windows in lab session
zsh
=== watcher (plain checkout) stdout
check: rearm-resurface
rc=0
=== ledger
(none - no attempt made)
Evidence: --claude-debug refused on fresh scout/ship/secondmate/batch and on exit

Source: --claude-debug refused on fresh scout/ship/secondmate/batch and on exit

=== $ bin/fm-spawn.sh se-two <demo> --scout --harness claude --claude-debug   (fresh scout)
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
error: --claude-debug applies to --relaunch only; turn it on for an existing worker with bin/fm-control.sh <id> relaunch --claude-debug
rc=
=== $ bin/fm-spawn.sh se-three <demo> --mode local-only --yolo off --harness claude --claude-debug   (fresh ship)
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
error: --claude-debug applies to --relaunch only; turn it on for an existing worker with bin/fm-control.sh <id> relaunch --claude-debug
rc=
=== $ bin/fm-spawn.sh se-four --secondmate --claude-debug   (fresh secondmate)
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
error: --claude-debug applies to --relaunch only; turn it on for an existing worker with bin/fm-control.sh <id> relaunch --claude-debug
rc=
=== $ bin/fm-spawn.sh a1=<demo> a2=<demo> --scout --claude-debug   (batch)
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
error: --claude-debug applies to --relaunch only; turn it on for an existing worker with bin/fm-control.sh <id> relaunch --claude-debug
rc=
=== $ bin/fm-control.sh se-scout exit --claude-debug
error: --harness, --model, --effort, --note, and --claude-debug apply to 'relaunch' only
rc=
=== task records created by these calls (expect none besides se-scout)
se-scout.meta
1
treehouse-state.json
treehouse-state.key
treehouse-state.lock
fresh scout rc=1
fresh ship rc=1
fresh secondmate rc=1
batch rc=1
control exit --claude-debug rc=1
Evidence: --claude-debug refused on a codex relaunch before checkpoint

Source: --claude-debug refused on a codex relaunch before checkpoint

=== fixture: shell-only window fm-se-scout recreated in the worktree (agent state: dead)
=== journal phase before: complete  ts: 2026-09-29T05:24:13Z
=== $ bin/fm-control.sh se-scout relaunch --harness codex --claude-debug --note ...
error: --claude-debug applies only to a claude relaunch; se-scout would be replaced on codex
rc=1
=== journal ts after (unchanged = refused before checkpoint): 2026-09-29T05:24:13Z
=== meta harness after: claude
Evidence: Teardown clears ledger, handled and control-exit files; keeps diagnostics

Source: Teardown clears ledger, handled and control-exit files; keeps diagnostics

=== before teardown: session-end state files
se-scout.claude-diagnostics.jsonl
se-scout.control-exit
=== $ bin/fm-teardown.sh se-scout
fm-captain-hold: origin se-scout has no completed captain-call inventory
REFUSED: scout task se-scout has not passed the captain-call completion gate.
Inventory its report and any visual review through bin/fm-captain-hold.sh before teardown.
=== after teardown: session-end state files
.session-end-handled-se-scout
.session-end-relaunch-se-scout
se-scout.claude-diagnostics.jsonl
se-scout.control-exit
=== meta present?
~/tmp/fm-lab.LdoTZs/state/se-scout.meta
=== $ bin/fm-captain-hold.sh complete se-scout --none ; bin/fm-teardown.sh se-scout
teardown: reaping leaked worktree process(es) for se-scout: 27979 81536 85758
teardown: force-killing leaked worktree process(es) for se-scout: 27979
A new version of treehouse is available: v2.3.1-fm.e58947de → v3.1.0
Run "treehouse update" to update

🌳 Worktree returned to pool.
teardown se-scout complete (window primary:fm-se-scout, worktree ~/tmp/fm-lab.LdoTZs/pool/.treehouse/demo-55eed5/1/demo)
Backlog: se-scout just finished (this home keeps no markdown backlog at ~/tmp/fm-lab.LdoTZs/data/backlog.md). Update ~/tmp/fm-lab.LdoTZs/data/backlog.md - move se-scout to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
=== after teardown: session-end state files
se-scout.claude-diagnostics.jsonl
=== meta present?
ls: ~/tmp/fm-lab.LdoTZs/state/se-scout.meta: No such file or directory
- Outcome: ⚠️ 1 info across 1 run (15m55s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ bin/fm-session-end-relaunch-lib.sh:222 - The tick treats a missing endpoint as eligible, but the intent only asks for a pane that 'returns to the shell' (dead, fm-control's exit_result=already-stopped). On tmux, fm-control relaunch always refuses a missing endpoint. fm_control_endpoint_absence_verdict returns unproven for tmux, do_exit dies, and stuck-crewmate-recovery/SKILL.md:45 says 'both exit and relaunch refuse'. Failing sequence: the tmux server or a lane window goes away. SIGHUP triggers Claude's graceful SessionEnd, so the record says session-end, and fm_backend_tmux_agent_state prints missing. Each tick then runs a relaunch that cannot succeed: checkpoint, note append, die, rollback. It emits auto-relaunch failed and repeats every 30 minutes, up to 3 times in any rolling 24 hours, indefinitely. tests/fm-session-end-relaunch.test.sh:test_missing_endpoint_relaunches asserts success on a tmux lane only because a recorder stands in for fm-control. The missing acceptance path exceeds the intent, and the defect lives inside it. The smallest honest remedy is to remove missing (accept dead only) and drop or invert that test. Keeping herdr-only missing would be an extension that needs your approval.
  • ⚠️ bin/fm-spawn.sh:771 - The intent asks for Claude's debug log on 'one idle worker'. For an existing idle worker, the only path is fm-control.sh &lt;id&gt; relaunch --claude-debug → fm-spawn --relaunch --claude-debug. The change also accepts --claude-debug on fresh ship spawns, scout spawns, and --secondmate spawns (usage lines 4-6), and no intent requirement needs those. A defect lives inside that extra surface. Batch dispatch (fm-spawn.sh a=repo b=repo --claude-debug) re-execs each pair with shared_args (fm-spawn.sh ~1469-1478), and shared_args does not forward --claude-debug. The flag is silently dropped, and no child launches with debug or diagnostics. Recommended narrower form: accept --claude-debug only together with --relaunch, and refuse it on the other forms.
  • ℹ️ bin/fm-session-end-relaunch-lib.sh:74 - The watcher touches .last-watcher-beat once at the top of each cycle. Then session_end_relaunch_tick blocks for up to FM_SESSION_END_TIMEOUT=300s inside that same cycle. The default WATCHER_STALE_GRACE is also 300 (fm_poll_derived_grace for POLL=15). If a relaunch runs close to its bound (a slow fm-spawn --relaunch plus the 90s launch wait and rollback), the beacon reaches the grace. fm-guard then prints the watcher-down banner, and fm-watch-arm refuses a re-arm with 'inspect or stop that watcher'. The watcher is alive, so this is a false alarm that invites someone to kill it mid-relaunch. Remedy options: keep the bound below the grace with margin, or refresh the beacon around the blocking call. This trades against the earlier review choice to widen the bound from 120 to 300, so confirm the direction.
  • ℹ️ bin/fm-timeout-lib.sh:224 - The merge commit 2f112f5 changes fm_exec_timed's owner-PID logic (an extra BASHPID guard). Neither parent has this change, so it goes beyond the recorded rebase decision (resolve conflicts, keep both behaviours). It only changes behaviour where BASHPID is unset (bash 3.2): a hard unbound-variable failure becomes owner=$$. This looks benign, and it may help the tasks-axi calls behind the new captain-hold check on /bin/bash. It is noted so the shared-helper change is visible, not hidden in a merge.

🔧 Fix applied.
1 info still open:

  • ℹ️ bin/fm-watch.sh:260 - The relaunch bound comes from the watcher's WATCHER_STALE_GRACE (FM_WATCHER_STALE_GRACE, then FM_GUARD_GRACE, then the poll-derived value). The consumers that show a false 'watcher down' banner read a different grace. bin/fm-guard.sh:56 and bin/fm-watch-arm.sh:109 use ${FM_GUARD_GRACE:-300}. docs/configuration.md:2361 documents FM_WATCHER_STALE_GRACE as a separate knob. Concrete case: FM_WATCHER_STALE_GRACE=900 with FM_GUARD_GRACE unset. The bound becomes 840s. fm-guard still uses 300s, so a slow relaunch can make fm-guard report the live watcher as down after 300s. That is the false alarm fix (3) was meant to remove. This happens only with a non-default override, and the user's instruction said to derive the bound from the watcher grace. A narrower remedy, if wanted: pass min(WATCHER_STALE_GRACE, ${FM_GUARD_GRACE:-300}) to fm_session_end_relaunch_scan. With default settings, all three values are 300, and the 60s margin holds.
⚠️ **Test** - 1 info
  • ℹ️ bin/fm-session-end-relaunch-lib.sh:283 - The automatic relaunch cannot succeed when the watcher runs from a no-mistakes gate worktree, even against a marked lab home. The library passes FM_STATE_OVERRIDE=$state to fm-control, and bin/fm-gate-refuse-lib.sh cancels the lab-home allowance when any FM_*_OVERRIDE is set. The result is 'auto-relaunch failed ... refusing fleet lifecycle from inside a no-mistakes gate worktree'. On a normal operator checkout this check does not apply, so production behaviour is not affected. The successful live passes ran the same target commit from a plain clone inside the lab. If future gate labs must drive this path from the gate worktree, pass FM_STATE_OVERRIDE only when the watcher itself was given one.
  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
An idle Claude worker stopped by SIGTERM records session-end: the pane prints 'Resume this session with:', returns to the shell, the busy record has event=session-end, and the agent state is dead ✅ pass live s1-signal-transcript.txt
The real watcher auto-relaunches that open lane within one cycle. The new Claude session is alive in the same worktree, the uncommitted notes.txt is kept, the session-end note is appended, the journal… ✅ pass live s1-auto-relaunch.txt (watcher run from a plain clone of 469b432 inside the lab)
A second session-end within 30 minutes is not relaunched. The watcher wakes once with 'auto-relaunch paused after 1 attempt in 1800s' and does not repeat on the next pass ✅ pass live s2-cap-holds.txt
fm-control relaunch --claude-debug launches claude --debug with CLAUDE_CODE_DIAGNOSTICS_FILE. The next stop (SIGHUP) is named in state/<id>.claude-diagnostics.jsonl as shutdown_signal SIGHUP ✅ pass live s3-claude-debug-names-signal.txt
A deliberate fm-control exit is not auto-relaunched: the busy record is retired, the control-exit marker is written, the watcher makes no attempt and the agent stays dead ✅ pass live s4-deliberate-exit-skipped.txt
A lane whose latest status is done is not auto-relaunched after a session-end ✅ pass live s5-done-lane-skipped.txt
A failed auto-relaunch is ledgered, wakes once with 'auto-relaunch failed after session-end: <detail>', and is not retried on the next watcher pass ✅ pass live s6-failed-relaunch-no-loop.txt (failure came from the gate check in the gate-worktree run)
A lane whose tmux window is gone (missing endpoint, session-end recorded) is not auto-relaunched ✅ pass live s7-missing-endpoint-skipped.txt
--claude-debug is refused with a clear error (rc=1, no task record created) on fresh scout, ship, secondmate and batch spawns, and on fm-control exit ✅ pass live s8-claude-debug-refused-outside-relaunch.txt
--claude-debug on a relaunch to a non-Claude harness (codex) is refused before the checkpoint, and the task record is unchanged ✅ pass live s9-claude-debug-refused-non-claude.txt
Teardown of the lane removes the session-end ledger, handled and control-exit files, and keeps the diagnostics file as evidence ✅ pass live s10-teardown-clears-session-end-files.txt
  • bash tests/fm-session-end-relaunch.test.sh (focused baseline, 12 ok)
  • bin/fm-lab-home.sh create $LAB plus a private tmux -L fm-lab server started with a clean env and TREEHOUSE_ROOT=$LAB/pool
  • bin/fm-brief.sh se-scout demo --scout then bin/fm-spawn.sh se-scout $LAB/projects/demo --scout --harness claude --effort low (real Claude worker)
  • kill -TERM &lt;claude pid&gt; on the idle worker; read the busy record and fm_backend_agent_state tmux primary:fm-se-scout
  • bin/fm-watch.sh from the gate worktree: relaunch refused by the gate check; then a second pass to confirm no retry
  • bin/fm-watch.sh from a plain clone of 469b432 inside the lab: auto-relaunch, then checks of meta, journal, ledger, wake queue, brief note, git status and pane
  • kill -TERM on the relaunched worker, then two watcher passes (30-minute cap)
  • bin/fm-control.sh se-scout relaunch --claude-debug --note ..., then kill -HUP, then grep for shutdown_signal in state/se-scout.claude-diagnostics.jsonl
  • bin/fm-control.sh se-scout exit, then two watcher passes (deliberate exit skip)
  • appended a done status line, kill -TERM, then a watcher pass (done skip)
  • tmux kill-window -t primary:fm-se-scout, then a watcher pass (missing endpoint skip)
  • bin/fm-spawn.sh ... --claude-debug for fresh scout, ship, secondmate and batch, and bin/fm-control.sh se-scout exit --claude-debug (refusals, rc=1)
  • bin/fm-control.sh se-scout relaunch --harness codex --claude-debug on a recreated shell window (refused before checkpoint)
  • bin/fm-captain-hold.sh complete se-scout --none then bin/fm-teardown.sh se-scout (state files cleared)
  • Lab cleanup: tmux -L fm-lab kill-server, rm -rf $LAB, worktree git status clean
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

The watcher reuses the recorded session-end event and the existing relaunch path, with a 30-minute and daily cap, and skips a deliberate exit or a paused or held task. Claude debug logging stays off unless --claude-debug is passed.
The watcher reuses the recorded session-end event and the existing relaunch path, with a 30-minute and daily cap, and skips a deliberate exit or a paused or held task. Claude debug logging stays off unless --claude-debug is passed.
@MrGTV-love
MrGTV-love force-pushed the fm/fm-idle-exit-relaunch-ship branch from 27169d7 to 0b7ee9a Compare September 29, 2026 00:52
@MrGTV-love
MrGTV-love merged commit a4d559e into main Sep 29, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant