Repository navigation
fix(bin): relaunch an in-flight worker when its session ends - #4
Merged
Merged
Conversation
The watcher reuses the recorded session-end event and the existing relaunch path, with a 30-minute and daily cap, and skips a deliberate exit or a paused or held task. Claude debug logging stays off unless --claude-debug is passed.
The watcher reuses the recorded session-end event and the existing relaunch path, with a 30-minute and daily cap, and skips a deliberate exit or a paused or held task. Claude debug logging stays off unless --claude-debug is passed.
MrGTV-love
force-pushed
the
fm/fm-idle-exit-relaunch-ship
branch
from
September 29, 2026 00:52
27169d7 to
0b7ee9a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
The captain's standing ask (2026-09-27): "we should not have any stoppages due to session limits." On the captain board (2026-09-28, card fm-idle-exit-relaunch) he chose option A: "build auto-restart plus one debug worker". Option A as he approved it: a small firstmate change where the existing session-end event triggers the existing relaunch path while the task is open, and Claude's debug log is enabled on one idle worker so the next stop names the signal. Result he approved: "A stopped lane restarts in minutes with nothing lost; the next stop tells us who sends the signal." Idle Claude Code worker sessions (2.1.283) take Claude's graceful-shutdown path after an unknown signal (SIGTERM, SIGHUP or orphan detection); the pane prints "Resume this session with:" and returns to the shell. One 01:23:03Z event ended several sessions at once. The proxy, usage limits, jetsam and the turn-end guard were ruled out. The per-worktree SessionEnd hook written by fm-spawn already records event=session-end in the task's busy-state file, but nothing acts on it, so the lane stays stopped until someone notices. bin/fm-control.sh relaunch already handles an already-stopped agent (exit_result=already-stopped).
What Changed
bin/fm-session-end-relaunch-lib.sh, runs from eachbin/fm-watch.shcycle. It looks for an open ship or scout whose busy record saysevent=session-endand whose pane shows no live agent. It then runsbin/fm-control.sh <id> relaunch, which keeps the recorded worktree and all uncommitted work.exitverb.fm-control.sh exitnow writes astate/<id>.control-exitmarker for the current busy generation, so the watcher can tell a deliberate stop.check: <id> auto-relaunched after session-endwake, afailedwake, or apausedwake. Teardown removes the new marker, ledger, and handled files.--claude-debugforbin/fm-control.sh <id> relaunchandbin/fm-spawn.sh --relaunch. It works only with a claude harness. It adds Claude's--debugto the launch and setsCLAUDE_CODE_DIAGNOSTICS_FILEtostate/<id>.claude-diagnostics.jsonl. Claude writes theshutdown_signalevent to that file, so the next stop shows which signal ended the session.fm_exec_timedinbin/fm-timeout-lib.shnow works whenBASHPIDis not set.AGENTS.mdand the docs (architecture, agent-control, configuration, turnend-guard, operational-home-layout skill) describe the new wakes, files, and grace limit. The newtests/fm-session-end-relaunch.test.shcovers this change, and the control and teardown tests are extended.Risk Assessment
✅ Low: The fix round follows the user's three instructions. Only
deadis eligible now, and the test checks that a missing endpoint gets no control call, no wake, and no ledger row. --claude-debug is refused before batch dispatch on every fresh spawn form. The bound is the grace minus 60s, and the beacon is touched before the blocking call. The new tests run the real code and would fail without the fixes. The only remaining note is a non-default configuration edge.Testing
The focused session-end unit test file passed as a baseline. Then I stood up a disposable lab home with a private tmux server, spawned a real Claude scout worker, and stopped it with real signals. The real watcher auto-relaunched the lane with nothing lost. Every guard (cap, deliberate exit, done, missing window, failed-attempt no-loop) and every --claude-debug rule held, and the debug diagnostics file named the SIGHUP that stopped the worker. The worker's Claude account was at its weekly usage limit, so the relaunched agent could not do model work. The runtime checks (agent alive, same worktree, file kept, note appended) prove the restart itself. The auto-relaunch passes ran the watcher from a plain clone of the target commit, because the gate check refuses them from a gate worktree. All evidence is CLI/TUI transcripts, since there is no graphical UI surface. The lab was fully removed afterward.
Evidence: Idle Claude worker stopped by SIGTERM: Resume line, session-end record, dead classifier
Source: Idle Claude worker stopped by SIGTERM: Resume line, session-end record, dead classifier
Evidence: Watcher auto-relaunch: same worktree, uncommitted file kept, note appended, exit_result=already-stopped
Source: Watcher auto-relaunch: same worktree, uncommitted file kept, note appended, exit_result=already-stopped
Evidence: 30-minute cap: one paused wake, no relaunch, no repeat wake
Source: 30-minute cap: one paused wake, no relaunch, no repeat wake
Evidence: --claude-debug relaunch; diagnostics file names SIGHUP
Source: --claude-debug relaunch; diagnostics file names SIGHUP
Evidence: Deliberate fm-control exit is not auto-relaunched
Source: Deliberate fm-control exit is not auto-relaunched
Evidence: Done lane is not auto-relaunched
Source: Done lane is not auto-relaunched
Evidence: Failed relaunch (gate refusal) wakes once and does not retry
Source: Failed relaunch (gate refusal) wakes once and does not retry
Evidence: Missing tmux window is not auto-relaunched
Source: Missing tmux window is not auto-relaunched
Evidence: --claude-debug refused on fresh scout/ship/secondmate/batch and on exit
Source: --claude-debug refused on fresh scout/ship/secondmate/batch and on exit
Evidence: --claude-debug refused on a codex relaunch before checkpoint
Source: --claude-debug refused on a codex relaunch before checkpoint
Evidence: Teardown clears ledger, handled and control-exit files; keeps diagnostics
Source: Teardown clears ledger, handled and control-exit files; keeps diagnostics
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-session-end-relaunch-lib.sh:222- The tick treats amissingendpoint as eligible, but the intent only asks for a pane that 'returns to the shell' (dead, fm-control's exit_result=already-stopped). On tmux, fm-control relaunch always refuses amissingendpoint. fm_control_endpoint_absence_verdict returnsunprovenfor tmux, do_exit dies, and stuck-crewmate-recovery/SKILL.md:45 says 'both exit and relaunch refuse'. Failing sequence: the tmux server or a lane window goes away. SIGHUP triggers Claude's graceful SessionEnd, so the record says session-end, and fm_backend_tmux_agent_state printsmissing. Each tick then runs a relaunch that cannot succeed: checkpoint, note append, die, rollback. It emitsauto-relaunch failedand repeats every 30 minutes, up to 3 times in any rolling 24 hours, indefinitely. tests/fm-session-end-relaunch.test.sh:test_missing_endpoint_relaunches asserts success on a tmux lane only because a recorder stands in for fm-control. Themissingacceptance path exceeds the intent, and the defect lives inside it. The smallest honest remedy is to removemissing(acceptdeadonly) and drop or invert that test. Keeping herdr-onlymissingwould be an extension that needs your approval.bin/fm-spawn.sh:771- The intent asks for Claude's debug log on 'one idle worker'. For an existing idle worker, the only path isfm-control.sh <id> relaunch --claude-debug→fm-spawn --relaunch --claude-debug. The change also accepts --claude-debug on fresh ship spawns, scout spawns, and --secondmate spawns (usage lines 4-6), and no intent requirement needs those. A defect lives inside that extra surface. Batch dispatch (fm-spawn.sh a=repo b=repo --claude-debug) re-execs each pair with shared_args (fm-spawn.sh ~1469-1478), and shared_args does not forward --claude-debug. The flag is silently dropped, and no child launches with debug or diagnostics. Recommended narrower form: accept --claude-debug only together with --relaunch, and refuse it on the other forms.bin/fm-session-end-relaunch-lib.sh:74- The watcher touches .last-watcher-beat once at the top of each cycle. Then session_end_relaunch_tick blocks for up to FM_SESSION_END_TIMEOUT=300s inside that same cycle. The default WATCHER_STALE_GRACE is also 300 (fm_poll_derived_grace for POLL=15). If a relaunch runs close to its bound (a slow fm-spawn --relaunch plus the 90s launch wait and rollback), the beacon reaches the grace. fm-guard then prints the watcher-down banner, and fm-watch-arm refuses a re-arm with 'inspect or stop that watcher'. The watcher is alive, so this is a false alarm that invites someone to kill it mid-relaunch. Remedy options: keep the bound below the grace with margin, or refresh the beacon around the blocking call. This trades against the earlier review choice to widen the bound from 120 to 300, so confirm the direction.bin/fm-timeout-lib.sh:224- The merge commit 2f112f5 changes fm_exec_timed's owner-PID logic (an extra BASHPID guard). Neither parent has this change, so it goes beyond the recorded rebase decision (resolve conflicts, keep both behaviours). It only changes behaviour where BASHPID is unset (bash 3.2): a hard unbound-variable failure becomes owner=$$. This looks benign, and it may help the tasks-axi calls behind the new captain-hold check on /bin/bash. It is noted so the shared-helper change is visible, not hidden in a merge.🔧 Fix applied.
1 info still open:
bin/fm-watch.sh:260- The relaunch bound comes from the watcher's WATCHER_STALE_GRACE (FM_WATCHER_STALE_GRACE, then FM_GUARD_GRACE, then the poll-derived value). The consumers that show a false 'watcher down' banner read a different grace. bin/fm-guard.sh:56 and bin/fm-watch-arm.sh:109 use ${FM_GUARD_GRACE:-300}. docs/configuration.md:2361 documents FM_WATCHER_STALE_GRACE as a separate knob. Concrete case: FM_WATCHER_STALE_GRACE=900 with FM_GUARD_GRACE unset. The bound becomes 840s. fm-guard still uses 300s, so a slow relaunch can make fm-guard report the live watcher as down after 300s. That is the false alarm fix (3) was meant to remove. This happens only with a non-default override, and the user's instruction said to derive the bound from the watcher grace. A narrower remedy, if wanted: pass min(WATCHER_STALE_GRACE, ${FM_GUARD_GRACE:-300}) to fm_session_end_relaunch_scan. With default settings, all three values are 300, and the 60s margin holds.bin/fm-session-end-relaunch-lib.sh:283- The automatic relaunch cannot succeed when the watcher runs from a no-mistakes gate worktree, even against a marked lab home. The library passes FM_STATE_OVERRIDE=$state to fm-control, and bin/fm-gate-refuse-lib.sh cancels the lab-home allowance when any FM_*_OVERRIDE is set. The result is 'auto-relaunch failed ... refusing fleet lifecycle from inside a no-mistakes gate worktree'. On a normal operator checkout this check does not apply, so production behaviour is not affected. The successful live passes ran the same target commit from a plain clone inside the lab. If future gate labs must drive this path from the gate worktree, pass FM_STATE_OVERRIDE only when the watcher itself was given one.bash tests/fm-session-end-relaunch.test.sh(focused baseline, 12 ok)bin/fm-lab-home.sh create $LABplus a privatetmux -L fm-labserver started with a clean env and TREEHOUSE_ROOT=$LAB/poolbin/fm-brief.sh se-scout demo --scoutthenbin/fm-spawn.sh se-scout $LAB/projects/demo --scout --harness claude --effort low(real Claude worker)kill -TERM <claude pid>on the idle worker; read the busy record andfm_backend_agent_state tmux primary:fm-se-scoutbin/fm-watch.shfrom the gate worktree: relaunch refused by the gate check; then a second pass to confirm no retrybin/fm-watch.shfrom a plain clone of 469b432 inside the lab: auto-relaunch, then checks of meta, journal, ledger, wake queue, brief note, git status and panekill -TERMon the relaunched worker, then two watcher passes (30-minute cap)bin/fm-control.sh se-scout relaunch --claude-debug --note ..., thenkill -HUP, then grep for shutdown_signal in state/se-scout.claude-diagnostics.jsonlbin/fm-control.sh se-scout exit, then two watcher passes (deliberate exit skip)appended adonestatus line,kill -TERM, then a watcher pass (done skip)tmux kill-window -t primary:fm-se-scout, then a watcher pass (missing endpoint skip)bin/fm-spawn.sh ... --claude-debugfor fresh scout, ship, secondmate and batch, andbin/fm-control.sh se-scout exit --claude-debug(refusals, rc=1)bin/fm-control.sh se-scout relaunch --harness codex --claude-debugon a recreated shell window (refused before checkpoint)bin/fm-captain-hold.sh complete se-scout --nonethenbin/fm-teardown.sh se-scout(state files cleared)Lab cleanup:tmux -L fm-lab kill-server,rm -rf $LAB, worktreegit statusclean✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.