fix(bin): confirm a spawned agent started and clear pending shell input before relaunch - #37
Merged
Merged
Conversation
…mpt before relaunch A launch typed into a pane can arrive cut short - on macOS a line over the 1024-byte canonical-input limit typed before the shell's line editor runs is truncated - leaving the shell at a continuation prompt with no agent, while fm-spawn.sh still reported success. The next relaunch then typed into the same open quote. On tmux, where the agent-state classifier proves an agent from the pane's own processes, every launch now waits for the agent to read alive, retries once in place after clearing the shell's pending input, and otherwise fails, closing a fresh spawn's endpoint and rolling back its record. A relaunch clears the adopted shell's pending input before typing anything.
…ane as pi-launcher Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
This PR delivers item 1 of 2 from the intent below: a spawn that reports success on a garbled launch, and relaunch clearing a stuck shell prompt before typing.
Item 2 - the alarm for a steer doorbell stuck unsent in a live worker's composer - is not in this PR and follows in a separate PR.
CI finding: how a started Pi reads under tmux
The first CI round failed
tests/fm-session-start.test.sh: its fake tmux reported a respawned Pi secondmate's foreground command asnodewith no foreground-process data, so the new launch confirmation could not name an agent and waited out its window.Before changing anything, the real Pi pane identity was checked.
Pi is not installed on the validation machine, so the recorded live evidence was used: real Pi under tmux reports
#{pane_current_command}aspi-launcherwith foreground processespi-signedandpi(pi 0.82.0 on 2026-08-03 and pi 0.84.4 on 2026-09-06,docs/verification/runtime-backends.md"Agent liveness name sources"), which classifies alive.The Pi engine's other recorded shape - kernel name
node, argv0pi- was then checked live in a private tmux server: the pane reportednode, and the realfm_backend_agent_state tmuxread the pane alive from its foreground process.So production already recognises a started Pi; the test stub was the unfaithful part.
The choice: the stub now reports
pi-launcherafter the respawn, and the launch confirmation was not loosened.The same round also fixed this PR's own test to accept bash's bare
>continuation prompt as well as zsh'sdquote cmdsubst>.Intent
"always make sure your lanes are filled with work" (2026-09-19). Standing lane rule (2026-09-17): firstmate's own work earns a lane when "they prevent firstmate from working properly"; "I cant get work done if firstmate is not shipshape".
Two recorded defects, both the same failure: text typed into a worker's terminal pane does not land, and nothing notices, so a worker silently never starts or silently stops.
A spawn reports success on a garbled launch (recorded 2026-09-18):
2026-09-18: fm-spawn.sh reported
spawned blu-orgunit-lock-live-tenant-walkthroughbut no agent ever started. The launch command typed into the fresh tmux pane arrived garbled mid-string ("...bin/fm-operational-in" followed directly by "env -u CURSOR_AGENT ..." from the start of the command), leaving zsh at adquote cmdsubst quote>continuation prompt. A followingfm-control.sh relaunchtyped its command into that same open quote and correctly reported that no agent came up; clearing the prompt with C-c and relaunching worked.Two defects: (1) spawn reports success without confirming an agent is running - same family as fm-spawn-accepts-missing-harness, but here the harness exists and the typed command itself was corrupted; (2) neither spawn nor relaunch clears a shell continuation prompt before typing, so one bad launch poisons every retry in that endpoint.
A steer's doorbell sits queued and unsent (recorded 2026-09-18):
A steer's doorbell can sit QUEUED and unsent in a live worker's composer, and every re-ring queues
behind it, so the escalation ladder never reaches the worker. Observed 2026-09-18 on
blu-3192-app-web-consume-lock-signals.
WHAT WAS SEEN:
"Press up to edit queued messages" and "ctrl+x ctrl+s to send now".
bin/fm-control.sh interruptcleared the composer; a fresh send then landed immediately and theworker acknowledged and resumed. Nothing was lost.
WHY THE EXISTING SAFETY NET DID NOT CATCH IT. The durable-inbox design is correct and is why no
work was lost: the doorbell is explicitly NOT treated as delivery proof, and the watcher re-rings
an unacknowledged message and escalates a stuck one. But a re-ring is another doorbell typed into
the SAME composer - which is still stuck - so each retry queues behind the last. The ladder climbs
a wall. The stale alarm did eventually fire, which is how firstmate found it, but the diagnosis
came from reading the pane by eye, not from anything the system reported.
Note the asymmetry worth preserving: for a DEAD agent the watcher already reports the right thing -
"unread firstmate instruction ... the worker's agent has exited or its endpoint is missing, so the
doorbell was not typed". There is no equivalent for a LIVE agent whose composer will not accept the
submit, which is the harder and quieter case.
THE CHANGE, in outline rather than prescription: the composer classifier in bin/fm-composer-lib.sh
already distinguishes empty / pending / pending-unproven / unknown, and the away daemon already
refuses to inject into anything but a confirmed-empty composer. That knowledge is not reaching the
steer path's retry logic. A doorbell whose composer reads pending across successive attempts is a
distinct condition - text stuck, agent alive - and should be reported as that rather than retried
indefinitely, so a supervisor is told "this worker cannot receive messages" instead of inferring it.
PROOF OBLIGATIONS - each needs a mutant that reds BY NAME:
that condition, not a generic stale wake.
holds queued text that WILL submit, is not alarmed on. This is the hard part - OpenCode keeps
queued text visible while working, and fm_composer_queued_enter_verdict exists precisely because
visible text alone does not prove a swallowed Enter. Do not turn a normal busy pane into noise.
whether an automatic interrupt is safe.
RELATED, same root family: fm-supervisor-ghost-injection, merged 2026-09-18, fixed the case where
an UNDELIVERABLE escalation was typed into a composer anyway and surfaced hours later. That fix
covers the away daemon's own escalations. It does NOT cover this path - a steer's doorbell to a
worker - so do not assume it is already handled.
What Changed
bin/fm-spawn.shnow confirms a launch on tmux, wherefm_backend_agent_statecan prove an agent from the pane's processes. It polls until the agent reads alive. If the endpoint still reads agent-free, it sends Ctrl+C, retypes the launch once, and re-polls. A second miss appends afailed:status line and exits non-zero. A fresh spawn closes its endpoint and rolls back its record. A relaunch keeps its endpoint and record. Poll count and interval are tunable viaFM_SPAWN_LAUNCH_POLLSandFM_SPAWN_LAUNCH_POLL_INTERVAL. Other backends launch unconfirmed.--relaunchnow sends Ctrl+C to the adopted shell once, before the first line it types (the worktreecd, or theGOTMPDIRexport). This means a garbled or unclosed earlier launch, such as adquote cmdsubst quote>continuation prompt, no longer swallows the replacement command. If the clear itself fails, the relaunch refuses instead of typing into pending input.fm_backend_launch_confirmableinbin/fm-backend.sh, which returns true only for tmux. Renamesrovo_endpoint_cleanuptospawn_launch_endpoint_cleanupso the launch-failure path shares it. Documents the contract in thefm-spawn.shheader anddocs/agent-control.md. Addstests/fm-spawn-launch-confirm.test.shand updates existing spawn, harness and secondmate tests, plus the shared fixtures, to fit the new confirmation and clearing steps.Risk Assessment
Testing
I ran the new launch-confirm test and the existing spawn tests the change touched. The three new behaviours pass: a spawn whose launch never started fails and closes its pane; a spawn clears a cut launch's continuation prompt and starts the agent; a relaunch clears the prompt before typing. One failure in fm-backend.test.sh (the symlinked-prefix Treehouse lock case) also fails identically at the base commit, so it is not caused by this change. fm-backlog-atomicity could not run because
timeoutis missing here. The steer-doorbell alarm, defect 2 of the intent, is not in the diff. The user declined that finding, so I left it untested rather than failing the change.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-spawn.sh:3549- The intent names two required fixes; only the first (spawn launch confirmation and Ctrl+C clear) is implemented. Defect 2, a steer's doorbell stuck in a live worker's composer, has no code in this diff. The intent requires: 'A live worker whose composer holds unsent text across N re-rings raises a distinct alarm naming that condition, not a generic stale wake.' It also requires a refusal test that an ordinary busy worker, whose queued text will submit, is not alarmed on, using fm_composer_queued_enter_verdict, and a mutant for each that reds by name. The diff touches only bin/fm-backend.sh and bin/fm-spawn.sh, plus tests and docs. It does not touch bin/fm-composer-lib.sh, the steer/doorbell path, or the watcher re-ring and escalation logic. Whether to drop this part of the scope or add it to this change is a decision for the author.bin/fm-spawn.sh:3585- The single in-place retry only fires when the endpoint reads 'dead' after the full poll window (60 x 0.5s = 30s by default). If a slow-starting agent has not yet become the pane's foreground process by then, the state reads dead. spawn_confirm_launch then sends Ctrl+C and types the launch again, which can interrupt the starting agent or run a second launch. This is a narrow race, and the retry is a deliberate part of the design.✅ **Test** - passed
✅ No issues found.
bash tests/fm-spawn-launch-confirm.test.shbash tests/fm-backend.test.sh(one failure, also present at base 70fc080)bash tests/fm-spawn-worktree-settle.test.shbash tests/fm-secondmate-harness.test.shbash tests/fm-trace-context-spawn.test.shbash tests/fm-backlog-atomicity.test.sh(could not run:timeoutcommand not found on this machine)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.