Skip to content

fix(bin): confirm a spawned agent started and clear pending shell input before relaunch - #37

Merged
marano merged 2 commits into
mainfrom
fm/fm-pane-delivery-batch
Sep 19, 2026
Merged

marano merged 2 commits into
mainfrom
fm/fm-pane-delivery-batch

Conversation

@marano

@marano marano commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Scope

This PR delivers item 1 of 2 from the intent below: a spawn that reports success on a garbled launch, and relaunch clearing a stuck shell prompt before typing.
Item 2 - the alarm for a steer doorbell stuck unsent in a live worker's composer - is not in this PR and follows in a separate PR.

CI finding: how a started Pi reads under tmux

The first CI round failed tests/fm-session-start.test.sh: its fake tmux reported a respawned Pi secondmate's foreground command as node with no foreground-process data, so the new launch confirmation could not name an agent and waited out its window.
Before changing anything, the real Pi pane identity was checked.
Pi is not installed on the validation machine, so the recorded live evidence was used: real Pi under tmux reports #{pane_current_command} as pi-launcher with foreground processes pi-signed and pi (pi 0.82.0 on 2026-08-03 and pi 0.84.4 on 2026-09-06, docs/verification/runtime-backends.md "Agent liveness name sources"), which classifies alive.
The Pi engine's other recorded shape - kernel name node, argv0 pi - was then checked live in a private tmux server: the pane reported node, and the real fm_backend_agent_state tmux read the pane alive from its foreground process.
So production already recognises a started Pi; the test stub was the unfaithful part.
The choice: the stub now reports pi-launcher after the respawn, and the launch confirmation was not loosened.
The same round also fixed this PR's own test to accept bash's bare > continuation prompt as well as zsh's dquote cmdsubst>.

Intent

"always make sure your lanes are filled with work" (2026-09-19). Standing lane rule (2026-09-17): firstmate's own work earns a lane when "they prevent firstmate from working properly"; "I cant get work done if firstmate is not shipshape".

Two recorded defects, both the same failure: text typed into a worker's terminal pane does not land, and nothing notices, so a worker silently never starts or silently stops.

  1. A spawn reports success on a garbled launch (recorded 2026-09-18):
    2026-09-18: fm-spawn.sh reported spawned blu-orgunit-lock-live-tenant-walkthrough but no agent ever started. The launch command typed into the fresh tmux pane arrived garbled mid-string ("...bin/fm-operational-in" followed directly by "env -u CURSOR_AGENT ..." from the start of the command), leaving zsh at a dquote cmdsubst quote> continuation prompt. A following fm-control.sh relaunch typed its command into that same open quote and correctly reported that no agent came up; clearing the prompt with C-c and relaunching worked.
    Two defects: (1) spawn reports success without confirming an agent is running - same family as fm-spawn-accepts-missing-harness, but here the harness exists and the typed command itself was corrupted; (2) neither spawn nor relaunch clears a shell continuation prompt before typing, so one bad launch poisons every retry in that endpoint.

  2. A steer's doorbell sits queued and unsent (recorded 2026-09-18):
    A steer's doorbell can sit QUEUED and unsent in a live worker's composer, and every re-ring queues
    behind it, so the escalation ladder never reaches the worker. Observed 2026-09-18 on
    blu-3192-app-web-consume-lock-signals.

WHAT WAS SEEN:

  • The instruction itself was delivered correctly - the durable inbox record existed and was intact.
  • The doorbell line was typed into the worker's pane and never submitted. The pane read
    "Press up to edit queued messages" and "ctrl+x ctrl+s to send now".
  • The worker was ALIVE, idle, and had no idea a message was waiting. It simply stopped.
  • bin/fm-control.sh interrupt cleared the composer; a fresh send then landed immediately and the
    worker acknowledged and resumed. Nothing was lost.

WHY THE EXISTING SAFETY NET DID NOT CATCH IT. The durable-inbox design is correct and is why no
work was lost: the doorbell is explicitly NOT treated as delivery proof, and the watcher re-rings
an unacknowledged message and escalates a stuck one. But a re-ring is another doorbell typed into
the SAME composer - which is still stuck - so each retry queues behind the last. The ladder climbs
a wall. The stale alarm did eventually fire, which is how firstmate found it, but the diagnosis
came from reading the pane by eye, not from anything the system reported.

Note the asymmetry worth preserving: for a DEAD agent the watcher already reports the right thing -
"unread firstmate instruction ... the worker's agent has exited or its endpoint is missing, so the
doorbell was not typed". There is no equivalent for a LIVE agent whose composer will not accept the
submit, which is the harder and quieter case.

THE CHANGE, in outline rather than prescription: the composer classifier in bin/fm-composer-lib.sh
already distinguishes empty / pending / pending-unproven / unknown, and the away daemon already
refuses to inject into anything but a confirmed-empty composer. That knowledge is not reaching the
steer path's retry logic. A doorbell whose composer reads pending across successive attempts is a
distinct condition - text stuck, agent alive - and should be reported as that rather than retried
indefinitely, so a supervisor is told "this worker cannot receive messages" instead of inferring it.

PROOF OBLIGATIONS - each needs a mutant that reds BY NAME:

  • A live worker whose composer holds unsent text across N re-rings raises a distinct alarm naming
    that condition, not a generic stale wake.
  • Refusal test that must stay green: an ordinary busy worker mid-turn, whose composer legitimately
    holds queued text that WILL submit, is not alarmed on. This is the hard part - OpenCode keeps
    queued text visible while working, and fm_composer_queued_enter_verdict exists precisely because
    visible text alone does not prove a swallowed Enter. Do not turn a normal busy pane into noise.
  • Recovery is not automated by this card. Report the condition; a human or a later card decides
    whether an automatic interrupt is safe.

RELATED, same root family: fm-supervisor-ghost-injection, merged 2026-09-18, fixed the case where
an UNDELIVERABLE escalation was typed into a composer anyway and surfaced hours later. That fix
covers the away daemon's own escalations. It does NOT cover this path - a steer's doorbell to a
worker - so do not assume it is already handled.

What Changed

  • bin/fm-spawn.sh now confirms a launch on tmux, where fm_backend_agent_state can prove an agent from the pane's processes. It polls until the agent reads alive. If the endpoint still reads agent-free, it sends Ctrl+C, retypes the launch once, and re-polls. A second miss appends a failed: status line and exits non-zero. A fresh spawn closes its endpoint and rolls back its record. A relaunch keeps its endpoint and record. Poll count and interval are tunable via FM_SPAWN_LAUNCH_POLLS and FM_SPAWN_LAUNCH_POLL_INTERVAL. Other backends launch unconfirmed.
  • --relaunch now sends Ctrl+C to the adopted shell once, before the first line it types (the worktree cd, or the GOTMPDIR export). This means a garbled or unclosed earlier launch, such as a dquote cmdsubst quote> continuation prompt, no longer swallows the replacement command. If the clear itself fails, the relaunch refuses instead of typing into pending input.
  • Adds fm_backend_launch_confirmable in bin/fm-backend.sh, which returns true only for tmux. Renames rovo_endpoint_cleanup to spawn_launch_endpoint_cleanup so the launch-failure path shares it. Documents the contract in the fm-spawn.sh header and docs/agent-control.md. Adds tests/fm-spawn-launch-confirm.test.sh and updates existing spawn, harness and secondmate tests, plus the shared fixtures, to fit the new confirmation and clearing steps.
  • This branch covers only the spawn and relaunch defect. It does not change the steer doorbell path or the composer classifier for a stuck queued doorbell.

Risk Assessment

⚠️ Medium: The spawn-confirmation half is well-bounded and covered by a new test. The doorbell-stuck-in-composer alarm the intent requires is absent, so half of the stated acceptance criteria is unmet.

Testing

I ran the new launch-confirm test and the existing spawn tests the change touched. The three new behaviours pass: a spawn whose launch never started fails and closes its pane; a spawn clears a cut launch's continuation prompt and starts the agent; a relaunch clears the prompt before typing. One failure in fm-backend.test.sh (the symlinked-prefix Treehouse lock case) also fails identically at the base commit, so it is not caused by this change. fm-backlog-atomicity could not run because timeout is missing here. The steer-doorbell alarm, defect 2 of the intent, is not in the diff. The user declined that finding, so I left it untested rather than failing the change.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Spawn whose launch never started fails and closes its pane instead of reporting success ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a spawn whose launch never started fails and closes its pane'
Spawn clears a cut launch's shell continuation prompt (Ctrl+C) and the agent then starts ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a spawn clears a cut launch's continuation prompt and starts the agent'
Relaunch clears a stuck continuation prompt before typing its launch ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a relaunch clears a continuation prompt before typing its launch'
Live worker whose composer holds unsent doorbell text across N re-rings raises a distinct alarm, and an ordinary busy worker is not alarmed on ⏸️ untested no The change does not implement this. The author declined review finding review-1 (defect 2). Adding it means a change to bin/fm-composer-lib.sh and the watcher re-ring logic, which the author has to de…

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 issues (1 error, 1 warning)
  • 🚨 bin/fm-spawn.sh:3549 - The intent names two required fixes; only the first (spawn launch confirmation and Ctrl+C clear) is implemented. Defect 2, a steer's doorbell stuck in a live worker's composer, has no code in this diff. The intent requires: 'A live worker whose composer holds unsent text across N re-rings raises a distinct alarm naming that condition, not a generic stale wake.' It also requires a refusal test that an ordinary busy worker, whose queued text will submit, is not alarmed on, using fm_composer_queued_enter_verdict, and a mutant for each that reds by name. The diff touches only bin/fm-backend.sh and bin/fm-spawn.sh, plus tests and docs. It does not touch bin/fm-composer-lib.sh, the steer/doorbell path, or the watcher re-ring and escalation logic. Whether to drop this part of the scope or add it to this change is a decision for the author.
  • ⚠️ bin/fm-spawn.sh:3585 - The single in-place retry only fires when the endpoint reads 'dead' after the full poll window (60 x 0.5s = 30s by default). If a slow-starting agent has not yet become the pane's foreground process by then, the state reads dead. spawn_confirm_launch then sends Ctrl+C and types the launch again, which can interrupt the starting agent or run a second launch. This is a narrow race, and the retry is a deliberate part of the design.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Spawn whose launch never started fails and closes its pane instead of reporting success ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a spawn whose launch never started fails and closes its pane'
Spawn clears a cut launch's shell continuation prompt (Ctrl+C) and the agent then starts ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a spawn clears a cut launch's continuation prompt and starts the agent'
Relaunch clears a stuck continuation prompt before typing its launch ✅ pass live tests/fm-spawn-launch-confirm.test.sh: 'ok - a relaunch clears a continuation prompt before typing its launch'
Live worker whose composer holds unsent doorbell text across N re-rings raises a distinct alarm, and an ordinary busy worker is not alarmed on ⏸️ untested no The change does not implement this. The author declined review finding review-1 (defect 2). Adding it means a change to bin/fm-composer-lib.sh and the watcher re-ring logic, which the author has to de…
  • bash tests/fm-spawn-launch-confirm.test.sh
  • bash tests/fm-backend.test.sh (one failure, also present at base 70fc080)
  • bash tests/fm-spawn-worktree-settle.test.sh
  • bash tests/fm-secondmate-harness.test.sh
  • bash tests/fm-trace-context-spawn.test.sh
  • bash tests/fm-backlog-atomicity.test.sh (could not run: timeout command not found on this machine)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

marano and others added 2 commits September 19, 2026 13:00
…mpt before relaunch

A launch typed into a pane can arrive cut short - on macOS a line over the
1024-byte canonical-input limit typed before the shell's line editor runs is
truncated - leaving the shell at a continuation prompt with no agent, while
fm-spawn.sh still reported success. The next relaunch then typed into the
same open quote.

On tmux, where the agent-state classifier proves an agent from the pane's own
processes, every launch now waits for the agent to read alive, retries once in
place after clearing the shell's pending input, and otherwise fails, closing a
fresh spawn's endpoint and rolling back its record. A relaunch clears the
adopted shell's pending input before typing anything.
…ane as pi-launcher

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@marano
marano merged commit 2f8f630 into main Sep 19, 2026
16 checks passed
@marano
marano deleted the fm/fm-pane-delivery-batch branch September 19, 2026 18:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant