fix(bin): let a waiting worker spend no turns until it is answered - #48
Merged
Merged
Conversation
tiago-peixoto
force-pushed
the
fm/firstmate-worker-waits-cost-no-tokens
branch
from
September 14, 2026 16:09
d7dae48 to
e082672
Compare
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. - fm-spawn no longer prints a missing-helper error on an early abort. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is already fixed on main by "drive no-mistakes with one foreground call, not a background poll"; this branch takes that text as is and only adds the regression test main shipped without. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto
force-pushed
the
fm/firstmate-worker-waits-cost-no-tokens
branch
from
September 14, 2026 20:41
e082672 to
4b7bd3f
Compare
tiago-peixoto
added a commit
that referenced
this pull request
Sep 17, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. - fm-spawn no longer prints a missing-helper error on an early abort. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is already fixed on main by "drive no-mistakes with one foreground call, not a background poll"; this branch takes that text as is and only adds the regression test main shipped without. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto
added a commit
that referenced
this pull request
Sep 17, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. - fm-spawn no longer prints a missing-helper error on an early abort. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is already fixed on main by "drive no-mistakes with one foreground call, not a background poll"; this branch takes that text as is and only adds the regression test main shipped without. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto
added a commit
that referenced
this pull request
Sep 18, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. - fm-spawn no longer prints a missing-helper error on an early abort. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is already fixed on main by "drive no-mistakes with one foreground call, not a background poll"; this branch takes that text as is and only adds the regression test main shipped without. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto
added a commit
that referenced
this pull request
Sep 18, 2026
…48) fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. The two senders that report the result classified it by matching the text of the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the deferred: line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the deferred: line out of the output rather than assuming it came first. The competing Waiting-section wait-shape that used to ride this patch is dropped in favor of upstream brief text.
tiago-peixoto
added a commit
that referenced
this pull request
Sep 21, 2026
…48) fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. The two senders that report the result classified it by matching the text of the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the deferred: line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the deferred: line out of the output rather than assuming it came first. The competing Waiting-section wait-shape that used to ride this patch is dropped in favor of upstream brief text.
tiago-peixoto
added a commit
that referenced
this pull request
Sep 21, 2026
…s Pi change. They fail with the same tests on main at the base commit 09dc7b3 (CI run 35613893238). Two tests that failed here but not visibly on main (fm-kimi-harness and fm-spawn-compact-adviser-disable-remote) were in main's shard 9, which stopped at an actionlint download error before any test ran. Two regressions came in when the fork's own commits were rebased onto upstream on Sep 21. As you asked, I kept the previous agent's two partial fixes. I checked each one against the logs and locally, and made no other changes. 1. bin/fm-spawn.sh: restored the short staged launch line. Upstream kunchenguid#4994 (a452a79) writes the full launch command to a private file and types only `. <launch-file>` into the pane, because typed lines over about 1,024 bytes get cut off. The fork's #57 (a32fce8) put the old `spawn_send_literal "$T" "$LAUNCH"` back while resolving a merge, so it typed the whole command again. That broke fm-claude-trust, fm-backend-orca, fm-kimi-harness, fm-spawn-dispatch-profile, both fm-spawn-compact-adviser-disable suites, fm-remote-secondmate-trace-context and fm-remote-secondmate-parent-binding. The fix is one line that restores kunchenguid#4994's `spawn_send_literal "$T" ". $(shell_quote "$LAUNCH_FILE")"` and keeps #57's `SPAWN_LAUNCH_SENT=1`. 2. tests/fm-remote-reply.test.sh: the fixture also resolves the `default` key. Upstream kunchenguid#3764 added two decision lines with no key (`needs-decision [at=...]: which base branch?`), and these count under the key `default`. The fork's #48 made the automatic recovery repost wait while the mate has any open decision, so "the one automatic recovery repost was not sent". I confirmed this by printing the open decisions at that point: only `default needs-decision which base branch?` was open. Only the fixture's setup changed; every assertion is unchanged, and #48's wait-while-open rule still applies. Verification on macOS: all 8 spawn test files and fm-remote-reply pass through bin/fm-test-run.sh. With the spawn line reverted, fm-claude-trust fails with the same message as CI ("the launch command did not carry the brief the worker must read"). bash -n and shellcheck -S warning pass on both files. No Pi code was touched
tiago-peixoto
added a commit
that referenced
this pull request
Sep 21, 2026
* fix(pi): stop nested Pi CLI from replacing a live session binding A short-lived child such as fm-spawn's pi --help probe was treated as lock-owned through ancestry and overwrote both markers with a pid that died immediately, causing false supervision alarms. * no-mistakes(review): Remove duplicate Pi turn-end marker regression tests * no-mistakes(review): Point Pi marker verification doc at remaining regression suite * no-mistakes(review): Test Pi self-lock marker binding; drop dead turn-end ownership code * no-mistakes(ci): These four failing shards are not caused by this PR's Pi change. They fail with the same tests on main at the base commit 09dc7b3 (CI run 35613893238). Two tests that failed here but not visibly on main (fm-kimi-harness and fm-spawn-compact-adviser-disable-remote) were in main's shard 9, which stopped at an actionlint download error before any test ran. Two regressions came in when the fork's own commits were rebased onto upstream on Sep 21. As you asked, I kept the previous agent's two partial fixes. I checked each one against the logs and locally, and made no other changes. 1. bin/fm-spawn.sh: restored the short staged launch line. Upstream kunchenguid#4994 (a452a79) writes the full launch command to a private file and types only `. <launch-file>` into the pane, because typed lines over about 1,024 bytes get cut off. The fork's #57 (a32fce8) put the old `spawn_send_literal "$T" "$LAUNCH"` back while resolving a merge, so it typed the whole command again. That broke fm-claude-trust, fm-backend-orca, fm-kimi-harness, fm-spawn-dispatch-profile, both fm-spawn-compact-adviser-disable suites, fm-remote-secondmate-trace-context and fm-remote-secondmate-parent-binding. The fix is one line that restores kunchenguid#4994's `spawn_send_literal "$T" ". $(shell_quote "$LAUNCH_FILE")"` and keeps #57's `SPAWN_LAUNCH_SENT=1`. 2. tests/fm-remote-reply.test.sh: the fixture also resolves the `default` key. Upstream kunchenguid#3764 added two decision lines with no key (`needs-decision [at=...]: which base branch?`), and these count under the key `default`. The fork's #48 made the automatic recovery repost wait while the mate has any open decision, so "the one automatic recovery repost was not sent". I confirmed this by printing the open decisions at that point: only `default needs-decision which base branch?` was open. Only the fixture's setup changed; every assertion is unchanged, and #48's wait-while-open rule still applies. Verification on macOS: all 8 spawn test files and fm-remote-reply pass through bin/fm-test-run.sh. With the spawn line reverted, fm-claude-trust fails with the same message as CI ("the launch command did not carry the brief the worker must read"). bash -n and shellcheck -S warning pass on both files. No Pi code was touched
tiago-peixoto
added a commit
that referenced
this pull request
Sep 25, 2026
…48) fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. The two senders that report the result classified it by matching the text of the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the deferred: line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the deferred: line out of the output rather than assuming it came first. The competing Waiting-section wait-shape that used to ride this patch is dropped in favor of upstream brief text.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot now spends no model turns until something actually changes.
Three contracts, each reproduced in an isolated Herdr lab with real Pi workers and the fleet's own scripts, then pinned by regression tests.
needs-decision:orblocked:ends its turn and takes no further turn until firstmate's answer arrives.fm-sendanswer still wakes it.This branch has been rebased onto current
main(efaca09, which restores the fleet'sfm/*no-mistakes exemption).Contract 3's definition-of-done half is now solved on
mainby 01d4c86, "drive no-mistakes with one foreground call, not a background poll", so that commit's text is taken as is and this branch no longer editsbin/fm-dod-lib.sh.What this branch adds there is the regression test that commit shipped without.
Tracks upstream kunchenguid#4228 (workers spend full-context model turns while they wait on gates, open decisions, and re-rings).
Contract 1 also addresses kunchenguid#3061 (a parked worker self-polling while its decision is open).
Open PR kunchenguid#2678 briefs workers never to wait in a foreground blocking sleep, preferring a tracked background job or brief polls between other work.
This change deliberately does the opposite and waits inside one blocking shell command, because its goal is zero model turns while waiting, and every background check or poll is a model turn.
Related upstream: kunchenguid#1095 (sleep-and-poll turns between
no-mistakes axi statuscalls) is the closest match to contract 3, and kunchenguid#3057 (a backgrounded pipeline call leaves the run parked at a gate with nobody listening) is the constraint contract 3 is built around.Diagnosis
Evidence came from the real Pi work-account session logs first, then from the lab.
Contract 1: workers polled their own inbox
paused:heavy-slot waits ran up to 44 sleep-and-check turns.Right after
needs-decision:real workers mostly already took no turns; the turns that did happen came from supervisor holding notes and automatic nudges, which is contract 2.Contract 2: supervisors poked waiting workers
Contract 3: the pipeline wait was a polling loop by instruction
bin/fm-dod-lib.shtold no-mistakes workers to "background the drive call and pollno-mistakes axi statusfrom a separate call".mainhas since fixed the instruction itself; this branch keeps the lab evidence below because it is what measured the cost, and adds the missing regression test.Naming the wait, not only forbidding the wrong ones
Forbidding a backgrounded poll is not enough on its own.
The brief told a worker what not to do without naming the wait it may actually use, and Claude Code can refuse a sleep-then-check command while pointing at
run_in_background, which is the one shape a waiting worker must not take.The Waiting section now names the blocking foreground
untilloop for that harness.The probes behind it are recorded in
docs/verification/runtime-backends.md.Change
bin/fm-brief.sh: deletes the "natural checkpoint" clause and adds one# Waitingsection to the ship and scout briefs.It tells the worker to end its turn after
needs-decision:orblocked:, and to wait on external state with ONE blocking command (no-mistakes axi run/respond --wait,gh pr checks --watch, or anuntilloop), never backgrounded in order to poll it.It bounds that command per harness, from each vendor's own code (recorded in
docs/verification/runtime-backends.md):timeoutof at most 2700 seconds, because Pi's bash tool has none by default; 2700 stays under the watcher's 3600-second busy-turn bound.timeoutof 600000 ms, because the default is 2 minutes.write_stdinpolls of up to 300000 ms.It then names the foreground
untilloop as the sanctioned wait for Claude Code.bin/fm-classify-lib.sh:status_own_open_decisions, the open decisions a task raised itself, excluding the reservedpending-reply-keys a supervisor library raises about it.bin/fm-send.sh:--automatic.While the target task has its own open decision or blocker, an automatic send writes nothing, rings nothing, prints one
deferred:line, and exits 4.The caller keeps its retry state.
--automaticand treats exit 4 as "retry later", not failure:bin/fm-bootstrap.sh);bin/fm-config-push.sh);bin/fm-config-inherit-lib.sh);bin/fm-secondmate-reconcile.sh).deferred:*, which discards the exit status the contract is actually stated in.fm-sendrunsbin/fm-guard.shas a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is.The banner lands ahead of the
deferred:line, so the match fell through and a waiting mate was reported as a failed send with the banner as the reason.This is why the first CI run of this branch failed
Behavior portable serial 1intests/fm-secondmate-sync.test.shwhile the same suite passed locally, where the fixtures' primary checkout is on its default branch and no banner is printed.bin/fm-bootstrap.shandbin/fm-config-inherit-lib.shnow classify on fm-send's exit status and select thedeferred:line out of the output instead of assuming it came first.bin/fm-secondmate-reconcile.shalready branched on exit 4 and is unchanged.bin/fm-pending-reply-lib.sh: the pending-reply recovery repost stays unattempted while the mate has its own open decision.AGENTS.mdsection 7: one sentence, "while a worker's decision or blocker is open, send it only the answer"..agents/skills/bootstrap-diagnostics: explains the newNUDGE_SECONDMATES: ... deferred:line and says to fold the re-read into the answer.bin/fm-spawn.sh: an abort before any lease was armed no longer printsspawn_return_abort_lease: command not found, because the EXIT trap is installed before that helper is defined; found while reproducing.The doorbell is unchanged: it is typed only when a record lands, and the watcher re-rings only an unacknowledged deliberate record, so a worker with an empty inbox is never rung.
A deliberate answer to a waiting worker still re-rings until acknowledged, which is what should wake it.
Automatic paths enumerated
Every sender into a worker or secondmate that acts on its own schedule, and what each does now:
The away daemon, stale, paused, and captain-call wakes all go to firstmate, never the worker.
Deliberate senders are unchanged:
fm-sendanswers,fm-secondmate-restart, the backlog-handoff receiver wake, the stow cascade,fm-control, and the spawn launch.Wedge backstops
None of these wakes the worker; each wakes firstmate:
FM_BUSY_TURN_MAX_SECS, default 3600 seconds), which flags a turn that runs too long, including a blocking wait that never returns;FM_TASK_INBOX_RING_MAXattempts;OPEN DECISIONSsection of every wake drain, which keeps an unanswered decision in front of firstmate until it is answered.How this layers with upstream neighbours
That bounds the supervisor side of an open decision; this change bounds the worker side, so the two compose without touching the same code.
fm-procevent-when, a generic supervisor-side primitive this change neither uses nor alters.whenwatch rings the worker on each pipeline-state change, and its DoD tells the worker to end its turn and wait for that ring.It rewrites the same
bin/fm-dod-lib.shparagraph this fork'smainhas already rewritten, so only one of them should land.The landed approach keeps the worker inside the drive call, so the call's own return is the signal and no watch has to observe a transition; that sidesteps the lost-transition race feat(bin): watcher-rung pipeline-state waits for no-mistakes spawns kunchenguid/firstmate#3979's own review raised, where a parked state reached while the watch is restarting never fires and the worker stalls.
If a pipeline-state ring does land later, it is an automatic sender and should pass
--automatic, so it also leaves a worker waiting on its own decision alone.Lab evidence
An isolated Herdr lab (
fm-herdr-lab.shsession, neverdefault), one home per variant, realfm-spawn,fm-send, andfm-secondmate-reconcilefrom each revision, real Pi 0.85.1 workers.Old =
mainat bda9903, new = this branch.The model was
cursor/composer-2.5through pi-cursor-sdk, because the Codex quota was exhausted; that extension runs tools on Cursor's side and replays them into Pi, so Pi's bashtimeoutdoes not bind it.The pipeline was a stand-in
no-mistakeswith axi's documented hold semantics that logged every call and published its outcome a fixed time after firstmate's validation steer.Turn counts come from the Pi session logs; pipeline calls come from the stand-in's own timestamped log.
Lab noise: both old and new workers spent some calls inspecting the stand-in to find its outcome marker, and one old-code worker ran a search across the home directory; later runs confine lab agents to their worktree.
Decision wait (contract 1)
needs-decision, before any answerfm-send --resolve-keydone: committed 55ad6d6done: committed 80ad98dThe lab worker had nothing to be unsure about, so it does not reproduce the checkpoint polling the real logs show; it proves the new brief keeps the zero and the answer still wakes.
Automatic nudge to a waiting secondmate (contract 2)
The harness opened
needs-decision [key=lab-name]for a real idle Pi secondmate, then ran the realfm-secondmate-reconcile.sh notifybackstop.fm-send --resolve-key lab-namesent: mate-old orphan_in_flightdeferred: mate-new orphan_in_flightdone [corr=...]Pipeline wait to an outcome (contract 3)
Outcome published 780 s after the validation steer.
axi run --wait 15s, thenaxi run --wait 8m &in the background, then 30axi statuspolls every 30 sdone: PR ... checks green16 s after the outcomeaxi run --wait 44m, re-run once with no status call in betweendone: PR ... checks green4 s after the outcomePipeline parked at a gate (kunchenguid#3057)
The stand-in's
axi runreturned a parked review gate (awaiting_agent: parked, oneauto-fixfinding) 300 s after the steer.New code, from the stand-in's log:
The worker was still inside its blocking call when the gate came back, and answered it 11 s later with no status call in between, so the run never sat parked with nobody listening.
Old code also answered this gate, after an
axi run --wait 2m, anaxi status, adaemon status, a secondaxi run --wait 10m, and one moreaxi status.The run was stopped during the fix round for a machine-wide quiet window, after the gate was answered; over the same span the old worker made 55 model calls and the new one 26.
All lab sessions were torn down by the helper; its tripwire was gone afterwards.
Tests
tests/fm-send-inbox.test.sh: an automatic send to a task with an open decision exits 4, writes no record, and types nothing; the keyed answer still lands and rings; a reserved pending-reply key does not defer; an explicit-target automatic send is refused.tests/fm-pending-reply.test.sh: recovery is not attempted while the mate has an open decision, and sends once afterresolved.tests/fm-secondmate-reconcile.test.sh: the notify defers with no inbox record and no cooldown; a queued request stays queued; after the answer it is delivered.tests/fm-secondmate-sync.test.sh: the bootstrap nudge defers, prints no "send failed", and keeps its marker.A second case puts the fixture's primary checkout on a feature branch, which is what made CI differ from a local run, and asserts up front that the guard banner really does precede the
deferred:line before checking that the nudge is still reported as deferred, so the case cannot pass vacuously.tests/fm-task-inbox.test.sh: the watcher never rings a worker waiting on its decision.tests/fm-brief.test.sh: ship, scout, and secondmate briefs carry the waiting contract and no checkpoint polling, name the wait a Claude Code worker may use, and keep the one-foreground-call definition of donemainintroduced.tests/fm-spawn-dispatch-profile.test.sh: an early spawn refusal prints no shell error.tests/fm-remote-reply.test.sh: its escalation scenario now answers the mate's earlier open decision and blocker before exercising the recovery repost, which would otherwise correctly wait.bin/fm-lint.shandbin/fm-doc-audience-check.shpass.Upstream
Upstream https://github.com/kunchenguid/firstmate
mainhas the same behavior, read on 2026-09-11:bin/fm-dod-lib.shline 236: "So background the drive call and pollno-mistakes axi statusfrom a separate call". This fork'smainhas since fixed that line; upstream has not.bin/fm-brief.shline 210: "and at any natural checkpoint when you are unsure - list $INBOX_DIR/*.msg".bin/fm-send.shhas no automatic-send guard,bin/fm-secondmate-reconcile.shline 527 sends the notify fire-and-forget whatever the mate's open decisions, and upstream has no blocking-wait guidance in the brief.Nothing was filed upstream.