Skip to content

fix(bin): let a waiting worker spend no turns until it is answered - #48

Merged
tiago-peixoto merged 1 commit into
mainfrom
fm/firstmate-worker-waits-cost-no-tokens
Sep 14, 2026
Merged

tiago-peixoto merged 1 commit into
mainfrom
fm/firstmate-worker-waits-cost-no-tokens

Conversation

@tiago-peixoto

@tiago-peixoto tiago-peixoto commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Summary

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot now spends no model turns until something actually changes.
Three contracts, each reproduced in an isolated Herdr lab with real Pi workers and the fleet's own scripts, then pinned by regression tests.

  1. A worker that appends needs-decision: or blocked: ends its turn and takes no further turn until firstmate's answer arrives.
  2. No automatic sender wakes a worker while its own decision or blocker is open; firstmate's deliberate fm-send answer still wakes it.
  3. External waits sit inside one blocking shell command bounded by the harness's command ceiling, instead of a sleep-then-check loop of model turns, and the worker stays attached so it acts on whatever the wait returns.

This branch has been rebased onto current main (efaca09, which restores the fleet's fm/* no-mistakes exemption).
Contract 3's definition-of-done half is now solved on main by 01d4c86, "drive no-mistakes with one foreground call, not a background poll", so that commit's text is taken as is and this branch no longer edits bin/fm-dod-lib.sh.
What this branch adds there is the regression test that commit shipped without.

Tracks upstream kunchenguid#4228 (workers spend full-context model turns while they wait on gates, open decisions, and re-rings).

Contract 1 also addresses kunchenguid#3061 (a parked worker self-polling while its decision is open).

Open PR kunchenguid#2678 briefs workers never to wait in a foreground blocking sleep, preferring a tracked background job or brief polls between other work.
This change deliberately does the opposite and waits inside one blocking shell command, because its goal is zero model turns while waiting, and every background check or poll is a model turn.

Related upstream: kunchenguid#1095 (sleep-and-poll turns between no-mistakes axi status calls) is the closest match to contract 3, and kunchenguid#3057 (a backgrounded pipeline call leaves the run parked at a gate with nobody listening) is the constraint contract 3 is built around.

Diagnosis

Evidence came from the real Pi work-account session logs first, then from the lab.

Contract 1: workers polled their own inbox

  • Trigger: the ship and scout brief told workers to list the inbox "at any natural checkpoint when you are unsure".
  • Masking condition: most checkpoints find an empty inbox, so each extra turn looks like diligence; the doorbell path works either way, which hid that the checkpoint clause is never needed.
  • Symptom: 612 inbox listings in the sampled logs, 481 of them unprompted and only 131 after a doorbell; paused: heavy-slot waits ran up to 44 sleep-and-check turns.

Right after needs-decision: real workers mostly already took no turns; the turns that did happen came from supervisor holding notes and automatic nudges, which is contract 2.

Contract 2: supervisors poked waiting workers

  • Trigger: six automatic senders nudge a secondmate on their own schedule, with no regard for its open decisions.
  • Masking condition: each nudge is fire-and-forget or retried by marker, so each looks harmless in isolation.
  • Symptom: the doorbell starts a full model turn, often a long one, in a worker whose only job was to wait for an answer.

Contract 3: the pipeline wait was a polling loop by instruction

  • Trigger: bin/fm-dod-lib.sh told no-mistakes workers to "background the drive call and poll no-mistakes axi status from a separate call".
  • Masking condition: the no-mistakes skill gives the same advice for harnesses with short command limits, so the loop looked sanctioned.
  • Symptom: status polls for the whole length of a fix round, which can run around thirty minutes, times up to three rounds; with a model that sleeps in a separate turn per poll, one model turn per poll; and a worker that backgrounds the call and ends its turn leaves the run parked at the next gate (Workers background a pipeline call and end their turn, so the run parks at a gate with nobody listening kunchenguid/firstmate#3057).

main has since fixed the instruction itself; this branch keeps the lab evidence below because it is what measured the cost, and adds the missing regression test.

Naming the wait, not only forbidding the wrong ones

Forbidding a backgrounded poll is not enough on its own.
The brief told a worker what not to do without naming the wait it may actually use, and Claude Code can refuse a sleep-then-check command while pointing at run_in_background, which is the one shape a waiting worker must not take.
The Waiting section now names the blocking foreground until loop for that harness.
The probes behind it are recorded in docs/verification/runtime-backends.md.

Change

  • bin/fm-brief.sh: deletes the "natural checkpoint" clause and adds one # Waiting section to the ship and scout briefs.
    It tells the worker to end its turn after needs-decision: or blocked:, and to wait on external state with ONE blocking command (no-mistakes axi run/respond --wait, gh pr checks --watch, or an until loop), never backgrounded in order to poll it.
    It bounds that command per harness, from each vendor's own code (recorded in docs/verification/runtime-backends.md):
    • Pi: a timeout of at most 2700 seconds, because Pi's bash tool has none by default; 2700 stays under the watcher's 3600-second busy-turn bound.
    • Claude Code: its maximum timeout of 600000 ms, because the default is 2 minutes.
    • Codex: empty write_stdin polls of up to 300000 ms.
    • Elsewhere: 10 minutes.
      It then names the foreground until loop as the sanctioned wait for Claude Code.
  • bin/fm-classify-lib.sh: status_own_open_decisions, the open decisions a task raised itself, excluding the reserved pending-reply- keys a supervisor library raises about it.
  • bin/fm-send.sh: --automatic.
    While the target task has its own open decision or blocker, an automatic send writes nothing, rings nothing, prints one deferred: line, and exits 4.
    The caller keeps its retry state.
  • Every automatic sender now passes --automatic and treats exit 4 as "retry later", not failure:
    • the bootstrap local nudge, its retry, and the remote nudge (bin/fm-bootstrap.sh);
    • the config-push remote nudge (bin/fm-config-push.sh);
    • the config re-read pointer (bin/fm-config-inherit-lib.sh);
    • the reconcile notify and its queued requests (bin/fm-secondmate-reconcile.sh).
  • The two senders that report the result classified it by matching the send's captured output against deferred:*, which discards the exit status the contract is actually stated in.
    fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is.
    The banner lands ahead of the deferred: line, so the match fell through and a waiting mate was reported as a failed send with the banner as the reason.
    This is why the first CI run of this branch failed Behavior portable serial 1 in tests/fm-secondmate-sync.test.sh while the same suite passed locally, where the fixtures' primary checkout is on its default branch and no banner is printed.
    bin/fm-bootstrap.sh and bin/fm-config-inherit-lib.sh now classify on fm-send's exit status and select the deferred: line out of the output instead of assuming it came first.
    bin/fm-secondmate-reconcile.sh already branched on exit 4 and is unchanged.
  • bin/fm-pending-reply-lib.sh: the pending-reply recovery repost stays unattempted while the mate has its own open decision.
  • AGENTS.md section 7: one sentence, "while a worker's decision or blocker is open, send it only the answer".
  • .agents/skills/bootstrap-diagnostics: explains the new NUDGE_SECONDMATES: ... deferred: line and says to fold the re-read into the answer.
  • bin/fm-spawn.sh: an abort before any lease was armed no longer prints spawn_return_abort_lease: command not found, because the EXIT trap is installed before that helper is defined; found while reproducing.

The doorbell is unchanged: it is typed only when a record lands, and the watcher re-rings only an unacknowledged deliberate record, so a worker with an empty inbox is never rung.
A deliberate answer to a waiting worker still re-rings until acknowledged, which is what should wake it.

Automatic paths enumerated

Every sender into a worker or secondmate that acts on its own schedule, and what each does now:

Path Before After
Bootstrap local nudge, retry, remote nudge sends defers, marker kept
Config-push remote re-read nudge sends defers, retry kept
Config re-read pointer sends defers, marker kept
Reconcile notify and process-requests sends defers, request kept, no cooldown started
Pending-reply recovery repost sends not attempted until the decision closes
Watcher inbox re-ring ladder re-rings only unhandled deliberate records unchanged; now pinned by a test that it never rings a waiting worker with an empty inbox

The away daemon, stale, paused, and captain-call wakes all go to firstmate, never the worker.
Deliberate senders are unchanged: fm-send answers, fm-secondmate-restart, the backlog-handoff receiver wake, the stow cascade, fm-control, and the spawn launch.

Wedge backstops

None of these wakes the worker; each wakes firstmate:

  • the watcher's busy-turn bound (FM_BUSY_TURN_MAX_SECS, default 3600 seconds), which flags a turn that runs too long, including a blocking wait that never returns;
  • the stale and captain-call bounded wakes;
  • the inbox escalation after the re-ring ladder's FM_TASK_INBOX_RING_MAX attempts;
  • the OPEN DECISIONS section of every wake drain, which keeps an unanswered decision in front of firstmate until it is answered.

How this layers with upstream neighbours

Lab evidence

An isolated Herdr lab (fm-herdr-lab.sh session, never default), one home per variant, real fm-spawn, fm-send, and fm-secondmate-reconcile from each revision, real Pi 0.85.1 workers.
Old = main at bda9903, new = this branch.
The model was cursor/composer-2.5 through pi-cursor-sdk, because the Codex quota was exhausted; that extension runs tools on Cursor's side and replays them into Pi, so Pi's bash timeout does not bind it.
The pipeline was a stand-in no-mistakes with axi's documented hold semantics that logged every call and published its outcome a fixed time after firstmate's validation steer.
Turn counts come from the Pi session logs; pipeline calls come from the stand-in's own timestamped log.
Lab noise: both old and new workers spent some calls inspecting the stand-in to find its outcome marker, and one old-code worker ran a search across the home directory; later runs confine lab agents to their worktree.

Decision wait (contract 1)

Window after needs-decision, before any answer After fm-send --resolve-key
old 416 s, 0 model calls woke, 6 model calls, done: committed 55ad6d6
new 417 s, 0 new turns (1 call is the closing message of the raising turn) woke, 6 model calls, done: committed 80ad98d

The lab worker had nothing to be unsure about, so it does not reproduce the checkpoint polling the real logs show; it proves the new brief keeps the zero and the answer still wakes.

Automatic nudge to a waiting secondmate (contract 2)

The harness opened needs-decision [key=lab-name] for a real idle Pi secondmate, then ran the real fm-secondmate-reconcile.sh notify backstop.

Reconcile notify 242 s after, decision open After fm-send --resolve-key lab-name
old sent: mate-old orphan_in_flight woke: 36 model calls, 52 tool calls, including an unrequested backlog edit 24 model calls
new deferred: mate-new orphan_in_flight 0 model calls woke: 41 model calls, recorded the answer, replied done [corr=...]

Pipeline wait to an outcome (contract 3)

Outcome published 780 s after the validation steer.

no-mistakes calls during the wait model calls reported
old axi run --wait 15s, then axi run --wait 8m & in the background, then 30 axi status polls every 30 s 17 done: PR ... checks green 16 s after the outcome
new axi run --wait 44m, re-run once with no status call in between 8 done: PR ... checks green 4 s after the outcome

Pipeline parked at a gate (kunchenguid#3057)

The stand-in's axi run returned a parked review gate (awaiting_agent: parked, one auto-fix finding) 300 s after the steer.
New code, from the stand-in's log:

1789153466 axi run --intent ... --wait 44m
1789153547 axi run --intent ... --wait 44m
1789153690 (gate published)
1789153701 axi respond --action fix --findings r1 --wait 44m

The worker was still inside its blocking call when the gate came back, and answered it 11 s later with no status call in between, so the run never sat parked with nobody listening.
Old code also answered this gate, after an axi run --wait 2m, an axi status, a daemon status, a second axi run --wait 10m, and one more axi status.
The run was stopped during the fix round for a machine-wide quiet window, after the gate was answered; over the same span the old worker made 55 model calls and the new one 26.

All lab sessions were torn down by the helper; its tripwire was gone afterwards.

Tests

  • tests/fm-send-inbox.test.sh: an automatic send to a task with an open decision exits 4, writes no record, and types nothing; the keyed answer still lands and rings; a reserved pending-reply key does not defer; an explicit-target automatic send is refused.
  • tests/fm-pending-reply.test.sh: recovery is not attempted while the mate has an open decision, and sends once after resolved.
  • tests/fm-secondmate-reconcile.test.sh: the notify defers with no inbox record and no cooldown; a queued request stays queued; after the answer it is delivered.
  • tests/fm-secondmate-sync.test.sh: the bootstrap nudge defers, prints no "send failed", and keeps its marker.
    A second case puts the fixture's primary checkout on a feature branch, which is what made CI differ from a local run, and asserts up front that the guard banner really does precede the deferred: line before checking that the nudge is still reported as deferred, so the case cannot pass vacuously.
  • tests/fm-task-inbox.test.sh: the watcher never rings a worker waiting on its decision.
  • tests/fm-brief.test.sh: ship, scout, and secondmate briefs carry the waiting contract and no checkpoint polling, name the wait a Claude Code worker may use, and keep the one-foreground-call definition of done main introduced.
  • tests/fm-spawn-dispatch-profile.test.sh: an early spawn refusal prints no shell error.
  • tests/fm-remote-reply.test.sh: its escalation scenario now answers the mate's earlier open decision and blocker before exercising the recovery repost, which would otherwise correctly wait.

bin/fm-lint.sh and bin/fm-doc-audience-check.sh pass.

Upstream

Upstream https://github.com/kunchenguid/firstmate main has the same behavior, read on 2026-09-11:

  • bin/fm-dod-lib.sh line 236: "So background the drive call and poll no-mistakes axi status from a separate call". This fork's main has since fixed that line; upstream has not.
  • bin/fm-brief.sh line 210: "and at any natural checkpoint when you are unsure - list $INBOX_DIR/*.msg".
  • bin/fm-send.sh has no automatic-send guard, bin/fm-secondmate-reconcile.sh line 527 sends the notify fire-and-forget whatever the mate's open decisions, and upstream has no blocking-wait guidance in the brief.

Nothing was filed upstream.

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.
- fm-spawn no longer prints a missing-helper error on an early abort.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is already fixed on main by "drive
no-mistakes with one foreground call, not a background poll"; this branch
takes that text as is and only adds the regression test main shipped without.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.
@tiago-peixoto
tiago-peixoto force-pushed the fm/firstmate-worker-waits-cost-no-tokens branch from e082672 to 4b7bd3f Compare September 14, 2026 20:41
@tiago-peixoto
tiago-peixoto merged commit a40f9e6 into main Sep 14, 2026
15 checks passed
tiago-peixoto added a commit that referenced this pull request Sep 17, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.
- fm-spawn no longer prints a missing-helper error on an early abort.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is already fixed on main by "drive
no-mistakes with one foreground call, not a background poll"; this branch
takes that text as is and only adds the regression test main shipped without.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto added a commit that referenced this pull request Sep 17, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.
- fm-spawn no longer prints a missing-helper error on an early abort.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is already fixed on main by "drive
no-mistakes with one foreground call, not a background poll"; this branch
takes that text as is and only adds the regression test main shipped without.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto added a commit that referenced this pull request Sep 18, 2026
A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.
- fm-spawn no longer prints a missing-helper error on an early abort.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is already fixed on main by "drive
no-mistakes with one foreground call, not a background poll"; this branch
takes that text as is and only adds the regression test main shipped without.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.
tiago-peixoto added a commit that referenced this pull request Sep 18, 2026
…48)

fm-send --automatic defers (exit 4, nothing written or rung) while the
target has an open decision or blocker of its own; every automatic sender
passes it and keeps its retry state, and the pending-reply recovery waits
the same way.

The two senders that report the result classified it by matching the text of
the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh
as a supervision warning, and that guard prints its worktree-tangle banner
whenever the primary checkout is on a feature branch, which is exactly what a
CI pull-request checkout is. The banner lands ahead of the deferred: line,
so the match fell through and a waiting mate was reported as a failed send,
with the banner as the reason. Both senders now classify on fm-send's exit
status, which is the contract the deferral is actually stated in, and select
the deferred: line out of the output rather than assuming it came first.

The competing Waiting-section wait-shape that used to ride this patch is
dropped in favor of upstream brief text.
tiago-peixoto added a commit that referenced this pull request Sep 21, 2026
…48)

fm-send --automatic defers (exit 4, nothing written or rung) while the
target has an open decision or blocker of its own; every automatic sender
passes it and keeps its retry state, and the pending-reply recovery waits
the same way.

The two senders that report the result classified it by matching the text of
the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh
as a supervision warning, and that guard prints its worktree-tangle banner
whenever the primary checkout is on a feature branch, which is exactly what a
CI pull-request checkout is. The banner lands ahead of the deferred: line,
so the match fell through and a waiting mate was reported as a failed send,
with the banner as the reason. Both senders now classify on fm-send's exit
status, which is the contract the deferral is actually stated in, and select
the deferred: line out of the output rather than assuming it came first.

The competing Waiting-section wait-shape that used to ride this patch is
dropped in favor of upstream brief text.
tiago-peixoto added a commit that referenced this pull request Sep 21, 2026
…s Pi change. They fail with the same tests on main at the base commit 09dc7b3 (CI run 35613893238). Two tests that failed here but not visibly on main (fm-kimi-harness and fm-spawn-compact-adviser-disable-remote) were in main's shard 9, which stopped at an actionlint download error before any test ran. Two regressions came in when the fork's own commits were rebased onto upstream on Sep 21. As you asked, I kept the previous agent's two partial fixes. I checked each one against the logs and locally, and made no other changes. 1. bin/fm-spawn.sh: restored the short staged launch line. Upstream kunchenguid#4994 (a452a79) writes the full launch command to a private file and types only `. <launch-file>` into the pane, because typed lines over about 1,024 bytes get cut off. The fork's #57 (a32fce8) put the old `spawn_send_literal "$T" "$LAUNCH"` back while resolving a merge, so it typed the whole command again. That broke fm-claude-trust, fm-backend-orca, fm-kimi-harness, fm-spawn-dispatch-profile, both fm-spawn-compact-adviser-disable suites, fm-remote-secondmate-trace-context and fm-remote-secondmate-parent-binding. The fix is one line that restores kunchenguid#4994's `spawn_send_literal "$T" ". $(shell_quote "$LAUNCH_FILE")"` and keeps #57's `SPAWN_LAUNCH_SENT=1`. 2. tests/fm-remote-reply.test.sh: the fixture also resolves the `default` key. Upstream kunchenguid#3764 added two decision lines with no key (`needs-decision [at=...]: which base branch?`), and these count under the key `default`. The fork's #48 made the automatic recovery repost wait while the mate has any open decision, so "the one automatic recovery repost was not sent". I confirmed this by printing the open decisions at that point: only `default needs-decision which base branch?` was open. Only the fixture's setup changed; every assertion is unchanged, and #48's wait-while-open rule still applies. Verification on macOS: all 8 spawn test files and fm-remote-reply pass through bin/fm-test-run.sh. With the spawn line reverted, fm-claude-trust fails with the same message as CI ("the launch command did not carry the brief the worker must read"). bash -n and shellcheck -S warning pass on both files. No Pi code was touched
tiago-peixoto added a commit that referenced this pull request Sep 21, 2026
* fix(pi): stop nested Pi CLI from replacing a live session binding

A short-lived child such as fm-spawn's pi --help probe was treated as
lock-owned through ancestry and overwrote both markers with a pid that
died immediately, causing false supervision alarms.

* no-mistakes(review): Remove duplicate Pi turn-end marker regression tests

* no-mistakes(review): Point Pi marker verification doc at remaining regression suite

* no-mistakes(review): Test Pi self-lock marker binding; drop dead turn-end ownership code

* no-mistakes(ci): These four failing shards are not caused by this PR's Pi change. They fail with the same tests on main at the base commit 09dc7b3 (CI run 35613893238). Two tests that failed here but not visibly on main (fm-kimi-harness and fm-spawn-compact-adviser-disable-remote) were in main's shard 9, which stopped at an actionlint download error before any test ran. Two regressions came in when the fork's own commits were rebased onto upstream on Sep 21. As you asked, I kept the previous agent's two partial fixes. I checked each one against the logs and locally, and made no other changes. 1. bin/fm-spawn.sh: restored the short staged launch line. Upstream kunchenguid#4994 (a452a79) writes the full launch command to a private file and types only `. <launch-file>` into the pane, because typed lines over about 1,024 bytes get cut off. The fork's #57 (a32fce8) put the old `spawn_send_literal "$T" "$LAUNCH"` back while resolving a merge, so it typed the whole command again. That broke fm-claude-trust, fm-backend-orca, fm-kimi-harness, fm-spawn-dispatch-profile, both fm-spawn-compact-adviser-disable suites, fm-remote-secondmate-trace-context and fm-remote-secondmate-parent-binding. The fix is one line that restores kunchenguid#4994's `spawn_send_literal "$T" ". $(shell_quote "$LAUNCH_FILE")"` and keeps #57's `SPAWN_LAUNCH_SENT=1`. 2. tests/fm-remote-reply.test.sh: the fixture also resolves the `default` key. Upstream kunchenguid#3764 added two decision lines with no key (`needs-decision [at=...]: which base branch?`), and these count under the key `default`. The fork's #48 made the automatic recovery repost wait while the mate has any open decision, so "the one automatic recovery repost was not sent". I confirmed this by printing the open decisions at that point: only `default needs-decision which base branch?` was open. Only the fixture's setup changed; every assertion is unchanged, and #48's wait-while-open rule still applies. Verification on macOS: all 8 spawn test files and fm-remote-reply pass through bin/fm-test-run.sh. With the spawn line reverted, fm-claude-trust fails with the same message as CI ("the launch command did not carry the brief the worker must read"). bash -n and shellcheck -S warning pass on both files. No Pi code was touched
tiago-peixoto added a commit that referenced this pull request Sep 25, 2026
…48)

fm-send --automatic defers (exit 4, nothing written or rung) while the
target has an open decision or blocker of its own; every automatic sender
passes it and keeps its retry state, and the pending-reply recovery waits
the same way.

The two senders that report the result classified it by matching the text of
the send's captured output against deferred:*. fm-send runs bin/fm-guard.sh
as a supervision warning, and that guard prints its worktree-tangle banner
whenever the primary checkout is on a feature branch, which is exactly what a
CI pull-request checkout is. The banner lands ahead of the deferred: line,
so the match fell through and a waiting mate was reported as a failed send,
with the banner as the reason. Both senders now classify on fm-send's exit
status, which is the contract the deferral is actually stated in, and select
the deferred: line out of the output rather than assuming it came first.

The competing Waiting-section wait-shape that used to ride this patch is
dropped in favor of upstream brief text.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant