Skip to content

fix(send): hoist busy-queued read-back verdict into the backend dispatch layer - #1

Merged
willchen95 merged 121 commits into
mainfrom
fm/fm-sendcore-busy-queued-verdict-p1
Aug 13, 2026
Merged

willchen95 merged 121 commits into
mainfrom
fm/fm-sendcore-busy-queued-verdict-p1

Conversation

@willchen95

Copy link
Copy Markdown
Owner

Intent

Hoist the busy-queued read-back verdict from caller-side into the fm_backend_send_text_submit dispatch layer so the daemon's inject_msg also benefits from proof-carrying busy-queued delivery. After each backend function returns, capture the verdict; for 'pending' (proven composer state) check if the pane is busy AND the typed text appears in a bounded capture, then return 'queued-busy'. For 'unknown' (unreadable composer), never upgrade — pass through the raw verdict as a delivery failure per captain standing rule: prefer false-full over false-fast, never claim delivery without proof. The hoisted read-back works on herdr only (native agent-state busy detection); correct docs/comments accordingly. inject_msg accepts 'queued-busy' alongside 'empty'. fm-send.sh accepts 'queued-busy'. Tests: test_inject_msg_accepts_queued_busy for daemon consumer, test_send_text_submit_busy_queued_readback for dispatch-layer conversion (pending->queued-busy with busy+capture-match, and unknown pass-through). All tests must pass.

What Changed

  • fm_backend_send_text_submit in bin/fm-backend.sh now captures each backend's submit verdict and performs the busy-queued read-back itself: a pending verdict (composer-proven text) is upgraded to queued-busy only when the pane is busy and the typed text appears in a bounded capture probe, while an unknown verdict (unreadable composer) is never rescued and passes through as a delivery failure. The read-back applies to herdr only, which has native agent-state busy detection, and docs/comments were scoped accordingly.
  • Consumers accept the new verdict instead of re-implementing the rescue: bin/fm-send.sh treats queued-busy alongside empty as a successful submit, and the supervise daemon's inject_msg logs and accepts queued-busy as a message queued for the next agent turn. New tests test_send_text_submit_busy_queued_readback (dispatch-layer pending→queued-busy conversion and unknown pass-through) and test_inject_msg_accepts_queued_busy (daemon consumer) cover both sides.
  • The branch also advances its base across ~114 upstream commits already merged to main (remote secondmates, procevent sources, agent lifecycle control, stow memory tiers, and related tests/docs), which accounts for the bulk of the diff stat; the busy-queued verdict hoist is this branch's own change.

Risk Assessment

✅ Low: The fix round closes the round-1 false-delivery race at the adapter boundary exactly as recommended (herdr emits final 'pending' only when composer-proven from a busy baseline, and only that verdict is rescued), all four consumers handle the new verdicts safely, and the behavioral tests reproduce the original race and prove the rescue never runs for unproven verdicts.

Testing

Ran the three affected suites (fm-backend, fm-daemon, fm-backend-herdr; 303 assertions, all passing, including both intent-named tests), then manually demonstrated the intent end-to-end at the fm-send.sh CLI surface with the herdr fakebin: a busy pane holding the queued message now delivers with queued-busy (rc=0) while the pre-change code refused the identical scenario, and unknown/stale-capture scenarios correctly remain loud delivery failures.

Evidence: fm-send queued-busy CLI transcript (new vs base commit, plus negative cases)
# End-to-end CLI demonstration: busy-queued proof-carrying delivery via fm-send.sh

Fixture: the herdr fakebin convention from tests/fm-backend-herdr.test.sh (call-numbered
canned responses, every `herdr` invocation logged), a herdr crew target recorded in
state meta (`window=default:w1:p2`, `backend=herdr`). Same scenario in every run:
the crewmate pane is BUSY (`agent_status: working`) when fm-send types the message,
Enter is swallowed by the busy composer, and per-attempt composer reads positively
prove the typed text sits in the composer (herdr verdict: composer-proven `pending`).

## 1. NEW code (353ad8a): busy pane + this message visible in capture → delivered

Dispatch-layer read-back sees busy + the typed text in a bounded capture and upgrades
the composer-proven `pending` to `queued-busy`; fm-send accepts it as delivery.

`` `
$ bin/fm-send.sh default:w1:p2 "status update: please prioritize the flaky login test"
rc=0
`` `

herdr call log (dispatch-layer probe visible as the trailing `agent get` + `pane read`):

`` `
pane | send-text | w1:p2 | status update: please prioritize the flaky login test
agent | get | w1:p2                      <- pre-Enter baseline: working (busy)
pane | send-keys | w1:p2 | enter
pane | read | w1:p2 | --lines 200 | ansi <- composer read: typed text proven (pending)
pane | send-keys | w1:p2 | enter         <- submit-only retry (never retypes)
pane | read | w1:p2 | --lines 200 | ansi <- composer read: still pending, retries exhausted
agent | get | w1:p2                      <- HOISTED read-back: busy probe -> working
pane | read | w1:p2 | --lines 200        <- HOISTED read-back: capture holds the message
`` `

## 2. OLD code (base b0ad61e): identical scenario → refused

The exact same fixture against the pre-change bin/ (extracted with `git archive b0ad61e bin`):

`` `
$ oldbin/bin/fm-send.sh default:w1:p2 "status update: please prioritize the flaky login test"
error: text not submitted to default:w1:p2 (delivery unconfirmed; verdict=pending;
  tried explicit target 'default:w1:p2' matched .../state/crew-1.meta; backend=herdr)
OLD rc=1
`` `

## 3. NEW code, negative: busy pane but the capture holds only a stale different message → refused

`` `
$ bin/fm-send.sh default:w1:p2 "status update: please prioritize the flaky login test"
error: text not submitted to default:w1:p2 (delivery unconfirmed; verdict=pending; ...)
STALE rc=1
`` `

## 4. NEW code, forbidden-behavior check: unreadable composer (verdict=unknown) → never upgraded

Empty canned responses (baseline/agent state unreadable): the raw `unknown` verdict passes
through as a loud delivery failure — never rescued to queued-busy.

`` `
$ bin/fm-send.sh default:w1:p2 "status update: please prioritize the flaky login test"
error: text not submitted to default:w1:p2 (delivery unconfirmed; verdict=unknown;
  tried explicit target 'default:w1:p2' matched .../state/crew-1.meta; backend=herdr)
rc=1
`` `

Daemon consumer coverage: `test_inject_msg_accepts_queued_busy` (tests/fm-daemon.test.sh)
proves `inject_msg` returns success and logs "inject queued: supervisor pane was busy,
message accepted for next turn (verdict=queued-busy)" when the dispatch layer reports
queued-busy — see fm-daemon.test.log line "ok - inject_msg: accepts queued-busy verdict".
Evidence: fm-backend suite log (dispatch-layer read-back test)
ok - fm_backend_name: FM_BACKEND env > config/backend > default tmux
ok - fm_backend_detect: no markers -> undetected, HERDR_ENV=1 -> herdr, $TMUX -> tmux, CMUX_WORKSPACE_ID -> cmux, nested combinations resolve innermost-first
ok - fm_backend_detect: falls back to __CFBundleIdentifier=com.cmuxterm.app when CMUX_WORKSPACE_ID is absent (signal bundle-id; foreign bundle ids rejected)
ok - fm_backend_detect: the cmux fallback signals are macOS-only (inert on a non-Darwin uname)
ok - fm_backend_detect: an inherited cmux bundle id never outranks $TMUX or HERDR_ENV (tmux/herdr-inside-cmux false positive absorbed)
ok - fm_backend_detect: ancestry fallback matches the lsappinfo-resolved (bundle-id) cmux app pid in the parent chain
ok - fm_backend_detect: ancestry fallback matches a bundle-shaped cmux comm path at any install location when lsappinfo cannot resolve a pid
ok - fm_backend_detect: ancestry fallback stops undetected at launchd (a reparented tmux server never reaches cmux)
ok - fm_backend_name: a fallback-detected cmux prints a NOTICE naming the fallback signal; the primary-marker notice is unchanged
ok - fm_backend_name: auto-detect selects herdr or cmux (loud notice) or tmux (silent, including nested tmux-in-herdr/tmux-in-cmux)
ok - fm_backend_name: an explicit FM_BACKEND or config/backend setting always wins over runtime auto-detection, including an ambient cmux marker
ok - fm_backend_validate: implemented adapters accepted, unknown and blocked codex-app backends refused loudly
ok - zsh: shell-portable backend matching skipped (zsh not found)
ok - bash: fm_backend_source recognizes known backends and rejects unknown ones
ok - fm_backend_validate_spawn: all implemented lifecycle backends are spawn-supported
ok - fm_meta_get / fm_backend_of_meta: read key=value, default backend to tmux
ok - fm_backend_resolve_selector: session:window literal, exact task id first, legacy fm-<id> label fallback, ad hoc bare name via tmux list-windows
ok - fm_backend_of_selector: exact task ids, legacy fm-<id> labels, and matching explicit targets inherit metadata backend
ok - fm-send.sh: explicit tmux targets are verified; text types once and submits with Enter
ok - fm_backend_send_text_submit: busy-queued read-back proves this message (distinctive middle), never a stale digest
ok - fm-peek.sh: capture-pane invocation and output are byte-identical old vs new
ok - fm-spawn.sh: a project reached through a symlinked prefix (e.g. macOS /tmp -> /private/tmp) does not trip the isolation guard's false refusal
ok - fm-teardown.sh: treehouse return remains compatible while tmux cleanup uses exact selectors
ok - fm-spawn.sh --backend bogus is refused loudly
ok - fm-spawn.sh --backend codex-app is refused
ok - fm-spawn.sh honors FM_BACKEND and refuses an unimplemented value loudly
ok - fm-spawn.sh: an explicit --backend tmux resolves silently and writes no backend= (missing means tmux)
ok - fm-spawn.sh: explicit --backend tmux wins over an ambient HERDR_ENV=1 auto-detect marker
ok - fm-spawn.sh: auto-detect resolves nested tmux-in-herdr to tmux and stays silent end to end
Evidence: fm-daemon suite log (inject_msg accepts queued-busy)
ok - fm-afk-start.sh fails before daemon startup when the afk flag cannot be written
ok - fm-afk-start.sh ignores stale pidfile-only live pids
ok - fm-afk-start.sh reclaims stale daemon locks whose live pid identity no longer matches
ok - supervise daemon state root is scoped by FM_HOME
ok - routine signal self-handles
ok - captain-relevant status verbs escalate
ok - check + unknown escalate; heartbeat self-handles
ok - transient stale self-handles and records a persistence marker
ok - enriched stale wedges bypass status absorption without disturbing busy workers
ok - stale + terminal status escalates immediately
ok - paused reasons with captain phrases remain pause-classified
ok - handle_wake on a paused stale records a pause marker, drops the wedge marker, and does not escalate
ok - handle_wake records a declared pause from a routine signal for long-cadence rechecks
ok - a terminal signal clears pause and stale tracking across both supervisors
ok - housekeeping migrates a normal-watcher's declared pause into daemon tracking
ok - housekeeping clears an already-resumed watcher pause across both supervisors
ok - housekeeping seeds pause tracking from status without a watcher marker
ok - persistent stale escalates after threshold and clears its marker
ok - resumed (busy) stale clears its marker without escalating
ok - housekeeping re-surfaces a stale declared pause on the long cadence and resets its window
ok - housekeeping clears a paused marker whose pane became busy again, without escalating
ok - housekeeping clears a paused marker once the crew is no longer declaring the pause
ok - housekeeping moves an existing stale marker to pause before wedge escalation
ok - housekeeping clears tracking when a crew leaves pause
ok - persistent herdr stale resolves the target from metadata and escalates
ok - herdr idle busy-footer stale clears through capture corroboration
ok - resumed herdr stale clears through backend-aware busy state
ok - persistent Orca stale resolves the terminal from metadata
ok - multiple escalations flush as a single batched digest
ok - batch flush measures max-delay from the first append, not the last
ok - catch-all scan escalates a missed terminal once, not twice
ok - handle_wake routes routine->self and captain->escalate
ok - INJECT_SKIP forces self-handle, bypassing captain-relevant classification
ok - is_wake_reason distinguishes watcher wake reasons from singleton-status stdout
ok - terminal-stale escalate removes its marker so housekeeping does not re-escalate
ok - captain signal escalate marks seen so the catch-all scan does not re-fire
ok - _collapse_newlines replaces newlines with literal separator
ok - afk flag absent: daemon does not inject, buffer preserved
ok - busy-guard defers injection when supervisor pane is busy
ok - marker detection: marker -> stay afk, no marker -> exit afk
ok - /afk invocation is exempt from afk exit (no self-cancel)
ok - should_exit_afk returns false when afk is not active
ok - strip_injection_marker removes the sentinel marker cleanly
ok - pane_input_pending detects partial input on the cursor line
ok - pane_input_pending: a blank unidentified cursor row defers (strict container-proof rule)
ok - pane_input_pending: only proven empty agent prompts pass
ok - fm_tmux_composer_state: a bare shell prompt ($/%/#/>) reads unknown, never empty (dead-shell injection safety)
ok - fm_tmux_composer_state: a bordered composer box and bare agent glyphs (❯/›) still read empty
ok - fm_tmux_composer_state: only matching edge borders form a composer box
ok - pane_input_pending preserves bright placeholder-like drafts in styled captures
ok - classify_signal dedupes against the catch-all scan seen marker
ok - classify_stale dedupes against the signal path seen marker
ok - AFK nonterminal working:+merged keeps wedge aging and re-escalates at bound
ok - genuine done: and merge-check events still escalate
ok - pane_input_pending: an idle bordered composer is NOT pending (afk-invx-i5)
ok - pane_input_pending: text inside a bordered composer is still pending
ok - submit-ACK confirms a submit when the composer returns to a bordered-empty box
ok - submit-ACK reports pending on a persistently swallowed Enter (type-once)
ok - max-defer on an empty stuck pane types once, alarms, and preserves the buffer
ok - max-defer flushes and clears the buffer on an empty bordered pane
ok - max-defer on a pending composer alarms without typing
ok - normal flush clears a stale wedge marker
ok - below MAX_DEFER: no inject, no alarm, buffer preserved
ok - max-defer does not flush or alarm while afk is inactive
ok - library mode: sourcing the daemon defaults FM_WEDGE_ALARM_EXEC to discard (no test can fire a real notification)
ok - wake helpers replace inherited notifier overrides with the safe recorder
ok - the discard seam suppresses every notifier, including command: (fires nothing)
ok - direct notifier helpers honor the discard seam, including command:
ok - osascript channel routes through the notifier seam with the summary (never a real notification)
ok - herdr channel routes through the notifier seam with the summary (never a real notification)
ok - command channel runs the captain command with the summary on $1 and on stdin
ok - command channel failures redact configured commands while logging their exit status
ok - unknown channel directives are redacted while the alarm keeps running
ok - off disables every active alert regardless of directive position (marker and tmux flash are unaffected)
ok - auto resolves to the macOS osascript notifier on Darwin (default-on)
ok - auto on a non-macOS platform selects no built-in OS channel (the marker or a configured command carries it)
ok - config/wedge-alarm selects every configured channel and skips comment and blank lines
ok - a failing channel logs and falls back to the next channel, never crashing the alarm
ok - a hung notifier is bounded, logged, and falls through to the next channel
ok - a backgrounded command notifier remains bounded until its process group is reaped
ok - a hung notifier override is bounded, logged, and proceeds to the next channel
ok - daemon shutdown stops and reaps the active notifier process group
ok - inject_wedge_alarm writes the marker AND emits the active alert even with no tmux status-line (herdr backend)
ok - in-process wedge throttle prevents alert spam when the marker cannot persist
ok - fm-send exits non-zero on a confirmed swallow, zero on a clean submit
ok - fm-send exits non-zero when initial text send fails
ok - fm-send exits non-zero unless delivery is proven empty
ok - discover_supervisor_backend: override > TMUX_PANE > HERDR_ENV+HERDR_PANE_ID > tmux fallback
ok - discover_supervisor_target: override > TMUX_PANE > herdr '<session>:<pane-id>' composition > firstmate:0 fallback
ok - pane_is_busy: herdr native busy_state='busy' short-circuits without a capture fallback
ok - primary busy guard isolates rendered signatures by detected harness
ok - pane_is_busy: omitted backend defaults to tmux for Grok's isolated fallback
ok - pane_input_pending: dispatches through fm_backend_composer_state for backend=herdr
ok - inject_msg: herdr busy-guard defers before ever attempting a submit
ok - inject_msg: herdr composer-guard defers before ever attempting a submit
ok - inject_msg: herdr pane-gone check defers before any busy/composer/submit call
ok - inject_msg: dispatches busy-guard/composer-guard/submit through the herdr backend and succeeds on a confirmed empty composer
ok - inject_msg: defers on a dead-shell/unreadable composer (unknown), never typing the escalation into a shell
ok - inject_msg: accepts queued-busy verdict as proof-carrying delivery (hoisted read-back confirmed)
ok - inject_msg: unrecognized composer states defer by default
Evidence: fm-backend-herdr suite log (pending-unproven + end-to-end rescue)
ok - fm_backend_herdr_version_check: accepts the current protocol (14)
ok - fm_backend_herdr_version_check: refuses an old protocol loudly
ok - fm_backend_herdr_version_check: refuses loudly when herdr is not installed
ok - fm_backend_herdr_workspace_label: a primary home (no marker) resolves to 'firstmate'
ok - fm_backend_herdr_workspace_label: a secondmate home (.fm-secondmate-home) resolves to '2ndmate-<id>'
ok - fm_backend_herdr_workspace_label: trims whitespace around the marker's secondmate id
ok - fm_backend_herdr_workspace_label: an empty marker file falls back to the primary label 'firstmate'
ok - fm_backend_herdr_workspace_label: two different secondmate homes get two different, non-colliding labels
ok - fm_backend_herdr_cli: sets HERDR_SESSION AND appends a trailing --session flag on every call
ok - fm_backend_herdr_launcher_identity: a firstmate not running inside herdr has no launcher workspace to inherit
ok - fm_backend_herdr_launcher_identity: HERDR_ENV=1 without a pane id selects the backend but binds no parent
ok - fm_backend_herdr_launcher_identity: resolves the launcher's exact workspace even when a same-labeled workspace sorts first
ok - fm_backend_herdr_launcher_identity: refuses a launcher pane that names a different herdr session
ok - fm_backend_herdr_launcher_identity: refuses a claimed pane without exact server identity
ok - fm_backend_herdr_launcher_identity: refuses a launcher pane whose injected socket belongs to another herdr server
ok - fm_backend_herdr_launcher_identity: refuses when the launcher's own pane no longer resolves
ok - fm_backend_herdr_launcher_identity: refuses when the launcher's pane and tab disagree about their workspace
ok - fm_backend_herdr_launcher_identity: refuses when the launcher's workspace is gone from its own session
ok - fm_backend_herdr_workspace_ensure: places a worker in the launcher's exact workspace, not the first same-labeled one
ok - fm_backend_herdr_workspace_ensure: refuses to guess between two same-labeled home workspaces
ok - fm_backend_herdr_workspace_ensure: a --secondmate container resolves that home's own workspace, not the launcher's
ok - fm_backend_herdr_container_ensure: surfaces the exact ambiguous-placement refusal instead of a generic failure
ok - fm_backend_herdr_container_ensure: version-gates, starts the server, ensures the firstmate workspace, echoes session:workspace_id + the seeded default tab id
ok - fm_backend_herdr_container_ensure: reuses an existing firstmate workspace without recreating it, and reports no seeded default tab (adopted, not created)
ok - fm_backend_herdr_container_ensure: workspace create passes --no-focus
ok - fm_backend_herdr_container_ensure: creates the workspace under the SECONDMATE home's own label, not 'firstmate'
ok - fm_backend_herdr_create_task: prunes exactly the seeded default tab container_ensure identified, once the first real task tab exists
ok - herdr repeated spawn/teardown: one persistent firstmate workspace reused, zero orphans, default tab pruned, create ran once
ok - fm_backend_herdr_create_task: an ADOPTED workspace's pre-existing tab is never pruned (the created-vs-adopted gate)
ok - fm_backend_herdr_create_task: the label-collision startup-workspace scenario (2026-07-02 incident) leaves the captain's live tab untouched
ok - fm_backend_herdr_workspace_prune_seeded_default_tab: refuses to close the seeded default tab when its pane reports a working agent (defense in depth)
ok - fm_backend_herdr_create_task: refuses a duplicate tab label (herdr's own tab create has no uniqueness check)
ok - fm_backend_herdr_create_task: a same-labeled tab with a live (even idle) registered agent still refuses exactly as before
ok - fm_backend_herdr_create_task: scans every same-labeled tab and refuses if any duplicate is live
ok - fm_backend_herdr_create_task: closes and replaces a same-labeled tab whose pane is dead (pane_not_found)
ok - fm_backend_herdr_create_task: closes and replaces a same-labeled tab whose pane is alive but hosts no registered agent (a restored plain shell)
ok - fm_backend_herdr_create_task: closes every confirmed same-labeled husk only after creating the replacement
ok - fm_backend_herdr_create_task: refuses success when a preexisting husk tab remains after replacement
ok - fm_backend_herdr_create_task: refuses (fail-safe) rather than guessing when the duplicate's agent state cannot be classified confidently
ok - fm_backend_herdr_create_task: creates the replacement tab BEFORE closing the husk tab, never the reverse
ok - fm_backend_herdr_create_task: creates a tab and parses tab_id/pane_id from the JSON response, prunes nothing when no seeded tab id is given
ok - fm_backend_herdr_create_task: tab create passes --no-focus
ok - herdr presentation: a home that set nothing gets the projection by default at or above the floor
ok - herdr presentation: an unconfigured home below the floor falls back flat with one naming warning
ok - herdr presentation: an unreadable client release falls back flat instead of guessing
ok - herdr presentation: a deliberate opt-in is never silently downgraded below the floor
ok - herdr presentation: an explicit off opts the home out
ok - herdr presentation: an unrecognized value warns and follows the default instead of failing a spawn
ok - herdr presentation: the below-floor warning is one per home per release, not one per spawn
ok - herdr presentation: warning marker publication is atomic, symlink-safe, and fails visible
ok - herdr presentation: client and selected server floors compose conservatively without overriding explicit opt-in
ok - herdr presentation floor: every measured release, and each signal alone, classifies correctly
ok - herdr presentation floor: either signal alone can carry an above verdict, and each divergence is real
ok - herdr presentation: config parsing separates a deliberate choice from an unconfigured default
ok - herdr presentation journal: atomically publishes one non-authoritative 128-bit correlator and refuses overwrite
ok - herdr presentation journal: version 2 binds exact home/endpoint/parent identities and advances atomically
ok - herdr presentation create: exact response IDs yield one normal task pane with no workspace-close authority
ok - herdr presentation create: concurrent same-label tabs are never prune targets
ok - herdr presentation focus: snapshot requires one exact active workspace and tab
ok - herdr presentation focus: exact pane close restores the exact prior workspace and tab
ok - herdr presentation focus: cleanup refuses rather than close the captain's active tab
ok - herdr presentation focus: pane close fails when exact focus restoration fails
ok - herdr presentation reclaim: live agent state at the close boundary refuses mutation
ok - herdr presentation cleanup: emptying close behind focus ends the exact shell without a move or focus change
ok - herdr presentation cleanup: emptying close before focus moves the doomed workspace to the end and ends its exact shell
ok - herdr presentation cleanup: emptying close with the focused workspace last skips the move
ok - herdr presentation cleanup: emptying close of the last workspace skips the move
ok - herdr presentation cleanup: a non-emptying close stays plain with no proof, move, or signal
ok - herdr presentation cleanup: no-move plain close requires structured pane removal
ok - herdr presentation cleanup: an ambiguous workspace layout falls back to the plain close
ok - herdr presentation cleanup: a failed repositioning move falls back to the plain close with a warning
ok - herdr presentation cleanup: a pane with a live foreground process falls back to the plain close
ok - herdr presentation cleanup: a transient prompt helper settles into the pane-death path instead of the plain close
tests/fm-backend-herdr.test.sh: line 1838: 2563564 Killed                     bash -c 'trap "" HUP; sleep 300'
ok - herdr presentation cleanup: a SIGHUP-surviving shell is escalated to SIGKILL before giving up
tests/fm-backend-herdr.test.sh: line 1876: 2563930 Killed                     bash -c 'trap "" HUP; sleep 300'
ok - herdr presentation cleanup: a failed pane-death close falls back to the plain close
tes

... [3061 bytes truncated] ...

space, never a sibling home's
ok - fm_backend_herdr_parse_target: splits '<session>:<pane_id>' on the FIRST colon (pane_id itself contains one)
ok - fm_backend_herdr_normalize_key: Enter/Escape/C-c map to herdr's verified enter/escape/ctrl+c
ok - fm_backend_herdr_capture: calls 'pane read <pane> --source recent --lines N' with the session set
ok - fm_backend_herdr_capture: works around the verified small-N '--lines' bug by over-fetching and trimming locally
ok - fm_backend_herdr_capture: ensures the session and preserves pane read failure
ok - fm_backend_herdr_send_key: normalizes the key and targets the right pane
ok - fm_backend_herdr_kill: calls pane close and stays best-effort on failure
ok - fm_backend_herdr_current_path: reads pane foreground_cwd (the live running process), not the frozen creation-time cwd
ok - fm_backend_herdr_busy_state: working -> busy
ok - fm_backend_herdr_busy_state: done -> idle, blocked -> idle (surfaced like a stale pane, not suppressed as busy)
ok - fm_backend_herdr_busy_state: unparseable/absent agent state reports unknown, the regex-fallback cue
ok - fm_backend_herdr_composer_state: a bare '❯' composer row reads empty
ok - fm_backend_herdr_composer_state: bright placeholder-like text stays pending rather than being mistaken for an idle ghost
ok - fm_backend_herdr_composer_state: real composer text reads pending
ok - fm_backend_herdr_composer_state: a slash-command popup's argument-hint placeholder still reads pending (the incident fix)
ok - fm_backend_herdr_composer_state: reports unknown when the pane cannot be captured
ok - fm_backend_herdr_composer_state: reports unknown for bare shell prompts with no composer row
ok - fm_backend_herdr_composer_state: a native idle Pi separator composer reads empty
ok - fm_backend_herdr_composer_state: real Pi composer text remains pending
ok - fm_backend_herdr_composer_state: an incomplete lower Pi separator cannot inherit a stale empty row
ok - fm_backend_herdr_composer_state: Pi separators never authorize working, non-Pi, unreadable, or over-tall targets
ok - fm_backend_herdr_composer_state: a real-claude unbordered '❯' prompt row (no border box in view) reads empty
ok - fm_backend_herdr_composer_state: a real-claude unbordered '❯ <text>' prompt row reads pending
ok - fm_backend_herdr_composer_state: a live unbordered prompt row below a stale bordered decorative box still wins (not misread as the box's own row)
ok - fm_backend_herdr_composer_state: claude's dim prompt-suggestion ghost (the overnight wedge shape) reads empty
ok - fm_backend_herdr_composer_state: real typed text on the same claude prompt row still reads pending
ok - fm_backend_herdr_composer_state: grok's dark-truecolor placeholder (the TRUECOLOR gap) reads empty
ok - fm_backend_herdr_composer_state: grok's real bright typed input still reads pending
ok - fm_backend_herdr_composer_state: a real-codex unbordered '›' prompt row reads empty
ok - fm_backend_herdr_composer_state: a faint real-codex ghost suggestion reads empty
ok - fm_backend_herdr_composer_state: non-faint codex prompt text still reads pending
ok - fm_backend_herdr_wait_for_working: reports 'busy' immediately on the first poll, without spending the rest of the budget
ok - fm_backend_herdr_wait_for_working: a slow transition landing on a later sample within one window is still caught (robust against the 'slow transition' failure direction)
ok - fm_backend_herdr_wait_for_working: spreads six samples across the full budget endpoint without a final trailing sleep
ok - fm_backend_herdr_send_text_submit: applies the herdr minimum confirmation budget before polling agent-state
ok - fm_backend_herdr_wait_for_working: reports 'idle' (readable, genuinely not yet working) when 'busy' never appears
ok - fm_backend_herdr_wait_for_working: reports 'unknown' (a hard read failure, not a timing race) only when EVERY poll in the window fails
ok - fm_backend_herdr_wait_for_working: treats blocked as submit-active for confirmation without changing watcher busy-state semantics
ok - fm_backend_herdr_send_text_submit: reports 'empty' once agent_status reports working after one Enter, without ever reading the composer
ok - fm_backend_herdr_send_text_submit: reports 'pending-unproven' when agent_status never reports working after retried Enters (swallowed; not eligible for the busy-queued rescue)
ok - fm_backend_herdr_send_text_submit: a busy baseline whose composer reads stay ambiguous exhausts to 'pending-unproven', never the rescue-eligible 'pending'
ok - fm_backend_herdr_send_text_submit: a slash-command popup's placeholder fill on Enter #1 never flips agent_status to working, so it does not short-circuit as submitted; Enter #2 is retried and lands it
ok - fm_backend_herdr_send_text_submit: a post-Enter blocked state confirms delivery without retrying into the prompt
ok - fm_backend_herdr_send_text_submit: preexisting working is not accepted as submit proof when the composer still holds the message
ok - fm_backend_send_text_submit (herdr): an idle-baseline swallow with a concurrent turn starting before the busy probe stays a delivery failure instead of a false queued-busy
ok - fm_backend_send_text_submit (herdr): the composer-proven busy-baseline pending still receives the busy-queued rescue end to end
ok - fm_backend_herdr_send_text_submit: confirms submission via native agent-state alone, immune to a codex-style dynamic idle-tip composer that would have misread as 'pending' under the old composer-based confirmation
ok - fm_backend_herdr_composer_state: a faint real-codex dynamic idle-tip composer row reads empty
ok - fm_backend_composer_state (herdr): the pre-injection empty-box guard still refuses a genuinely non-empty composer, unaffected by the submit-confirmation change
ok - fm_backend_herdr_send_text_submit: a slow transition landing on a later sample within one Enter's budget is confirmed WITHOUT sending a needless extra Enter
ok - fm_backend_herdr_send_text_submit: reports 'send-failed' when the literal send-text call itself errors
ok - fm_backend_herdr_send_text_submit: reports 'unknown' when the post-Enter agent-get read fails (never retries past an unreadable target)
ok - fm_backend_validate: herdr is a known backend (P2)
ok - fm_backend_busy_state: tmux (no native primitive) always reports unknown, preserving the P1 regex-only path
error: unknown backend 'bogus' (known: tmux herdr zellij orca cmux)
ok - fm_backend_composer_state dispatches every backend to its named thin classifier, unknown for unrecognized backends
ok - fm-peek/fm-send: explicit stale targets matching metadata use the recorded backend
ok - fm_backend_herdr_normalize_event routes through the shared record with an empty from_status
ok - fm_backend_herdr_escalation_marker keys the dedupe marker exactly like the watcher's .stale-<key>
ok - fm_backend_herdr_apply_transition: blocked dedupe starts only after explicit commit
ok - fm_backend_herdr_apply_transition: a working edge clears the marker so the next ->blocked re-escalates
ok - fm_backend_herdr_clear_transition removes task-owned dedupe state
ok - fm_backend_herdr_apply_transition: idle/done (defer) and unknown/empty (fallback) take no fast action
ok - fm_backend_herdr_wait_transition: a home with no herdr panes falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: below-capability protocol/schema falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: reconnect level-reconcile returns an uncommitted blocked pane
ok - fm_backend_herdr_wait_transition: subscribes before reconnect level-reconcile
ok - fm_backend_herdr_wait_transition: a still-blocked, already-escalated pane is not re-delivered on reconnect
ok - fm_backend_herdr_wait_transition: a streamed ->blocked edge returns the record sub-poll
ok - fm_backend_herdr_wait_transition: streamed working clears the marker, idle/done are deferred (clean timeout)
ok - fm_backend_herdr_wait_transition: a reader/subscribe failure falls back to polling (rc 2)
ok - fm_backend_herdr_wait_transition: Bash 3.2-safe bad-ack path closes fd 9 and removes its FIFO
ok - fm_backend_herdr_wait_transition: stock macOS Bash clean timeout closes fd 9 and returns 1

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ bin/fm-backend.sh:771 - The hoisted rescue assumes verdict 'pending' means proven composer state (per intent), but the herdr adapter emits 'pending' from two weaker sub-paths: the idle-baseline path (bin/backends/herdr.sh:2750) returns 'pending' after exhausting Enter retries without ever reading the composer, and the busy-baseline path collapses the classifier's 'pending-unproven' (ambiguous composer) into the same final 'pending'. The dispatch-layer capture match scans the last 80 pane lines including scrollback, not the composer. Concrete false-delivery sequence: idle herdr pane; digest typed; every Enter swallowed (popup) with all agent-state polls idle; backend returns 'pending'; a concurrent writer (daemon wake or another fm-send) starts a turn in the sub-second gap before the busy probe; probe reads busy and the capture shows the text still sitting unsent in the composer; verdict upgrades to queued-busy and fm-send marks the pending-reply delivered ('do not resend') for text that was never queued. This contradicts the intent's own premise ("for 'pending' (proven composer state)") and the captain rule 'never claim delivery without proof', and skips the tmux adapter's documented rule that ambiguous pending never receives the busy-queue conversion. Recommended durable boundary: have the herdr adapter distinguish composer-proven pending (busy baseline plus classifier-proven text) from retries-exhausted-idle and pending-unproven, and let the dispatch rescue fire only on the proven variant.

🔧 Fix: emit herdr pending only composer-proven; unproven never rescued
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-backend.test.sh — 29 ok / 0 failures, including intent-named test_send_text_submit_busy_queued_readback (pending→queued-busy on busy+capture-match; unknown pass-through; stale-digest, idle-pane, and absent-text negative cases; conclusive verdicts untouched)
  • bash tests/fm-daemon.test.sh — 100 ok / 0 failures, including intent-named test_inject_msg_accepts_queued_busy (daemon inject_msg treats queued-busy as proof-carrying delivery)
  • bash tests/fm-backend-herdr.test.sh — 174 ok / 0 failures, including test_dispatch_herdr_proven_busy_baseline_pending_upgrades_to_queued_busy (end-to-end dispatch rescue), test_dispatch_herdr_idle_swallow_with_concurrent_turn_never_upgrades_to_queued_busy, and the pending-unproven exhaustion tests
  • Manual end-to-end CLI run: bin/fm-send.sh default:w1:p2 &#34;status update: ...&#34; against the herdr fakebin with a busy pane, swallowed Enter, and composer-proven text → rc=0 via queued-busy; herdr call log shows the hoisted busy probe + bounded capture
  • Before/after contrast: identical fixture against pre-change bin/ extracted from base commit b0ad61e → rc=1 verdict=pending (refused), proving the change enables busy-queued delivery
  • Negative CLI runs on new code: stale capture without this message → rc=1 verdict=pending (no false rescue); unreadable composer → rc=1 verdict=unknown (never upgraded)
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix: silence SC2329 for indirect subshell mocks in busy-queued test
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits July 27, 2026 12:28
* feat: add verified pi-signed adapter

* no-mistakes(review): Correct pi-signed maintainer verification date

* no-mistakes(review): Correct remaining pi-signed verification dates

* no-mistakes(review): Preserve authoritative pi-signed runtime identity

* no-mistakes(document): Document pi-signed shared adapter semantics

* no-mistakes: apply CI fixes
* fix(pi): rearm watcher across same-process session transitions

Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement
as well as terminal quit. The primary watcher extension latched a module-level
stopping flag on every shutdown, so a replacement session in the same process
could not arm monitoring until Pi restarted.

Own arm authority per session generation so only the active live generation
may start, stop, or rearm the child. Replacement sessions can arm again without
restarting Pi, stale prior-generation callbacks cannot mutate the active cycle,
and real quit still blocks late rearm.

* no-mistakes(review): Preserve Pi generation isolation and exit cleanup

* no-mistakes(document): Correct Pi watcher transition documentation
* Consume quota-axi pace signals in dispatch profile array selection.

Add quota-array-dispatch as the single owner of the pace-aware candidate
choice, keep AGENTS.md to the intake boundary and load trigger, and cover
the acceptance cases with sanitized schemaVersion 3 fixtures.

* no-mistakes(review): Stop and report genuine quota dispatch ties

* no-mistakes(document): Document quota pace freshness and uncertainty
…nguid#1171)

* fix(grok): adapt Stop continuation to runtime capability

* no-mistakes(review): Reject ambiguous Grok Stop payloads

* no-mistakes(review): Reject duplicate Grok fields and accept spaced tmux sessions

* no-mistakes(review): Enforce exact tmux cleanup selectors

* no-mistakes(test): Fix historical tmux fixture and validate Grok Stop

* no-mistakes: apply CI fixes
* fix(brief): make DOD scaffolding parse-safe on stock macOS Bash 3.2

fm-brief.sh built each Definition-of-done block and the not-enabled
Herdr declaration with `VAR=$(cat <<EOF ... EOF)`. On Bash 3.2 (macOS
/bin/bash) the lexer scans for the command substitution's closing `)`
textually and tracks quote state through the heredoc body, so a single
apostrophe, unbalanced quote, or unbalanced paren in that prose breaks
parsing of the whole script. Every ship-brief scaffold (no-mistakes,
direct-PR, local-only) failed with `unexpected EOF while looking for
matching )`. Bash 4+ parses it fine, so the breakage stayed invisible
everywhere except stock macOS.

Replace all four command-substitution heredocs with
`IFS= read -r -d '' VAR <<EOF || true`. That removes the `$(...)`
wrapper and the entire defect class regardless of future prose, and
preserves the variable expansion the direct-PR and local-only bodies
need. `read` keeps the heredoc's trailing newline that `$(...)` used to
strip, so trim one newline to keep every generated brief byte-identical
to prior output.

Guard the structure, not one historical phrase: a new test rejects any
heredoc nested in a command substitution anywhere in fm-brief.sh, where
the old assertion pinned a single apostrophe phrase and so missed the
reintroduction. Extend the stock-macOS Bash CI job from parsing one
script to the whole maintained shell surface (bin/*.sh,
bin/backends/*.sh, tests/*.sh), matching bin/fm-lint.sh's canonical file
set so parse scope and lint scope cannot drift apart.

* no-mistakes(review): Captain: harden Bash structure and inventory guards

* no-mistakes(document): Align stock macOS Bash contributor checks

* no-mistakes(lint): Suppress deliberate SC2016 literal fixture warnings
* fix(test): pin teardown tmux baseline to historical kill selectors

merge-base HEAD main collapses to HEAD after the exact-selector change
lands on the default branch, so the old teardown fixture was accidentally
exercising current exact targets. Resolve a content-historical permissive
tmux adapter from first-parent history and force that post-squash topology
inside the conformance case so main and feature branches keep the same
old-vs-new contract.

* no-mistakes(lint): Suppress intentional literal-pattern ShellCheck warnings
…id#1197)

Cut the runtime skill to the compact pace-aware selection procedure plus
minimum owner pointers. Keep every distinct decision rule and move expanded
acceptance scenarios to deterministic fixture ownership assertions.

Size: 170/1374/10187 -> 63/544/4068 (about 63%/60%/60% reduction).
…1219)

* Inherit config/backend into secondmate homes with deliberate-override preservation

Add backend to the shared inheritable config allowlist so launch, locked
bootstrap, and config-push converge a primary pin into secondmate homes as each
home local future-spawn default. Track last-inherited bytes in a private state
provenance marker so deliberate per-home overrides survive present and absent
primary convergence, keep --backend and FM_BACKEND stronger, and extend the
existing inheritance tests plus docs and skill claims.

* no-mistakes(review): Preserve equal unprovenanced backend overrides

* no-mistakes(review): Preserve symlink overrides and verify spawn precedence

* no-mistakes(review): Snapshot backend inheritance for consistent provenance

* no-mistakes(review): Simplify backend inheritance to primary-authoritative convergence

* no-mistakes(document): Document inherited backend override preservation

* fix: restore primary-authoritative backend inheritance after document regression

The document step reintroduced provenance and deliberate per-home override
semantics after review had simplified config/backend to plain primary-authoritative
allowlist membership. Restore the primary-always-wins path: present overwrites,
absent removes, no provenance marker, and docs/tests match that contract.

* no-mistakes(review): Add divergent backend precedence regression fixtures

* no-mistakes(document): Document backend inheritance contract
* fix(pi): remove Calm's exclusive Pi upper-version ceiling

tests/fm-calm-pi-extension.test.sh gated on a closed PI_COMPAT_VERSIONS
allowlist ("0.81.1 0.82.0") that refused any other installed Pi, and docs
described that range as "supported" rather than verified evidence. The
Calm CHANGELOG shows no API introduced at either version, so there is no
evidence for a real minimum; the presentation adapters already probe the
exact method they patch rather than checking a version.

Replace the allowlist with dated version evidence that never rejects a
newer Pi, and make each presentation adapter degrade independently with
a diagnostic if a future Pi removes its API, instead of the whole Calm
extension failing to load. Rewrite the feasibility doc's "Pi 0.81.1
through 0.82.0" phrasing to state it as verified evidence, not a
ceiling.

* no-mistakes(review): Probe missing Calm adapter exports safely

* no-mistakes(document): Document Calm's unbounded Pi compatibility
…enguid#1204)

* fix(guard): allow session-local todo tools in the primary

The delegation-shape guard denied TaskCreate and TaskUpdate because their
normalized names contain the `task` stem. Those tools write only the harness's
session-local todo list, which has no executor: it spawns no agent, allocates
no worktree, registers no schedule, and starts nothing that outlives the
session. That is not the unaccounted work the guard exists to stop, so the stem
match was a false positive, and the deny text told the primary to run
bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry.

Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than
widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only
observe or stop existing work. Both lists stay exact-name so neither can widen
by substring.

Tests cover the two allowed names and six near-miss names that a substring or
shortened-stem widening would release; both mutations were watched red.

* no-mistakes(review): drop session-local todo tools from recommended deny list

* no-mistakes: apply CI fixes
…claude pid (kunchenguid#1206)

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid

fm_harness_ancestry_pid() previously returned the first ancestor process
whose command matched a verified harness name. Claude Code's Stop hook
fires as a bg-spare worker several levels below the session's actual
lock-owning claude process (hook shell -> claude bg-spare ->
claude bg-pty-host -> claude -> claude(lock)), so the first match was
the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self()
then never matched state/.lock, and the Claude Stop auto-arm silently
treated its own primary session as an unrelated live owner and never
armed the watcher.

The walk now keeps going past a claude-named match, looking for a still
more ancestral claude-named match, and stops the instant a non-match
follows an already-found match (bounding it to a contiguous run rather
than the literal ancestry top, so an unrelated claude-named process
further up the real process tree is never mistaken for part of this
session's own nested chain). Every other harness keeps the original
first-match-wins behavior, since e.g. Pi's shared signed-wrapper
ancestry actually holds the session at the inner engine pid, not an
outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper
bg-spare chain.

* no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim

* no-mistakes: apply CI fixes
* fix: confirm watcher startup on MSYS

* no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test

* no-mistakes(review): validate OpenCode ready timeout, make uname cache internal
…d#1195)

* fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates

Crewmate panes are created by a long-lived tmux/herdr daemon that does not
inherit firstmate's current environment. When firstmate runs under a non-default
CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare
`claude` in the crewmate pane fell back to the default ~/.claude store and
launched unauthenticated, blocking the crewmate before it could do any work.

fm-spawn now prefixes the claude launch with firstmate's own resolved
CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config
store firstmate is authenticated with. An unset value is the single-store
default and adds no prefix; non-claude harnesses are unaffected.

Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set,
omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test
helper so launch assertions no longer depend on the developer's environment.

* no-mistakes: apply CI fixes
…guid#1233)

* fix: preserve dispatch harness identity

* no-mistakes(review): Fix Grok counterfactual tuple validation

* no-mistakes(document): Scope dispatch authentication to selected tuple

* fix: restore dispatch instruction budget

* no-mistakes(review): Scope dispatch authentication after candidate selection
* fix(bin): handle dash-leading harness process names (#2)

* fix: handle dash-leading harness process names

* no-mistakes(review): Make dash-leading harness regression hermetic

* fix: preserve secondmate reply routes across relative homes

Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved.

Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory.

* no-mistakes(review): Preserve absolute overrides and normalize relative durable paths

* no-mistakes(review): Normalize relative home before deriving durable paths

* no-mistakes(document): Document relative durable-path normalization

* no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization

* no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck
* Add internal status skill

* no-mistakes(document): register /status skill in documentation-audiences inventory

* no-mistakes(lint): replace grep|wc -l with grep -c in status skill test

* test: silence literal status skill patterns

* Refactor bearings default to chat-only

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* no-mistakes(review): Align project removal preflight with approved exception

* no-mistakes(document): Align project removal documentation with approved exception

* fix: restore removal test byte-for-byte and preserve the default sentence

tests/fm-instruction-owners.test.sh had been changed to assert different
text; restore it byte-for-byte to origin/main. project-management SKILL.md's
Remove section now keeps the exact default "Never issue a raw removal
command from Firstmate." sentence that test still asserts, immediately
followed by the already-approved captain-operation-or-scope exception, so
the default and the exception both stay explicit and consistent.

* no-mistakes(document): Align project-write boundary documentation
…henguid#1275)

* Route project intake through secondmate scopes

* no-mistakes(test): Guard all main-home project registry mutations

* no-mistakes(document): Consolidate secondmate routing documentation

* no-mistakes: apply CI fixes

* Restore new-project routing scope

* no-mistakes(document): Clarify secondmate routing for new-project intake

* no-mistakes: apply CI fixes
)

* fix: scope validation corrections by accepted behavior

* no-mistakes(review): Classify stale delivery evidence as an autonomous correction
…#1282)

* test: remove source-content assertions

* no-mistakes(review): Replace source assertions with runtime behavior coverage

* no-mistakes(review): Isolate Kimi task temp runtime coverage

* no-mistakes(document): Refresh test cleanup documentation

* no-mistakes: apply CI fixes
…#1286)

* fix(watch): bound how long a busy pane may run with no completed turn

A busy pane (backend busy state or the harness's rendered footer) was
unconditional, unbounded proof of liveness in every escalation path, so a
hung foreground tool call behind a busy signature could run for hours
undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic-
backtracking regex hung one bash call for 25h behind an unchanging
"Working..." footer).

FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may
run with no completed turn (state/<id>.turn-ended, or its spawn record
before any turn has completed). Past the bound, busy_turn_over_age routes
the pane through the existing wedge_timer_check, reusing the identical
stale reason, escalation counter, and demand-deep-inspection marker for
human inspection only - never an automatic interrupt, signal, or restart
of the worker or its tool process. A completed turn resets the age.

Reproduced end-to-end against the real installed Pi TUI: a foreground
`sleep 999999` bash call with no timeout renders the actual busy footer,
and two captures ~15s apart show the elapsed counter changing the pane
hash while the same turn stays unfinished. Running the pre-fix watcher
against the real captures showed it never starts a wedge timer no matter
how long the pane stays busy; the fixed watcher starts and escalates the
timer through the same mechanism, while the real hung process remained
untouched and alive throughout.

* no-mistakes(review): fix: parse enriched AFK stale reasons

* no-mistakes(review): fix: preserve enriched wedges during AFK supervision

* no-mistakes(review): fix: route all enriched AFK wedges

* no-mistakes(document): Clarify busy-turn age supervision documentation
…unchenguid#1261)

A name-by-name list of config/ entries silently stops ignoring any new or
home-local file placed there, which makes the working tree read as dirty and
blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md
already documents config/ as captain-private and gitignored as a category;
this makes .gitignore match that contract.
…al coverage (kunchenguid#1304)

The second assertion in fm-gitignore-config.test.sh (added by kunchenguid#1261) greps
.gitignore for a specific spelling of the config/ ignore pattern. It fails
on a semantically equivalent pattern like config/** and does not prove Git
actually ignores anything, per the completed source-content-test audit.

Replace it with a real git check-ignore control test on a generated
unrelated path, and strengthen the existing directory-coverage test with
generated unpredictable direct and nested config/ paths.
)

* Add bounded startup memory curation

* no-mistakes(review): Record reproducible stow verification evidence

* no-mistakes(review): Validate inherited secondmate stow evidence

* no-mistakes(document): Document editable startup-memory budget propagation
* fix(herdr): place workers in the launching agent's exact workspace

Herdr enforces no workspace-label uniqueness, and spawn resolved its
container by taking the FIRST workspace whose label matched the home
label. With two workspaces both labeled "firstmate", a worker launched
from the second one was created in the first, so it appeared in a
different space than the Firstmate the captain was watching.

Reproduced end to end on Herdr 0.7.5 protocol 17 by running the real
bin/fm-spawn.sh inside a launcher pane in the second "firstmate"
workspace: the worker landed in w1 while its launcher was in w2, with an
unrelated third workspace focused throughout, which also rules out any
dependence on the focused workspace.

Placement now binds to the launching process's own Herdr identity. Herdr
injects HERDR_PANE_ID, HERDR_SESSION, and HERDR_SOCKET_PATH into every
process it manages a pane for, and fm_backend_herdr_launcher_identity
resolves that pane's current owning tab and workspace live from Herdr,
cross-checking the pane against its tab and confirming the workspace
exists exactly once in the session. The injected HERDR_TAB_ID and
HERDR_WORKSPACE_ID are creation-time snapshots and are deliberately not
read as current identity. Labels are no longer placement authority.

A claimed parent identity that is unreadable, contradictory, stale, or
from another named session or Herdr server stops the spawn before any
worker endpoint exists, rather than degrading to a label search. A
launcher with no Herdr ancestry has no workspace to inherit and keeps
the per-home labeled container, which must now resolve to exactly one
workspace; two same-labeled candidates refuse instead of adopting
either. A --secondmate launch keeps standing up that home's own
workspace by design.

With presentation spaces enabled, the projected child is created and
bound under that same exact parent and anchors its ordering on it, so a
duplicated home label no longer makes the layout ambiguous. Projection,
focus restoration, restart binding, and quarantine rules are unchanged,
and children are never collapsed into the parent. tmux, Zellij, cmux,
Orca, and the away-mode daemon terminal were each inspected and are not
affected: none resolves a container by searching mutable labels.

tests/fm-backend-herdr-launcher-workspace-e2e.test.sh drives the real
spawn and teardown against an isolated Herdr lab, with its headline case
running fm-spawn.sh inside a real Herdr pane so the identity comes from
Herdr's own injection. The refusal matrix and the ordering anchor are
covered deterministically in tests/fm-backend-herdr.test.sh.

Eight existing real-Herdr suites inherited the developer terminal's own
Herdr pane into their isolated lab sessions, which the new cross-session
check correctly refuses. tests/herdr-test-safety.sh now owns
herdr_forget_inherited_pane and those suites call it, so what they assert
no longer depends on where they were launched from.

Two unrelated fixes found along the way. tests/fm-secondmate-harness.test.sh
had the same class of environment leak through CLAUDECODE, which outranks
PI_CODING_AGENT in bin/fm-harness.sh and made its pi-signed ancestry case
resolve "claude" whenever the suite ran inside Claude Code. And
fm-spawn.sh's usage() printed a fixed line range that had already been
truncating its own help mid-sentence.

* no-mistakes(review): Enforce exact Herdr launcher and projection identity

* no-mistakes(document): Document exact Herdr launcher workspace placement
* feat(calm): replace Pi's working row with an animated ship while Calm is on

While Calm is active and one logical agent run is under way, Calm now hides
Pi's built-in working row and renders a small two-row SSHHIP-derived boat in
its place. When Calm is off, Pi's stock working row is left untouched.

The presentation uses only public Pi extension API: setWorkingVisible(false)
plus a temporary setWidget() component whose render(width) owns the responsive
geometry and whose timer requests a TUI render. Visibility follows agent_start
through agent_settled, so the boat does not flicker between tool calls,
automatic continuations, retries, or compaction inside the same run, and
settle, abort, and failure all reach the same cleanup.

fm-calm.ts stays the sole owner of the presentation choice and the only caller
of setWorkingVisible(); the new lib owns the sprite geometry and widget.

* no-mistakes(review): Guarded Calm-off lifecycle visibility writes; focused tests pass

* no-mistakes(test): Fixed Calm E2E wait to include tmux scrollback

* no-mistakes(document): Document Calm working boat behavior

* no-mistakes: apply CI fixes

* feat(calm): slow the Calm boat, animate blue water, and make the sail directional

The boat now moves one column every 880ms while a bounded fixed-cell water phase
advances every 220ms, so the water ripples several times between boat steps and
the presentation reads as calm. One scheduler drives both clocks and disposing
the widget stops them together; ticks rather than wall-clock timestamps drive
every state change, so tests seek animation time exactly.

Colors are standard ANSI foreground codes instead of theme lookups: blue for
every water cell and yellow for the complete boat, each run closed with a
default-foreground reset so nothing bleeds into padding or later frames. ANSI
bytes never enter geometry, so visible width stays exact.

The mainsail is directional and trails aft of the mast: <| travelling right and
|> travelling left. Direction reverses the moment the boat lands on an endpoint,
so the endpoint frame already shows the new heading and no frame at or after a
bounce shows the previous sail.

* test(calm): wait for the Ctrl+O expansion redraw this block asserts

* docs(calm): record the revised working-presentation verification evidence

* no-mistakes(document): Fix Calm feasibility document EOF whitespace
…henguid#1349)

* fix(dispatch): scope candidate authentication to its own surface

A locally expired timestamp in one credential store was reported to the
captain as a sign-out, including for dispatch candidates that never read
that store. A `harness=pi, model=xai/grok-*` candidate authenticates
through Pi's own xAI credential, but the only Grok quota reading
available was gated on the standalone Grok CLI's separate token, whose
expiry clock drifts independently. The always-loaded intake rule then
turned that unreadable quota into a mandatory captain escalation.

Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts
that must not depend on agent memory: it resolves a tuple's
authentication surface from quota-axi's own emitted auth sources rather
than from a harness or model name, so another harness's CLI can never
gate a candidate that does not use it. A vendor CLI is launched only
when the tuple's own harness owns the credential store under test and a
non-destructive discovery command is registered for it, which today is
`grok models` alone. That probe runs at most once with stdin closed and
a hard timeout, reads its verdict from the first stdout line because the
command exits 0 either way, treats unrecognized output as indeterminate,
and never invokes login, logout, or the interactive TUI. Quota is read
at most twice, and unknown headroom never makes a candidate ineligible
on its own.

Update the dispatch procedure to match: usable authentication with
unmeasurable headroom stays eligible at lower preference with the
unknown disclosed, and stop-and-report is reserved for unresolved
authentication, an unresolved relationship, or malformed configuration.
Record that Grok's `credits.remaining` is a prepaid balance rather than
window headroom.

Gate quota-axi at 0.1.16 in bootstrap, the first build reporting
per-credential auth sources. A stale install previously passed the
presence check silently, which is why a fix published two days earlier
was still not in effect.

Replace the orphaned quota-array-dispatch fixtures, which encoded a
`provider: "xai"` shape the tool never emits and had no consumer, with
fixtures shaped like real 0.1.16 output that the new suite drives the
script against. The suite asserts the verdict and, separately, which
vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI
fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture
change selects the right tests instead of refusing.

* refactor(bootstrap): give the quota-axi floor one owner

The floor was stated twice - once in bootstrap's gate and once inline in
the auth preflight - so bumping it needed two edits that could drift.
Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the
existing tasks-axi library, and derive the comparison from the constant
so the number appears exactly once. Bootstrap turns a failing check into
the operator diagnostic; the preflight refuses to emit an unscoped
verdict. Map the new library to both consuming suites so a bump re-runs
them, and record that any usable source means the surface authenticates.

* no-mistakes(review): Captain: bound quota checks and removed Python dependency

* no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry

* no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing

* no-mistakes(review): Captain: reject malformed OpenCode model relationships

* no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation

* no-mistakes(document): Updated dispatch authentication documentation

* no-mistakes: apply CI fixes
…nchenguid#1350)

* feat(x-mode): reconcile promised public replies deterministically

A promised final reply in an X or Discord thread was only kept while the
primary remembered it. Compaction or restart erased that memory, so a typed
public-followup obligation could sit at pending-work after its PR merged and
the original thread never got its reply.

Make the promise durable state instead:

- bin/fm-public-followup-emit.sh reports a typed terminal work result (source
  home, work id, generation, outcome, safe deliverables, bounded public-safe
  text) into the owning home's private inbox. The event id is derived from
  that identity tuple, so duplicate reports and restart replay converge with
  no coordination, and nothing ever parses a free-form done: sentence.
- bin/fm-public-followup.sh registers a commitment, reconciles events through
  tasks-axi public-followup, and runs the idempotent delivery sequence
  (begin-delivery with the payload hash, post, record the posted receipt or a
  typed error) against the stored platform and opaque thread binding. A
  delivery interrupted between post and receipt refuses rather than risk a
  second public reply.
- Session start surfaces unresolved commitments from disk, the existing relay
  poll surfaces a new terminal-result set once, and teardown refuses while
  this home still owes a public reply for that exact work.

tasks-axi public-followup remains the only owner of the obligation state
machine, state/x-context/ the only owner of the private request context, and
fm-x-reply.sh the only thing that posts. Its new optional --receipt-file is
the one addition there, so a caller can record how many messages were sent.

A home that never opted into the myfirstmate relay gates out on a single
[ -f "$FM_HOME/.env" ] test: no tasks-axi call, no backlog or context scan,
no output, and no artifact. Evidence in docs/verification/public-followup.md.

* no-mistakes(review): Hardened public-followup reconciliation and ownership guards

* no-mistakes(review): Hardened typed terminal cleanup and receipt reconciliation

* no-mistakes(review): Automated typed-delivery cleanup and strict backlog validation

* no-mistakes(review): Fail-closed parent resolution and registration-safe delivery

* no-mistakes(review): Harden relay gating and validate secondmate bindings

* no-mistakes(review): Use owner-aware single-gate teardown protection

* no-mistakes(document): Correct public-followup documentation drift

* no-mistakes(lint): Quote done literals to fix ShellCheck warnings

* no-mistakes: apply CI fixes
…chenguid#1327)

* feat: add semantic busy-state contract owner and event writer

One owner (bin/fm-busy-lib.sh) for the captain-approved semantic
busy-state redesign: a per-task gen-bound record written only by
bin/fm-busy-event.sh, per-harness trusted-source classification with
explicit source attribution, busy/idle/unknown/dead semantics where
missing, malformed, stale, or untrusted semantic data is unknown -
never idle - and endpoint death is the only process-level override.
The Grok-only rendered-tail fallback and the standalone-Kimi
verification gate live behind the same classifier.

* feat: arm busy-state at spawn and convert Pi to the semantic extension path

fm-spawn arms the busy-state contract for converted adapters and seeds
busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed
per-task extension now reports agent_start -> busy and agent_settled ->
idle confirmed by ctx.isIdle(), covering auto-retries, compaction
retries, tool loops, and queued continuations, while turn_end stays a
wake notification touch. Teardown removes the new record, gen sidecar,
and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy ->
agent-settled idle with the marker still touched.

* feat: convert OpenCode to the semantic session.status plugin path

The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now
classifies from OpenCode's semantic session.status events - busy and
retry are active, idle is inactive - latched to the worker's own
session so a subagent child session can never clear the worker's busy
state. The session.idle marker touch stays a wake notification.
Teardown removes both the new and the legacy plugin filenames.
Live-verified on OpenCode 1.17.18 in a real TUI pane: seed ->
session-busy -> session-status-idle.

* feat: convert Claude to the full lifecycle hooks path

The per-task settings.local.json now wires UserPromptSubmit -> busy
and Stop, StopFailure, and SessionEnd -> idle, so API-error and
shutdown turn ends can never strand a busy record; Stop keeps the
turn-ended notification touch. A refused (stale-gen) event exits 0 and
stays silent so Claude's own lifecycle is never broken. Live-verified
on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch
prompt, Stop closes each turn, a mid-stream Escape interrupt fires no
closing hook, and the firstmate-controlled idle/fm-interrupt clear
resolves it.

* feat: gate Codex busy state behind verified semantic sources

The approved contract prefers Codex's app-server turn lifecycle with
capability negotiation and sanctions its lifecycle hooks as the
intermediate. Live probes on codex-cli 0.145.0 show neither is usable
for a pane worker: the app-server daemon is unreachable for a TUI
thread and refuses to start outside the managed standalone install,
and firstmate-written project hooks never fired (interactive with
directory trust granted, and exec, both with
--dangerously-bypass-hook-trust) while global hooks fired in the same
runs. Codex therefore classifies unknown codex-unverified behind an
explicit probe rather than falling back to idle or footer text, and
fm-spawn installs no unverified Codex wiring.

* feat: gate standalone Kimi busy state on live verification

Standalone Kimi has no installed binary here, so per the approved
contract its semantic path stays guarded and it classifies unknown
kimi-unverified rather than idle - and never from its locale-sensitive
moon-phase spinner, which the redesign forbids inventing as a state
source. The gate records the preferred source order (Wire prompt
request lifetime, which brackets a turn and reports cancellation, then
the documented hooks including Interrupt because Stop does not fire on
interrupts) and the exact evidence required to open it. Arming without
wiring would seed a busy record nothing could clear, so both land
together behind the same gate.

* feat: route busy consumers through the contract and drop the global OR

The watcher, crew-state reader, and away-mode daemon now decide busy
state through bin/fm-busy-lib.sh: only an exact busy verdict counts as
working, and unknown never becomes working or a silent idle, so a crew
whose semantic state is missing, malformed, stale, or unverified
surfaces instead of being absorbed. Crew-state reports the producing
source in its detail. The watcher's global OR regex default is gone;
Grok keeps its isolated fallback inside the contract. The daemon's
supervisor-pane reader stays rendered-text - that pane is not a
recorded task - but is now scoped to firstmate's own detected harness
instead of every vendor signature. Secondmate pending-reply
observation is deliberately unchanged and documented as a
delivery-confirmation signal, not task state.

* docs: point busy-state documentation at the single contract owner

Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as
the owner of what busy means, with per-adapter sources, the
unknown-never-idle rule, the endpoint-death override, and the two
rendered-text readers that deliberately stay outside the contract.
Replaces the stale regex-first prose in architecture, tmux-backend,
herdr-backend, and configuration; converts the harness-adapters
per-harness rows from UI signatures to the semantic source each
harness uses; and records the live verification evidence, including
why Codex and standalone Kimi stay unknown.

* fix: arm away-launch signal handlers before acquiring the lifecycle lock

fm_afk_launch_main acquired its lock and only then installed the EXIT,
INT, and TERM traps. A signal arriving in that window terminated the
process by default action and left the lock directory behind, which
blocks the next away-mode launch until the stale-owner reclaim path
clears it. The release helper only removes a lock this process owns,
so the handlers are now armed first. The accompanying test also killed
the child whether or not the lock had appeared and sampled cleanup the
instant wait returned; it now requires the lock, then allows a bounded
settle, so it proves the guarantee instead of racing it.

* test: align fleet, Kimi, lifecycle, and detection suites with the contract

The fleet snapshot and wake-daemon lifecycle fixtures now prove a
working crew through its own semantic busy-state record instead of
rendered pane text, which is what those consumers read. The Kimi
watcher test asserts the approved contract directly: a standalone Kimi
task classifies unknown rather than matching its moon-phase spinner,
while Grok's isolated fallback still classifies only Grok. The
pi-signed detection cases clear ambient harness markers, fixing a
pre-existing failure where the running session's own CLAUDECODE
outranked the fixture's marker.

* fix: stop teardown from deleting a project's own Codex hooks file

An intermediate revision wired Codex through a firstmate-written
<worktree>/.codex/hooks.json, and teardown removed it alongside the
other generated wiring. The Codex wiring was dropped when its probes
came back unverified, so that removal now targets a file firstmate
never creates - and a project may legitimately track its own
.codex/hooks.json, which teardown would then delete from a pooled
worktree.

* fix: keep busy-record parsing from disturbing its sourcing caller

The record parser split fields with set -- under a temporary noglob,
which clobbers a sourcing caller's positional parameters and restores
glob expansion even when the caller had disabled it. The watcher, the
daemon, and the crew-state reader all source this library, so it now
reads fields with read -a, which never globs and never touches caller
state.

* docs: state exactly which Claude hook paths were reproduced live

The busy-state record listed all four wired Claude hooks in the source
column, which could read as a claim that every one fired during the
pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are
wired from hook names confirmed present in the installed binary, but
the abnormal turn ends they cover were not reproduced.

* test: let reset_fakes own the crew-state busy-text fixture lifecycle

The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so
the variable's lifetime was owned by one test rather than by the
shared reset that every other fake already uses.

* no-mistakes(review): Fix semantic busy-state lifecycle races

* no-mistakes(review): Make busy-state retirement idempotent

* no-mistakes(review): Enforce semantic state boundaries for status and injection

* no-mistakes(review): Restore harness-scoped away-mode busy guard

* no-mistakes(document): Refresh semantic busy-state documentation

* no-mistakes: apply CI fixes
kunchenguid and others added 29 commits August 8, 2026 22:58
* feat(stow): tiered decaying memory with captain-gated offload to local excluded skills

Implement the captain-adopted /stow redesign from the v2 tiering report as
amended by the adoption decision:

- Per-entry trailing HTML-comment markers with three tiers named for their
  handling: pinned (no clock, no eviction), aging (stale after 30 days),
  perishable (stale after 7 days, mandatory checkable expiry condition).
- File-scoped defaults (captain.md and captain-shared.md pinned,
  learnings.md aging) with a self-describing legend line per file header.
- Reinforcement requires session evidence; re-reading memory never counts.
- Archive-not-delete: stale and budget-evicted entries move with provenance
  to the never-injected data/memory-archive.md; prune always means the cold
  tier, and a stale unique fact is never deleted.
- Captain-gated over-budget offload: staleness evaluated before scope, the
  sweep runs only when still over budget after decay and consolidation,
  proposals go through the receipt plus one durable captain-held backlog
  item, migration runs through the destination's normal path, and the
  memory entry leaves only once the destination is live.
- Offload destination per the adoption decision: a user-owned skill under
  .agents/skills/<freeform-name>/ excluded via the local .git/info/exclude,
  with the hard rule that stow never creates or writes a tracked skill.
- Five graduation moves, receipt verbs archived and proposed-offload, and
  the one-time non-destructive migration of unmarked legacy entries.

The public skills/stow/SKILL.md mirrors the generic parts (markers, decay,
archive exit, user-approved on-demand offload exit, migration) with no
firstmate-specific paths.

The load-bearing assumption that a git-excluded skill is still discovered
was verified empirically against Claude Code 2.1.226 (direct
.git/info/exclude scratch-repo test plus an in-repo ignored-probe test);
the dated evidence is recorded in docs/verification/stow-memory.md.

The graduation list's deletion move is deliberately narrowed to duplicates
already preserved by a stronger owner, reconciling the v2 report's retained
'deletion of a stale entry' wording with its own prune-always-archives
rule.

* no-mistakes(review): Persist legacy migration grace across stow passes

* no-mistakes(review): Enforce archival invariants and exempt default-pinned legacy entries

* no-mistakes(review): Clarify offload scope, archive placement, and marker boundaries

* no-mistakes(review): Enforce aging fallback and verify excluded skill loading

* no-mistakes(review): Fix stow decay, pinned offload, and archival safeguards

* no-mistakes(review): Preserve pinned entries, approvals, and archive provenance

* no-mistakes(review): Restrict stow mutations to editable memory files

* no-mistakes(review): Clarify skill destinations, collision checks, and migration legends

* no-mistakes(review): Resolve exclude paths for linked worktrees

* no-mistakes(review): Secure per-home excluded skill migration

* no-mistakes(test): Require explicit tier markers on new stow entries

* no-mistakes(test): Route missing shared legends to primary owner

* no-mistakes(document): Align stow documentation with tiered memory

* fix(stow): converge the pass on an over-budget home (dogfood D1-D3)

The dogfood run against a copy of the real over-budget home showed the
pass increasing the deficit from 624 to 1,107 estimated tokens and the
relief ladder provably unable to reach budget. Three skill-text fixes:

- D1: markers become single-token spellings (<!--a:DATE-->, <!--p:DATE-->,
  <!--P-->, <!--g-->), entries matching a pinned file default carry no
  marker, the per-file policy legend collapses to a one-line pointer
  naming the stow skill as the scheme owner, and marker/pointer bytes are
  explicitly counted content - roughly 76% less metadata cost on the
  dogfooded home's first installment.
- D2: the eviction rung gains a convergence precondition - total the
  eligible pool first, and when archiving all of it cannot reach budget,
  skip eviction entirely, archive nothing for budget reasons, and report
  the exempt pinned floor as the concrete inability in the final step.
- D3: budget eviction considers only dated aging entries; <!--g-->
  legacy-grace entries are ineligible until their grace cycle resolves,
  so eviction cannot cancel promised grace or invert against validation.

Public skill mirrors the D1 marker/pointer changes; D2/D3 are internal
because the public skill has no budget ladder.

* no-mistakes(test): Enforce evidence-only reinforcement during stow migration

* no-mistakes(document): Clarify stow receipt marker actions
* docs: add firstmate vision

* no-mistakes(test): Classify VISION.md as public product documentation

* no-mistakes(document): Restore approved one-file vision diff

* no-mistakes: apply CI fixes
* fix(spawn): force regular Pi TUI for crews

* no-mistakes(document): Documented Pi regular TUI launch mode
* fix(cmux): classify borderless Claude composer

* no-mistakes(review): Normalize cmux NBSP prompts across locales

* no-mistakes(document): Document cmux borderless Claude composer classification
…nchenguid#2091)

The public installer-facing stow skill scoped its classify-then-replace
discipline to TODO/BACKLOG items only, so findings routed to a memory file
had no stated rule against a blind append or a wholesale overwrite.

Step 6 now classifies every finding against the destination's current
contents as new, duplicate, superseding, or obsolete, and states the
considered replacement each classification implies. The outcomes follow the
tiered-memory contract already in the file: an obsolete entry is refreshed,
archived, or replaced in a way that preserves its fact, a duplicate folds
into the entry that already carries it, and a superseded body worth keeping
leaves through step 7's existing exits rather than a second recovery
mechanism.
* fix(watcher): resurface durable work after downtime

* no-mistakes(review): Make watcher rearm recovery durable and cursor-safe

* no-mistakes(review): Persist safe recovery markers across migration lock recovery

* no-mistakes(review): Retain stale lock when recovery marker publication fails

* no-mistakes(review): Preserve delivery-gap recovery and quarantine malformed markers

* no-mistakes(review): Serialize recovery consumption and report acknowledgment failures

* no-mistakes(review): Centralize recovery publication before clearing watcher evidence

* no-mistakes(review): Guarantee recovery evidence across queue and lock handoffs

* no-mistakes(review): Publish recovery evidence before durable wake commits

* no-mistakes(review): Replace recovery marker Perl dependency with Node

* no-mistakes(review): Keep interrupted wakes durable until handling acknowledgment

* no-mistakes(review): Add post-handling durable wake acknowledgements

* no-mistakes(review): Enforce post-handling acknowledgement across recovery and AFK return

* no-mistakes(review): Bind wake acknowledgements to recovery generations

* no-mistakes(review): Align wake regressions with generation-bound acknowledgements

* no-mistakes(document): Document durable re-arm recovery semantics

* no-mistakes(lint): Resolve ShellCheck warnings in recovery and watcher tests

* no-mistakes: apply CI fixes

* test(watcher): assert post-handling wake replay

* no-mistakes(review): Prevent successor loops and adopt legacy wake generations

* no-mistakes(review): Rearm durable wakes without recursive successor recovery

* no-mistakes(review): Align recovery tests with handling marker state

* no-mistakes(review): Delay handling transition until successor launch is established

* no-mistakes(review): Confirm wake handling only after successful prompt delivery

* no-mistakes(review): Acknowledge AFK wakes only after evidence publication

* no-mistakes(review): Prevent AFK wake loss before post-handling acknowledgement

* no-mistakes(document): Document durable wake acknowledgement semantics

* no-mistakes(lint): Suppress false positive for recovery action output

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* ci: add Windows Herdr automation spike

* ci: run Windows spike on its pull request

* fix: wait for Windows Herdr command output

* fix: run ANSI probe in pane shell

* ci: keep Windows Herdr spike manually triggered

* docs: clarify Windows Herdr spike verdict
* Add guided ahoy decision flow

* no-mistakes(document): Document guided Ahoy decision flow
* Harden stow memory budget policy

* Refine internal stow offload policy

* no-mistakes(review): Enforce shared-budget decisions and autonomous offload
…enguid#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes
…#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of kunchenguid#1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR kunchenguid#1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.
* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…unchenguid#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.
* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims
* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes
…id#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes kunchenguid#2109
…uid#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements
* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership
* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing
* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts
fm_backend_send_text_submit now wraps the backend-specific verdict with a
hoisted read-back: when the backend returns pending or unknown, it checks
whether the pane is provably busy with the typed text still visible in a
capture. If so, it returns queued-busy — a proof-carrying verdict proving
the message was queued for the next agent turn.

Previously only fm-send.sh had a caller-side read-back rescue; the daemon's
inject_msg accepted only exact empty and rejected every busy-pane queued
nudge with 'inject failed'. Now inject_msg (and fm-send.sh, fm-control.sh,
fm-spawn.sh) all benefit from the single hoisted read-back in the dispatch
layer.

Changes:
- fm-backend.sh: add fm_backend_send_condense helper; refactor
  fm_backend_send_text_submit to capture verdict, run read-back, emit
  queued-busy when busy-queued delivery is proven
- fm-supervise-daemon.sh: inject_msg accepts queued-busy alongside empty,
  logs 'inject queued' for the new verdict
- fm-send.sh: accept queued-busy as delivery in verdict case statement
- tests/fm-daemon.test.sh: add test_inject_msg_accepts_queued_busy,
  registered in runner

Evidence: data/loop-audit-2026-08-02.md (finding 1: injection into busy
pane was unconfirmable; backlog fm-sendcore-busy-queued-verdict-p1)
@willchen95
willchen95 merged commit e2a3634 into main Aug 13, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.