Skip to content

fix(afk): gate away-mode injection on a confirmed live agent - #577

Closed
zachlandes wants to merge 14 commits into
kunchenguid:mainfrom
zachlandes:fm/afk-guard-fix-p2
Closed

zachlandes wants to merge 14 commits into
kunchenguid:mainfrom
zachlandes:fm/afk-guard-fix-p2

Conversation

@zachlandes

Copy link
Copy Markdown
Contributor

Intent

Fix a verified safety hole: the away-mode supervise daemon could type an escalation digest (plus Enter) into a live shell, executing arbitrary commands when nobody is watching. Root cause: inject_msg (bin/fm-supervise-daemon.sh) gated injection ONLY on composer-emptiness, but the default starship/pure/spaceship shell prompt glyph U+276F (the heavy right-angle quotation mark) is byte-identical to the glyph claude draws for its own empty composer, so a supervisor pane whose agent had exited to a login shell read as an empty agent composer and was injected into. Fix: after the composer-empty check, inject_msg now also calls the pre-existing fm_backend_agent_alive probe (bin/fm-backend.sh; already used by the session-start secondmate-liveness sweep) and injects ONLY on a confident 'alive'. Deliberate decisions a diff-only reviewer would not know: (1) This FAILS CLOSED on purpose - both 'dead' (a bare shell) and 'unknown' (an unreadable pane, or a harness whose liveness is unverifiable for the backend, e.g. pi's generic node interpreter) defer, on the stated principle that a missed escalation is an inconvenience but a command typed into the captain's shell is not; deferred escalations are not lost (buffer + max-defer wedge alarm + afk-exit catch-up flush). A pi-based supervisor in away mode therefore stops receiving injected escalations - this is an accepted, documented consequence, not a regression to fix here. (2) I intentionally did NOT remove the glyph from any character class - that is a band-aid because an empty shell prompt and an empty agent composer are indistinguishable by content; the guard asks what process is running instead. (3) fm-send.sh shares the same content-blindness but is deliberate/firstmate-initiated (not autonomous) and a strict alive-gate there would break steering for every unknown-liveness backend (pi, zellij, orca, cmux); I analyzed and documented it with a recommended separate dead-only guard but deliberately did NOT change fm-send in this PR to keep the safety fix focused. Tests: three deterministic inject_msg units in tests/fm-daemon.test.sh plus a real-shell end-to-end Scenario D in tests/fm-afk-inject-e2e.test.sh, each of which fails against the pre-fix daemon (verified by neutering the gate: the suite exits non-zero with an explicit 'typed into a dead shell' / 'SHELL INJECTION' failure). Test-helper changes in tests/wake-helpers.sh add a pane_current_command reply to the fake tmux so the fake supervisor panes (which stand in for a live agent) report as alive. Docs: amended the one-owner composer-emptiness safety policy in docs/herdr-backend.md (new dated subsection) and corrected the now-inaccurate claim in docs/architecture.md that a bare dead-shell prompt cannot receive an escalation. This is a generic bug affecting every firstmate user with a default shell prompt, intended to go upstream.

What Changed

  • inject_msg in bin/fm-supervise-daemon.sh now probes fm_backend_agent_alive after the composer-empty check and injects only on a confident alive. An empty starship/pure prompt glyph (U+276F) is byte-identical to claude's empty-composer glyph, so a supervisor pane whose agent had exited to a login shell previously read as an empty composer and got the escalation digest typed in and submitted as shell commands. The guard fails closed: both dead and unknown (an unreadable pane, or a harness whose liveness the backend cannot attribute, e.g. pi's generic node) defer instead of injecting, so a pi-backed supervisor stops receiving injections in away mode - deferred escalations still survive via the buffer, the wedge alarm, and the afk-exit catch-up flush.
  • Because the guard can leave away mode permanently unable to deliver, new bin/fm-afk-canary-lib.sh derives and words that verdict in one place, and surfaces it on both the arm path (bin/fm-afk-launch.sh) and session start (bin/fm-bootstrap.sh, as an AFK_INJECTION_DISABLED: diagnostic). It probes the running daemon's own recorded target rather than the current pane, so a firstmate restarted into a new pane learns the daemon is still injecting elsewhere, and reports a vanished target instead of staying quiet.
  • The wedge alarm now records why the last inject deferred and reports that cause rather than always blaming a busy pane, and gains a separate debounced, latched pane-gone alarm (its own .subsuper-pane-gone marker, one alert per absence episode) that declares supervision DOWN when the supervisor pane vanishes. Captain-facing wording keys off the alarm kind, not the last recorded cause, so a buffer wedge can never render as "away mode is dead". Covered by three deterministic inject_msg units in tests/fm-daemon.test.sh and a real-shell Scenario D in tests/fm-afk-inject-e2e.test.sh, each of which fails against the pre-fix daemon; bin/fm-send.sh shares the same content-blindness and is documented in docs/herdr-backend.md but deliberately left unchanged here.

Risk Assessment

⚠️ Medium: A large, safety-critical change to the autonomous away-mode injection path that is well-tested and strictly improves on the pre-existing behavior, but it adds a fail-closed gate that fully disables away-mode escalation delivery for any supervisor whose harness a backend cannot attribute (pi, and any node-wrapped harness) — an accepted, documented regression — and ships with three deliberately deferred robustness gaps in the new alarm and canary paths.

Testing

Baseline test command had already passed; I re-ran the full suite (68 files, 0 failures) and then produced product-level before/after evidence for the safety fix. Using a real login shell displaying the default starship ❯ glyph, the end-to-end Scenario D shows the fixed daemon refuses to inject and defers the escalation, while a neutered (pre-fix) daemon typed the digest into and executed it in the live shell — the test caught this with an explicit SHELL INJECTION failure. The deterministic inject_msg units confirm the same at the function level, including fail-closed-on-unknown and the alive+empty positive path. I restored the daemon file and verified a clean working tree afterward. This is a CLI/daemon behavior change with no rendered UI surface, so evidence is CLI transcripts rather than screenshots.

Evidence: Evidence summary: before/after shell-injection guard

### A. Real-shell end-to-end (Scenario D) PRE-FIX (guard neutered): not ok - Scenario D: the escalation digest was typed into the live shell (found 'Supervisor escalate' in the pane) — SHELL INJECTION [e2e exit 1] WITH FIX: ok - Scenario D: a dead shell showing a bare starship ❯ prompt is refused — the escalation defers, nothing is typed into the shell [e2e exit 0] ### B. Deterministic inject_msg units PRE-FIX (neutered): not ok - inject_msg typed into a dead shell: send_text_submit ran despite agent_alive=dead (starship-glyph hole) WITH FIX: ok - defers on empty-composer dead shell (agent_alive=dead) ok - defers when agent liveness is unknown (fail closed) ok - injects only when agent alive and composer empty (positive path) ### C. Full suite: 68 test files, 0 failures, SUITE_EXIT=0

AFK away-mode shell-injection guard — evidence
==============================================

The hazard: the away-mode daemon typed an escalation digest + Enter into a
supervisor pane whenever the composer read EMPTY. A dead login shell showing the
default starship glyph U+276F (❯) is byte-identical to claude's empty composer,
so a pane whose agent exited read 'empty' and the digest was executed as a shell
command. Fix: inject only when fm_backend_agent_alive confirms a live agent.

### A. Real-shell end-to-end (Scenario D) — the actual product surface

PRE-FIX (agent-liveness guard neutered → gate on composer-emptiness only):
not ok - Scenario D: the escalation digest was typed into the live shell (found 'Supervisor escalate' in the pane) — SHELL INJECTION
  e2e exit: 1 (FAIL — digest was typed into and run by the live shell)

WITH FIX (guard restored):
ok - Scenario D: a dead shell showing a bare starship ❯ prompt is refused — the escalation defers, nothing is typed into the shell
  e2e exit: 0

### B. Deterministic inject_msg units

PRE-FIX (neutered):
not ok - inject_msg typed into a dead shell: send_text_submit ran despite agent_alive=dead even though the composer read empty (starship-glyph hole)
not ok - starship-glyph inject_msg subshell failed

WITH FIX:
ok - inject_msg: defers on an empty-composer dead shell (bare starship ❯ prompt, agent_alive=dead), closing the shell-injection hole
ok - inject_msg: defers when agent liveness is unknown (fail closed), even on a confirmed-empty composer
ok - inject_msg: injects only when the agent is confirmed alive and the composer is confirmed empty (positive path)

### C. Full suite
  68 test files, 0 failures, SUITE_EXIT=0
Evidence: Full test suite log (68 files, 0 failures)
tmux 3.7b
== tests/fm-afk-inject-e2e.test.sh ==
ok - Scenario A: partial input defers injection; digest arrives clean after idle
ok - Scenario B: swallowed Enter produces exactly one clean digest
ok - Scenario C: a normal captain status injects exactly one clean single-line sentinel digest
ok - Scenario D: a dead shell showing a bare starship ❯ prompt is refused — the escalation defers, nothing is typed into the shell
all e2e injection tests passed
== tests/fm-afk-inject-herdr-e2e.test.sh ==
ok - real herdr Scenario A: partial input defers injection; digest arrives clean after idle
ok - real herdr Scenario B: swallowed Enter (via the herdr shim) produces exactly one clean digest
ok - real herdr Scenario C: a normal captain status injects exactly one clean single-line sentinel digest
ok - real herdr Scenario D: a persistently pending composer raises the max-defer wedge alarm, preserves the buffer, and never crashes the daemon
all real-herdr afk injection e2e tests passed
== tests/fm-afk-launch.test.sh ==
ok - clear-stale: removes every away-session delivery artifact the one list names
ok - clear-stale: leaves the durable wake-queue intact (no pending work dropped)
ok - refresh: daemon already alive - stale artifacts preserved (current session's buffer kept)
ok - arm canary: a dead supervisor target warns in the foreground, names the wrong target, and still arms
ok - arm canary: an ambiguous verdict names BOTH its causes and confirms the target before accepting it
ok - arm canary: a live supervisor agent arms with no warning
ok - arm canary: a refresh probes the running daemon's target and says to stop away mode, not just re-arm it
ok - arm canary: a live daemon injecting into a vanished pane is reported loudly, not swallowed
ok - arm canary: an unresolvable pane with no live daemon stays quiet
ok - stop-ordering: daemon SIGTERM'd while .afk still present (flush is not a no-op)
ok - stop-ordering: .afk cleared last
ok - stop-ordering: daemon-terminal record removed
ok - stop identity: stale lock cannot signal an unrelated live process
ok - failed start: away flag and delivery artifacts roll back
ok - concurrent start: one serialized daemon terminal remains tracked
ok - launcher lock: incomplete publication receives initialization grace
ok - launcher signal: TERM exits and releases the lifecycle lock
fm-afk-launch: daemon launched in non-visible herdr workspace ws-partial (pane lab:pane-exact), supervising lab:captain
ok - herdr create: malformed response recovers durable exact ownership
fm-afk-launch: herdr create failed after returning exact ids; closing lab:pane-exact
fm-afk-launch: recorded terminal teardown is unconfirmed; preserving exact id
ok - herdr create error: unconfirmed exact id is persisted for reconciliation
fm-afk-launch: failed to run daemon in herdr pane lab:pane-exact; closing it
fm-afk-launch: recorded terminal teardown is unconfirmed; preserving exact id
ok - herdr run failure: unconfirmed exact id remains reconcilable
fm-afk-launch: failed to persist daemon terminal record; closing tmux:exact-session
ok - record failure: newly created terminal is closed by exact id
fm-afk-launch: daemon did not become ready; closing tmux:exact-session
ok - readiness failure: exact terminal and durable record roll back
fm-afk-launch: daemon did not become ready; closing tmux:exact-session
fm-afk-launch: recorded terminal teardown is unconfirmed; preserving exact id
ok - readiness failure: unconfirmed terminal retains its reconciliation id
ok - tmux absence: clean missing differs from transport probe failure
ok - native lifecycle: launcher owns state with no terminal
ok - native lifecycle: uniform stop clears state without closing a terminal
ok - native entry: launcher-prepared lifecycle state is not rewritten
fm-afk-launch: reconciling leaked daemon terminal tmux:exact-session
fm-afk-launch: recorded terminal teardown is unconfirmed; preserving exact id
ok - teardown failure: exact terminal record is preserved
ok - record publication: failed atomic rename preserves the complete prior record
fm-afk-launch: daemon terminal record is malformed; refusing to act on it
ok - record read: malformed record fails closed without acting on a partial id
fm-afk-launch: daemon terminal record is malformed; refusing to act on it
fm-afk-launch: malformed daemon terminal record; refusing to stop away mode
ok - stop: malformed terminal record preserves away state and fails closed
fm-afk-launch: failed to create detached tmux daemon session 'fm-afk-daemon-2134678404-23528-23822-1784063985'
ok - tmux launch: planned exact target is recorded before creation and removed on failure
fm-afk-launch: failed to create detached tmux daemon session 'fm-afk-daemon-3400271325-23564-7409-1784063985'
ok - tmux launch: unique names eliminate collision teardown
ok - stop validation: malformed record causes no daemon or state side effects
ok - launcher lock: incomplete metadata fails acquisition and releases lock
fm-afk-launch: failed to clear away-mode flag
fm-afk-launch: away mode stopped; terminal teardown remains recorded for retry
ok - stop state: away-flag removal failure is surfaced
fm-afk-launch: away-mode daemon did not exit after SIGTERM; preserving lifecycle state
ok - stop liveness: captured live daemon preserves lifecycle state after lock release
fm-afk-launch: daemon terminal record is malformed; refusing to act on it
fm-afk-launch: daemon terminal record is malformed; refusing to act on it
ok - refresh record: malformed terminal identity fails closed
fm-afk-launch: failed to clear stale away-mode artifacts
ok - clear failure: native entry aborts and restores prior state
fm-afk-launch: reconciling leaked daemon terminal tmux:exact-session
fm-afk-launch: terminal close command failed, but exact absence was confirmed
ok - confirmed absence: cleanup succeeds and removes the stale record
fm-afk-launch: rollback restoration incomplete; backup retained at /var/folders/4n/_bv20k5d26dgrq1t5y4l69940000gn/T//fm-afk-restore-fail.VPsgV8/state/.afk-launch-backup.wlwtsZ
ok - rollback restore: incomplete restoration retains its recovery backup
ok - rollback: every away-session artifact a failed arm can write is cleared
ok - flag failure: lifecycle aborts without active state
ok - herdr e2e: captain tab pane count unchanged after start (no split)
ok - herdr e2e: daemon launched in a separate non-visible workspace
ok - herdr e2e: daemon pane is NOT in the captain's tab
ok - herdr e2e: daemon terminal scoped to the lab session
ok - herdr e2e: captain tab pane count restored after stop
ok - herdr e2e: daemon workspace removed by exact id on stop
ok - herdr e2e: record + .afk cleared on stop
ok - tmux e2e: captain window pane count unchanged after start (no split-window)
ok - tmux e2e: daemon launched in a separate detached session
ok - tmux e2e: captain window pane count unchanged after stop
ok - tmux e2e: daemon session killed by exact id on stop
ok - tmux e2e: record + .afk cleared on stop
== tests/fm-arm-pretool-check.test.sh ==
ok - matrix A01: allow through all five entry forms
ok - matrix A02: allow through all five entry forms
ok - matrix A03: allow through all five entry forms
ok - matrix A04: allow through all five entry forms
ok - matrix A05: allow through all five entry forms
ok - matrix A06: allow through all five entry forms
ok - matrix A07: allow through all five entry forms
ok - matrix A08: allow through all five entry forms
ok - matrix A09: allow through all five entry forms
ok - matrix A10: allow through all five entry forms
ok - matrix A11: allow through all five entry forms
ok - matrix A12: allow through all five entry forms
ok - matrix A13: allow through all five entry forms
ok - matrix A14: allow through all five entry forms
ok - matrix A15: allow through all five entry forms
ok - matrix A16: allow through all five entry forms
ok - matrix A17: allow through all five entry forms
ok - matrix R01: allow through all five entry forms
ok - matrix R02: allow through all five entry forms
ok - matrix R03: allow through all five entry forms
ok - matrix R04: allow through all five entry forms
ok - matrix R05: allow through all five entry forms
ok - matrix R06: al

... [103640 bytes truncated] ...

ff plain stale identities for daemon-owned pause triage
== tests/fm-watcher-lock.test.sh ==
ok - simultaneous watcher starts leave exactly one live process
ok - fm_pid_identity is locale-invariant across LC_ALL/LC_TIME
ok - killed watcher stale lock is reclaimed
ok - live watcher lock with stale heartbeat is actionable
ok - guard banner leads when down with pending wakes (re-arm-after-drain) and stays silent when fresh
ok - concurrent fm_lock_try_acquire yields exactly one winner
ok - dead-pid stale lock is reclaimed by a single acquirer
ok - concurrent stale-lock steal yields exactly one winner
ok - live steal mutex is not reclaimed
ok - live-held lock is not stolen
ok - empty mid-acquire lock keeps a minimum grace
ok - late original claimant cannot claim a recreated lock
ok - paused mid-acquire claimant backs off to active stealer
ok - watch restart refuses to signal a reused pid
ok - watch restart reports a healthy peer without attaching to it
ok - watcher self-evicts when the lock pid no longer names it
ok - arm attaches to a live fresh watcher and exits only when that cycle ends
ok - arm starts+confirms a fresh watcher on a clean lock and self-heals a dead-pid lock (never healthy off a dead pid)
ok - arm cleans child watcher and temp output on HUP
ok - arm propagates an immediate watcher wake before confirmation
ok - arm attaches to a peer watcher after child stands down and exits when peer dies
watcher: lock held by live pid 5522 but heartbeat is stale for 837350992s (>300s); inspect or stop that watcher before re-arming.
ok - arm reports FAILED and exits non-zero when no fresh watcher can be confirmed
== tests/fm-x-mode.test.sh ==
ok - fm-x-poll is a hard no-op without a token (inert default)
ok - fm-x-poll treats an explicitly empty env token as configured
ok - fm-x-poll stays silent on HTTP 204 (the common case)
ok - fm-x-poll lets an explicitly empty relay env override .env
ok - fm-x-poll surfaces auth/config errors once and clears on recovery
ok - fm-x-poll stashes the question and prints the compact marker
ok - fm-x-poll preserves in_reply_to conversation context in the inbox
ok - fm-x-poll reports inbox commit failures without emitting a mention wake
ok - fm-x-poll requires a non-empty question before waking
ok - fm-x-poll rejects an unsafe request_id (path-traversal guard)
ok - fm-x-reply posts a request-bound answer and echoes only the request_id
ok - fm-x-reply accepts the reply via --text-file and stdin (safe, unexpanded)
ok - fm-x-reply exits non-zero on a non-2xx relay response
ok - fm-x-reply cleans up auth header temp files on interrupted posts
ok - fm-x-reply rejects missing arguments with a usage error
ok - fm-x-reply --help makes image support discoverable
ok - fm-x-reply rejects whitespace-only reply text
ok - fm-x-reply dry-run records the would-be reply and never posts
ok - fm-x-reply dry-run works without a token
ok - fm-x-reply honors FMX_DRY_RUN from .env
ok - fm-x-reply lets an explicitly empty dry-run env override .env
ok - fm-x-reply dry-run fails when it cannot record the preview
ok - fmx_split_thread: word-boundary, fence-aware, within-limit, numbered, lossless, capped
ok - fm-x-reply keeps a concise reply as a single unnumbered tweet
ok - fm-x-reply auto-splits a long reply into a numbered thread (texts[])
ok - fm-x-reply uses the Discord inbox platform budget instead of the X tweet budget
ok - fm-x-reply keeps numeric X requests on the X tweet budget
ok - fm-x-reply prefers an explicit relay-provided reply limit
ok - fm-x-reply clamps a below-floor max to 50 characters
ok - fm-x-reply posts a thread payload (texts[]) to the relay
ok - fm-x-reply --image posts an image object on answer
ok - fm-x-reply streams large image payloads outside curl argv
ok - fm-x-reply dry-run records compact image metadata for threaded replies
ok - fm-x-reply cleans image and payload temp files
ok - fm-x-reply --image rejects missing and unsupported image paths clearly
ok - fm-x-reply --followup posts to /connector/followup with the same request-bound body
ok - fm-x-reply maps a followup_unavailable follow-up 409 to exit 9
ok - fm-x-reply maps every follow-up 409 to exit 9 even without the marker
ok - fm-x-reply treats answer-endpoint 409 as a generic failure
ok - fm-x-reply --followup --image posts an image object
ok - fm-x-reply --followup is accepted in any position and leaves the answer path default
ok - fm-x-reply --followup dry-run marks the endpoint without changing the answer path
ok - fm-x-reply --followup auto-splits a long follow-up into a marked thread
ok - fm-x-reply followup dry-run keeps endpoint marker and compact image metadata
ok - fm-x-poll records the durable per-request reply context from the relay payload
ok - context registry retention is bounded to the seven-day follow-up window
ok - context registry rewrites preserve the first-seen timestamp
ok - context retention starts only when a live initial answer succeeds
ok - a delayed Discord follow-up stays one message after inbox cleanup via the durable registry
ok - an X follow-up over 280 still splits correctly after inbox cleanup
ok - every unresolved follow-up is refused before posting
ok - a partial registry platform combines with the relay's authoritative budget
ok - concurrent requests each recover their own platform/budget with no cross-overwrite
ok - fm-x-dismiss clears the durable per-request context (a dismissed mention gets no follow-up)
ok - fm-x-dismiss posts a request-bound dismiss and echoes only the request_id
ok - fm-x-dismiss dry-run records the would-be body and never posts
ok - fm-x-dismiss dry-run works without a token
ok - fm-x-dismiss exits non-zero on a non-2xx relay response
ok - fm-x-dismiss exits non-zero on a transport failure
ok - fm-x-dismiss rejects an unsafe request_id (path-traversal guard)
ok - fm-x-dismiss rejects missing or extra arguments with a usage error
ok - fm-x-link records and refreshes the X-request link without disturbing meta
ok - fm-x-link records Discord platform context so follow-ups keep the Discord budget
ok - fm-x-link resolves the platform by request_id so a post-cleanup link keeps the Discord budget
ok - fm-x-link warns loudly and the follow-up is held (not wrongly split) when the platform is unknown
ok - fm-x-link paired carry flags preserve a prior task's follow-up binding onto a successor
ok - fm-x-link recovery relink preserves Discord platform context after inbox drain
ok - fm-x-link rejects malformed or unpaired carry flags
ok - meta rewrites are independent of TMPDIR
ok - fm-x-link rejects unsafe ids, missing meta, and missing arguments
ok - fm-x-followup --check reports postable / not-linked correctly
ok - fm-x-followup --check prunes a link past the 7-day window
ok - fm-x-followup --check prunes a link that already reached the follow-up cap
ok - fm-x-followup posts a follow-up, increments the counter, and keeps the link under the cap
ok - fm-x-followup --final clears the link after one post regardless of the remaining count
ok - fm-x-followup clears the link once the third follow-up reaches the cap
ok - fm-x-followup --image forwards the attachment through fm-x-reply --followup
ok - fm-x-followup keeps the link and counter when the post fails
ok - fm-x-followup tombstones the link when a post-success counter write fails
ok - fm-x-followup treats a relay cap/window rejection as an already-exhausted link, not a retry
ok - fm-x-followup skips silently and clears the link past the 7-day window
ok - fm-x-followup is a no-op for a task with no X link
ok - fm-x-followup dry-run records the follow-up and increments the counter, keeping the link
ok - fm-x-followup dry-run --final clears the link just as a live post would
ok - fm-x-followup rejects malformed invocations
ok - bootstrap activates X mode from an .env token, idempotently
ok - bootstrap reports missing X-mode dependencies before arming
ok - bootstrap does not report X mode on when activation artifacts cannot be written
ok - bootstrap is inert without a non-empty .env token (non-X users unaffected)
ok - bootstrap cleans up X artifacts on opt-out and is silent once off
ok - bootstrap reports failed X artifact cleanup on opt-out
SUITE_EXIT=0
Evidence: Pre-fix (neutered) e2e — SHELL INJECTION failure
ok - Scenario A: partial input defers injection; digest arrives clean after idle
ok - Scenario B: swallowed Enter produces exactly one clean digest
ok - Scenario C: a normal captain status injects exactly one clean single-line sentinel digest
not ok - Scenario D: the escalation digest was typed into the live shell (found 'Supervisor escalate' in the pane) — SHELL INJECTION
Evidence: Fixed e2e — Scenario D refuses injection
ok - Scenario A: partial input defers injection; digest arrives clean after idle
ok - Scenario B: swallowed Enter produces exactly one clean digest
ok - Scenario C: a normal captain status injects exactly one clean single-line sentinel digest
ok - Scenario D: a dead shell showing a bare starship ❯ prompt is refused — the escalation defers, nothing is typed into the shell
all e2e injection tests passed
Evidence: Pre-fix (neutered) inject_msg units — failures
ok - fm-afk-start.sh fails before daemon startup when the afk flag cannot be written
ok - fm-afk-start.sh ignores stale pidfile-only live pids
ok - fm-afk-start.sh reclaims stale daemon locks whose live pid identity no longer matches
ok - supervise daemon state root is scoped by FM_HOME
ok - routine signal self-handles
ok - captain-relevant status verbs escalate
ok - check + unknown escalate; heartbeat self-handles
ok - transient stale self-handles and records a persistence marker
ok - stale + terminal status escalates immediately
ok - paused reasons with captain phrases remain pause-classified
ok - handle_wake on a paused stale records a pause marker, drops the wedge marker, and does not escalate
ok - handle_wake records a declared pause from a routine signal for long-cadence rechecks
ok - a terminal signal clears pause and stale tracking across both supervisors
ok - housekeeping migrates a normal-watcher's declared pause into daemon tracking
ok - housekeeping clears an already-resumed watcher pause across both supervisors
ok - housekeeping seeds pause tracking from status without a watcher marker
ok - persistent stale escalates after threshold and clears its marker
ok - resumed (busy) stale clears its marker without escalating
ok - housekeeping re-surfaces a stale declared pause on the long cadence and resets its window
ok - housekeeping clears a paused marker whose pane became busy again, without escalating
ok - housekeeping clears a paused marker once the crew is no longer declaring the pause
ok - housekeeping moves an existing stale marker to pause before wedge escalation
ok - housekeeping clears tracking when a crew leaves pause
ok - persistent herdr stale resolves the target from metadata and escalates
ok - herdr idle busy-footer stale clears through capture corroboration
ok - resumed herdr stale clears through backend-aware busy state
ok - persistent Orca stale resolves the terminal from metadata
ok - multiple escalations flush as a single batched digest
ok - batch flush measures max-delay from the first append, not the last
ok - catch-all scan escalates a missed terminal once, not twice
ok - handle_wake routes routine->self and captain->escalate
ok - INJECT_SKIP forces self-handle, bypassing captain-relevant classification
ok - is_wake_reason distinguishes watcher wake reasons from singleton-status stdout
ok - terminal-stale escalate removes its marker so housekeeping does not re-escalate
ok - captain signal escalate marks seen so the catch-all scan does not re-fire
ok - _collapse_newlines replaces newlines with literal separator
ok - afk flag absent: daemon does not inject, buffer preserved
ok - busy-guard defers injection when supervisor pane is busy
ok - marker detection: marker -> stay afk, no marker -> exit afk
ok - /afk invocation is exempt from afk exit (no self-cancel)
ok - should_exit_afk returns false when afk is not active
ok - strip_injection_marker removes the sentinel marker cleanly
ok - pane_input_pending detects partial input on the cursor line
ok - pane_input_pending: blank cursor line is not pending
ok - pane_input_pending: bare prompts are not pending (idle)
ok - fm_tmux_composer_state: a bare shell prompt ($/%/#/>) reads unknown, never empty (dead-shell injection safety)
ok - fm_tmux_composer_state: a bordered composer box and bare agent glyphs (❯/›) still read empty
ok - fm_tmux_composer_state: only matching edge borders form a composer box
ok - pane_input_pending honors FM_COMPOSER_IDLE_RE after border stripping
ok - classify_signal dedupes against the catch-all scan seen marker
ok - classify_stale dedupes against the signal path seen marker
ok - pane_input_pending: an idle bordered composer is NOT pending (afk-invx-i5)
ok - pane_input_pending: text inside a bordered composer is still pending
ok - submit-ACK confirms a submit when the composer returns to a bordered-empty box
ok - submit-ACK reports pending on a persistently swallowed Enter (type-once)
ok - max-defer on an empty stuck pane types once, alarms with the swallowed-Enter cause, and preserves the buffer
ok - max-defer flushes and clears the buffer on an empty bordered pane
ok - max-defer on a pending composer alarms without typing, naming the composer as the cause
ok - normal flush clears a stale wedge marker
ok - below MAX_DEFER: no inject, no alarm, buffer preserved
ok - max-defer does not flush or alarm while afk is inactive
ok - library mode: sourcing the daemon defaults FM_WEDGE_ALARM_EXEC to discard (no test can fire a real notification)
ok - wake helpers replace inherited notifier overrides with the safe recorder
ok - the discard seam suppresses every notifier, including command: (fires nothing)
ok - direct notifier helpers honor the discard seam, including command:
ok - osascript channel routes through the notifier seam with the summary (never a real notification)
ok - herdr channel routes through the notifier seam with the summary (never a real notification)
ok - command channel runs the captain command with the summary on $1 and on stdin
ok - command channel failures redact configured commands while logging their exit status
ok - unknown channel directives are redacted while the alarm keeps running
ok - off disables every active alert regardless of directive position (marker and tmux flash are unaffected)
ok - auto resolves to the macOS osascript notifier on Darwin (default-on)
ok - auto on a non-macOS platform selects no built-in OS channel (the marker or a configured command carries it)
ok - config/wedge-alarm selects every configured channel and skips comment and blank lines
ok - a failing channel logs and falls back to the next channel, never crashing the alarm
ok - a hung notifier is bounded, logged, and falls through to the next channel
ok - a backgrounded command notifier remains bounded until its process group is reaped
ok - a hung notifier override is bounded, logged, and proceeds to the next channel
/Users/zacharylandes/.no-mistakes/worktrees/01004ff68993/01KXFBCYZTCDTN4KY16GR41M1S/bin/fm-supervise-daemon.sh: line 752: 12893 Terminated: 15          sh -c 'sleep 30 & printf "%s" "$!" > "$1"; wait' sh "$child_file"
ok - daemon shutdown stops and reaps the active notifier process group
ok - inject_wedge_alarm writes the marker AND emits the active alert even with no tmux status-line (herdr backend)
ok - in-process wedge throttle prevents alert spam when the marker cannot persist
ok - fm-send exits non-zero on a confirmed swallow, zero on a clean submit
ok - fm-send exits non-zero when initial text send fails
ok - discover_supervisor_backend: override > TMUX_PANE > HERDR_ENV+HERDR_PANE_ID > tmux fallback
ok - discover_supervisor_target: override > TMUX_PANE > herdr '<session>:<pane-id>' composition > firstmate:0 fallback
ok - pane_is_busy: herdr native busy_state='busy' short-circuits without a capture fallback
ok - pane_is_busy: herdr falls back to the shared regex-over-capture reader when native busy_state is unknown
ok - pane_is_busy: herdr corroborates native idle with the shared regex-over-capture reader
ok - pane_is_busy: omitted backend arg defaults to tmux (pre-existing callers unaffected)
ok - pane_input_pending: dispatches through fm_backend_composer_state for backend=herdr
ok - inject_msg: herdr busy-guard defers before ever attempting a submit
ok - inject_msg: herdr composer-guard defers before ever attempting a submit
ok - inject_msg: herdr pane-gone check defers before any busy/composer/submit call
ok - inject_msg: dispatches busy-guard/composer-guard/submit through the herdr backend and succeeds on a confirmed empty composer
ok - inject_msg: defers on a dead-shell/unreadable composer (unknown), never typing the escalation into a shell
not ok - inject_msg typed into a dead shell: send_text_submit ran despite agent_alive=dead even though the composer read empty (starship-glyph hole)
not ok - starship-glyph inject_msg subshell failed
- Evidence: Fixed inject_msg units — pass (local file: /var/folders/4n/_bv20k5d26dgrq1t5y4l69940000gn/T/no-mistakes-evidence/01KXFBCYZTCDTN4KY16GR41M1S/fixed-fm-daemon-units.log) - Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (56m48s)

Pipeline

Updates from git push no-mistakes

... (7 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 2 infos

🔧 Fix: alarm the away captain when the supervisor pane vanishes
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-supervise-daemon.sh:958 - pane_gone_wedge_alarm calls inject_wedge_alarm with no buffer gate, so the alarm fires even when nothing is buffered - and it inherits inject_wedge_alarm's escalation-centric wording. Housekeeping's max-defer alarm (line 1004) gates on [ -s &#34;$state/.subsuper-escalations&#34; ]; this one does not. Worse, the buffer cannot grow while the pane is gone: the main-loop backoff continues past the wake handling that calls escalate_add, so a buffer that was empty when the pane vanished stays empty for the entire away window. In that case the captain's Notification Center alert reads "away-mode escalations WEDGED 300s undelivered (pane-gone)" (line 922), the durable marker's header repeats it (line 902) with an empty "Buffered items:" list beneath, and the ERROR log says "away-mode escalation undelivered 300s" - three statements about undelivered escalations that do not exist, repeated once per max-defer window for as long as the daemon lives. The number is also the pane's absence, not any escalation's age. Concrete case: crew work finishes and flushes, the buffer empties, then firstmate's pane dies; the captain wakes to a stack of alerts telling them escalations are stuck, goes looking for them, and finds none - while the actual message ("away mode is dead, restart it") is only implied by the (pane-gone) tag. Firing the alarm here is right and is what was asked for; the fix is to branch the summary, marker header, and ERROR line on cause=pane-gone (say the supervisor pane is gone and away mode can deliver nothing, and state the buffered count) rather than reusing the undelivered-escalation framing.
  • ℹ️ bin/fm-supervise-daemon.sh:1569 - When the target resolves again the main loop only resets PANE_GONE_SINCE; it never clears state/.subsuper-inject-wedged. The existing recovery path (line 1013) removes the marker only when housekeeping's max-defer escape flush succeeds, which requires the buffer to still be stuck past max_defer - so a pane-gone alarm raised on an empty or promptly-flushed buffer leaves the marker behind permanently. This is reachable whenever the target is a name rather than a unique pane id: FM_SUPERVISOR_TARGET_DEFAULT is firstmate:0, and any session:window target can be killed and recreated, as can a herdr session:pane. Concrete case: the supervisor window is closed and recreated during a 6-minute gap; the daemon alarms, then delivers normally for the rest of the night; on return, firstmate reads the stale marker as part of the afk-exit catch-up (AGENTS.md section 8) and reports a wedge that resolved hours earlier, naming a pane that is now healthy. Clearing the marker on the recovery edge when the last recorded cause was pane-gone would close it.

🔧 Fix: name supervision-down cause and retire pane-gone alarm marker
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-supervise-daemon.sh:999 - The pane-gone alarm re-fires once per max-defer window for the entire life of the daemon, on a condition that can never self-heal. inject_wedge_alarm rewrites the marker on every alarm (line 942), so its marker-age throttle (line 908) re-opens after max_defer, and WEDGE_ALARM_LAST_EPOCH (line 931) likewise re-permits wedge_alarm_notify after max_defer. The comment at line 976 calls this "one alarm per window", which is accurate — but the window repeats indefinitely. Unlike the pre-existing buffer wedge (a busy composer clears; a flush succeeds and clears the marker), a vanished pane never comes back: the supervisor target is a unique pane id, not a name. fm-afk-launch.sh captures $TMUX_PANE (e.g. %5) and passes it as FM_SUPERVISOR_TARGET, and the daemon's own discovery does the same, so FM_SUPERVISOR_TARGET_DEFAULT ('firstmate:0', the recreatable name the pane_gone_recovered comment cites) is only the unconfigured fallback. Concrete case: away mode armed at 23:00; the captain's terminal dies at 23:30 (the ordinary reason firstmate restarts elsewhere, and exactly the scenario this alarm was added for); the detached daemon alarms at 23:35 and every 300s until morning — ~96 Notification Center banners with the Basso sound, and ~96 pushes through a configured command: channel, which is the documented way to reach a captain who is away from the machine (docs/wedge-alarm.md). The captain cannot act on any of them remotely, and alert fatigue on a safety channel is the failure mode this whole alarm exists to prevent. Your instruction for this round was "one alert, respecting the existing alarm cadence, not every poll". Suggest firing the pane-gone alert once per absence EPISODE (PANE_GONE_SINCE already identifies the episode, and pane_gone_recovered already resets it), or backing it off geometrically — the durable marker and the ERROR log remain the record either way, so nothing is lost by not re-pushing.
  • ℹ️ bin/fm-supervise-daemon.sh:1016 - pane_gone_recovered decides whether to retire state/.subsuper-inject-wedged by reading the in-memory INJECT_LAST_DEFER_CAUSE, not by what actually wrote the marker — but that global is set to 'pane-gone' from two places, and only one of them writes a pane-gone marker. inject_msg's own target-exists check (line 1247) records cause=pane-gone whenever the pane dies between the main loop's check and housekeeping's flush. Concrete failure: a buffer is genuinely wedged behind a busy composer, so housekeeping raised an alarm and the on-disk marker describes 900s of undelivered escalations; the pane then blinks out for a few seconds; housekeeping's escalate_flush hits inject_msg, whose target-exists check fails and sets cause=pane-gone, but the max-defer escape's own age gate (line 1074) does not re-fire, so the composer-wedge marker is left untouched; the main loop then sees the pane missing (PANE_GONE_SINCE set) and, on the next poll, sees it back — and pane_gone_recovered deletes the composer-wedge marker, whose contract (line 1010) explicitly says such a marker "still describes a real undelivered buffer" and must be left alone. The buffer itself survives and housekeeping re-raises within one tick, so this self-corrects, but a captain returning inside that window loses the wedge report from the afk-exit catch-up. Gate the retirement on the marker's own recorded cause (it is written into the header/detail at lines 918-920) rather than on the process-global, or set a separate flag only in pane_gone_wedge_alarm's alarm branch.

🔧 Fix: latch pane-gone alarm to one alert per absence episode
2 warnings still open:

  • ⚠️ bin/fm-supervise-daemon.sh:1033 - The fire-once latch is set BEFORE the alert is attempted, and the alert has no success signal, so one transient channel failure silences the worst state for the whole away window. PANE_GONE_ALARMED=1 (line 1033) latches, then inject_wedge_alarm calls wedge_alarm_notify (line 981), which is best-effort by contract: every channel failure only logs and it always returns 0 ("Always returns 0: a channel failure can never abort inject_wedge_alarm"), and each channel runs under wedge_alarm_run_bounded's 10s timeout. Nothing re-fires afterwards - line 1025 returns early for the rest of the episode, and the pane never comes back on its own. Every OTHER wedge cause re-alarms once per max-defer window, so a failed push is effectively retried; pane-gone now gets exactly one attempt. Concrete failure: away mode armed at 23:00 with a command: ntfy/Slack channel (the documented way to reach a captain away from the machine, docs/wedge-alarm.md); the captain's terminal dies at 23:30; at 23:35 the single push hits a network blip (or osascript times out under load) and returns non-zero; wedge_alarm_via_command logs "command channel exited N" into a log nobody reads, PANE_GONE_ALARMED is already 1, and the captain gets no alert at all for a daemon that is now typing at nothing - the exact away-window silence this alarm was added to end. The marker and ERROR log survive, but both are synchronous surfaces the captain only sees once they are back, which is what round 5 established is not good enough. Suggest latching on a DELIVERED alert rather than an attempted one (have wedge_alarm_notify report whether any channel succeeded, leaving the latch clear when none did), or allowing a small bounded number of re-pushes per episode - either satisfies "one alert per episode, not one per window forever" without making a single dropped packet permanent.
  • ⚠️ bin/fm-supervise-daemon.sh:924 - The latched flag skips the marker-age throttle at line 924 but NOT the unconditional marker write at lines 959-965, so a pane-gone alarm overwrites whatever marker is already on disk - including a composer-wedge marker housekeeping wrote moments earlier - replacing its cause: line with pane-gone. pane_gone_recovered's on-disk cause check (line 1061) then matches and deletes it (line 1062). That is precisely the deletion the comment at lines 953-958 promises is impossible ("a marker describing a real undelivered buffer is never mistaken for a pane-gone alarm and deleted with it"): keying on the marker instead of the global closes the mid-tick-blink path, but not the path where the pane-gone alarm actually fires. Concrete sequence, all reachable: escalations wedge behind a busy composer past max_defer, so housekeeping (line 1104) writes a marker with cause: composer-not-empty describing 900s undelivered; the captain's terminal then dies; after a full max-defer window of absence pane_gone_wedge_alarm fires latched and rewrites that marker as cause: pane-gone; the target is a session:window name (FM_SUPERVISOR_TARGET_DEFAULT is firstmate:0, and pane_gone_recovered's own comment cites the killed-and-recreated window), so it resolves again and the composer-wedge marker is removed. The new test test_pane_gone_recovery_keeps_a_real_buffer_wedge_marker only passes because it hand-sets PANE_GONE_ALARMED=1 without ever calling pane_gone_wedge_alarm, so it pins a state the production path cannot reach. Impact is bounded - the buffer survives and housekeeping re-raises on the next tick - but for that window firstmate's afk-exit catch-up has no wedge report, and the re-raise is silent (the latched alarm set WEDGE_ALARM_LAST_EPOCH, so line 947 suppresses both the ERROR log and the active alert for up to max_defer). Suggest not rewriting an existing marker whose recorded cause is something else (leave it, or give the pane-gone alarm its own marker file) so recovery can only ever retire what it wrote.

🔧 Fix: give the pane-gone alarm its own marker file
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-supervise-daemon.sh:941 - inject_wedge_alarm selects its marker file and its throttle from $kind (line 928) but selects the captain-facing summary, the Notification Center banner title (WEDGE_ALARM_TITLE), and the ERROR log from $cause (line 941). Those two disagree on a reachable path, and the result is a false 'away mode is dead' alert to an away captain. Concrete failure: the main loop's fm_backend_target_exists check passes, the watcher runs for a poll interval, and the supervisor pane dies in that window; housekeeping's max-defer escape (line 1121) then calls escalate_flush -> inject_msg, whose own target-exists check fails and records cause=pane-gone (line 1247 via _inject_defer). escalate_flush returns non-zero, so housekeeping calls inject_wedge_alarm "$state" "$oldest" with the DEFAULT kind=buffer. Inside, the cause=pane-gone branch fires: WEDGE_ALARM_TITLE becomes 'firstmate: away-mode supervision DOWN', the active alert reads 'away-mode supervision DOWN Ns (pane-gone): supervisor pane gone, nothing can be delivered', and that text is written into .subsuper-inject-wedged (the BUFFER marker). This entirely bypasses the debounce pane_gone_wedge_alarm was built around - it requires a full max-defer window of CONTINUOUS absence before declaring supervision down, precisely because 'a pane that blinks out for one poll is a backend hiccup' (line 1003) - so one flaky tmux/herdr existence probe at flush time now produces the loudest verdict in the design. Per the new .agents/skills/afk/SKILL.md guidance ('The one verdict that is never an accepted degradation is a vanished pane... report it and fix it rather than continuing away mode'), firstmate would tear down a healthy away session on that alert. This is the same misattribution class the wedge-alarm and canary work spent four rounds eliminating. Fix: branch the summary/header/detail/err/title on $kind instead of $cause (line 941 -> if [ &#34;$kind&#34; = pane-gone ]). Nothing is lost: with kind=buffer and cause=pane-gone the buffer wording still names the tag - 'away-mode escalations WEDGED Ns undelivered (pane-gone)' - which is accurate, because the buffer IS stuck and IS still being retried, while the debounced, latched supervision-DOWN declaration stays where it belongs.
  • ℹ️ bin/fm-afk-launch.sh:353 - FM_AFK_LAUNCH_ARTIFACTS (line 64) was introduced this commit with the comment 'One list, so a new artifact cannot be added to the daemon and forgotten by the rollback that has to put it back' - but fm_afk_launch_restore_backup's rm -f at lines 353-358 still hardcodes the same four artifact names immediately above the loop that iterates the array. The drift the array exists to prevent is therefore only half-closed: a future fifth artifact added to FM_AFK_LAUNCH_ARTIFACTS but not to that rm -f would survive a rolled-back arm (the loop only restores what was in the backup, so an artifact the failed arm created fresh is never cleared). Iterate the array in the rm -f too, keeping .afk as the one explicit extra. (bin/fm-afk-start.sh's fm_afk_clear_stale_artifacts is a third hardcoded copy of the same list, in a different file - worth noting, but the launcher's internal duplication is the one that contradicts its own comment.)

🔧 Fix: key wedge alarm wording to alarm kind, not last cause
2 issues (1 warning, 1 info) still open:

  • ⚠️ docs/wedge-alarm.md:12 - The alarm-semantics section (lines 12-16) still attributes the supervision-DOWN wording, the fire-once latch, and the empty-buffer firing to the CAUSE pane-gone ("pane-gone is the one cause whose captain-facing wording is not about the buffer", "It also fires on an EMPTY buffer", "Every channel therefore says supervision is down"). This commit changed all three to key on the alarm KIND instead. On the reachable pane-died-mid-flush path - the main loop's fm_backend_target_exists check passes, the pane dies during housekeeping's max-defer escalate_flush, and inject_msg records cause=pane-gone (bin/fm-supervise-daemon.sh:1298) while the alarm housekeeping raises is kind=buffer - the channels say "away-mode escalations WEDGED 900s undelivered (pane-gone)", the alarm IS throttled once per max-defer window, and it IS gated on a non-empty buffer. That is the exact opposite of what these lines promise, and it is the behavior the new test test_buffer_wedge_alarm_never_declares_supervision_down pins. Line 109 of the same file, added in this commit, states the correct rule ("Which alarm fired is what the captain is told, never the last recorded cause"), so the file now contradicts itself: a maintainer reading top-down would reintroduce the false 'away mode is dead' alert this commit fixed. Line 6 carries a smaller version of the slip - it claims pane-gone comes only from the daemon's main loop rather than an inject attempt, but inject_msg's own target-exists check records it too. Rewrite lines 6 and 12-16 so the supervision-DOWN wording, the latch, and the empty-buffer firing are described as properties of the pane-gone ALARM (the debounced, latched one raised from the backoff path), with pane-gone named only as a cause tag that can also appear on a buffer wedge.
  • ℹ️ bin/fm-afk-start.sh:63 - FM_AFK_LAUNCH_ARTIFACTS (bin/fm-afk-launch.sh:64) was introduced this commit with the comment "One list, so a new artifact cannot be added to the daemon and forgotten by the rollback that has to put it back", and the rollback rm now iterates it. But fm_afk_clear_stale_artifacts - the fresh-arm clear that drops the PRIOR away session's delivery artifacts - still hardcodes the same four names, and it lives in bin/fm-afk-start.sh, which the launcher SOURCES, so it cannot reference the launcher's array. The two agree today (both gained .subsuper-pane-gone in this PR), but a fifth artifact added to FM_AFK_LAUNCH_ARTIFACTS alone would be backed up and rolled back correctly while surviving a fresh arm - resurfacing a prior session's buffered escalation or wedge marker under a new away session, which is the failure fm_afk_clear_stale_artifacts exists to prevent. Defining the list once in bin/fm-afk-start.sh (where the clear lives) and having bin/fm-afk-launch.sh iterate that same array would close the drift and shrink the footprint rather than grow it. Flagged as an aside in round 9 and deliberately scoped out of that fix; noting the residual, not re-litigating.

🔧 Fix: key wedge-alarm doc to alarm kind and unify afk artifact list
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-supervise-daemon.sh:200 - The WEDGE_ALARM_TITLE declaration comment still states the captain-facing banner title "follows the CAUSE" ("so it follows the CAUSE: a vanished supervisor pane means away-mode supervision is down"), but commit 1496b85 changed inject_wedge_alarm to select the summary, WEDGE_ALARM_TITLE, and the ERROR log from $kind instead — the same discriminator that picks the marker file and the throttle — specifically so a cause=pane-gone recorded by inject_msg during housekeeping's max-defer flush cannot render the supervision-DOWN verdict from what is actually a buffer wedge. docs/wedge-alarm.md was rewritten for this (lines 201-209: "WHICH one fired - its kind - is what every captain-facing surface says") and tests/fm-daemon.test.sh's test_buffer_wedge_alarm_never_declares_supervision_down pins it, so the code comment is now the only place in the repo asserting the superseded rule. inject_wedge_alarm's own header (line ~892) carries a milder version: it opens with "the ERROR log, the durable marker, and the active alert all report the recorded cause" before the next paragraph corrects it to the kind. Concrete failure: a maintainer reads the declaration comment, re-keys the branch at line 946 on $cause, and the reachable pane-died-mid-flush path (main-loop target-exists passes, pane dies during escalate_flush -> inject_msg records pane-gone at line 1298, housekeeping raises kind=buffer at line 1137) again titles an away captain's Notification Center banner "firstmate: away-mode supervision DOWN" for a buffer that is still stuck behind a pane the daemon keeps retrying — which per .agents/skills/afk/SKILL.md means tearing down a healthy away session. Rewrite lines 198-204 (and the opening sentence of the inject_wedge_alarm header) to say the title follows the alarm KIND, with the cause carried only as a diagnostic tag.
  • ℹ️ bin/fm-bootstrap.sh:609 - Noting the boundary of the new session-start diagnostic, not requesting a change. afk_injection_check gates on state/.afk, then fm_afk_canary_resolve falls back to discover_supervisor_target whenever no LIVE daemon owns the lock — so when away mode is armed but the daemon has died, bootstrap probes its own pane, finds firstmate obviously alive, and prints nothing. That is the right answer for what the canary measures (can the resolved supervisor be injected into), and AGENTS.md section 5 step 8 already covers the dead-daemon case by having /afk ensure the daemon is running on recovery, so nothing is actually missed. But docs/herdr-backend.md's "a cold or restarted session that re-enters afk from the flag alone still learns injection is off" reads broader than the check guarantees: the one away-mode failure a restarted session cannot learn from this line is the daemon simply not being there.

🔧 Fix: key wedge-alarm title comment to alarm kind, not cause
2 infos still open:

  • ℹ️ bin/fm-afk-start.sh:79 - fm_afk_clear_stale_artifacts changed from a single best-effort rm -f (which returns 0 essentially always) to a loop that returns 1 if ANY rm fails. In bin/fm-afk-launch.sh both call sites already test the return value and log "failed to clear stale away-mode artifacts", so there the new failure signal is an improvement. But fm_afk_start_main calls it bare inside an if body (bin/fm-afk-start.sh:159) and that file runs under set -eu with no set +e, so a non-zero return is not a caught error — errexit fires and the daemon start aborts before exec &#34;$FM_AFK_DAEMON&#34;, with no message at all (the rm's stderr goes to /dev/null and errexit is silent). Failure scenario: a direct, non-prepared bin/fm-afk-start.sh on a state dir holding an undeletable artifact (root-owned leftover from a prior run under a different user, or a read-only mount) previously logged nothing and started away mode anyway; now it exits 1 with no output and away mode silently does not arm. Reachability is narrow — the /afk skill always goes through the launcher with FM_AFK_STATE_PREPARED=1, so only a manual/legacy direct invocation takes this branch — but the silence is the problem, not the abort. Either echo the failure before returning non-zero, or guard the call site the way the launcher does.
  • ℹ️ bin/fm-supervise-daemon.sh:1081 - When the supervisor target resolves again, pane_gone_recovered deletes state/.subsuper-pane-gone and logs the recovery only to the daemon log. But the captain has, by then, already received the pane-gone active alert — a Notification Center banner titled "firstmate: away-mode supervision DOWN", and a push through any configured command: channel, telling them to stop away mode and arm it again. Nothing retracts it, and the marker (the one surface the afk-exit catch-up actually reads) is gone. Per .agents/skills/afk/SKILL.md:144, an absent .subsuper-pane-gone marker means "no live absence", so firstmate's catch-up reports all-clear for an alarm the captain demonstrably received. Failure scenario: the supervisor window is a session:window name (the FM_SUPERVISOR_TARGET_DEFAULT firstmate:0 fallback, or an explicit named target — the exact case pane_gone_recovered was written for), it is killed and recreated during a >300s gap overnight; the captain gets a 2am "supervision DOWN" push, comes back to a clean catch-up, and neither they nor firstmate can explain the alert without grepping state/.supervise-daemon.log, which no catch-up instruction points at. A one-line resolved note left behind on the recovery edge, or having the /afk catch-up read the daemon log's ERROR/recovery pair, would close it. Flagging for the record rather than requesting a change: this is the same best-effort-alert reach you already scoped out (docs/wedge-alarm.md "Known follow-ups"), and it belongs with the bounded-re-fire item there.
🔧 **Test** - 1 issue found → auto-fixed ✅
  • 🚨 tests failed with exit code 1
  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"

🔧 Fix: provide mkdir to x-mode fakebin, silence node ExperimentalWarning
✅ Re-checked - no issues remain.

  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"
  • bash tests/*.test.sh full suite: 68 files, 0 failures, exit 0
  • bash tests/fm-afk-inject-e2e.test.sh — real-shell Scenario D (genuine login shell with a bare starship ❯ prompt) passes: escalation defers, nothing typed into the shell
  • bash tests/fm-daemon.test.sh — the three safety-critical inject_msg units pass: defers on empty-composer dead shell (agent_alive=dead), defers on unknown liveness (fail closed), injects on alive+empty (positive path)
  • Pre-fix regression proof: neutered the liveness gate to liveness=alive (simulating gate-on-composer-only) and re-ran both files — Scenario D failed with 'SHELL INJECTION' (digest typed into and run by the live shell) and the starship-glyph unit failed with 'send_text_submit ran despite agent_alive=dead'; restored the file (git checkout) and confirmed clean tree + tests green again
⚠️ **Document** - 1 info
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

* The away-mode daemon gated injection solely on composer-emptiness, but
  the default starship/pure/spaceship shell prompt glyph U+276F (❯) is
  byte-identical to the glyph claude draws for its own empty composer.
* A supervisor pane whose agent had exited to a login shell therefore read
  as an empty agent composer, so the escalation digest (plus Enter) was
  typed into and executed by the live shell - arbitrary command execution
  at exactly the moment nobody is watching.
* inject_msg now confirms a live agent PROCESS via the existing
  fm_backend_agent_alive probe after the composer-empty check, and fails
  closed: dead (a bare shell) or unknown liveness defers instead of
  injecting. A deferred escalation stays buffered for the next cycle, the
  max-defer wedge alarm, or the afk-exit catch-up flush.
* Regression coverage: three deterministic inject_msg units and a real-shell
  e2e scenario (a bare starship ❯ prompt driven end-to-end), each failing
  against the pre-fix daemon.
* fm-send shares the same content-blindness but is deliberate and
  firstmate-initiated; documented for a separate dead-only guard.
@kunchenguid

kunchenguid commented Jul 15, 2026 •

Copy link
Copy Markdown
Owner

Thanks for the PR! This branch currently has a merge conflict with the base branch.

When you get a chance, please rebase onto (or merge) the latest base branch, resolve the conflict, and push. After that, checks will re-run and the PR will get looked at again.

Noted for firstmate#577 at b4ce454e.

@kunchenguid

Copy link
Copy Markdown
Owner

Automated reminder: this PR still looks blocked on a rebase or merge conflict fix.

If you are still interested, please rebase onto the current base branch, resolve the conflict, and push.

If I do not hear back, I may close this as inactive.

@kunchenguid

Copy link
Copy Markdown
Owner

I am closing this because it has been waiting on a rebase or merge-conflict fix since 2026-07-15, and I have not seen a comment or push since then.

If you still want to keep working on this, please reopen it or open a new PR and mention this one.

Happy to take another look when there is an update.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants