fix(afk): gate away-mode injection on a confirmed live agent - #577
zachlandes wants to merge 14 commits into
Conversation
* The away-mode daemon gated injection solely on composer-emptiness, but the default starship/pure/spaceship shell prompt glyph U+276F (❯) is byte-identical to the glyph claude draws for its own empty composer. * A supervisor pane whose agent had exited to a login shell therefore read as an empty agent composer, so the escalation digest (plus Enter) was typed into and executed by the live shell - arbitrary command execution at exactly the moment nobody is watching. * inject_msg now confirms a live agent PROCESS via the existing fm_backend_agent_alive probe after the composer-empty check, and fails closed: dead (a bare shell) or unknown liveness defers instead of injecting. A deferred escalation stays buffered for the next cycle, the max-defer wedge alarm, or the afk-exit catch-up flush. * Regression coverage: three deterministic inject_msg units and a real-shell e2e scenario (a bare starship ❯ prompt driven end-to-end), each failing against the pre-fix daemon. * fm-send shares the same content-blindness but is deliberate and firstmate-initiated; documented for a separate dead-only guard.
|
Thanks for the PR! This branch currently has a merge conflict with the base branch. When you get a chance, please rebase onto (or merge) the latest base branch, resolve the conflict, and push. After that, checks will re-run and the PR will get looked at again. Noted for firstmate#577 at |
|
Automated reminder: this PR still looks blocked on a rebase or merge conflict fix. If you are still interested, please rebase onto the current base branch, resolve the conflict, and push. If I do not hear back, I may close this as inactive. |
|
I am closing this because it has been waiting on a rebase or merge-conflict fix since 2026-07-15, and I have not seen a comment or push since then. If you still want to keep working on this, please reopen it or open a new PR and mention this one. Happy to take another look when there is an update. |
Intent
Fix a verified safety hole: the away-mode supervise daemon could type an escalation digest (plus Enter) into a live shell, executing arbitrary commands when nobody is watching. Root cause: inject_msg (bin/fm-supervise-daemon.sh) gated injection ONLY on composer-emptiness, but the default starship/pure/spaceship shell prompt glyph U+276F (the heavy right-angle quotation mark) is byte-identical to the glyph claude draws for its own empty composer, so a supervisor pane whose agent had exited to a login shell read as an empty agent composer and was injected into. Fix: after the composer-empty check, inject_msg now also calls the pre-existing fm_backend_agent_alive probe (bin/fm-backend.sh; already used by the session-start secondmate-liveness sweep) and injects ONLY on a confident 'alive'. Deliberate decisions a diff-only reviewer would not know: (1) This FAILS CLOSED on purpose - both 'dead' (a bare shell) and 'unknown' (an unreadable pane, or a harness whose liveness is unverifiable for the backend, e.g. pi's generic node interpreter) defer, on the stated principle that a missed escalation is an inconvenience but a command typed into the captain's shell is not; deferred escalations are not lost (buffer + max-defer wedge alarm + afk-exit catch-up flush). A pi-based supervisor in away mode therefore stops receiving injected escalations - this is an accepted, documented consequence, not a regression to fix here. (2) I intentionally did NOT remove the glyph from any character class - that is a band-aid because an empty shell prompt and an empty agent composer are indistinguishable by content; the guard asks what process is running instead. (3) fm-send.sh shares the same content-blindness but is deliberate/firstmate-initiated (not autonomous) and a strict alive-gate there would break steering for every unknown-liveness backend (pi, zellij, orca, cmux); I analyzed and documented it with a recommended separate dead-only guard but deliberately did NOT change fm-send in this PR to keep the safety fix focused. Tests: three deterministic inject_msg units in tests/fm-daemon.test.sh plus a real-shell end-to-end Scenario D in tests/fm-afk-inject-e2e.test.sh, each of which fails against the pre-fix daemon (verified by neutering the gate: the suite exits non-zero with an explicit 'typed into a dead shell' / 'SHELL INJECTION' failure). Test-helper changes in tests/wake-helpers.sh add a pane_current_command reply to the fake tmux so the fake supervisor panes (which stand in for a live agent) report as alive. Docs: amended the one-owner composer-emptiness safety policy in docs/herdr-backend.md (new dated subsection) and corrected the now-inaccurate claim in docs/architecture.md that a bare dead-shell prompt cannot receive an escalation. This is a generic bug affecting every firstmate user with a default shell prompt, intended to go upstream.
What Changed
inject_msginbin/fm-supervise-daemon.shnow probesfm_backend_agent_aliveafter the composer-empty check and injects only on a confidentalive. An empty starship/pure prompt glyph (U+276F) is byte-identical to claude's empty-composer glyph, so a supervisor pane whose agent had exited to a login shell previously read as an empty composer and got the escalation digest typed in and submitted as shell commands. The guard fails closed: bothdeadandunknown(an unreadable pane, or a harness whose liveness the backend cannot attribute, e.g. pi's genericnode) defer instead of injecting, so a pi-backed supervisor stops receiving injections in away mode - deferred escalations still survive via the buffer, the wedge alarm, and the afk-exit catch-up flush.bin/fm-afk-canary-lib.shderives and words that verdict in one place, and surfaces it on both the arm path (bin/fm-afk-launch.sh) and session start (bin/fm-bootstrap.sh, as anAFK_INJECTION_DISABLED:diagnostic). It probes the running daemon's own recorded target rather than the current pane, so a firstmate restarted into a new pane learns the daemon is still injecting elsewhere, and reports a vanished target instead of staying quiet..subsuper-pane-gonemarker, one alert per absence episode) that declares supervision DOWN when the supervisor pane vanishes. Captain-facing wording keys off the alarm kind, not the last recorded cause, so a buffer wedge can never render as "away mode is dead". Covered by three deterministicinject_msgunits intests/fm-daemon.test.shand a real-shell Scenario D intests/fm-afk-inject-e2e.test.sh, each of which fails against the pre-fix daemon;bin/fm-send.shshares the same content-blindness and is documented indocs/herdr-backend.mdbut deliberately left unchanged here.Risk Assessment
Testing
Baseline test command had already passed; I re-ran the full suite (68 files, 0 failures) and then produced product-level before/after evidence for the safety fix. Using a real login shell displaying the default starship ❯ glyph, the end-to-end Scenario D shows the fixed daemon refuses to inject and defers the escalation, while a neutered (pre-fix) daemon typed the digest into and executed it in the live shell — the test caught this with an explicit SHELL INJECTION failure. The deterministic inject_msg units confirm the same at the function level, including fail-closed-on-unknown and the alive+empty positive path. I restored the daemon file and verified a clean working tree afterward. This is a CLI/daemon behavior change with no rendered UI surface, so evidence is CLI transcripts rather than screenshots.
Evidence: Evidence summary: before/after shell-injection guard
### A. Real-shell end-to-end (Scenario D) PRE-FIX (guard neutered): not ok - Scenario D: the escalation digest was typed into the live shell (found 'Supervisor escalate' in the pane) — SHELL INJECTION [e2e exit 1] WITH FIX: ok - Scenario D: a dead shell showing a bare starship ❯ prompt is refused — the escalation defers, nothing is typed into the shell [e2e exit 0] ### B. Deterministic inject_msg units PRE-FIX (neutered): not ok - inject_msg typed into a dead shell: send_text_submit ran despite agent_alive=dead (starship-glyph hole) WITH FIX: ok - defers on empty-composer dead shell (agent_alive=dead) ok - defers when agent liveness is unknown (fail closed) ok - injects only when agent alive and composer empty (positive path) ### C. Full suite: 68 test files, 0 failures, SUITE_EXIT=0Evidence: Full test suite log (68 files, 0 failures)
Evidence: Pre-fix (neutered) e2e — SHELL INJECTION failure
Evidence: Fixed e2e — Scenario D refuses injection
Evidence: Pre-fix (neutered) inject_msg units — failures
/var/folders/4n/_bv20k5d26dgrq1t5y4l69940000gn/T/no-mistakes-evidence/01KXFBCYZTCDTN4KY16GR41M1S/fixed-fm-daemon-units.log) - Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (56m48s)Pipeline
Updates from git push no-mistakes
... (7 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 Fix: alarm the away captain when the supervisor pane vanishes
2 issues (1 warning, 1 info) still open:
bin/fm-supervise-daemon.sh:958- pane_gone_wedge_alarm calls inject_wedge_alarm with no buffer gate, so the alarm fires even when nothing is buffered - and it inherits inject_wedge_alarm's escalation-centric wording. Housekeeping's max-defer alarm (line 1004) gates on[ -s "$state/.subsuper-escalations" ]; this one does not. Worse, the buffer cannot grow while the pane is gone: the main-loop backoffcontinues past the wake handling that calls escalate_add, so a buffer that was empty when the pane vanished stays empty for the entire away window. In that case the captain's Notification Center alert reads "away-mode escalations WEDGED 300s undelivered (pane-gone)" (line 922), the durable marker's header repeats it (line 902) with an empty "Buffered items:" list beneath, and the ERROR log says "away-mode escalation undelivered 300s" - three statements about undelivered escalations that do not exist, repeated once per max-defer window for as long as the daemon lives. The number is also the pane's absence, not any escalation's age. Concrete case: crew work finishes and flushes, the buffer empties, then firstmate's pane dies; the captain wakes to a stack of alerts telling them escalations are stuck, goes looking for them, and finds none - while the actual message ("away mode is dead, restart it") is only implied by the(pane-gone)tag. Firing the alarm here is right and is what was asked for; the fix is to branch the summary, marker header, and ERROR line on cause=pane-gone (say the supervisor pane is gone and away mode can deliver nothing, and state the buffered count) rather than reusing the undelivered-escalation framing.bin/fm-supervise-daemon.sh:1569- When the target resolves again the main loop only resets PANE_GONE_SINCE; it never clears state/.subsuper-inject-wedged. The existing recovery path (line 1013) removes the marker only when housekeeping's max-defer escape flush succeeds, which requires the buffer to still be stuck past max_defer - so a pane-gone alarm raised on an empty or promptly-flushed buffer leaves the marker behind permanently. This is reachable whenever the target is a name rather than a unique pane id: FM_SUPERVISOR_TARGET_DEFAULT isfirstmate:0, and any session:window target can be killed and recreated, as can a herdr session:pane. Concrete case: the supervisor window is closed and recreated during a 6-minute gap; the daemon alarms, then delivers normally for the rest of the night; on return, firstmate reads the stale marker as part of the afk-exit catch-up (AGENTS.md section 8) and reports a wedge that resolved hours earlier, naming a pane that is now healthy. Clearing the marker on the recovery edge when the last recorded cause was pane-gone would close it.🔧 Fix: name supervision-down cause and retire pane-gone alarm marker
2 issues (1 warning, 1 info) still open:
bin/fm-supervise-daemon.sh:999- The pane-gone alarm re-fires once per max-defer window for the entire life of the daemon, on a condition that can never self-heal. inject_wedge_alarm rewrites the marker on every alarm (line 942), so its marker-age throttle (line 908) re-opens after max_defer, and WEDGE_ALARM_LAST_EPOCH (line 931) likewise re-permits wedge_alarm_notify after max_defer. The comment at line 976 calls this "one alarm per window", which is accurate — but the window repeats indefinitely. Unlike the pre-existing buffer wedge (a busy composer clears; a flush succeeds and clears the marker), a vanished pane never comes back: the supervisor target is a unique pane id, not a name. fm-afk-launch.sh captures $TMUX_PANE (e.g. %5) and passes it as FM_SUPERVISOR_TARGET, and the daemon's own discovery does the same, so FM_SUPERVISOR_TARGET_DEFAULT ('firstmate:0', the recreatable name the pane_gone_recovered comment cites) is only the unconfigured fallback. Concrete case: away mode armed at 23:00; the captain's terminal dies at 23:30 (the ordinary reason firstmate restarts elsewhere, and exactly the scenario this alarm was added for); the detached daemon alarms at 23:35 and every 300s until morning — ~96 Notification Center banners with the Basso sound, and ~96 pushes through a configuredcommand:channel, which is the documented way to reach a captain who is away from the machine (docs/wedge-alarm.md). The captain cannot act on any of them remotely, and alert fatigue on a safety channel is the failure mode this whole alarm exists to prevent. Your instruction for this round was "one alert, respecting the existing alarm cadence, not every poll". Suggest firing the pane-gone alert once per absence EPISODE (PANE_GONE_SINCE already identifies the episode, and pane_gone_recovered already resets it), or backing it off geometrically — the durable marker and the ERROR log remain the record either way, so nothing is lost by not re-pushing.bin/fm-supervise-daemon.sh:1016- pane_gone_recovered decides whether to retire state/.subsuper-inject-wedged by reading the in-memory INJECT_LAST_DEFER_CAUSE, not by what actually wrote the marker — but that global is set to 'pane-gone' from two places, and only one of them writes a pane-gone marker. inject_msg's own target-exists check (line 1247) records cause=pane-gone whenever the pane dies between the main loop's check and housekeeping's flush. Concrete failure: a buffer is genuinely wedged behind a busy composer, so housekeeping raised an alarm and the on-disk marker describes 900s of undelivered escalations; the pane then blinks out for a few seconds; housekeeping's escalate_flush hits inject_msg, whose target-exists check fails and sets cause=pane-gone, but the max-defer escape's own age gate (line 1074) does not re-fire, so the composer-wedge marker is left untouched; the main loop then sees the pane missing (PANE_GONE_SINCE set) and, on the next poll, sees it back — and pane_gone_recovered deletes the composer-wedge marker, whose contract (line 1010) explicitly says such a marker "still describes a real undelivered buffer" and must be left alone. The buffer itself survives and housekeeping re-raises within one tick, so this self-corrects, but a captain returning inside that window loses the wedge report from the afk-exit catch-up. Gate the retirement on the marker's own recorded cause (it is written into the header/detail at lines 918-920) rather than on the process-global, or set a separate flag only in pane_gone_wedge_alarm's alarm branch.🔧 Fix: latch pane-gone alarm to one alert per absence episode
2 warnings still open:
bin/fm-supervise-daemon.sh:1033- The fire-once latch is set BEFORE the alert is attempted, and the alert has no success signal, so one transient channel failure silences the worst state for the whole away window. PANE_GONE_ALARMED=1 (line 1033) latches, then inject_wedge_alarm calls wedge_alarm_notify (line 981), which is best-effort by contract: every channel failure only logs and it always returns 0 ("Always returns 0: a channel failure can never abort inject_wedge_alarm"), and each channel runs under wedge_alarm_run_bounded's 10s timeout. Nothing re-fires afterwards - line 1025 returns early for the rest of the episode, and the pane never comes back on its own. Every OTHER wedge cause re-alarms once per max-defer window, so a failed push is effectively retried; pane-gone now gets exactly one attempt. Concrete failure: away mode armed at 23:00 with acommand:ntfy/Slack channel (the documented way to reach a captain away from the machine, docs/wedge-alarm.md); the captain's terminal dies at 23:30; at 23:35 the single push hits a network blip (or osascript times out under load) and returns non-zero; wedge_alarm_via_command logs "command channel exited N" into a log nobody reads, PANE_GONE_ALARMED is already 1, and the captain gets no alert at all for a daemon that is now typing at nothing - the exact away-window silence this alarm was added to end. The marker and ERROR log survive, but both are synchronous surfaces the captain only sees once they are back, which is what round 5 established is not good enough. Suggest latching on a DELIVERED alert rather than an attempted one (have wedge_alarm_notify report whether any channel succeeded, leaving the latch clear when none did), or allowing a small bounded number of re-pushes per episode - either satisfies "one alert per episode, not one per window forever" without making a single dropped packet permanent.bin/fm-supervise-daemon.sh:924- Thelatchedflag skips the marker-age throttle at line 924 but NOT the unconditional marker write at lines 959-965, so a pane-gone alarm overwrites whatever marker is already on disk - including a composer-wedge marker housekeeping wrote moments earlier - replacing itscause:line withpane-gone. pane_gone_recovered's on-disk cause check (line 1061) then matches and deletes it (line 1062). That is precisely the deletion the comment at lines 953-958 promises is impossible ("a marker describing a real undelivered buffer is never mistaken for a pane-gone alarm and deleted with it"): keying on the marker instead of the global closes the mid-tick-blink path, but not the path where the pane-gone alarm actually fires. Concrete sequence, all reachable: escalations wedge behind a busy composer past max_defer, so housekeeping (line 1104) writes a marker withcause: composer-not-emptydescribing 900s undelivered; the captain's terminal then dies; after a full max-defer window of absence pane_gone_wedge_alarm fires latched and rewrites that marker ascause: pane-gone; the target is a session:window name (FM_SUPERVISOR_TARGET_DEFAULT isfirstmate:0, and pane_gone_recovered's own comment cites the killed-and-recreated window), so it resolves again and the composer-wedge marker is removed. The new test test_pane_gone_recovery_keeps_a_real_buffer_wedge_marker only passes because it hand-sets PANE_GONE_ALARMED=1 without ever calling pane_gone_wedge_alarm, so it pins a state the production path cannot reach. Impact is bounded - the buffer survives and housekeeping re-raises on the next tick - but for that window firstmate's afk-exit catch-up has no wedge report, and the re-raise is silent (the latched alarm set WEDGE_ALARM_LAST_EPOCH, so line 947 suppresses both the ERROR log and the active alert for up to max_defer). Suggest not rewriting an existing marker whose recorded cause is something else (leave it, or give the pane-gone alarm its own marker file) so recovery can only ever retire what it wrote.🔧 Fix: give the pane-gone alarm its own marker file
2 issues (1 warning, 1 info) still open:
bin/fm-supervise-daemon.sh:941- inject_wedge_alarm selects its marker file and its throttle from $kind (line 928) but selects the captain-facing summary, the Notification Center banner title (WEDGE_ALARM_TITLE), and the ERROR log from $cause (line 941). Those two disagree on a reachable path, and the result is a false 'away mode is dead' alert to an away captain. Concrete failure: the main loop's fm_backend_target_exists check passes, the watcher runs for a poll interval, and the supervisor pane dies in that window; housekeeping's max-defer escape (line 1121) then calls escalate_flush -> inject_msg, whose own target-exists check fails and records cause=pane-gone (line 1247 via _inject_defer). escalate_flush returns non-zero, so housekeeping calls inject_wedge_alarm "$state" "$oldest" with the DEFAULT kind=buffer. Inside, the cause=pane-gone branch fires: WEDGE_ALARM_TITLE becomes 'firstmate: away-mode supervision DOWN', the active alert reads 'away-mode supervision DOWN Ns (pane-gone): supervisor pane gone, nothing can be delivered', and that text is written into .subsuper-inject-wedged (the BUFFER marker). This entirely bypasses the debounce pane_gone_wedge_alarm was built around - it requires a full max-defer window of CONTINUOUS absence before declaring supervision down, precisely because 'a pane that blinks out for one poll is a backend hiccup' (line 1003) - so one flaky tmux/herdr existence probe at flush time now produces the loudest verdict in the design. Per the new .agents/skills/afk/SKILL.md guidance ('The one verdict that is never an accepted degradation is a vanished pane... report it and fix it rather than continuing away mode'), firstmate would tear down a healthy away session on that alert. This is the same misattribution class the wedge-alarm and canary work spent four rounds eliminating. Fix: branch the summary/header/detail/err/title on $kind instead of $cause (line 941 ->if [ "$kind" = pane-gone ]). Nothing is lost: with kind=buffer and cause=pane-gone the buffer wording still names the tag - 'away-mode escalations WEDGED Ns undelivered (pane-gone)' - which is accurate, because the buffer IS stuck and IS still being retried, while the debounced, latched supervision-DOWN declaration stays where it belongs.bin/fm-afk-launch.sh:353- FM_AFK_LAUNCH_ARTIFACTS (line 64) was introduced this commit with the comment 'One list, so a new artifact cannot be added to the daemon and forgotten by the rollback that has to put it back' - but fm_afk_launch_restore_backup'srm -fat lines 353-358 still hardcodes the same four artifact names immediately above the loop that iterates the array. The drift the array exists to prevent is therefore only half-closed: a future fifth artifact added to FM_AFK_LAUNCH_ARTIFACTS but not to that rm -f would survive a rolled-back arm (the loop only restores what was in the backup, so an artifact the failed arm created fresh is never cleared). Iterate the array in the rm -f too, keeping .afk as the one explicit extra. (bin/fm-afk-start.sh's fm_afk_clear_stale_artifacts is a third hardcoded copy of the same list, in a different file - worth noting, but the launcher's internal duplication is the one that contradicts its own comment.)🔧 Fix: key wedge alarm wording to alarm kind, not last cause
2 issues (1 warning, 1 info) still open:
docs/wedge-alarm.md:12- The alarm-semantics section (lines 12-16) still attributes the supervision-DOWN wording, the fire-once latch, and the empty-buffer firing to the CAUSEpane-gone("pane-goneis the one cause whose captain-facing wording is not about the buffer", "It also fires on an EMPTY buffer", "Every channel therefore says supervision is down"). This commit changed all three to key on the alarm KIND instead. On the reachable pane-died-mid-flush path - the main loop's fm_backend_target_exists check passes, the pane dies during housekeeping's max-defer escalate_flush, and inject_msg records cause=pane-gone (bin/fm-supervise-daemon.sh:1298) while the alarm housekeeping raises is kind=buffer - the channels say "away-mode escalations WEDGED 900s undelivered (pane-gone)", the alarm IS throttled once per max-defer window, and it IS gated on a non-empty buffer. That is the exact opposite of what these lines promise, and it is the behavior the new test test_buffer_wedge_alarm_never_declares_supervision_down pins. Line 109 of the same file, added in this commit, states the correct rule ("Which alarm fired is what the captain is told, never the last recorded cause"), so the file now contradicts itself: a maintainer reading top-down would reintroduce the false 'away mode is dead' alert this commit fixed. Line 6 carries a smaller version of the slip - it claimspane-gonecomes only from the daemon's main loop rather than an inject attempt, but inject_msg's own target-exists check records it too. Rewrite lines 6 and 12-16 so the supervision-DOWN wording, the latch, and the empty-buffer firing are described as properties of the pane-gone ALARM (the debounced, latched one raised from the backoff path), withpane-gonenamed only as a cause tag that can also appear on a buffer wedge.bin/fm-afk-start.sh:63- FM_AFK_LAUNCH_ARTIFACTS (bin/fm-afk-launch.sh:64) was introduced this commit with the comment "One list, so a new artifact cannot be added to the daemon and forgotten by the rollback that has to put it back", and the rollback rm now iterates it. But fm_afk_clear_stale_artifacts - the fresh-arm clear that drops the PRIOR away session's delivery artifacts - still hardcodes the same four names, and it lives in bin/fm-afk-start.sh, which the launcher SOURCES, so it cannot reference the launcher's array. The two agree today (both gained .subsuper-pane-gone in this PR), but a fifth artifact added to FM_AFK_LAUNCH_ARTIFACTS alone would be backed up and rolled back correctly while surviving a fresh arm - resurfacing a prior session's buffered escalation or wedge marker under a new away session, which is the failure fm_afk_clear_stale_artifacts exists to prevent. Defining the list once in bin/fm-afk-start.sh (where the clear lives) and having bin/fm-afk-launch.sh iterate that same array would close the drift and shrink the footprint rather than grow it. Flagged as an aside in round 9 and deliberately scoped out of that fix; noting the residual, not re-litigating.🔧 Fix: key wedge-alarm doc to alarm kind and unify afk artifact list
2 issues (1 warning, 1 info) still open:
bin/fm-supervise-daemon.sh:200- The WEDGE_ALARM_TITLE declaration comment still states the captain-facing banner title "follows the CAUSE" ("so it follows the CAUSE: a vanished supervisor pane means away-mode supervision is down"), but commit 1496b85 changed inject_wedge_alarm to select the summary, WEDGE_ALARM_TITLE, and the ERROR log from $kind instead — the same discriminator that picks the marker file and the throttle — specifically so a cause=pane-gone recorded by inject_msg during housekeeping's max-defer flush cannot render the supervision-DOWN verdict from what is actually a buffer wedge. docs/wedge-alarm.md was rewritten for this (lines 201-209: "WHICH one fired - its kind - is what every captain-facing surface says") and tests/fm-daemon.test.sh's test_buffer_wedge_alarm_never_declares_supervision_down pins it, so the code comment is now the only place in the repo asserting the superseded rule. inject_wedge_alarm's own header (line ~892) carries a milder version: it opens with "the ERROR log, the durable marker, and the active alert all report the recorded cause" before the next paragraph corrects it to the kind. Concrete failure: a maintainer reads the declaration comment, re-keys the branch at line 946 on $cause, and the reachable pane-died-mid-flush path (main-loop target-exists passes, pane dies during escalate_flush -> inject_msg records pane-gone at line 1298, housekeeping raises kind=buffer at line 1137) again titles an away captain's Notification Center banner "firstmate: away-mode supervision DOWN" for a buffer that is still stuck behind a pane the daemon keeps retrying — which per .agents/skills/afk/SKILL.md means tearing down a healthy away session. Rewrite lines 198-204 (and the opening sentence of the inject_wedge_alarm header) to say the title follows the alarm KIND, with the cause carried only as a diagnostic tag.bin/fm-bootstrap.sh:609- Noting the boundary of the new session-start diagnostic, not requesting a change. afk_injection_check gates on state/.afk, then fm_afk_canary_resolve falls back to discover_supervisor_target whenever no LIVE daemon owns the lock — so when away mode is armed but the daemon has died, bootstrap probes its own pane, finds firstmate obviously alive, and prints nothing. That is the right answer for what the canary measures (can the resolved supervisor be injected into), and AGENTS.md section 5 step 8 already covers the dead-daemon case by having /afk ensure the daemon is running on recovery, so nothing is actually missed. But docs/herdr-backend.md's "a cold or restarted session that re-enters afk from the flag alone still learns injection is off" reads broader than the check guarantees: the one away-mode failure a restarted session cannot learn from this line is the daemon simply not being there.🔧 Fix: key wedge-alarm title comment to alarm kind, not cause
2 infos still open:
bin/fm-afk-start.sh:79- fm_afk_clear_stale_artifacts changed from a single best-effortrm -f(which returns 0 essentially always) to a loop that returns 1 if ANY rm fails. In bin/fm-afk-launch.sh both call sites already test the return value and log "failed to clear stale away-mode artifacts", so there the new failure signal is an improvement. But fm_afk_start_main calls it bare inside anifbody (bin/fm-afk-start.sh:159) and that file runs underset -euwith noset +e, so a non-zero return is not a caught error — errexit fires and the daemon start aborts beforeexec "$FM_AFK_DAEMON", with no message at all (the rm's stderr goes to /dev/null and errexit is silent). Failure scenario: a direct, non-preparedbin/fm-afk-start.shon a state dir holding an undeletable artifact (root-owned leftover from a prior run under a different user, or a read-only mount) previously logged nothing and started away mode anyway; now it exits 1 with no output and away mode silently does not arm. Reachability is narrow — the /afk skill always goes through the launcher with FM_AFK_STATE_PREPARED=1, so only a manual/legacy direct invocation takes this branch — but the silence is the problem, not the abort. Either echo the failure before returning non-zero, or guard the call site the way the launcher does.bin/fm-supervise-daemon.sh:1081- When the supervisor target resolves again, pane_gone_recovered deletes state/.subsuper-pane-gone and logs the recovery only to the daemon log. But the captain has, by then, already received the pane-gone active alert — a Notification Center banner titled "firstmate: away-mode supervision DOWN", and a push through any configuredcommand:channel, telling them to stop away mode and arm it again. Nothing retracts it, and the marker (the one surface the afk-exit catch-up actually reads) is gone. Per .agents/skills/afk/SKILL.md:144, an absent .subsuper-pane-gone marker means "no live absence", so firstmate's catch-up reports all-clear for an alarm the captain demonstrably received. Failure scenario: the supervisor window is a session:window name (the FM_SUPERVISOR_TARGET_DEFAULTfirstmate:0fallback, or an explicit named target — the exact case pane_gone_recovered was written for), it is killed and recreated during a >300s gap overnight; the captain gets a 2am "supervision DOWN" push, comes back to a clean catch-up, and neither they nor firstmate can explain the alert without grepping state/.supervise-daemon.log, which no catch-up instruction points at. A one-line resolved note left behind on the recovery edge, or having the /afk catch-up read the daemon log's ERROR/recovery pair, would close it. Flagging for the record rather than requesting a change: this is the same best-effort-alert reach you already scoped out (docs/wedge-alarm.md "Known follow-ups"), and it belongs with the bounded-re-fire item there.🔧 **Test** - 1 issue found → auto-fixed ✅
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"🔧 Fix: provide mkdir to x-mode fakebin, silence node ExperimentalWarning
✅ Re-checked - no issues remain.
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"bash tests/*.test.shfull suite: 68 files, 0 failures, exit 0bash tests/fm-afk-inject-e2e.test.sh— real-shell Scenario D (genuine login shell with a bare starship ❯ prompt) passes: escalation defers, nothing typed into the shellbash tests/fm-daemon.test.sh— the three safety-critical inject_msg units pass: defers on empty-composer dead shell (agent_alive=dead), defers on unknown liveness (fail closed), injects on alive+empty (positive path)Pre-fix regression proof: neutered the liveness gate toliveness=alive(simulating gate-on-composer-only) and re-ran both files — Scenario D failed with 'SHELL INJECTION' (digest typed into and run by the live shell) and the starship-glyph unit failed with 'send_text_submit ran despite agent_alive=dead'; restored the file (git checkout) and confirmed clean tree + tests green againdocs/scripts.md:47- docs/scripts.md (the bin/ toolbelt inventory) is missing five scripts that pre-date this change: fm-lint.sh, fm-install-shellcheck.sh (added in fix: enforce deterministic ShellCheck parity #481), fm-cd-pretool-check.sh, fm-cd-command-policy.mjs (added in feat: guard primary shells from persistent cd commands #483), and fm-transition-lib.sh (added in feat(herdr): escalate blocked panes immediately #472). I added the row for fm-afk-canary-lib.sh, which this change introduced, but deliberately left the pre-existing five alone: backfilling unrelated scripts would expand a focused away-mode safety fix's documentation diff with work from three other PRs. Needs a human call on whether to fix them here or in a separate docs cleanup.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.