Skip to content

feat(watcher): absorb wakes only when the crew is provably working - #126

Merged
kunchenguid merged 2 commits into
mainfrom
fm/watcher-absorb-provable-w4
Jun 28, 2026
Merged

kunchenguid merged 2 commits into
mainfrom
fm/watcher-absorb-provable-w4

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Jun 28, 2026 •

Copy link
Copy Markdown
Owner

Intent

Rework the watcher's wake classification to "absorb-only-when-provably-working" so a crewmate that finishes (or otherwise stops and waits) is never silently swallowed.

The bug: the no-verb triage path (a bare turn-end, a working: note, a non-terminal stale) was benign by default and surfaced only when an accompanying status carried a captain-relevant verb. A crew that finished but reported through interactive pane menus - writing no done: status - had its final turn-end absorbed, so firstmate was never woken and the finish was missed (~45 minutes in the incident that motivated this).

The fix inverts the rule: a no-verb turn-end or non-terminal stale is absorbed ONLY when the crew shows positive evidence it is still working - its no-mistakes run for its branch is in an actively-running step, OR its pane shows the harness busy signature - and is surfaced otherwise (it may be done, waiting on a decision, or wedged).

Key decisions and tradeoffs (so review does not flag them as mistakes):

  • Added crew_is_provably_working and signal_crew_provably_working to the shared classifier bin/fm-classify-lib.sh, reusing bin/fm-crew-state.sh for the run-step/pane verdict rather than duplicating run-step logic. FM_CREW_STATE_BIN is a documented test override, consistent with the repo's existing FM_*_OVERRIDE style.
  • In bin/fm-watch.sh the costly provably-working read (it may make a bounded no-mistakes call) runs ONLY on the no-verb, non-afk signal path via short-circuit ordering (afk_present, then captain-verb, then provably-working), and for stale only on first sight of a distinct stale hash - never on every poll - so per-wake triage stays cheap.
  • A not-provably-working non-terminal stale now surfaces immediately instead of waiting out the wedge timer; a provably-working non-terminal stale still absorbs and keeps the existing wedge-timer backstop, so the low-churn behavior during long validations is preserved.
  • The afk/away-mode path is intentionally left UNCHANGED to avoid regressing afk: the watcher stays one-shot under state/.afk and skips the provably-working read entirely; the daemon (bin/fm-supervise-daemon.sh) keeps its own bounded-latency stale backstop and its classify functions were deliberately not modified.

What Changed

  • bin/fm-classify-lib.sh: adds the crew_is_provably_working predicate (reusing fm-crew-state.sh) and its signal_crew_provably_working wrapper.
  • bin/fm-watch.sh: the no-verb signal path surfaces a wake whose crew is not provably working; the non-terminal stale path surfaces immediately when not provably working, else absorbs with the wedge timer.
  • Tests: classifier unit tests plus behavioral fm-watch.sh runs covering every required semantic (mid-pipeline absorb, finished/parked surface, no-running-pipeline idle surface, busy absorb, captain-verb surface), a queue-safety test for the new immediate-surface stale path, and a hermetic fake fm-crew-state.sh test helper.
  • Docs: AGENTS.md section 8 and the architecture/configuration/scripts docs describe the new rule.

Risk Assessment

Low: the change is confined to the watcher's no-verb triage path. Detection is unchanged, the enqueue-before-suppressor ordering is preserved, the afk daemon path is untouched, and the only new failure mode is an occasional cheap false-positive peek (by design) instead of a swallowed finish.

Testing

Validated through the no-mistakes pipeline: review, test, document, lint, and push all completed with no findings. The full local suite (21 behavior suites) and shellcheck bin/*.sh tests/*.sh are green.

The pr step skipped (the no-mistakes daemon's GitHub auth), so this PR was created manually; GitHub Actions CI runs here.

Pipeline

Updates from git push no-mistakes

This change was validated through the no-mistakes pipeline (run 01KW806XEDYN5Q1V9SRPZCWVQ0). All gating steps passed with no findings; the pr step skipped because the no-mistakes daemon's GitHub auth was unavailable, so this PR was opened manually.

✅ intent - passed

✅ No issues found.

✅ rebase - passed

✅ No issues found.

✅ review - passed

✅ No issues found.

✅ test - passed

✅ No issues found.

✅ document - passed

✅ Synced watcher documentation (AGENTS.md section 8, docs/architecture.md, docs/configuration.md, docs/scripts.md).

✅ lint - passed

✅ No issues found.

✅ push - passed

✅ No issues found.

⏭️ pr - skipped (no-mistakes daemon GitHub auth unavailable; PR opened manually)
⏭️ ci - skipped (follows the pr step; GitHub Actions CI runs on this PR)

The no-verb triage path (a bare turn-end, a working: note, a non-terminal
stale) used to be benign by default and surfaced only on a captain-relevant
status verb. A crew that finished but reported through interactive pane menus
(no done: status) had its final turn-end absorbed, so firstmate was never
woken and the finish was missed.

Invert the rule: absorb a no-verb turn-end or non-terminal stale ONLY when the
crew shows positive evidence it is still working - its no-mistakes run for its
branch is in an actively-running step, or its pane shows the harness busy
signature. Otherwise surface it so firstmate peeks (done, waiting, or wedged).

- fm-classify-lib.sh: add crew_is_provably_working (reuses fm-crew-state.sh,
  no run-step duplication) and signal_crew_provably_working; FM_CREW_STATE_BIN
  override for tests.
- fm-watch.sh: signal path surfaces a no-verb wake whose crew is not provably
  working (costly check runs only on the no-verb, non-afk path); non-terminal
  stale surfaces immediately when not provably working, else absorbs with the
  wedge timer (run-step read only on first sight of a stale hash).
- afk path unchanged: the watcher stays one-shot and skips the provably-working
  read; the daemon keeps its bounded-latency stale backstop.
- tests: cover every required semantic (mid-pipeline absorb, finished/parked
  surface, no-running-pipeline idle surface, busy absorb, captain-verb surface)
  as classifier unit tests and behavioral watcher runs; queue-safety test for
  the new immediate-surface stale path.
- AGENTS.md section 8: document absorb-only-when-provably-working.
@kunchenguid
kunchenguid merged commit 81c94db into main Jun 28, 2026
4 of 5 checks passed
@kunchenguid
kunchenguid deleted the fm/watcher-absorb-provable-w4 branch June 28, 2026 21:21
mark-swift added a commit to mark-swift/firstmate that referenced this pull request Jun 29, 2026
…unchenguid#126) (#4)

* feat(watcher): absorb wakes only when the crew is provably working

The no-verb triage path (a bare turn-end, a working: note, a non-terminal
stale) used to be benign by default and surfaced only on a captain-relevant
status verb. A crew that finished but reported through interactive pane menus
(no done: status) had its final turn-end absorbed, so firstmate was never
woken and the finish was missed.

Invert the rule: absorb a no-verb turn-end or non-terminal stale ONLY when the
crew shows positive evidence it is still working - its no-mistakes run for its
branch is in an actively-running step, or its pane shows the harness busy
signature. Otherwise surface it so firstmate peeks (done, waiting, or wedged).

- fm-classify-lib.sh: add crew_is_provably_working (reuses fm-crew-state.sh,
  no run-step duplication) and signal_crew_provably_working; FM_CREW_STATE_BIN
  override for tests.
- fm-watch.sh: signal path surfaces a no-verb wake whose crew is not provably
  working (costly check runs only on the no-verb, non-afk path); non-terminal
  stale surfaces immediately when not provably working, else absorbs with the
  wedge timer (run-step read only on first sight of a stale hash).
- afk path unchanged: the watcher stays one-shot and skips the provably-working
  read; the daemon keeps its bounded-latency stale backstop.
- tests: cover every required semantic (mid-pipeline absorb, finished/parked
  surface, no-running-pipeline idle surface, busy absorb, captain-verb surface)
  as classifier unit tests and behavioral watcher runs; queue-safety test for
  the new immediate-surface stale path.
- AGENTS.md section 8: document absorb-only-when-provably-working.

* no-mistakes(document): Sync watcher documentation

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
leo1oel added a commit to leo1oel/nemo that referenced this pull request Jun 29, 2026
…rking (#15)

Port upstream kunchenguid#107 + kunchenguid#126 to the herdr fork: always-on in-bash wake triage.

The watcher now classifies every wake in bash and absorbs the benign majority
(advance the suppression marker, log to state/.watch-triage.log, keep blocking)
instead of waking firstmate's LLM for each. It queues and exits only on an
actionable wake, so firstmate re-arms once per actionable event instead of once
per wake - eliminating the quiet-stretch churn during a long crew validation.

kunchenguid#126's inversion is the safety property: a no-verb wake (a bare turn-end, a
working: note, a non-terminal stale) is absorbed ONLY while the crew shows
positive evidence it is still working - its no-mistakes run for its branch is in
an actively-running step, or its pane shows the harness busy signature, read via
bin/fm-crew-state.sh. A crew that stopped with no running pipeline and no busy
pane is surfaced immediately, so a finish reported only through interactive pane
menus (no done: status) is never swallowed.

- bin/fm-classify-lib.sh (new): the shared classifier - captain-relevant verb
  set, signal/stale/heartbeat predicates, scan_captain_relevant_statuses, and the
  provably-working predicate (crew_is_provably_working, reusing fm-crew-state.sh;
  FM_CREW_STATE_BIN override for tests). Sourced by BOTH the watcher and the
  away-mode daemon so the overlapping policy cannot drift.
- bin/fm-watch.sh: absorb gates on the signal / stale / heartbeat paths; the
  costly provably-working read runs only on the no-verb, non-afk path. A
  not-provably-working non-terminal stale surfaces at once; a provably-working one
  absorbs and starts a wedge timer that still escalates past FM_STALE_ESCALATE_SECS.
  Heartbeats absorb unless the fleet-scan backstop finds an unsurfaced
  captain-relevant status. While state/.afk exists the watcher reverts to one-shot
  (every wake surfaced for the daemon) and skips the provably-working read.
- bin/fm-supervise-daemon.sh: sources fm-classify-lib.sh and drops its now-duplicate
  CAPTAIN_RE_DEFAULT / last_status_line / status_is_captain_relevant / window_to_task;
  its housekeeping heartbeat scan uses the shared scan_captain_relevant_statuses.
- bin/fm-watch-arm.sh: comment updated for the actionable-wake behavior.

herdr adaptation: the stale path keeps our handle=/fm-backend read meta loop
(not upstream's window=/tmux capture-pane) and w="fm-<id>" labels; window_to_task
maps "fm-<id>" cleanly. Busy detection stays the fork's Claude-only regex.

Tests: tests/fm-watch-triage.test.sh (new, 11 cases, self-contained per the
fork's harness convention with a fake herdr pane-read + FM_CREW_STATE_BIN stub)
covers the classifier predicates and the behavioral absorb/surface contract
(captain-verb surface, no-verb absorb-when-working, no-verb surface-when-stopped,
terminal stale, non-terminal stale absorb-then-wedge, immediate stopped-crew
surface, heartbeat absorb + backstop). tests/fm-wake-queue.test.sh gains a
queue-safety test that a provably-working stale never enters the durable queue.
AGENTS.md sections 2/5/8, the afk skill, and docs synced. All 14 suites green;
shellcheck clean.
JTInventory added a commit to JTInventory/firstmate that referenced this pull request Jun 30, 2026
* feat(x-mode): add X mention completion follow-ups (kunchenguid#113)

* feat(x-mode): X-mention completion follow-up flow

Acknowledge an actionable X mention first, do the work, then post one
follow-up reply when it completes.

- fm-x-reply.sh: add --followup mode posting to the relay's
  /connector/followup endpoint; reuses thread-split, payload shape,
  dry-run (with a self-describing endpoint marker), and never-inline
  safety. Answer path unchanged.
- fm-x-link.sh: link a spawned task to its originating mention via
  x_request/x_request_ts in state/<id>.meta (atomic, preserves other
  lines).
- fm-x-followup.sh: --check detection plus post-and-clear on terminal
  completion; honors the 24h window (skip+prune past it), keeps the link
  on a failed post for retry.
- fm-x-lib.sh: shared meta link get/set/clear helpers.
- Docs: fmx-respond reads as one ack-first -> act -> follow-up flow;
  AGENTS.md §14 + supervision pointer document the link, completion
  follow-up, and 24h public-safe window.
- Tests: cover --followup endpoint/payload/dry-run, link, and the
  followup helper; shellcheck clean.

* no-mistakes(review): Captain, fix atomic X meta rewrites

* no-mistakes(document): Document X completion follow-ups

* feat(x-mode): dismiss skipped X mentions through the relay (kunchenguid#120)

* feat(x-mode): dismiss skipped mentions at the relay

The relay now exposes POST /connector/dismiss: acknowledge a pending
mention without replying - it drops the request, posts nothing, and stops
re-offering it. Wire firstmate to use it on the skip path so a deliberately
unanswered mention no longer churns every poll and times out to the relay's
"offline" auto-reply.

- bin/fm-x-dismiss.sh: new client modeled on fm-x-reply.sh. POSTs
  {request_id} (no body) to /connector/dismiss with the bearer; echoes the
  request_id on 2xx, exits non-zero on non-2xx/transport failure. Honors
  FMX_DRY_RUN (records the would-be POST to state/x-outbox/ with an
  endpoint:"dismiss" marker, posts nothing) and rejects unsafe request_ids.
- fmx-respond skill: the skip path now calls bin/fm-x-dismiss.sh before
  clearing the inbox file; answer and follow-up paths unchanged.
- AGENTS.md section 14: documents that a skipped mention is dismissed at the
  relay, not just locally cleared.
- tests: dismiss posts {request_id} to /connector/dismiss with the bearer
  and echoes it; dry-run records and posts nothing; non-2xx and transport
  failures exit non-zero; unsafe id and bad args rejected.

* chore(no-mistakes): run the bash suite directly as the test step

The test step had no configured test command, so it delegated to an agent;
that agent-driven run crashed the no-mistakes daemon mid-step on this repo.
Configure commands.test to run the firstmate behavior suite deterministically
instead, mirroring .github/workflows/ci.yml: iterate every tests/*.test.sh,
run each, and fail the step if any exits non-zero. This removes the agent from
the test step entirely (no crash) and makes the gate's test baseline match CI.
Same pattern myfirstmate uses (commands.test: mix deps.get && mix test).

* no-mistakes(review): Fix X dismiss docs and gate preflight

* no-mistakes(document): Document X dismiss and gate tests

* feat(watcher): absorb wakes only when the crew is provably working (kunchenguid#126)

* feat(watcher): absorb wakes only when the crew is provably working

The no-verb triage path (a bare turn-end, a working: note, a non-terminal
stale) used to be benign by default and surfaced only on a captain-relevant
status verb. A crew that finished but reported through interactive pane menus
(no done: status) had its final turn-end absorbed, so firstmate was never
woken and the finish was missed.

Invert the rule: absorb a no-verb turn-end or non-terminal stale ONLY when the
crew shows positive evidence it is still working - its no-mistakes run for its
branch is in an actively-running step, or its pane shows the harness busy
signature. Otherwise surface it so firstmate peeks (done, waiting, or wedged).

- fm-classify-lib.sh: add crew_is_provably_working (reuses fm-crew-state.sh,
  no run-step duplication) and signal_crew_provably_working; FM_CREW_STATE_BIN
  override for tests.
- fm-watch.sh: signal path surfaces a no-verb wake whose crew is not provably
  working (costly check runs only on the no-verb, non-afk path); non-terminal
  stale surfaces immediately when not provably working, else absorbs with the
  wedge timer (run-step read only on first sight of a stale hash).
- afk path unchanged: the watcher stays one-shot and skips the provably-working
  read; the daemon keeps its bounded-latency stale backstop.
- tests: cover every required semantic (mid-pipeline absorb, finished/parked
  surface, no-running-pipeline idle surface, busy absorb, captain-verb surface)
  as classifier unit tests and behavioral watcher runs; queue-safety test for
  the new immediate-surface stale path.
- AGENTS.md section 8: document absorb-only-when-provably-working.

* no-mistakes(document): Sync watcher documentation

* feat: add grok crewmate harness support (kunchenguid#143)

* feat(harness): add grok (Grok Build) as a verified crewmate adapter

Empirically verified against grok 0.2.73 and encoded across the machinery:

- fm-harness.sh: detect grok via GROK_AGENT=1 env marker (grok does not set
  CLAUDECODE) and `grok` command-name ancestry.
- fm-spawn.sh: grok launch template (`grok --always-approve "$(cat BRIEF)"`,
  fully autonomous, no permission gate) and a turn-end Stop hook. grok only
  loads project hooks after a manual folder-trust grant, so the hook is a
  single firstmate-owned global hook (~/.grok/hooks/fm-turn-end.json, always
  trusted) that is a guarded no-op unless the workspace holds a per-task
  .fm-grok-turnend pointer; fm-spawn drops that gitignored pointer naming
  state/<id>.turn-ended. Hook stays outside the worktree, needs no trust grant.
- fm-watch.sh + fm-tmux-lib.sh: grok busy signature `Ctrl+c:cancel` (the
  mid-turn cancel hint; ASCII, present iff a turn runs).
- harness-adapters skill: grok facts section (busy, exit=Ctrl+Q x2,
  interrupt=Ctrl+C, skill invocation /<skill>, resume) and /no-mistakes form.

Gating question confirmed: grok invokes /no-mistakes and drives a real
no-mistakes axi run, so grok is usable for no-mistakes-mode tasks. End-to-end
verified through fm-spawn: autonomous launch past the dir picker into the
worktree, brief processed, busy->idle and turn-end signal detected, fm-send
steer lands, clean Ctrl+Q exit and teardown. config/crew-harness is left
unchanged; this only makes grok available as a verified option.

* no-mistakes(review): Captain, harden Grok hook lifecycle

* no-mistakes(review): Captain, make Grok harness test executable

* no-mistakes(review): Captain, bound Grok pointer reads

* no-mistakes(test): Captain, harden crew-state and watcher-lock timing

* no-mistakes(document): Document Grok harness support

* feat(harness): split secondmate harness configuration (kunchenguid#144)

* feat(harness): split secondmate harness and inherit primary config into secondmate homes

Add config/secondmate-harness so secondmates can run on a different adapter
than crewmates. fm-harness.sh gains a `secondmate` mode resolving the chain
config/secondmate-harness -> config/crew-harness -> own; `crew` mode is
unchanged. fm-spawn resolves a --secondmate launch through that mode (durable:
every respawn re-resolves), while an explicit per-spawn harness arg still wins
and the unverified-adapter guard still holds.

Add a generic, extensible inheritable-config mechanism (fm-config-inherit-lib.sh)
that pushes the primary's declared LOCAL config into each secondmate home's
config/ at secondmate spawn and on the bootstrap secondmate sweep. Exactly one
item is wired today: config/crew-harness, so a secondmate's own crewmates use
the primary's setting. Primary-authoritative (re-pushed every convergence,
mirrors absence); config/secondmate-harness is deliberately not inherited since
secondmates never spawn secondmates. config/ is gitignored, so this is a copy
separate from the tracked-files fast-forward.

Update AGENTS.md (layout, bootstrap, harness, spawn), the harness-adapters
skill, docs/scripts.md, and .gitignore. New tests cover secondmate resolution
and fallback, spawn/respawn honoring config/secondmate-harness, config
propagation on spawn and sweep, the unverified-adapter guard, and backward
compatibility.

* no-mistakes(review): Surface inherited config propagation failures

* no-mistakes(review): Harden inherited config propagation

* no-mistakes(review): Document literal harness inheritance requirement

* no-mistakes(document): Document secondmate harness config

* feat(backlog): default backlog operations to tasks-axi (kunchenguid#145)

* feat(backlog): default to tasks-axi backend

* no-mistakes(document): Sync backlog backend docs

* fix(spawn): set per-task GOTMPDIR so interrupted Go builds don't leak /tmp (#36)

* fix(spawn): set per-task GOTMPDIR so interrupted Go builds don't leak /tmp

Go's GOTMPDIR is unset, so every go build/test creates numbered /tmp/go-build*
dirs. Go cleans them on a clean exit but LEAVES THEM when interrupted (signal,
timeout, OOM, full disk), accumulating and filling the disk over time.

Give each task its own temp root at /tmp/fm-<id>/ with Go's build temp nested at
gotmp/. fm-spawn creates the dir (Go won't mkdir GOTMPDIR), exports GOTMPDIR into
the crewmate pane so the agent and child processes inherit it, and records
tasktmp= in meta. fm-teardown reads tasktmp= and removes the whole root on
cleanup, deterministically.

GOTMPDIR (not TMPDIR) is the targeted knob: TMPDIR is too broad (affects every
program's temp). The nested root is extensible: other per-task temp can live
under /tmp/fm-<id>/ later.

Backward compat: tasks spawned before this change have no tasktmp= in meta;
teardown tolerates the empty value as a no-op. The daily fm-disk-cleanup.sh cron
remains a safety net for any pre-fix stray dirs.

* fix(tests): silence SC2016 for literal grep -F patterns in fm-gotmp test

The structural grep -F assertions deliberately match literal $TASK_TMP in the
fm-spawn source; add per-line shellcheck disable=SC2016 (the codebase's existing
pattern, e.g. bin/fm-spawn.sh) so CI lint passes.

* no-mistakes(document): docs: document tasktmp= meta field for per-task GOTMPDIR

---------

Co-authored-by: e-jung <8334081+e-jung@users.noreply.github.com>

* fix: accept landed squash-merged PR heads (kunchenguid#149)

* fix(teardown): accept landed squash-merge PR heads

* no-mistakes(document): Document teardown landing behavior

* no-mistakes: apply CI fixes

* fix(test): pass explicit teardown git identity

* feat(dispatch): add dynamic crew profiles (kunchenguid#154)

* feat(dispatch): add dynamic crew profiles

* no-mistakes(review): Captain, document dispatch profile inheritance

* no-mistakes(review): Captain, guard stale dispatch inheritance

* no-mistakes(document): Sync dispatch profile docs

* no-mistakes: apply CI fixes

* fix: harden crew dispatch profile enforcement (kunchenguid#159)

* Harden crew dispatch profile enforcement

* no-mistakes(document): Captain, synced crew dispatch docs

* feat: add live secondmate config push (kunchenguid#161)

* feat(config): add live secondmate config push

* no-mistakes(document): Document config push behavior

* no-mistakes(lint): Clean changed shell lint

* no-mistakes: apply CI fixes

* feat: support image attachments in X replies (kunchenguid#162)

* feat(x): add image attachments to reply helpers

* no-mistakes(review): Stream X image replies safely

* no-mistakes(review): Captain, clean X reply temp tracking

* no-mistakes(document): Document X reply image support

* Harden cleanup and image payload limits

* no-mistakes(review): Captain, validate spawn task IDs

* no-mistakes(document): Document X image cap

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: e-jung <e-jung@users.noreply.github.com>
Co-authored-by: e-jung <8334081+e-jung@users.noreply.github.com>
vipentti pushed a commit to vipentti/firstmate that referenced this pull request Aug 5, 2026
…unchenguid#126)

* feat(watcher): absorb wakes only when the crew is provably working

The no-verb triage path (a bare turn-end, a working: note, a non-terminal
stale) used to be benign by default and surfaced only on a captain-relevant
status verb. A crew that finished but reported through interactive pane menus
(no done: status) had its final turn-end absorbed, so firstmate was never
woken and the finish was missed.

Invert the rule: absorb a no-verb turn-end or non-terminal stale ONLY when the
crew shows positive evidence it is still working - its no-mistakes run for its
branch is in an actively-running step, or its pane shows the harness busy
signature. Otherwise surface it so firstmate peeks (done, waiting, or wedged).

- fm-classify-lib.sh: add crew_is_provably_working (reuses fm-crew-state.sh,
  no run-step duplication) and signal_crew_provably_working; FM_CREW_STATE_BIN
  override for tests.
- fm-watch.sh: signal path surfaces a no-verb wake whose crew is not provably
  working (costly check runs only on the no-verb, non-afk path); non-terminal
  stale surfaces immediately when not provably working, else absorbs with the
  wedge timer (run-step read only on first sight of a stale hash).
- afk path unchanged: the watcher stays one-shot and skips the provably-working
  read; the daemon keeps its bounded-latency stale backstop.
- tests: cover every required semantic (mid-pipeline absorb, finished/parked
  surface, no-running-pipeline idle surface, busy absorb, captain-verb surface)
  as classifier unit tests and behavioral watcher runs; queue-safety test for
  the new immediate-surface stale path.
- AGENTS.md section 8: document absorb-only-when-provably-working.

* no-mistakes(document): Sync watcher documentation
delaanthonio added a commit to delaanthonio/firstmate that referenced this pull request Aug 23, 2026
* fix(send): settle codex skill popups before submit (#103)

* fix(send): settle codex $skill popup before submit

A `$<skill>` invocation (e.g. $no-mistakes) opens codex's $-autocomplete
popup; submitting too fast lets it swallow the Enter so the invocation
never lands - biting every pipeline trigger to a codex crew/secondmate.
Mirror the existing `/` slash-popup handling: give a `$...` message the
1.2s popup-settle before the (retried) Enter, but scope it to
harness=codex (read from the target's meta) so `$`-prefixed plain text
("$5/month", "$HOME") to claude/opencode/pi keeps the 0.3s fast path. An
explicit session:window target has no meta -> harness unknown -> treated
as non-codex. The retried Enter in fm_tmux_submit_core still backs the
settle up; the `/` case, --key path, marker, and meta-resolution
contract are unchanged.

Record the codex $-popup fact in the harness-adapters skill, and add a
per-harness settle-selection test (codex $ ->1.2, claude $ ->0.3,
explicit $ ->0.3, any / ->1.2, plain ->0.3).

* no-mistakes(document): Sync fm-send popup docs

* feat(supervise): add deterministic crew state helper (#104)

* feat(supervise): deterministic crew current-state helper

Add bin/fm-crew-state.sh, a deterministic one-line read of a crew's
CURRENT state that ends the stale-status-line friction in supervision.
The status log is an append-only wake-event log, so tail -1 reports the
last event, not the current state: after firstmate resolves a
needs-decision/blocked and the crew silently resumes, the log stays
stale. The helper takes the active no-mistakes run-step as the source of
truth (attributed to the crew's branch, via axi status then the run
list), reconciles a possibly-stale needs-decision/blocked log line
against it (flagging it superseded when the run has resumed), and falls
back to the pane busy-signature + status log when no run is active.
Handles scout/secondmate kinds and torn-down/missing crews gracefully.

Reuses fm-tmux-lib.sh busy detection and meta parsing; fully
deterministic (run-step/pane/log reads only, no heuristics, no LLM).

Wire AGENTS.md sections 7 and 8 to read current state via the helper and
to stop inferring state from a tail of the status log; the status file
stays the wake-event log. Watcher core, status-file/brief contracts, and
wake/dedup machinery are unchanged.

Add tests/fm-crew-state.test.sh covering run-step authority, superseded
stale log lines, genuine-parked, cross-branch attribution, pane and
status-log fallback, scout skip, and torn-down/missing-meta cases.

* no-mistakes(review): Fix dead-window crew state fallback

* no-mistakes(review): Harden crew state liveness checks

* no-mistakes(review): Harden crew state gate handling

* no-mistakes(review): Harden crew state gate parsing

* no-mistakes(document): Document crew state helper

* no-mistakes: apply CI fixes

* fix(crew-state): keep run-step authoritative over a closed pane

The dead-window guard ran before the no-mistakes run lookup, so a
finished crew whose agent had exited and closed its window reported
'unknown' instead of its authoritative run-step state (e.g. done) - the
normal gap between a crew completing and teardown, and a regression
against the helper's core principle of judging by the run-step, not the
shell. Move the pane-readability guard into the no-run fallback path so
it runs only after the run-step lookup: a crew with a run reports its
run-step state regardless of pane liveness, while a genuinely runless
dead window still reports unknown rather than trusting a stale log.

Replace the test that encoded the regression with two that pin the
correct behavior: a closed pane with a terminal run reports done, and a
closed pane with an active run reports working.

* no-mistakes(document): Sync crew-state docs

* feat(bin): absorb benign watcher wakes (#107)

* feat(watch): always-on wake triage absorbs benign wakes in bash

Factor the afk daemon's classify_signal/classify_stale/heartbeat-scan into a
shared bin/fm-classify-lib.sh used by both the always-on watcher and the daemon.
fm-watch.sh now loops internally: it classifies each wake and absorbs the benign
majority (working: signals, bare turn-ended, non-terminal stale, no-change
heartbeat) by advancing the suppression marker and logging, without queuing or
exiting; it writes the durable queue and exits only on an actionable wake
(captain-relevant signal, any check, terminal stale, a non-terminal stale that
persists past FM_STALE_ESCALATE_SECS, or the heartbeat fleet-scan backstop). So
firstmate re-arms once per actionable event instead of once per wake.

Safety preserved: singleton lock + beacon (touched every poll including while
absorbing), durable queue for actionable wakes, kind=secondmate stale-skip,
per-task check polling, heartbeat backoff, bounded wedge latency. While state/.afk
exists the daemon owns triage and the watcher reverts to one-shot, so the two
never double-triage. AGENTS.md section 8 + afk skill updated; tests added.

* no-mistakes(review): Captain, repair watcher triage edge cases

* no-mistakes(document): Sync watcher triage docs

* no-mistakes: apply CI fixes

* feat: firstmate listens and replies on X (inert-by-default client) (#87)

The firstmate side of "listen on X and reply", shipped inside the repo so every
user has it but INERT until they opt in (a non-empty FMX_PAIRING_TOKEN in .env).
Purely additive: the watcher backbone (fm-watch.sh, fm-watch-arm.sh,
fm-wake-lib.sh) and the afk daemon (fm-supervise-daemon.sh, afk skill) are
untouched.

- bin/fm-x-poll.sh: one short-poll of GET /connector/poll; hard no-op without a
  token; requires non-empty text; stashes the full mention (incl. in_reply_to
  conversation context) to state/x-inbox/<id>.json behind a path-traversal guard;
  prints "x-mention <id>" (or a rate-limited "x-mode-error ...") for the watcher.
- bin/fm-x-reply.sh: POST /connector/answer; bearer token via a 0600 header file
  (never argv); reply via --text-file/stdin so mention text is never inlined into
  a shell command. Long replies auto-split into a premium-independent numbered
  "(k/n)" thread (codepoint-aware, capped); single tweet stays unnumbered. Wire:
  {request_id, text}, plus {texts:[...]} for a thread. FMX_DRY_RUN previews to
  state/x-outbox and posts nothing.
- bin/fm-x-lib.sh: .env/env config (token, relay default, dry-run, max chars,
  thread cap; env wins over .env) and the thread splitter.
- bin/fm-bootstrap.sh: .env-presence activation - drop the check shim + 30s
  cadence config on opt-in, remove on opt-out, idempotent, silent off.
- .agents/skills/fmx-respond: public-safe answer playbook - drain inbox, judge
  follow-up worthiness (skip pure acks), conversation continuity via in_reply_to,
  concise by default, dry-run aware.
- AGENTS.md section 14 + README/CONTRIBUTING/docs; tests/fm-x-mode.test.sh
  (hermetic: fake curl, real jq).

* fix: clarify X-mode owner mention handling (#109)

* fix(fmx): X-mode mentions are captain instructions - act, then reply

Owner-only relay routing means every routed mention is from the firstmate's
own owner, so fmx-respond no longer frames the asker as a stranger and may
address them as captain. Enabling X mode is the standing authorization, so
replies are composed and posted autonomously with no per-reply confirmation;
dry-run stays the only non-posting path.

A mention's request is a real captain instruction to act on, not merely to
reply to: the drain loop now classifies each mention as an actionable request
(run firstmate's normal lifecycle - intake, backlog, dispatch, scout, ship -
then report the outcome), a question (answer from fleet state), or a pure
acknowledgment (skip). The public channel keeps the yolo carve-out: anything
destructive, irreversible, or security-sensitive is flagged to the captain
through the trusted channel first, never executed straight from a mention.

Public-safety (outcomes only, no secrets/internals), the untrusted-in_reply_to
caution, and the skip-acknowledgment judgment remain intact. AGENTS.md §14
reflects the owner identity, autonomous answering, and act-on-requests carve-out.

* no-mistakes(review): Captain: clarify X-mode consent safeguards

* no-mistakes(review): Captain: document X-mode action consent

* no-mistakes(document): Captain, sync X-mode docs

* fix: recover safe fleet sync drift (#111)

* fix(fleet-sync): self-heal safe detached drift, loudly flag stuck clones

A pooled clone that drifts off its default branch was silently skipped
forever by both the post-merge teardown sync and bootstrap fleet-sync,
falling further behind on every merge with only an easy-to-miss skip line.

- Auto-recover the one safe case: a clean, detached HEAD that is an
  ancestor of origin/<default> and whose <default> is free to check out
  is re-attached and fast-forwarded, reported as 'recovered:'.
- Every other off-default state (dirty, non-default named branch,
  detached with unique commits, diverged) is left untouched and reported
  as a quantified 'STUCK: ... N commits behind ... - needs attention'
  warning instead of a quiet drift. Nothing is forced, stashed, or discarded.
- Relay the new recovered:/STUCK: outcomes through bootstrap FLEET_SYNC
  lines; document both in AGENTS.md section 3.
- Add tests/fm-fleet-sync.test.sh covering recover, every stuck variant,
  ordinary fast-forward, already-current, local-only/no-origin skips, the
  whole-fleet form, and the bootstrap relay.

* no-mistakes(review): Guard detached recovery from diverged local defaults

* no-mistakes(document): Sync fleet refresh docs

* no-mistakes: apply CI fixes

* feat(x-mode): add X mention completion follow-ups (#113)

* feat(x-mode): X-mention completion follow-up flow

Acknowledge an actionable X mention first, do the work, then post one
follow-up reply when it completes.

- fm-x-reply.sh: add --followup mode posting to the relay's
  /connector/followup endpoint; reuses thread-split, payload shape,
  dry-run (with a self-describing endpoint marker), and never-inline
  safety. Answer path unchanged.
- fm-x-link.sh: link a spawned task to its originating mention via
  x_request/x_request_ts in state/<id>.meta (atomic, preserves other
  lines).
- fm-x-followup.sh: --check detection plus post-and-clear on terminal
  completion; honors the 24h window (skip+prune past it), keeps the link
  on a failed post for retry.
- fm-x-lib.sh: shared meta link get/set/clear helpers.
- Docs: fmx-respond reads as one ack-first -> act -> follow-up flow;
  AGENTS.md §14 + supervision pointer document the link, completion
  follow-up, and 24h public-safe window.
- Tests: cover --followup endpoint/payload/dry-run, link, and the
  followup helper; shellcheck clean.

* no-mistakes(review): Captain, fix atomic X meta rewrites

* no-mistakes(document): Document X completion follow-ups

* feat(x-mode): dismiss skipped X mentions through the relay (#120)

* feat(x-mode): dismiss skipped mentions at the relay

The relay now exposes POST /connector/dismiss: acknowledge a pending
mention without replying - it drops the request, posts nothing, and stops
re-offering it. Wire firstmate to use it on the skip path so a deliberately
unanswered mention no longer churns every poll and times out to the relay's
"offline" auto-reply.

- bin/fm-x-dismiss.sh: new client modeled on fm-x-reply.sh. POSTs
  {request_id} (no body) to /connector/dismiss with the bearer; echoes the
  request_id on 2xx, exits non-zero on non-2xx/transport failure. Honors
  FMX_DRY_RUN (records the would-be POST to state/x-outbox/ with an
  endpoint:"dismiss" marker, posts nothing) and rejects unsafe request_ids.
- fmx-respond skill: the skip path now calls bin/fm-x-dismiss.sh before
  clearing the inbox file; answer and follow-up paths unchanged.
- AGENTS.md section 14: documents that a skipped mention is dismissed at the
  relay, not just locally cleared.
- tests: dismiss posts {request_id} to /connector/dismiss with the bearer
  and echoes it; dry-run records and posts nothing; non-2xx and transport
  failures exit non-zero; unsafe id and bad args rejected.

* chore(no-mistakes): run the bash suite directly as the test step

The test step had no configured test command, so it delegated to an agent;
that agent-driven run crashed the no-mistakes daemon mid-step on this repo.
Configure commands.test to run the firstmate behavior suite deterministically
instead, mirroring .github/workflows/ci.yml: iterate every tests/*.test.sh,
run each, and fail the step if any exits non-zero. This removes the agent from
the test step entirely (no crash) and makes the gate's test baseline match CI.
Same pattern myfirstmate uses (commands.test: mix deps.get && mix test).

* no-mistakes(review): Fix X dismiss docs and gate preflight

* no-mistakes(document): Document X dismiss and gate tests

* feat(watcher): absorb wakes only when the crew is provably working (#126)

* feat(watcher): absorb wakes only when the crew is provably working

The no-verb triage path (a bare turn-end, a working: note, a non-terminal
stale) used to be benign by default and surfaced only on a captain-relevant
status verb. A crew that finished but reported through interactive pane menus
(no done: status) had its final turn-end absorbed, so firstmate was never
woken and the finish was missed.

Invert the rule: absorb a no-verb turn-end or non-terminal stale ONLY when the
crew shows positive evidence it is still working - its no-mistakes run for its
branch is in an actively-running step, or its pane shows the harness busy
signature. Otherwise surface it so firstmate peeks (done, waiting, or wedged).

- fm-classify-lib.sh: add crew_is_provably_working (reuses fm-crew-state.sh,
  no run-step duplication) and signal_crew_provably_working; FM_CREW_STATE_BIN
  override for tests.
- fm-watch.sh: signal path surfaces a no-verb wake whose crew is not provably
  working (costly check runs only on the no-verb, non-afk path); non-terminal
  stale surfaces immediately when not provably working, else absorbs with the
  wedge timer (run-step read only on first sight of a stale hash).
- afk path unchanged: the watcher stays one-shot and skips the provably-working
  read; the daemon keeps its bounded-latency stale backstop.
- tests: cover every required semantic (mid-pipeline absorb, finished/parked
  surface, no-running-pipeline idle surface, busy absorb, captain-verb surface)
  as classifier unit tests and behavioral watcher runs; queue-safety test for
  the new immediate-surface stale path.
- AGENTS.md section 8: document absorb-only-when-provably-working.

* no-mistakes(document): Sync watcher documentation

* feat: add grok crewmate harness support (#143)

* feat(harness): add grok (Grok Build) as a verified crewmate adapter

Empirically verified against grok 0.2.73 and encoded across the machinery:

- fm-harness.sh: detect grok via GROK_AGENT=1 env marker (grok does not set
  CLAUDECODE) and `grok` command-name ancestry.
- fm-spawn.sh: grok launch template (`grok --always-approve "$(cat BRIEF)"`,
  fully autonomous, no permission gate) and a turn-end Stop hook. grok only
  loads project hooks after a manual folder-trust grant, so the hook is a
  single firstmate-owned global hook (~/.grok/hooks/fm-turn-end.json, always
  trusted) that is a guarded no-op unless the workspace holds a per-task
  .fm-grok-turnend pointer; fm-spawn drops that gitignored pointer naming
  state/<id>.turn-ended. Hook stays outside the worktree, needs no trust grant.
- fm-watch.sh + fm-tmux-lib.sh: grok busy signature `Ctrl+c:cancel` (the
  mid-turn cancel hint; ASCII, present iff a turn runs).
- harness-adapters skill: grok facts section (busy, exit=Ctrl+Q x2,
  interrupt=Ctrl+C, skill invocation /<skill>, resume) and /no-mistakes form.

Gating question confirmed: grok invokes /no-mistakes and drives a real
no-mistakes axi run, so grok is usable for no-mistakes-mode tasks. End-to-end
verified through fm-spawn: autonomous launch past the dir picker into the
worktree, brief processed, busy->idle and turn-end signal detected, fm-send
steer lands, clean Ctrl+Q exit and teardown. config/crew-harness is left
unchanged; this only makes grok available as a verified option.

* no-mistakes(review): Captain, harden Grok hook lifecycle

* no-mistakes(review): Captain, make Grok harness test executable

* no-mistakes(review): Captain, bound Grok pointer reads

* no-mistakes(test): Captain, harden crew-state and watcher-lock timing

* no-mistakes(document): Document Grok harness support

* feat(harness): split secondmate harness configuration (#144)

* feat(harness): split secondmate harness and inherit primary config into secondmate homes

Add config/secondmate-harness so secondmates can run on a different adapter
than crewmates. fm-harness.sh gains a `secondmate` mode resolving the chain
config/secondmate-harness -> config/crew-harness -> own; `crew` mode is
unchanged. fm-spawn resolves a --secondmate launch through that mode (durable:
every respawn re-resolves), while an explicit per-spawn harness arg still wins
and the unverified-adapter guard still holds.

Add a generic, extensible inheritable-config mechanism (fm-config-inherit-lib.sh)
that pushes the primary's declared LOCAL config into each secondmate home's
config/ at secondmate spawn and on the bootstrap secondmate sweep. Exactly one
item is wired today: config/crew-harness, so a secondmate's own crewmates use
the primary's setting. Primary-authoritative (re-pushed every convergence,
mirrors absence); config/secondmate-harness is deliberately not inherited since
secondmates never spawn secondmates. config/ is gitignored, so this is a copy
separate from the tracked-files fast-forward.

Update AGENTS.md (layout, bootstrap, harness, spawn), the harness-adapters
skill, docs/scripts.md, and .gitignore. New tests cover secondmate resolution
and fallback, spawn/respawn honoring config/secondmate-harness, config
propagation on spawn and sweep, the unverified-adapter guard, and backward
compatibility.

* no-mistakes(review): Surface inherited config propagation failures

* no-mistakes(review): Harden inherited config propagation

* no-mistakes(review): Document literal harness inheritance requirement

* no-mistakes(document): Document secondmate harness config

* feat(backlog): default backlog operations to tasks-axi (#145)

* feat(backlog): default to tasks-axi backend

* no-mistakes(document): Sync backlog backend docs

* fix(spawn): set per-task GOTMPDIR so interrupted Go builds don't leak /tmp (#36)

* fix(spawn): set per-task GOTMPDIR so interrupted Go builds don't leak /tmp

Go's GOTMPDIR is unset, so every go build/test creates numbered /tmp/go-build*
dirs. Go cleans them on a clean exit but LEAVES THEM when interrupted (signal,
timeout, OOM, full disk), accumulating and filling the disk over time.

Give each task its own temp root at /tmp/fm-<id>/ with Go's build temp nested at
gotmp/. fm-spawn creates the dir (Go won't mkdir GOTMPDIR), exports GOTMPDIR into
the crewmate pane so the agent and child processes inherit it, and records
tasktmp= in meta. fm-teardown reads tasktmp= and removes the whole root on
cleanup, deterministically.

GOTMPDIR (not TMPDIR) is the targeted knob: TMPDIR is too broad (affects every
program's temp). The nested root is extensible: other per-task temp can live
under /tmp/fm-<id>/ later.

Backward compat: tasks spawned before this change have no tasktmp= in meta;
teardown tolerates the empty value as a no-op. The daily fm-disk-cleanup.sh cron
remains a safety net for any pre-fix stray dirs.

* fix(tests): silence SC2016 for literal grep -F patterns in fm-gotmp test

The structural grep -F assertions deliberately match literal $TASK_TMP in the
fm-spawn source; add per-line shellcheck disable=SC2016 (the codebase's existing
pattern, e.g. bin/fm-spawn.sh) so CI lint passes.

* no-mistakes(document): docs: document tasktmp= meta field for per-task GOTMPDIR

---------

Co-authored-by: e-jung <8334081+e-jung@users.noreply.github.com>

* fix: accept landed squash-merged PR heads (#149)

* fix(teardown): accept landed squash-merge PR heads

* no-mistakes(document): Document teardown landing behavior

* no-mistakes: apply CI fixes

* fix(test): pass explicit teardown git identity

* feat(dispatch): add dynamic crew profiles (#154)

* feat(dispatch): add dynamic crew profiles

* no-mistakes(review): Captain, document dispatch profile inheritance

* no-mistakes(review): Captain, guard stale dispatch inheritance

* no-mistakes(document): Sync dispatch profile docs

* no-mistakes: apply CI fixes

* fix: harden crew dispatch profile enforcement (#159)

* Harden crew dispatch profile enforcement

* no-mistakes(document): Captain, synced crew dispatch docs

* feat: add live secondmate config push (#161)

* feat(config): add live secondmate config push

* no-mistakes(document): Document config push behavior

* no-mistakes(lint): Clean changed shell lint

* no-mistakes: apply CI fixes

* feat: support image attachments in X replies (#162)

* feat(x): add image attachments to reply helpers

* no-mistakes(review): Stream X image replies safely

* no-mistakes(review): Captain, clean X reply temp tracking

* no-mistakes(document): Document X reply image support

* fix(teardown): make landed PR detection robust (#167)

* fix(teardown): make landed-check robust when no pr= was ever recorded

fm-teardown.sh's squash-merge landed-check already falls back to
discovering a merged PR by branch name when state/<id>.meta has no
recorded pr=, but nothing guaranteed pr=/pr_head= actually got
recorded on a yolo-authorized merge - the "checks green" trigger that
normally runs fm-pr-check.sh never fires on repos with no PR CI, so a
merge done via a bare `gh-axi pr merge` silently skips it.

Add bin/fm-pr-merge.sh as the one path for merging a task's PR: it
always runs fm-pr-check.sh first, so pr=/pr_head= land in meta as part
of the merge itself regardless of any CI signal. Document both the
existing branch-name discovery fallback and the new merge path in
AGENTS.md, and add regression coverage for the no-pr=-recorded landed
scenario and for fm-pr-merge.sh's record-then-merge behavior.

* no-mistakes(review): Guard PR merges on task metadata

* no-mistakes(document): Document PR merge wrapper

* no-mistakes: apply CI fixes

* fix: parse PR merge URLs for gh-axi (#168)

* Fix fm-pr-merge.sh to parse PR URLs for gh-axi

gh-axi pr merge expects a PR number and --repo, not a full GitHub URL.
Parse the URL, default to --squash when no merge method is passed, and
fail fast on malformed URLs. Tests cover parsing, defaults, and refusal.

* no-mistakes(review): Harden PR merge validation

* no-mistakes(review): Harden PR merge URL guards

* no-mistakes(document): Document PR merge URL handling

* no-mistakes(lint): Clean shell lint

* feat(bin): pin secondmate model and effort (#180)

* feat: pin secondmate model/effort in config/secondmate-harness

Extend config/secondmate-harness's format to an optional
"<harness> [<model>] [<effort>]" line so a secondmate can be durably
locked to a concrete model/effort in the same file, without adding a
new config file. A bare harness-only file behaves exactly as before.

fm-harness.sh gains secondmate-model/secondmate-effort accessors;
fm-spawn.sh populates MODEL/EFFORT from them on every secondmate spawn
(including respawns) unless the caller passed an explicit --model/--effort.

* no-mistakes(review): Fix secondmate override pin precedence

* no-mistakes(document): Document secondmate harness pins

* feat(bin): add runtime backend interface (#183)

* feat(bin): extract tmux runtime behind a backend interface (P1)

Add bin/fm-backend.sh (selection, meta helpers, selector resolution,
dispatch) and bin/backends/tmux.sh (the tmux adapter), then route
fm-send.sh, fm-peek.sh, fm-watch.sh, fm-spawn.sh, and fm-teardown.sh
through them. Every default tmux command sequence, meta shape, and
printed output stays byte-identical: missing backend= still means
tmux, and a default spawn never writes backend=tmux.

Adds a --backend flag (tmux-only for now) and FM_BACKEND/config/backend
selection, refusing any unimplemented backend loudly. Names the
watcher's poll loop as the default event-source implementation over the
backend's pull primitives, per the herdr-addendum's events-as-the-core-
abstraction direction, without changing its behavior.

Verification: fake-tmux/treehouse old-vs-new command-log conformance
tests for send/peek/spawn/teardown, a real-tmux smoke test for the
adapter, and the full existing suite passing unmodified (bar two
fixture-only additions in fm-gotmp.test.sh for the new sibling
scripts).

* no-mistakes(review): Captain, harden backend baseline resolution

* no-mistakes(review): Captain, ignore and document backend config

* no-mistakes(review): Captain, make backend tests executable

* no-mistakes(document): Sync runtime backend documentation

* feat(bin): add experimental Herdr runtime backend (#186)

* feat(bin): add experimental herdr runtime backend (P2)

Implements bin/backends/herdr.sh (session-provider adapter, D3: treehouse
stays the worktree provider) wired through fm-backend.sh's dispatch, with
--backend herdr / FM_BACKEND=herdr / config/backend selection, a
version/protocol gate at spawn, semantic busy-state detection via herdr's
agent.get (fm-watch.sh and fm-crew-state.sh consult it before falling back to
the existing tmux pane-regex path), and label-based recovery discovery.

Container shape (D4) decided empirically: tab-per-task in one "firstmate"
workspace, mirroring tmux's one-session-many-windows model.

Found and fixed two real herdr v0.7.1 bugs during verification: `pane read
--lines N` returns empty for small N (worked around by over-fetching and
trimming locally), and `pane get`'s cwd field is frozen at pane-creation time
(fixed to read foreground_cwd instead, needed for fm-spawn's worktree-
discovery poll after `treehouse get`). Also fixed a pre-existing bug in
tests/fm-backend.test.sh's old-vs-new fixture that was silently missing
fm-backend.sh/bin/backends/ from the old bin/ shim.

Full empirical verification, the D4 decision evidence, and a real end-to-end
run (spawn/steer/peek/done/merge-local/teardown, including confirming
teardown refuses before the merge) are recorded in docs/herdr-backend.md.
The entire existing tmux conformance suite stays green.

* no-mistakes(review): Fix Herdr supervision recovery gaps

* no-mistakes(review): Captain, fix Herdr stale recovery gaps

* no-mistakes(review): Document Herdr composer primitive candidate

* no-mistakes(review): Captain, harden Herdr stale recovery and tests

* no-mistakes(document): Sync herdr backend docs

* no-mistakes: apply CI fixes

* feat(bin): auto-detect runtime backend (#188)

* feat(bin): auto-detect runtime backend from HERDR_ENV/TMUX markers

fm_backend_name now falls through to runtime auto-detection between
config/backend and the hard tmux default: a firstmate running natively
inside herdr (HERDR_ENV=1) now spawns crewmates into herdr by default,
mirroring how harness detection already works in fm-harness.sh. Nesting
resolves innermost-first (tmux wins over a nested herdr pane). Explicit
--backend/FM_BACKEND/config/backend settings always win over detection.
Selecting herdr via auto-detect prints a loud stderr notice; auto-detecting
tmux stays silent so the unconfigured default path is unchanged.

* no-mistakes(review): Captain, pin tmux tests and backend docs

* no-mistakes(document): Sync backend autodetect docs

* no-mistakes: apply CI fixes

* feat(stow): add operational memory capture (#197)

* feat(stow): add operational-memory learnings convention and /stow skill

Add data/learnings.md as the fleet-local operational-learnings home,
a knowledge-routing table in AGENTS.md, and a user-invocable /stow
skill that sweeps a session for uncaptured durable knowledge and
files it to the right disk home before a reset.

* no-mistakes(review): Fix stow backlog note command

* no-mistakes(document): Document stow memory routing

* fix(tests): protect herdr smoke cleanup from default sessions (#199)

* fix(tests): stop real-herdr smoke tests from ever killing the default session

Both fm-backend-herdr-smoke.test.sh and fm-backend-autodetect-smoke.test.sh
tore down their isolated throwaway HERDR_SESSION via a bare/inline-prefixed
`herdr server stop`, which is unscoped and resolves ambiently. On this herdr
client, that ambient resolution silently falls back to whatever server is
already running instead of the requested session - it killed the captain's
live default herdr server twice in production (2026-07-02), once from each
smoke test's cleanup trap.

Add tests/herdr-test-safety.sh with herdr_safe_stop_and_delete: it uses the
explicit-by-name `herdr session stop/delete <name>` form (never the ambient
`server stop`) and, before that, a read-only hard guard
(herdr_refuse_if_default) that re-queries `herdr session list --json` and
refuses outright if the target is literally "default", not found, or flagged
default:true. Fails closed on any ambiguity. Verified empirically against a
real isolated session: refuses on default/nonexistent/empty names without
ever calling stop, and correctly tears down a genuine isolated session while
leaving the default session's workspace state byte-identical before and
after.

* no-mistakes(review): guard herdr delete with fresh check

* no-mistakes(document): Document herdr smoke cleanup safety

* feat(bin): route herdr secondmates into per-home workspaces (#200)

* feat(bin): give each secondmate its own labeled herdr workspace

Give each secondmate its own labeled herdr workspace, and land crewmates
spawned from a secondmate home in that secondmate's own space, instead of
every firstmate home (primary and all secondmates) sharing one "firstmate"
workspace.

bin/backends/herdr.sh: replace the constant FM_BACKEND_HERDR_WORKSPACE_LABEL
with fm_backend_herdr_workspace_label(), resolved fresh from FM_HOME on every
call. The primary (no .fm-secondmate-home marker) still resolves to
"firstmate" - byte-identical to every pre-existing task's recorded label, no
forced migration. A secondmate home resolves to "firstmate-<secondmate-id>".
Every workspace-scoped path (find/ensure, tab create + duplicate check,
list-live recovery, pane-for-tab) uses this same resolution, so recovery and
duplicate checks stay scoped to each home's own space. Workspace and tab
create now pass --no-focus unconditionally (verified: neither focuses by
default once a workspace exists; --no-focus is defense in depth against the
one bootstrap edge case where the very first workspace in a session auto-
focuses).

Also fixes a session-targeting bug found while verifying this empirically:
HERDR_SESSION (env var, exported or inline-prefixed) is not reliably honored
by herdr 0.7.1 CLI subcommands once another herdr server is already running -
it silently falls back to whatever server IS running. fm_backend_herdr_cli
wraps every herdr invocation with both HERDR_SESSION and a trailing
--session <name> flag (verified to route correctly in every case tried),
fixing this for the whole adapter, not just the new label-scoped calls.

bin/fm-spawn.sh: a --secondmate spawn is launched BY the primary's own
process, whose FM_HOME still names the primary at that point. The herdr case
arm now shadows FM_HOME to the secondmate's own home (PROJ_ABS) for just the
two calls that resolve/create the workspace and tab, restored automatically
afterward (bash's temporary-assignment-before-a-command form works for shell
functions too). A crewmate/scout spawned FROM a secondmate's own fm-spawn.sh
process needs no such glue - its own FM_HOME already names it.

Tests: extended tests/fm-backend-herdr.test.sh (per-home label resolution,
--no-focus, --session flag, workspace-find/list-live scoping) and
tests/fm-backend-herdr-smoke.test.sh (a secondmate-shaped home's workspace
label, list-live scoping, restart stability in the multi-workspace shape).
Added tests/fm-backend-herdr-workspace-per-home-e2e.test.sh: the mandatory
isolated E2E, driving real bin/fm-spawn.sh/fm-teardown.sh - a primary-shaped
home into "firstmate", a --secondmate spawn into its own labeled space, a
crewmate spawned FROM that secondmate-shaped home landing in the same space
(this exact path had never run before), teardown closing only the right tab,
and list-live recovery seeing only each home's own tabs. All ten assertions
passed on the real binary; the default herdr session's own workspace state
was confirmed byte-identical before and after every real-herdr test run in
this change.

docs/herdr-backend.md: rewrote "Task container shape" for the workspace-per-
home design (label derivation, the --secondmate FM_HOME-shadow wrinkle, focus
behavior, label-collision/adopt-don't-duplicate semantics, no-forced-
migration), added "Session targeting: the --session flag, not HERDR_SESSION
alone", extended "ID stability" to the multi-workspace shape, and documented
the new E2E test.

* no-mistakes(review): Clarify herdr focus docs

* no-mistakes(review): Captain, clarify herdr server session docs

* no-mistakes(document): Document herdr per-home spaces

* fix(backends): rename Herdr secondmate workspace labels (#203)

* Rename herdr secondmate workspace prefix to 2ndmate-

The primary home keeps the firstmate label; secondmate homes now
resolve to 2ndmate-<id> so the herdr spaces sidebar is unambiguous.
Tests and docs updated; pre-rename workspaces can be aligned with
herdr workspace rename.

* no-mistakes(review): Clarify herdr workspace migration behavior

* no-mistakes(document): Herdr docs label alignment

* feat(bin): add unified session start digest (#201)

* feat(bin): collapse session start into one command

Add bin/fm-session-start.sh, composing fm-lock.sh, fm-bootstrap.sh, and
fm-wake-drain.sh into one ordered digest (lock, bootstrap diagnostics,
wake queue, context files, fleet state) instead of six-plus separate
turns. Lock now runs before bootstrap's mutating sweeps, closing a race
where a second concurrent session could mutate shared state before
discovering the lock was held. A lock refusal prints a loud read-only
banner, skips every mutating step via a new opt-in
FM_BOOTSTRAP_DETECT_ONLY flag on fm-bootstrap.sh, and still completes
the read-only-safe digest.

Add fm_backend_target_exists to fm-backend.sh as a shared, read-only,
never-side-effecting per-task endpoint-liveness primitive for both the
tmux and herdr backends.

Rewrite AGENTS.md sections 3 and 5 around the single command and add
tests/fm-session-start.test.sh.

* no-mistakes(review): Harden session-start read-only guidance

* no-mistakes(review): Suppress read-only tangle repair guidance

* no-mistakes(review): Include orphan status logs

* no-mistakes(review): Captain: make session-start test executable

* no-mistakes(review): Captain: clarify status tail guidance

* no-mistakes(document): Sync session-start docs

* no-mistakes(lint): Shell lint clean

* fix(tests): avoid shellcheck boolean chain

* fix(bin): corroborate herdr idle crew state (#207)

* fix(bin): corroborate herdr idle agent_status with the pane's own text

crew_pane_is_busy trusted a bare `idle` verdict from herdr's agent.get
outright, skipping the tail-regex corroboration unknown already gets.
agent.get reports generation state only (working while the model streams
a turn), so it reads idle for a crew blocked on its own long foreground
no-mistakes run - even though the pane still shows the busy banner the
whole time. Combined with the no-mistakes CLI's 10-run attribution cap,
this made a genuinely working herdr crew read as not provably working,
triggering an immediate stale wake instead of absorb-then-escalate.

* no-mistakes(document): Align herdr busy-state docs

* fix(backends): reuse the herdr firstmate workspace instead of leaking one per spawn (#202)

* fix(backends): reuse the herdr firstmate workspace instead of leaking one per spawn

Every herdr-backed crewmate left an orphaned `firstmate`-labelled workspace
behind, one per task, because `fm_backend_herdr_workspace_find` never matched
the existing workspace: its jq filter used `--arg label ... $label`, and
`label` is a reserved keyword in jq (label/break), so the filter was a compile
error. The error was swallowed by `2>/dev/null`, the find returned empty on
every call, and `workspace_ensure` took the create path each spawn, minting a
fresh workspace. The same collision silently disabled the create-task
duplicate-label check and the bare-selector tab lookup.

Rename the jq variable to `$want` in all three affected filters so reuse,
duplicate detection, and bare-selector lookup work. With reuse restored the
single `firstmate` workspace is persistent (like tmux's session) and teardown
correctly leaves it in place, closing only the task's pane/tab.

Also prune the default tab (label "1") herdr auto-creates inside a freshly
created workspace, best-effort, so the workspace holds only real task tabs.

Add stateful-fake-CLI tests that replay repeated spawn/teardown cycles and
assert one reused workspace, zero orphans, the default tab pruned, and
`workspace create` invoked exactly once. Verified against the real herdr binary
too: the pre-fix code failed the smoke idempotency check (minted wE then wF in
an isolated session); the fix passes it.

Document the workspace lifecycle, the jq-keyword pitfall, the default-tab
prune, and the project-labelled-workspace anomaly (not adapter-created) in
docs/herdr-backend.md.

* docs(herdr): trim workspace-lifecycle addition to current-state facts

The docs/herdr-backend.md convention documents current behavior, not
history - narrative belongs in the PR/commit message. Tightened the
workspace-leak and default-tab-prune write-up down to the operative
facts (the jq reserved-keyword guard, when pruning is safe, and the
persistence caveat), and corrected the CLI-facts table rows to match
the corrected prune timing.

* fix(backends): defer herdr default-tab prune until a real task tab exists

Closing a workspace's LAST tab deletes the whole workspace on real
herdr (verified). Pruning the auto-created default tab right after
workspace create closed the workspace's only tab at that point,
destroying the just-created workspace on every single spawn instead of
reusing it - the fake-CLI unit tests didn't model this real-herdr
behavior, so they passed while the real-herdr smoke test failed with
"container_ensure is not idempotent".

Move the prune into fm_backend_herdr_create_task, right after the
first real task tab is added to a freshly created workspace, when
closing the default tab alongside it is safe. Update the fake-CLI unit
test to match the corrected timing.

Also fix a smoke-test-only bug this surfaced: the test's second
create_task call reused a $CONTAINER captured before the first task
was killed, rather than re-running container_ensure like real
fm-spawn.sh always does immediately before every create_task call - so
once the workspace (correctly) disappeared after its last tab closed,
the stale reference no longer named a live workspace.

* test(backends): guard against jq --arg names colliding with jq keywords

Regression guard for the workspace-leak bug this PR fixes: a jq
--arg/--argjson variable named after a jq reserved keyword (e.g.
label) is a compile error on jq <= 1.6, and this adapter's
2>/dev/null silently turns that into an empty result instead of a
visible failure. Greps bin/ for the pattern so a future violation
fails loudly here instead of silently misbehaving on an older jq.

* docs(herdr): fix per-home staleness and drop contributor-specific example

The workspace-lifecycle write-up hardcoded "the firstmate workspace"
as if the label were always the fixed constant, stale against the
per-home labeling documented earlier in this file (primary: firstmate,
secondmate: 2ndmate-<id>). Rephrased per-home throughout, and pointed
the "workspace this adapter did not derive" case at the existing
Label-derivation section instead of a separate anomaly writeup.

Dropped the "Anomaly: a workspace labelled with a project name"
section - the python-teslemetry-stream example was a contributor's own
environment, not current adapter fact, and it repeated a now-incorrect
FM_BACKEND_HERDR_WORKSPACE_LABEL constant claim. Replaced with one
generic sentence already covered by the corrected wording above.

Also fixed the jq-reserved-keyword guard test's file reference, which
named tests/fm-backend.test.sh when the test actually lives in
tests/fm-backend-herdr.test.sh.

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* fix(backends): make Herdr default-tab pruning safe (#215)

* fix(backends): make the herdr default-tab prune provably safe

The default-tab prune could close a pane holding a LIVE agent: an
ADOPTED workspace (found pre-existing by label match) was pruned using
the same tab-count/label heuristic as a freshly created one, and herdr
derives a workspace's displayed label from its cwd basename when no
explicit --label is given. A captain launching herdr directly inside a
directory named "firstmate" produces a workspace that looks identical,
by label alone, to firstmate's own container - so the very next spawn
adopted the captain's own live workspace and closed their live pane
27ms after creating its task tab (2026-07-02 incident).

The fix is structural: fm_backend_herdr_workspace_ensure now captures
the seeded default tab's id straight from its own `workspace create`
response, only when it just created the workspace. That id threads
through fm_backend_herdr_container_ensure to fm_backend_herdr_create_task,
which is the only function allowed to prune it - an adopted workspace's
caller always passes an empty seeded-tab-id, so create_task never
re-derives "prunable" from a tab's label or count. Defense in depth:
the prune also refuses a tab whose pane reports a working agent.

Covered by new unit tests (adopted-never-prunes, created-prunes-exactly,
the exact label-collision incident shape) and a new isolated real-herdr
E2E test that reproduces the incident against the pre-fix code and shows
it fixed, plus the normal happy path.

* no-mistakes(document): Sync herdr prune docs

* no-mistakes: apply CI fixes

* fix(brief): remove apostrophe breaking bash -n and guard bash 3.2 set -u in spawn (#173)

* fix(brief): remove apostrophe that broke bash -n on fm-brief.sh

The no-mistakes DOD heredoc, built via VAR=$(cat <<EOF ... EOF), had an
unescaped apostrophe in "no-mistakes' own guidance". Nesting a heredoc
inside $(...) makes bash track quote state through the body, so the lone
apostrophe broke parsing of the rest of the script (bin/fm-brief.sh:211),
making the default no-mistakes ship path fail outright. Audited the other
two $(cat <<EOF...EOF) blocks (direct-PR, local-only) for the same class
of bug; none found. Added tests/fm-brief.test.sh as a regression guard.

* no-mistakes(test): fix(spawn): guard empty shared_args under bash 3.2 set -u

* no-mistakes(document): docs(contributing): list tests/fm-brief.test.sh in the test suite inventory

* fix(test): silence shellcheck SC2034/SC2100 in fm-brief.test.sh

Drop the unused out= capture (redirect to /dev/null instead) and quote
the id= assignment so shellcheck stops reading the hyphenated id value
as an arithmetic expression.

* feat: add experimental zellij runtime backend (#217)

* feat(backends): add experimental zellij runtime backend (P3)

Implements bin/backends/zellij.sh on the P1 dispatcher + P2 herdr precedent:
one zellij session, one tab per task, treehouse stays the worktree provider.
Wired through fm-backend.sh/fm-spawn.sh so fm-send/fm-peek/fm-watch/
fm-crew-state/fm-teardown work generically with zero changes to those scripts.

Empirically verified against real zellij 0.44.0: every "gaps to verify" item
from the design report, plus real findings the report missed - new-tab always
steals focus (mitigated with a restore call), zellij action always exits 0
even against a dead target, every pane op needs an explicit --pane-id, and
pane_cwd never tracks a subshell's own cd (treehouse get's exact shape) so
worktree-path discovery uses an active pwd-probe instead of passive JSON
polling. Findings and the full real-CLI evidence log are in
docs/zellij-backend.md.

Full real E2E cycle passed: spawn a real claude crewmate, accept its trust
dialog, steer it, receive done, confirm teardown refuses before merge, merge
local-only, confirm teardown then succeeds and the zellij tab is gone - all
in a scratch FM_HOME against a uniquely-named isolated zellij session, never
touching the real "firstmate" session or the live fleet.

Existing tmux and herdr conformance suites stay green; the two P1-era tests
asserting zellij was unimplemented now assert that of orca instead.

* no-mistakes(review): captain, guard zellij paste payloads

* no-mistakes(review): Captain, guard zellij pane readiness

* no-mistakes(review): Captain, harden zellij dead-target handling

* no-mistakes(review): Captain, harden zellij teardown and tests

* no-mistakes(review): Captain, harden zellij target validation

* no-mistakes(document): Document zellij backend

* docs: document Orca backend adapter contract (#209)

* docs: specify Orca backend adapter contract

* docs: require Orca window target alias

* no-mistakes(test): Captain: handle empty arrays under nounset

* no-mistakes(document): Document Orca backend proposal

---------

Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>

* feat(backends): add Orca primitive backend support (#210)

* feat(backends): add Orca adapter primitives

* no-mistakes(review): Gate Orca from task spawning

* no-mistakes(document): Document Orca backend limits

* fix(backends): stop mapping Orca Escape to interrupt

* no-mistakes(document): Document Orca primitive key support

* no-mistakes: apply CI fixes

* fix: drop CI-gate rewrite and normalize shared array-guard hunks

---------

Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>

* fix(backends): submit herdr slash commands reliably (#223)

* fix(backends): herdr slash-command submit verification false-positives on popup autocomplete

Two grok/herdr crewmates left /no-mistakes fully typed but unsubmitted for
minutes while fm-send exited 0. Live-reproduced against real grok 0.2.82:
the herdr adapter's submit verification declared success on ANY pane content
change after Enter, but an argument-taking slash command's first Enter only
closes the popup and expands the composer into an argument-hint placeholder
(/compact -> /compact compaction instructions) rather than submitting - a
real, visible change that isn't a submission. A second Enter is required.

fm_backend_herdr_composer_state replaces the delta check with a structural
read of the composer's own row (located by border-glyph shape, since herdr
exposes no cursor-row primitive), mirroring what cursor_y gives the tmux
adapter. A popup-close-with-placeholder-fill still reads pending, so the
retry loop now correctly sends the needed second Enter instead of stopping
early. The tmux backend was unaffected (its cursor-row read already handled
this correctly, verified side by side against the same live repro).

* no-mistakes(document): Sync herdr submit docs

* feat(skills): publish public stow skill with internal-skill hiding (#221)

* feat(skills): hide agent-only skills from installer discovery, add public stow

Mark every .agents/skills/* skill metadata.internal: true so the
skills.sh installer (npx skills add) hides them from discovery - all
assume a live firstmate home and are meaningless elsewhere. This is
inert to firstmate's own harness skill loader.

Add a new, fully standalone skills/stow for non-firstmate users: sweep
a conversation for durable knowledge and file it into whatever notes
convention the host project/user already has, asking once and
remembering the answer when ambiguous. No shared code with the
internal stow by design.

Document the two-tier layout in README and CONTRIBUTING.

* no-mistakes(review): Tighten public stow tracker routing

* no-mistakes(document): Sync skill docs

* fix(skills): make public stow's undone-next-steps routing local-first

Supersede the earlier ask-before-tracker-write tightening: the
standalone stow no longer treats an issue tracker as a routing option
at all based on inference (git remote, .github/ presence, etc).
Undone next steps always land in a local file by default - an
existing TODO/BACKLOG/NOTES file, or a freshly created local scratch
file otherwise. A tracker (or any other external system) is only ever
used when the user has explicitly said so, this session or as a
previously recorded standing preference.

* no-mistakes(document): Sync skill documentation

* feat(skills): clarify public stow resume and fallback routing (#225)

* feat(skills): resume pointer, routing tiers, default notes file for public stow

Adds a copy-pasteable resume pointer to the safe-to-end verdict so a
new session can pick the work back up cold, states the explicit three-tier
routing priority (explicit instruction > existing local convention > default
NOTES.md fallback) instead of leaving it implicit, and names NOTES.md as the
top-level discoverable default instead of letting each agent improvise a
location. Also removes a leftover internal-tooling word from the maintainer
comment so the public file carries zero internal vocabulary.

* fix(skills): make the public stow default fallback private and gitignored

Supersedes the earlier NOTES.md/ask-once split: the tier-3 default fallback
(no existing local convention fits) is now .stow-notes.md, gitignored so it
never lands unprompted in a shared/committed file. Because it's private, it's
safe for every finding-kind including user preferences, so the separate
ask-once carve-out for personal material is no longer needed there. Tier 2
(an already-established tracked convention) is unchanged and is the only
tier that still writes into a shared file. Step 7's resume pointer now flags
when notes landed in the private fallback and that they can be promoted into
a shared file later.

* fix(skills): split the public stow private fallback by scope, use git exclude

Supersedes the single .stow-notes.md fallback: user preferences (cross-project
by nature) now default to a host-local ~/.stow/notes.md instead of being
siloed into one project's repo. Project-scoped findings still default to
.stow-notes.md at the project root, but it's now kept out of git via the
local-only .git/info/exclude instead of the tracked .gitignore, so the
fallback is truly zero-shared-footprint: nothing lands in a tracked file and
nothing is left for the user to review or commit.

* fix(skills): keep the public stow private fallback sandbox-safe (current dir only)

Supersedes the home-file split: a home-directory path fails for agents
sandboxed to their current working directory, so the tier-3 default now
stays a single .stow-notes.md at the project root for every finding-kind,
including user preferences. Tier 2's user-level memory file is now framed
as a bonus when accessible, never assumed or required. Step 7 gains a
caveat when a preference lands in the project-local fallback: it applies to
this project only, and the user can copy it into their own global memory
file if they want it to follow them everywhere.

* fix(skills): use a current-directory .gitignore for the public stow fallback

Supersedes .git/info/exclude: that mechanism resolves outside the working
directory in a linked worktree, breaking the sandbox-safety guarantee for
exactly the setup this fleet uses everywhere. Switch to an ordinary
.gitignore file in the current directory instead - always in-directory
regardless of worktree layout - creating or appending a .stow-notes.md
line, left uncommitted for the user. If that write itself fails, the skill
still creates .stow-notes.md and tells the user to ignore it manually
rather than blocking. Also scopes the "never writes outside the current
directory" guarantee precisely to tier 3: tier 2 is exempted since it only
targets a destination the user's own existing convention already
established, which can legitimately be a user-level file outside the
project.

* no-mistakes(review): Clarify stow routing precedence

* no-mistakes(review): Clarify stow fallback metadata boundary

* no-mistakes(review): Guard tracked stow fallback

* no-mistakes(review): Clarify stow fallback verdict

* no-mistakes(document): Sync stow skill docs

* no-mistakes(lint): Public skill lint cleanup

* docs: clarify note hygiene guidance (#226)

* docs: add note-hygiene rule to backlog format section

Backlog and task notes accumulate volatile specifics that drift and
mislead; capture the general principle so every firstmate user avoids
trusting a stale note over the authoritative source.

* no-mistakes(review): Clarify note hygiene schema exemptions

* no-mistakes(document): Clarify note-hygiene docs

* revert: drop out-of-scope stow/architecture doc edits

Keep this PR's diff scoped to the AGENTS.md note-hygiene addition only;
skills/stow/SKILL.md, .agents/skills/stow/SKILL.md, and docs/architecture.md
are under separate close review and must not change out of band here.

This reverts commit ed2d73205b01e24ba8ffd40c523a24993f6eb1a5.

* feat(backends): add Orca task lifecycle support (#228)

* feat(backends): add Orca task lifecycle support

* fix(backends): harden Orca spawn lifecycle

* no-mistakes(review): Fix Orca lifecycle cleanup gaps

* no-mistakes(review): Release Orca worktrees without paths

* no-mistakes(review): Guard Orca spawn abort cleanup

* no-mistakes(review): Allow partial Orca child cleanup

* no-mistakes(review): Fix Orca selector and cleanup leaks

* no-mistakes(review): Preserve pathless Orca cleanup metadata

* no-mistakes(review): Harden Orca spawn and teardown lifecycle

* no-mistakes(review): Enforce Orca scout report gate

* no-mistakes(review): Harden Orca teardown path validation

* no-mistakes(review): Harden Orca capture errors

* no-mistakes(review): Harden Orca JSON cleanup validation

* no-mistakes(test): Fix zellij scout teardown fixture

* no-mistakes(document): Document Orca lifecycle support

* Harden Orca runtime and submit verification

* no-mistakes(review): Captain, preserve Orca current-tail verification

* no-mistakes(document): Sync Orca lifecycle docs

* no-mistakes(lint): Captain, silence deliberate ShellCheck

---------

Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>

* fix: use PR head for review diffs (#229)

* Fix fm-review-diff to compare PR head when pr= is recorded

After no-mistakes fix rounds push to the open PR, the crewmate worktree
branch can lag the authoritative PR head. When meta records pr=, resolve
the compare ref from reachable pr_head= or refs/pull/<n>/head before
diffing against the fetched authoritative base; fall back to the local
branch with a loud warning when the PR head cannot be resolved.

Add behavior tests for pr_head resolution, fetch, unchanged no-pr path,
and unreachable-PR fallback.

* no-mistakes(review): Make review-diff test executable

* no-mistakes(document): Document PR-head review diffs

* no-mistakes: apply CI fixes

* fix: make backend matching shell-portable (#230)

* Make fm-backend.sh backend-name matching shell-portable for zsh

Replace word-split-dependent for-loops in fm_backend_is_known() and
fm_backend_validate_spawn() with case-based membership tests so sourcing
the library from zsh no longer falsely rejects known backends.

Add zsh/bash regression coverage in tests/fm-backend.test.sh.

* no-mistakes(review): Fix zsh backend loading and validation

* no-mistakes(document): Document backend membership portability

* fix(backends): make herdr respawn idempotent (#231)

* fix(backends): make herdr respawn idempotent against restored-layout husks

herdr persists and restores its session layout (workspaces/tabs/panes)
across a server restart, so a restored fm-<id> task tab comes back a
husk - a dead pane, or a plain agent-less shell - which fm-spawn.sh's
duplicate-tab guard refused unconditionally, forcing manual pane closes
after every restart.

fm_backend_herdr_create_task now classifies an existing same-labeled
tab's pane conservatively (dead/no-agent/live/unknown) and
closes-and-replaces only a confirmed husk, always creating the
replacement tab before closing the old one so a husk that is a
workspace's only tab is never at risk of taking the whole workspace
down with it. A genuinely live agent, or anything not confidently
classifiable, still refuses exactly as before.

* no-mistakes(review): Harden herdr duplicate respawn guard

* no-mistakes(review): Enforce herdr husk cleanup postcondition

* no-mistakes(document): sync herdr respawn docs

* no-mistakes: apply CI fixes

* fix: absorb stale wakes during active validation (#233)

* fix(watcher): stop stale_is_terminal from ignoring an active run-step

A crewmate's status log gets no new entry once firstmate hands it to a
no-mistakes validation (the sparse status-reporting contract), so the
log's last line can stay a pre-validation "done:" (or needs-decision/
blocked) leftover for the run's entire duration. fm-watch.sh's
stale_is_terminal only reads that raw last line - it has no run-step
awareness - so it kept surfacing a stale pane as immediately terminal
every time it went quiet for two polls, no matter how actively the
pipeline was validating (confirmed live against fm-herdr-respawn-idem,
whose status log's last line is literally "done: ..." while its
no-mistakes run-step reads "validating (running)"). crew_is_provably_working
now gets a chance to override a stale captain-relevant log line on a
new stale hash, exactly as it already did for a non-captain-relevant one.

Also fixed a separate, independently-confirmed dead code path in
fm-crew-state.sh: its cross-branch run-attribution fallback shelled out
to `no-mistakes axi` (bare) expecting a runs[N]{...} TOON table that the
real CLI (v1.32.2) never emits - verified the axi surface exposes only
abort/logs/respond/run/status. Replaced it with the real top-level
`no-mistakes runs` listing.

* no-mistakes(document): Sync watcher stale docs

* no-mistakes(lint): Fix stale fake run-list variables

* docs: clarify no-mistakes evidence commit handling (#232)

* docs: no-mistakes evidence commits in crew branches are intentional

* no-mistakes(review): Scope evidence guidance to project repos

* no-mistakes(document): Document evidence commit policy

* fix: tighten Orca backend parsing (#237)

* fix: tighten orca parser coverage

* no-mistakes(document): Sync Orca backend docs

* no-mistakes(lint): Captain, lint clean

---------

Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>

* docs: add backend setup guides and slim README (#238)

* docs: make README pointer-first, add per-backend setup guides

Trim the Quick Start and How It Works walls of prose down to overview plus
pointers, relocating every removed sentence's content into docs/architecture.md,
docs/configuration.md, or the relevant backend doc. Add docs/tmux-backend.md as
the reference-backend setup guide, and add a Setup section to each experimental
backend doc (herdr, zellij, Orca) covering prerequisites, selection, first run,
watching/attaching, verification, and limitations. Record the README convention
in CONTRIBUTING.md.

* no-mistakes(review): Clarify backend setup docs

* no-mistakes(review): Clarify tmux secondmate support

* no-mistakes(document): Align backend documentation

* docs: stop telling users to run fm-spawn.sh for backend selection

fm-spawn.sh is firstmate-internal; a user never runs it directly. Rephrase
every user-facing backend-selection sentence across the tmux/herdr/zellij/orca
guides and docs/configuration.md to present the actual user mechanisms: a
local config/backend file, FM_BACKEND at launch, or telling the first mate in
chat. Internal-reference mentions of the --backend flag (fm-spawn.sh usage
notes, test coverage lists) are left as mechanics, not user instructions.

* Document herdr dual license in docs and install hint (#239)

Add AGPL-3.0-or-later/commercial licensing note to herdr-backend Setup
and the missing-binary error in bin/backends/herdr.sh.

* feat: support multiple X-mode follow-ups (#241)

* feat(x-mode): raise X follow-up cap to 3 within a 7-day window

Matches the relay's parallel contract change: fm-x-link.sh now records a
follow-up counter (with --carry-count to preserve it across a re-link onto
a successor task), fm-x-followup.sh posts up to three follow-ups per
mention within a 7-day window instead of one within 24h, clearing the
link on --final, cap exhaustion, window lapse, or a distinguishable relay
rejection (fm-x-reply.sh HTTP 409 -> exit 9) rather than treating that
rejection as a retryable failure. AGENTS.md and the fmx-respond skill are
updated to keep usage disciplined: spend follow-ups only on genuine
milestones, always finish with --final.

* no-mistakes(review): Update follow-up dry-run docs

* no-mistakes(review): Captain: refresh X follow-up docs

* no-mistakes(review): Captain: harden X follow-up relinks

* no-mistakes(review): Captain: harden follow-up state persistence

* no-mistakes(document): Document X follow-up carryover

* feat(backends): add experimental cmux runtime backend (#246)

* feat(backends): add cmux runtime backend (experimental)

Session-provider-only adapter for cmux (bin/backends/cmux.sh), mirroring
zellij/herdr structurally, wired into fm-backend.sh and fm-spawn.sh with
--secondmate refused for now. Verified against the real cmux 0.64.17 app:
send does not auto-submit, cwd is creation-time-frozen (zellij-shape,
pwd-marker-probe workaround), close-surface refuses on a workspace's last
surface (falls back to close-workspace), workspace ids do not survive a
relaunch, and the control socket defaults to cmuxOnly access (requires a
one-time password-mode setup, documented in docs/cmux-backend.md). Also
found and fixed a live bug during development: read-screen fails on a
surface that has never been written to, so liveness now uses list-panes
instead. Fake-CLI unit suite (40 tests), a real-binary smoke test, and a
full spawn/steer/peek/done/merge/teardown E2E pass against a real claude
crewmate all pass, including the popup/second-Enter regression class.

* no-mistakes(review): Harden cmux recovery and password parsing

* no-mistakes(review): Harden cmux capture failure handling

* no-mistakes(review): Mark cmux test scripts executable

* no-mistakes(review): Scope cmux workspaces and teardown

* no-mistakes(review): Captain, honor cmux password config override

* no-mistakes(review): Captain, hash cmux home labels

* no-mistakes(document): Sync cmux backend docs

* feat(agents): add firstmate coding guidelines skill (#248)

* Add firstmate-coding-guidelines skill (AGENTS.md diet PR 0)

Encodes the knowledge-placement decision tree, one-owner rule, and
inline-stub pattern from the diet analysis so future contributions stop
adding conditional detail inline. AGENTS.md gets one section-13 trigger
line; fm-brief.sh's REPO argument has no reliable signal for "this is
firstmate's own repo", so the load instruction goes in CONTRIBUTING.md's
Development section instead of the scaffold.

* no-mistakes(review): Captain, align tracked-material trigger scope

* no-mistakes(document): Sync coding guidelines docs

* no-mistakes(lint): Fix Markdown style issues

* fix: add turn-end supervision guard (#249)

* feat: structural Stop-hook backstop for primary turn-end supervision

fm-guard.sh is pull-based: it only warns when some other supervision
script happens to run, so a primary session that ends a turn without
re-arming the watcher and then runs no further fleet-touching command
can sit blind for hours (the 2026-07-04 incident this fixes).

Add bin/fm-turnend-guard.sh, a Claude Code Stop hook registered in the
tracked .claude/settings.json, that fires on every primary turn end and
blocks (exit 2, verified empirically to force continuation) when work
is in flight with no fresh watcher beacon. It never blocks more than
once per turn, using Claude Code's own stop_hook_active loop-guard
field, and scopes itself to the actual primary checkout only (inert in
crewmate/scout worktrees and secondmate homes).

Factor the shared "in-flight but no live watcher" predicate out of
fm-guard.sh into bin/fm-supervision-lib.sh so the pull-based banner and
the push-based hook can never drift on what "unhealthy" means.

Document the verified Stop-hook mechanism and scoping in
docs/turnend-guard.md, add a harness-adapters note, and cover the
predicate and hook with tests/fm-turnend-guard.test.sh.

* no-mistakes(review): Respect active home in turnend guard

* no-mistakes(review): Require live watcher for turn-end guard

* no-mistakes(review): Captain: portable turn-end timing

* no-mistakes(document): Sync turn-end guard documentation

* feat(backends): auto-detect cmux runtime (#250)

* feat(backends): auto-detect cmux runtime from CMUX_WORKSPACE_ID

Wires cmux into fm_backend_detect the same way herdr already is: a
firstmate process running inside a cmux-spawned terminal now spawns
new tasks into cmux by default, no config needed. Verified from cmux's
own shipped source that CMUX_WORKSPACE_ID/CMUX_SURFACE_ID/CMUX_SOCKET_PATH
are unconditionally, no…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant