fix: detect Git and centralize backend configuration - #445
Merged
Merged
Conversation
Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3).
The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements.
Keigyoku
pushed a commit
to Keigyoku/firstmate
that referenced
this pull request
Jul 11, 2026
* docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation (cherry picked from commit bc558c6)
Keigyoku
pushed a commit
to Keigyoku/firstmate
that referenced
this pull request
Jul 11, 2026
* docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation (cherry picked from commit bc558c6)
Keigyoku
added a commit
to Keigyoku/firstmate
that referenced
this pull request
Jul 11, 2026
…, kunchenguid#435 deferred) (#8) * feat: add harness-aware supervision (kunchenguid#367) * Add harness-aware supervision * no-mistakes(review): Captain, harden watcher supervision regressions * no-mistakes(review): Captain, harden watcher supervision cadence * no-mistakes(review): Harden watcher supervision ownership * no-mistakes(review): Captain, harden Pi extension marker * no-mistakes(review): Captain, harden Pi supervision restart checks * no-mistakes(review): Harden watcher ownership checks * no-mistakes(review): Captain, harden Pi supervision loading * no-mistakes(review): Captain, require Pi guard extension loading * no-mistakes(review): Captain, harden watcher supervision recovery * no-mistakes(test): Fix fm-send baseline log filtering * no-mistakes(document): Sync harness supervision docs * no-mistakes: apply CI fixes (cherry picked from commit 090483e) * fix: split X-mode replies by platform (kunchenguid#369) * fix: make x replies split by platform * no-mistakes(review): Captain: preserve Discord recovery relink context * no-mistakes(test): Captain: keep split markers outside fences * no-mistakes(document): Sync X-mode reply docs (cherry picked from commit 4790bc5) * fix: make stow memory writes inspect before update (kunchenguid#372) * docs: make stow inspect-then-update * no-mistakes(review): Remove unsupported archive-body guidance * no-mistakes(review): Clarify stow read-before-write exception * no-mistakes(test): Require archive-body for stow task notes * no-mistakes(document): Sync stow memory docs * no-mistakes(lint): Silence ShellCheck source warning (cherry picked from commit af5361e) * fix(watcher): wait when arm attaches to a healthy watcher (kunchenguid#375) * fix: attach-and-wait when arm finds a healthy watcher Grok and Claude re-arm after every turn with work in flight. When a watcher was already healthy, fm-watch-arm exited immediately with watcher: healthy, which completed the harness background task and injected an empty false wake. Attach to the live identity-matched holder instead, stay until that cycle ends, then exit 0 so notify fires for a real end-of-cycle. The peer-startup-race path uses the same contract. --restart and the started path are unchanged. * no-mistakes(review): Gate restart watcher peer attach * no-mistakes(document): Sync watcher arm docs (cherry picked from commit 7302a95) * feat(pi): simplify primary session launch (kunchenguid#386) * docs(readme): reformat Quick Start and recommend Grok equally with Claude Code * no-mistakes(review): Captain: align harness launch guidance * no-mistakes(review): Captain, clarify Pi supervised launch * no-mistakes(review): Captain, document Pi first-launch bridge * feat(pi): track primary watcher extension for plain-pi launch Move Pi's primary watcher bridge from a generated state/ file to a tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end guard extension already works: self-hashing version, project-local auto-discovery after one-time Pi trust. This drops the state/-generation step and dual -e requirement from the happy path, so Pi's Quick Start launch becomes plain 'pi', the same friction class as 'claude' and 'grok --trust'. - bin/fm-pi-watch-extension.sh is removed; nothing generates the extension anymore since it is committed. - fm-session-start.sh and fm-supervision-instructions.sh resolve the watcher extension path from FM_ROOT instead of state/, and the session-start diagnostic now points at restarting plain pi after trust, with -e as a documented fallback. - fm-spawn.sh points Pi secondmate launches at the tracked extension path in the secondmate home instead of generating a state/ copy. - README Quick Start Pi block is now just 'pi' plus a trust note. - Tests, docs, and the harness-adapters skill updated to match. * fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016 (cherry picked from commit 1b5a9f5) * feat(supervision): prevent unsafe watcher-arm commands (kunchenguid#387) * feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that denies a primary shell command backgrounding, piping, or bundling the watcher arm/checkpoint, or force-killing the watcher process broadly - the exact shapes that silently took Grok's supervision down. Wires it into all five verified harnesses (grok, claude, codex, opencode, pi), each validated empirically against the real harness. Also fixes a grok 0.2.93 regression discovered during that validation: the existing turnend-guard Stop hook's bare root variable broke grok's own variable pre-substitution and silently no-op'd the hook. * no-mistakes(review): Harden watcher arm validation * no-mistakes(review): Harden arm guard metacharacter checks * no-mistakes(review): Harden nested shell arm guard * fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then A && B || C is not if-then-else; C can run when A is true. Replace both occurrences of the quote-state early-continue with an explicit if/then. (cherry picked from commit 766f772) * fix(pi): restore primary watcher supervision lifecycle (kunchenguid#397) * fix Pi primary supervision lifecycle * no-mistakes(document): Synchronize Pi primary extension documentation (cherry picked from commit 80ecc8d) * fix: keep persistent secondmates out of the main backlog (kunchenguid#398) * fix secondmate backlog guidance * no-mistakes(review): Require reasons for captain backlog holds * no-mistakes(test): Document secondmate handoff skill requirement * fix secondmate teardown reminder * no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract (cherry picked from commit 7525a37) * fix(backlog-handoff): move full item blocks including indented bodies (kunchenguid#401) * fix(backlog-handoff): move full item blocks including indented bodies fm-backlog-handoff only moved the checklist header line, so multi-line item bodies were left orphaned in the source backlog and never reached the secondmate. Move the full block (header plus indented body lines) atomically, treating body membership by indentation so lines like ## Intent stay with the item, and add regression coverage. * no-mistakes(review): Captain: preserve EOF handoff terminators * no-mistakes(review): treat blank lines inside item bodies as movable body * no-mistakes(document): sync backlog-handoff docs with full-block move behavior (cherry picked from commit 600075e) * feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (kunchenguid#402) * guard Herdr lab lifecycle in briefs * no-mistakes(review): Fix Herdr lab helper and provisioning safety * no-mistakes(review): Captain, harden Herdr lab lifecycle safety * no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range * no-mistakes(review): reject leading options in Herdr lab run guard * no-mistakes(review): strip leading non-alnum in Herdr lab name generator * no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag * no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab * no-mistakes: apply CI fixes (cherry picked from commit 4bc0824) * fix(watcher): classify arm-command seatbelt by execution position (kunchenguid#403) * fix watcher arm command policy * no-mistakes(review): Harden watcher command policy parsing * no-mistakes(review): Captain: harden watcher policy parsing * no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter * no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps * no-mistakes(review): fail closed on loop-wrapped broad watcher kills * no-mistakes(document): sync docs for watcher-arm command-position policy (cherry picked from commit 22b1d71) * fix: reconcile existing AGENTS.md safely (kunchenguid#405) * fix(agents-md): inject self-governance section into existing AGENTS.md fm-ensure-agents-md.sh only appended the canonical "## Maintaining this file" section on skeleton create or CLAUDE.md promotion, so an existing AGENTS.md that lacked it exited unchanged and forced hand-copying the wording during a rollout across existing projects. Call the already- idempotent ensure_maintenance_section on the existing-AGENTS.md paths and report whether the file changed; a re-run and an already-complete file stay byte-identical. Also fixes kunchenguid#389: refuse a case-variant real memory file (e.g. a lowercase agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase literal target dangles once the tree lands on a case-sensitive filesystem. Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and CLAUDE.md-promotion regressions still pass. Docs updated to match. * no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency * no-mistakes(review): Preserve CRLF during maintenance-section injection * no-mistakes(review): Captain: harden dangling-symlink regression coverage * no-mistakes(document): Document agent-memory injection outcomes (cherry picked from commit 96244f4) * feat: support project-less secondmate homes (kunchenguid#409) * feat(secondmate): support project-less homes via --no-projects fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit --no-projects signal to scaffold, seed, and register a secondmate home whose subject is the firstmate repo itself (no clones). The signal is mutually exclusive with a project list; omitting both still fails loudly so an accidental omission is never a silent project-less seed. The registry line renders an empty projects: field, which spawn and the snapshot already tolerate. Docs updated in the secondmate-provisioning skill and both script headers. * no-mistakes(review): Captain: document project-less secondmate flow * no-mistakes(review): Captain: refuse project-less reseeding of populated homes * fix(seed): fail closed on unreadable project data * no-mistakes(review): Captain: reject stale projectful charters * no-mistakes(review): Captain: fail closed on unsafe project paths * no-mistakes(review): Captain: validate project-less charter clone sections * no-mistakes(document): Document project-less secondmate seeding (cherry picked from commit 08453f9) * fix: delegate backlog handoffs to tasks-axi (kunchenguid#411) * wip(handoff): record verified delegation design + tasks-axi mv blocker No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically move a blocked-by-linked item set across backlogs (deadlocks both orders, no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked pending a tasks-axi connected-set mv enhancement; note captures the verified design, semantics, test/CI/doc changes, and resume checklist. * refactor(handoff): delegate the item move to tasks-axi mv fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog format and the source of the PR kunchenguid#401 body-orphaning drift. Delete it and delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic multi-id), the single owner of the format: a connected set (blocker plus dependents) moves together with blocked-by preserved, item blocks stay byte-exact, and destination section placement holds. The helper keeps only the fleet-level validation tasks-axi cannot know - secondmate-home resolution, the seeded-home safety checks, the In-flight refusal, and idempotent per-key reporting - and is atomic: on any move failure nothing moves. Tests: fm-backlog-handoff.test.sh keeps PR kunchenguid#401's regression matrix but now exercises the delegated path and skips cleanly when tasks-axi is absent; the two whole-file fixtures move to tasks-axi's canonical whitespace. The lifecycle-e2e and safety move-cases gain the same skip guard. CI installs tasks-axi so the delegated path is exercised. Docs state that config/backlog-backend=manual governs firstmate's own hand-editing, not this validated helper, which delegates fleet-wide because bootstrap requires tasks-axi on PATH. Remove the now-redundant WIP design note. * no-mistakes(review): Captain: harden atomic backlog handoffs * no-mistakes(review): Captain: enforce queued-only backlog handoffs * no-mistakes(review): Captain: harden handoff section parsing * no-mistakes(document): Document delegated backlog handoffs * no-mistakes(lint): Silence ShellCheck source diagnostics (cherry picked from commit 31afb8c) * fix: ignore secondmate home marker during sync (kunchenguid#417) * fix: gitignore the secondmate home marker bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into every seeded secondmate home. A secondmate home is a worktree of the firstmate repo, so any plain `git status --porcelain` dirtiness check counted the untracked marker and the home read as dirty forever: fleet-sync reported it STUCK and the local fast-forward convergence sweeps risked leaving it stale on firstmate updates. Add .fm-secondmate-home to the tracked .gitignore so the marker is invisible to every dirtiness check uniformly, without weakening fleet-sync's deliberate untracked-counting for project clones. Convergence chicken-and-egg: existing homes predate the fix and it only arrives by fast-forward. The already-present marker-tolerant ff-skip (ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate, and spawn pre-launch) advances such a home past the fix commit, after which .gitignore takes over - no hand intervention. Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home reading clean, an existing marker-only home converging then reading clean, and a genuinely dirty home still skipping. * no-mistakes(review): Captain: document standalone-clone update path * no-mistakes(document): Document secondmate marker migration (cherry picked from commit 171207c) * fix(composer): prevent dead-shell message injection (kunchenguid#416) * fix(composer): stop reading dead-shell prompts as empty agent composers Consolidate composer empty/pending/unknown classification into one shared owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This replaces four drifting copies of the glyph decision. Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now classified unknown (a dead shell, unsafe for injection), not empty. It is only empty inside a bordered composer box (the harness's own prompt). Agent glyphs ❯ (claude) and › (codex) read empty either way. The away-mode injector (inject_msg) now requires an affirmatively-empty composer, deferring on pending or unknown, so an escalation can never be typed into (or executed by) a pane whose agent exited to its login shell. Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner; per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay green. * no-mistakes(review): Captain: harden composer safety checks * no-mistakes(test): Stabilize Herdr prune safety setup * no-mistakes(document): Document composer injection safety * no-mistakes(lint): Clean composer safety lint * no-mistakes: apply CI fixes (cherry picked from commit a955a05) * feat(watcher): add paused external-wait supervision (kunchenguid#421) * feat(watcher): add paused/awaiting-external crew state A crew (or firstmate steering it) can declare a deliberate wait on a known external dependency with a paused: <reason> status. Both the always-on watcher and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh vocabulary instead of tripping the possible-wedge stale escalation, and re-surface it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused distinctly. A crew that goes idle without declaring a pause classifies exactly as before. Docs and brief scaffold state lists updated; tests colocated. * no-mistakes(review): Captain: fix paused-state transitions * init * no-mistakes(review): Captain: fix paused-state supervision transitions * no-mistakes(review): Captain: fix paused supervision handoffs * no-mistakes(review): Reconcile paused supervision markers * no-mistakes(review): Captain: prioritize paused states over captain relevance * no-mistakes(review): Captain: preserve paused-working wedge timer * no-mistakes(review): Captain: honor configured pause verb in briefs * no-mistakes(test): Captain: fix AFK paused watcher handoff * no-mistakes(document): Document declared external waits * no-mistakes(lint): Clean paused-state lint --------- Co-authored-by: fmtest <fmtest@example.invalid> (cherry picked from commit 7788fa3) * fix: preserve X-mode follow-up platform limits (kunchenguid#425) * fix(x-mode): make follow-up platform splitting immune to link ordering A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280 chars because fm-x-link only learned the platform from the inbox payload, and the fmx-respond ack path can drain that inbox file before the task is linked. A link recorded after cleanup silently lost the platform and the splitter defaulted to the X 280-char budget. Make platform resolution ordering-proof: - fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a new fmx_request_relay_context helper (POST /connector/request-context) when neither the inbox payload nor carry flags carry it. The request_id survives the inbox drain, so a post-cleanup link still learns the right split budget. Best-effort: no token/curl or a non-2xx relay degrades to the loud warning below rather than a silent X default. - fm-x-link warns loudly when no platform source resolves, so the loss is never silent. - The fmx-respond procedure now orders link-before-inbox-cleanup so the fast local path stays correct without a relay round-trip. Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE message even when linked after inbox cleanup, and an unresolvable platform warns loudly instead of splitting silently. docs/configuration.md documents the request-context lookup. The relay endpoint is the companion durable change (see done status); until it ships, the link-before-cleanup reorder keeps the normal path correct. * no-mistakes(document): Document X-mode platform recovery (cherry picked from commit 5f808cc) * fix(composer): handle ANSI ghost text safely (kunchenguid#429) * fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor Away-mode injection wedged all night on the primary claude-on-herdr pane: the herdr composer classifier never stripped generic dim ghost text (only a narrow codex bold-wrapped byte-pattern check), so claude's rotating prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI pane read preserves - read as real pending input and every escalation deferred (6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s). Consolidate ghost extraction into one fleet-wide ANSI-aware owner, fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX, default 128, dark-theme assumption). Both ANSI-capable backends route through it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a dark box border dropped with the ghost does not lose the composer shape. This also closes the documented grok TRUECOLOR placeholder gap by the same mechanism (harness-adapters skill note updated). Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle) and the incident write-up are in docs/herdr-backend.md; deterministic regressions feed the exact captured bytes through the real classifiers (tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text (never a realistic real-input color) are corrected to a bright 38;2;224;222;244, preserving the truecolor payload-skip parser intent. * no-mistakes(review): Preserve dark shell prompt safety * no-mistakes(review): Harden erased shell prompt classification * no-mistakes(document): Document shared composer ghost extraction * no-mistakes(lint): Normalize tmux comment punctuation (cherry picked from commit 0eaf293) * fix(spawn): make tmux window handling robust under non-default config (kunchenguid#134) (cherry picked from commit 6d90240) * test: isolate session-start suite from ambient harness markers (kunchenguid#432) * fix(session-start): isolate harness env markers in suite runner Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in run_session_start so ambient interactive shells cannot override the suite's fake ps harness (local-vs-CI split on the pi supervision case). * no-mistakes(document): Correct Pi marker documentation (cherry picked from commit 3e3dff6) * fix: complete brief help and consolidate documentation (kunchenguid#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage (cherry picked from commit 492c937) * fix: detect Git and centralize backend configuration (kunchenguid#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation (cherry picked from commit bc558c6) * feat(daemon): add backend-independent wedge alerts (kunchenguid#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior (cherry picked from commit 52241a5) * fix(successor): carry reply-platform context through X relinks Integration reconciliation: upstream's platform-aware fm-x-link.sh requires carried reply context on relink, so the successor's X-link carry now forwards the predecessor's x_platform/x_reply_max_chars, defaulting a pre-platform link to the X budget instead of halting the watchdog. * no-mistakes(review): Captain, restore cursor/hermes tmux detection * no-mistakes(test): captain: fix watcher test regressions * test(turnend-guard): restore PR #7 project-dir fallback coverage Rebase reconciliation: re-port the unset-CLAUDE_PROJECT_DIR fallback and physical-identity tests that the Pi logical-run test replay had displaced, on top of the logical-run suite. * no-mistakes(test): Hermeticize turnend guard harness tests * no-mistakes(document): Sync session-start doc comments --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: fmtest <fmtest@example.invalid> Co-authored-by: Pierre Marais <pierremarais67@gmail.com>
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Jul 12, 2026
…te automation (#7) * feat(backends): add experimental cmux runtime backend (#246) * feat(backends): add cmux runtime backend (experimental) Session-provider-only adapter for cmux (bin/backends/cmux.sh), mirroring zellij/herdr structurally, wired into fm-backend.sh and fm-spawn.sh with --secondmate refused for now. Verified against the real cmux 0.64.17 app: send does not auto-submit, cwd is creation-time-frozen (zellij-shape, pwd-marker-probe workaround), close-surface refuses on a workspace's last surface (falls back to close-workspace), workspace ids do not survive a relaunch, and the control socket defaults to cmuxOnly access (requires a one-time password-mode setup, documented in docs/cmux-backend.md). Also found and fixed a live bug during development: read-screen fails on a surface that has never been written to, so liveness now uses list-panes instead. Fake-CLI unit suite (40 tests), a real-binary smoke test, and a full spawn/steer/peek/done/merge/teardown E2E pass against a real claude crewmate all pass, including the popup/second-Enter regression class. * no-mistakes(review): Harden cmux recovery and password parsing * no-mistakes(review): Harden cmux capture failure handling * no-mistakes(review): Mark cmux test scripts executable * no-mistakes(review): Scope cmux workspaces and teardown * no-mistakes(review): Captain, honor cmux password config override * no-mistakes(review): Captain, hash cmux home labels * no-mistakes(document): Sync cmux backend docs * feat(agents): add firstmate coding guidelines skill (#248) * Add firstmate-coding-guidelines skill (AGENTS.md diet PR 0) Encodes the knowledge-placement decision tree, one-owner rule, and inline-stub pattern from the diet analysis so future contributions stop adding conditional detail inline. AGENTS.md gets one section-13 trigger line; fm-brief.sh's REPO argument has no reliable signal for "this is firstmate's own repo", so the load instruction goes in CONTRIBUTING.md's Development section instead of the scaffold. * no-mistakes(review): Captain, align tracked-material trigger scope * no-mistakes(document): Sync coding guidelines docs * no-mistakes(lint): Fix Markdown style issues * fix: add turn-end supervision guard (#249) * feat: structural Stop-hook backstop for primary turn-end supervision fm-guard.sh is pull-based: it only warns when some other supervision script happens to run, so a primary session that ends a turn without re-arming the watcher and then runs no further fleet-touching command can sit blind for hours (the 2026-07-04 incident this fixes). Add bin/fm-turnend-guard.sh, a Claude Code Stop hook registered in the tracked .claude/settings.json, that fires on every primary turn end and blocks (exit 2, verified empirically to force continuation) when work is in flight with no fresh watcher beacon. It never blocks more than once per turn, using Claude Code's own stop_hook_active loop-guard field, and scopes itself to the actual primary checkout only (inert in crewmate/scout worktrees and secondmate homes). Factor the shared "in-flight but no live watcher" predicate out of fm-guard.sh into bin/fm-supervision-lib.sh so the pull-based banner and the push-based hook can never drift on what "unhealthy" means. Document the verified Stop-hook mechanism and scoping in docs/turnend-guard.md, add a harness-adapters note, and cover the predicate and hook with tests/fm-turnend-guard.test.sh. * no-mistakes(review): Respect active home in turnend guard * no-mistakes(review): Require live watcher for turn-end guard * no-mistakes(review): Captain: portable turn-end timing * no-mistakes(document): Sync turn-end guard documentation * feat(backends): auto-detect cmux runtime (#250) * feat(backends): auto-detect cmux runtime from CMUX_WORKSPACE_ID Wires cmux into fm_backend_detect the same way herdr already is: a firstmate process running inside a cmux-spawned terminal now spawns new tasks into cmux by default, no config needed. Verified from cmux's own shipped source that CMUX_WORKSPACE_ID/CMUX_SURFACE_ID/CMUX_SOCKET_PATH are unconditionally, non-overridably injected into every terminal surface it spawns, and that cmux's own CLI treats CMUX_WORKSPACE_ID as its own ambient-target fallback - the same role $TMUX/HERDR_ENV play for their backends. CMUX_WORKSPACE_ID is checked last (after $TMUX and HERDR_ENV=1) since cmux is a terminal application, not a nestable multiplexer. Socket auth (config/cmux-socket-password) stays required regardless of how the backend was selected; the existing spawn refusal now also names the config/backend=tmux / --backend tmux opt-out for a caller who never explicitly chose cmux. A live env dump inside a real cmux terminal was not obtained safely on the shared dev machine (documented in docs/cmux-backend.md); this rests on the source read instead, mirroring this doc's existing verified-from-source precedent. * no-mistakes(review): Fix cmux autodetect docs and tests * no-mistakes(document): Document cmux auto-detection * fix(afk): support herdr away-mode injection (#251) * fix(afk): make the away-mode daemon backend-aware for herdr bin/fm-supervise-daemon.sh discovered its supervisor pane and injected via raw tmux calls only, so /afk failed outright on a herdr-based fleet (TMUX_PANE unset, firstmate:0 fallback unresolvable). Discovery now resolves backend (tmux|herdr) and target independently, mirroring fm-backend.sh's own runtime auto-detection, with an explicit FM_SUPERVISOR_BACKEND override alongside the existing FM_SUPERVISOR_TARGET. zellij/orca refuse loudly at startup instead of misapplying tmux primitives. Injection (pane-exists probe, busy-guard, composer-guard, verified submit) now dispatches through bin/fm-backend.sh's generic primitives, adding a new fm_backend_composer_state dispatcher; the tmux path is byte-identical to before. Also fixes a pre-existing bug in fm_backend_target_exists's herdr arm (missing --session, so it silently misrouted once more than one herdr server was running) found while verifying this end to end against a real isolated herdr session. Classification, batching, max-defer, the marker contract, locks, and wake-queue handling are unchanged - this is a transport-layer fix. * no-mistakes(review): Corroborate Herdr idle busy state * no-mistakes(review): Stabilize Herdr daemon startup wait * no-mistakes(review): Captain, route cmux composer and update AFK docs * no-mistakes(document): Document AFK supervisor backend support * docs(agents): move X-mode procedures out of AGENTS (#253) * docs(agents): collapse X-mode section 14 into fmx-respond/docs pointers AGENTS.md diet PR 1 of 3 (agentsmd-diet-s2 report, move-plan items 1-2). Replaces section 14's "Answering"/"Completion follow-up"/"Conversations"/ "Length and threads"/"Preview / dry-run" blocks (54 lines) and the "Mechanism" narrative (6 lines) with two short pointers: fmx-respond (section 13) for the procedure, docs/configuration.md "X mode (.env)" for the wire protocol. Net -55 lines in AGENTS.md. Destination edits landed first, deletions second (q4 discipline): - docs/configuration.md: added the "purely additive, watcher untouched" guarantee that AGENTS.md's Mechanism block stated but configuration.md did not. - fmx-respond/SKILL.md: added the x-mode-error wake boundary (report as a blocker, do not load this skill), the --image flag for replies and follow-ups, the "images are for real artifacts, not prose" rule, and the dry-run compact-image-marker behavior - none of these were previously in the skill even though AGENTS.md described them, so they were genuine gaps, not pre-existing duplication. Also made the skill's own "Completion follow-up" section the sole, full owner of that procedure instead of deferring to AGENTS.md section 14 for substance that no longer lives there (two internal cross-references updated to point at section 8's terminal-wake trigger and the skill's own section instead). Mechanical line-by-line audit of every removed AGENTS.md line: Mechanism block (6 lines removed): - bootstrap artifact-writing description -> already owned by docs/configuration.md "X mode (.env)" (locked-bootstrap paragraph) - check-shim/poll mechanism description -> already owned by docs/configuration.md same section - missing-deps/x-mode-error diagnostic description -> already owned by docs/configuration.md ("Relay auth or config problems...") plus bin/fm-x-poll.sh's own header comment for the missing-curl/jq mechanics - opt-out artifact removal description -> already owned by docs/configuration.md same section - "purely additive, no edit to fm-watch.sh/fm-watch-arm.sh/fm-wake-lib.sh/ afk daemon" guarantee -> MOVED to docs/configuration.md (added in this PR; this fact had no other home before) Answering/Completion follow-up/Conversations/Length and threads/ Preview-dry-run blocks (54 lines removed): - x-mention wake -> load fmx-respond: already owned by section 13's existing trigger line (unchanged) and restated in the new pointer - x-mode-error wake -> report as blocker, don't load fmx-respond: MOVED to fmx-respond/SKILL.md (added in this PR) - inbox-draining, classification, acting, reply composition, submission, cleanup-on-success/failure: already owned by fmx-respond/SKILL.md "Procedure" section (unchanged, pre-existing) - owner-only routing / captain-as-asker framing: already owned by fmx-respond/SKILL.md "The asker is your own captain" section - standing X-mode authorization / autonomous posting / dry-run as only non-posting path: already owned by fmx-respond/SKILL.md same section - acknowledge-first -> act -> follow-up shape, three-case classification: already owned by fmx-respond/SKILL.md "A request to act on" section - destructive/irreversible/security-sensitive escalation guardrail: already owned by fmx-respond/SKILL.md "Public channel..." section and Procedure step 2c - dismiss-instead-of-reply for pure acknowledgments, relay re-offer prevention, dry-run honoring: already owned by fmx-respond/SKILL.md Procedure steps 2b/2c/2e-skip and docs/configuration.md - public-safety bar (no task ids/internals/captain-private/secrets): already owned by fmx-respond/SKILL.md "The reply is public" section - never-inline-into-shell-command / --text-file or stdin: already owned by fmx-respond/SKILL.md Procedure step 2e and Notes - --image flag for replies (formats, base64, no-inline guarantee): MOVED to fmx-respond/SKILL.md Procedure step 2e (added in this PR - this was not previously in the skill) - fm-x-link field names (x_request=, x_request_ts=, x_followups=): already owned by AGENTS.md section 2's state/<id>.meta field list (untouched, out of scope for this PR) and fmx-respond/SKILL.md - carry-count/carry-ts relink behavior, three-follow-up budget, milestone sparingness, --check/--text-file posting, connector/followup wire detail, --final clearing, cap/window graceful degradation: already owned by fmx-respond/SKILL.md "Completion follow-up" section (now sole owner) and docs/configuration.md wire-protocol paragraphs - --image flag for follow-ups: MOVED to fmx-respond/SKILL.md "Completion follow-up" section (added in this PR - genuine gap) - "failed task still gets an honest final follow-up": already owned by fmx-respond/SKILL.md "Completion follow-up" section - FMX_DRY_RUN whole-loop previewability: already owned by fmx-respond/SKILL.md "Dry-run / preview mode" section - in_reply_to conversation continuity, untrusted-thread handling, follow-up worthiness judgment, relay-owned self-reply guard/cap: already owned by fmx-respond/SKILL.md "The direct ask is the captain's" section and Notes (one bullet is a verbatim match) - concise-by-default / no hand-numbered threads: already owned by fmx-respond/SKILL.md "Voice" section - auto-split behavior, char/tweet caps, premium-independence, wire shape ({text}/{text,texts}): behavior already owned by fmx-respond/SKILL.md Voice section; exact defaults and wire shape already owned by docs/configuration.md; "premium-independent" mechanics already owned by bin/fm-x-reply.sh's own header comment - "images are for real artifacts, not prose": MOVED to fmx-respond/ SKILL.md "Voice" section (added in this PR - genuine gap) - image-on-thread wire behavior: already owned by docs/configuration.md; reinforced in fmx-respond/SKILL.md's new --image note - dry-run POST-body shape, endpoint marker, truthy-value definition, jq-only dependency, end-to-end testability, x-outbox inspection: already owned by fmx-respond/SKILL.md "Dry-run / preview mode" section (several near-verbatim matches) and docs/configuration.md wire detail - dry-run compact image marker: MOVED to fmx-respond/SKILL.md "Dry-run / preview mode" section (added in this PR - genuine gap) Section 8's terminal-wake completion-follow-up trigger (the one fact required to survive inline) is untouched and already present; the new section 14 pointer references it instead of restating it. Nothing outside section 14 (plus the two destination files) is touched. Full test suite green, including all 74 fm-x-mode.test.sh checks. * no-mistakes(review): Preserve X-linked follow-up triggers * no-mistakes(review): Fix x-mode error trigger * no-mistakes(document): Docs cross-reference synchronized * no-mistakes(lint): Clean Markdown lint pass * docs: trim duplicated harness guidance (#255) * docs(agents): trim section 4 harness/secondmate duplication AGENTS.md diet PR 2 of 3 (data/agentsmd-diet-s2/report.md, move-plan items 3-4; redundancy item 2 folded into item 3). Removed the five claude/codex/grok/pi/opencode model/effort-flag bullets from section 4 - byte-for-byte duplicated by harness-adapters' "Launch profile axes" table (which is already a superset: it carries verified CLI versions per adapter that the AGENTS.md bullets lacked). Replaced with a one-line pointer; the skill is already loaded before every spawn per section 4's own closing trigger, so no new trigger was needed. Moved the config/secondmate-harness model/effort pin-format detail (the `<harness> [<model>] [<effort>]` line format, the secondmate-model/secondmate-effort accessors, back-compat, and the durability-across-respawn behavior) into secondmate-provisioning, which is already a mandatory load at every secondmate lifecycle touchpoint. Added the destination content to the skill first, then replaced the AGENTS.md paragraph with a 3-line pointer. Mechanical audit - every removed line's new home: - 5 harness bullets (claude/codex/grok/pi/opencode model+effort flags, per-harness max-omission rationale) -> already present in harness-adapters SKILL.md's "Launch profile axes" table (lines 53-59), confirmed fact-by-fact before deleting. - "config/secondmate-harness may also pin..." paragraph (pin format, bare-harness back-compat, secondmate-model/secondmate-effort accessors, per-spawn override precedence, respawn durability, secondmate-only scope) -> secondmate-provisioning SKILL.md's "Charter and seed" section, added verbatim before this trim. - The following paragraph (inheritable config: crew-dispatch.json, crew-harness, backlog-backend) is untouched - out of scope for this PR, still inline. - The bootstrap CREW_DISPATCH effort-mismatch diagnostic sentence is untouched - not part of the five-bullet duplication, stays inline. No script changes. Section 4 shrinks from 104 to 87 lines (958 -> 901 total AGENTS.md lines) with zero facts lost: every fact is reachable through harness-adapters or secondmate-provisioning, both already mandatory loads at the relevant lifecycle points. * no-mistakes(document): Align secondmate skill triggers * no-mistakes(lint): Markdown style clean * fix: anchor turn-end Stop hook to project root (#256) * Fix turn-end Stop hook to use CLAUDE_PROJECT_DIR path Claude Code runs hook commands via /bin/sh from the session cwd, so the bare relative bin/fm-turnend-guard.sh path fails when cwd is not the repo root. Anchor the command with "$CLAUDE_PROJECT_DIR"/bin/fm-turnend-guard.sh instead; verified CLAUDE_PROJECT_DIR is set on Stop hooks in Claude Code 2.1.201. Document the cwd caveat and add a settings.json regression test. * no-mistakes(document): Document Stop hook path anchoring * docs: trim firstmate agent guidance duplication (#258) * docs: trim AGENTS.md redundancy (diet PR 3/3) Consolidates five duplicated passages to a single owner each, per data/agentsmd-diet-s2/report.md redundancy items c3-c7: - Inheritable-config propagation mechanism: owned by section 3 (where the sweep runs); sections 4 and 7 keep compact references. Section 4 retains its one genuinely unique fact (crew-harness inherit-vs-fallback semantics), just no longer restates the propagation mechanism itself. - Landed-work definition: owned by section 7's ship-teardown detail (PR-containment mechanics, pr= discovery fallback); section 1's hard rule #3 keeps the rule plus a three-case summary and a pointer. - Backend meta-field enumeration: owned by docs/configuration.md ("Runtime backend", already comprehensive including cmux) and each backend's own doc; AGENTS.md keeps only the fields common to every task plus a pointer. - Dropped one redundant restatement of "silence is correct while waiting" in section 8. - Worktree-tangle guard explanation: owned by section 8 (already the fuller, cross-referenced version); section 3's TANGLE bullet keeps the remediation action and points at section 8 for the why. Also adds two captain-requested single-sentence rules: invoke bin/ scripts by absolute $FM_ROOT path after any cd away from the home, and a backend spawn refusal must be surfaced to the captain rather than silently worked around by switching backends. AGENTS.md: 901 -> 889 lines, 112355 -> 108560 bytes. * no-mistakes(review): Clarify post-cd bin invocation guidance * no-mistakes(document): Sync AGENTS trim docs * no-mistakes(lint): Fix Markdown line style * feat(backends): improve cmux detection and socket-mode guidance (#259) * feat(backends): cmux detection fallbacks and socket-mode matrix Workstream A: cmux's bundled claude wrapper strips every CMUX_* env var on its passthrough path (reproduced live 2026-07-04, cmux 0.64.17), so a claude-harness firstmate inside a cmux tab has no CMUX_WORKSPACE_ID. fm_backend_detect now falls back - macOS-only, only when the primary marker is absent - to __CFBundleIdentifier=com.cmuxterm.app and then a process ancestry walk resolved by bundle id (lsappinfo) plus a bundle-shaped ps comm match. Innermost-first ordering is unchanged and absorbs the tmux-inside-cmux bundle-id false positive; the auto-detect NOTICE names the winning fallback signal. Workstream B: the five socketControlMode values were traced through cmux source (commit 9c91710e3f58): off/cmuxOnly can never admit an external CLI, automation admits same-user clients with no secret (0600 socket only), password needs the auth handshake, allowAll opens the socket to every local user (0666). Automation mode is now the documented recommendation; the adapter's refusals name every viable mode, classify Invalid password as unauth, and the launch-timeout message names the off-mode possibility. Docs carry the wrapper-strip empirical record, the fallback contract and authority split, and the full mode matrix with rationale; tests cover the new detection paths, the nested false positive, and the refusal wording. * no-mistakes(review): Document cmux fallback detection * no-mistakes(review): Update cmux architecture docs * no-mistakes(document): Align cmux backend docs * fix(backends): scope zellij tabs by firstmate home (#252) * fix(backends): home-scope zellij tab titles to close cross-home collision gap Zellij's one shared "firstmate" session has no per-home split and enforces no tab-name uniqueness, so two firstmate homes with colliding task ids could send/peek/close each other's tabs - the same gap a no-mistakes review gate caught for cmux (docs/cmux-backend.md). Ports that fix: every new tab is created with a home-scoped title (fm-<home-label>-<id>), and every list/find/recover/kill path scopes matches to this home's own tag. A tab spawned before this change still matches via its old untagged bare title, but only when unambiguous - two live tabs sharing a bare title refuse rather than guessing which one is ours. Factors the home-label/hash derivation shared with cmux into bin/fm-backend-hometag-lib.sh so the two adapters can't drift. * no-mistakes(review): Fix zellij child teardown home tag * no-mistakes(review): Fix zellij teardown and selector scoping * no-mistakes(document): Sync zellij home-scope docs * fix: sync project clones after merged PR wakes (#293) * fix(fleet-sync): auto-sync on merged-PR wake, accept project name fm-fleet-sync.sh's single-project form failed on a bare project name ("not a directory"), forcing hand-typed full paths (4 manual runs in one incident). It now resolves a bare name or projects/<name> against the home's projects dir. AGENTS.md now encodes the trigger: a wake whose status reports a merged PR for a project cloned in this home runs fleet-sync for that project as part of handling the wake, so a secondmate-reported merge does not leave the primary's clone stale until the next session start or teardown. * no-mistakes(review): Fix fleet-sync project name shadowing * no-mistakes(document): sync fleet-sync docs * fix: canonicalize spawn worktree path checks (#294) * fix(spawn): canonicalize worktree-isolation guard against symlinked project prefixes fm-spawn.sh compared a logical PROJ_ABS against the physically-resolved pane cwd every backend reports, so a project reached through a symlinked prefix (e.g. macOS's /tmp -> /private/tmp) could trip the isolation guard's false refusal before treehouse ever moved the pane. Canonicalize once into PROJ_ABS_REAL and compare against that everywhere instead. * no-mistakes(review): Canonicalize spawn cwd comparisons * no-mistakes(document): Refresh symlinked spawn docs * docs: add Orca operator skill (#276) * docs: add Orca operator skill * no-mistakes(document): Document Orca checklist --------- Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net> * fix: surface green PRs during CI monitoring (#297) * fix(crew-state): detect green-PR CI monitoring, escalate repeat wedges fm-crew-state.sh's ci step never distinguishes "still waiting on checks" from "checks green, waiting on merge" via axi status alone, since a repo that defers merge to the captain keeps the ci step at status=running for the whole monitor phase. Read the ci step's own log tail (axi logs) for the checks-passed marker and surface done instead of a false "validating (running)" - verified against the real PR #252 run's ci.log. The watcher's wedge timer can re-escalate the same stale pane forever without ever signaling that it is a repeat; track a per-pane consecutive escalation count and add a demand-deep-inspection marker to the wake payload once it crosses a threshold, so the supervisor can no longer dismiss each one as an isolated, still-validating pane. Also clarify the ship-brief's checks-green line: it is owed at the CI-ready return point, not after the background monitor-until-merge loop finishes. * no-mistakes(review): Captain, distinguish pending no-checks CI marker * no-mistakes(review): Harden CI relapse handling * no-mistakes(review): Block stale done during fixing * no-mistakes(review): Captain, tighten CI status gating * no-mistakes(review): Captain, harden stale CI green handling * no-mistakes(review): Captain, recognize ranged CI rearm markers * no-mistakes(document): Sync crew-state supervision docs * fix(teardown): recover provably stale git index locks (#296) * fix(teardown): recover from a stale worktree git index.lock A crew process killed mid-git-operation can leave a stale .git/worktrees/<wt>/index.lock behind, making fm-teardown.sh's `treehouse return --force` fail closed. On that failure, retry once after a short wait (the owning process may be exiting), then remove the lock and retry once more only when it is provably stale: old enough by mtime and lsof shows no live holder on the lock or the worktree itself. A lock that isn't provably stale is left in place and the original failure still surfaces. * no-mistakes(review): Harden teardown lock refusal paths * no-mistakes(review): Harden stale-lock teardown safety rechecks * no-mistakes(review): Harden stale teardown lock checks * no-mistakes(document): Document teardown lock recovery * feat(bin): encode project AGENTS.md authoring bar with canonical self-governance section (#307) * Encode project AGENTS authoring bar * no-mistakes(review): Captain, centralize CLAUDE promotion governance * no-mistakes(review): make ensure_maintenance_section idempotent-success, drop || true guards * no-mistakes(review): separate appended maintenance section on newline-less CLAUDE.md promotion * no-mistakes(review): assert maintenance heading present before separator check in test * no-mistakes(document): sync docs with AGENTS.md authoring bar and self-governance --------- Co-authored-by: fmtest <fmtest@example.invalid> * feat(skills): add captain-invocable bearings status-report skill (#300) * Add captain-invocable bearings skill Generates a pick-up-where-I-left-off status report from live fleet state to data/status-report-<YYYY-MM-DD>.md plus a concise chat summary. Read-mostly procedure: reads backlog, per-task crew state via bin/fm-crew-state.sh, open PRs via gh-axi, scout reports, pending decisions, and date-gated queued work; composes the exemplar's sections (TL;DR, Check first, Landed, In flight, Plans, Decisions pending, Date-gated/queued); never tears down, merges, or mutates task state as a side effect. * no-mistakes(document): docs: list new /bearings skill in README built-in skills table * fix(watcher): make PID identity locale-invariant (#285) * fix(watcher): pin LC_ALL=C in fm_pid_identity for locale-invariant identity ps's lstart date format follows the caller's LC_TIME/LC_ALL. The watcher records its process identity under one locale, but arm/guard/turn-end re-read it under the machine's ambient locale. On a non-C locale (e.g. ko_KR) the two strings differ only in the date portion, so fm_watcher_lock_matches_pid / fm_watcher_healthy reject a genuinely live watcher - breaking fm-watch-arm.sh, fm-guard.sh, and fm-turnend-guard.sh on every non-C-locale machine. Pin LC_ALL=C on that one ps call so the write and read sides agree regardless of machine locale, matching the LC_ALL=C determinism the file already uses elsewhere. Add a colocated regression test asserting fm_pid_identity is locale-invariant across exported LC_ALL/LC_TIME. * no-mistakes(document): Document watcher PID identity coverage * docs: document codex app backend contract (#222) * docs: reconcile Codex App backend contract * no-mistakes(document): Sync backend docs * docs: clarify Codex Desktop bridge blocker * no-mistakes(document): Align Codex App backend docs * no-mistakes(test): Captain, stabilize watcher self-eviction test cadence * no-mistakes(document): Document Codex App backend contract * no-mistakes(document): Captain, document blocked codex-app coverage * docs: make Codex App contract doc authoritative * no-mistakes(document): Align Codex App backend docs * docs: redact local Codex App smoke paths --------- Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net> * docs: add Codex Desktop coordination skill (#275) * docs: add Codex App coordination skill * no-mistakes(review): Captain, mark Codex App skill agent-only * no-mistakes(document): Document Codex App backend boundary * no-mistakes(document): Captain, document Codex Desktop backend boundary * no-mistakes(lint): Captain, lint clean * no-mistakes(document): Document Codex Desktop boundaries * docs: narrow Codex App skill playbook * fix(afk): stop herdr escalation redelivery loop (#317) * fix(afk): recognize unbordered herdr composer rows to stop escalation redelivery loop fm_backend_herdr_composer_state only recognized bordered composer rows (the grok shape). Real claude and codex render their live input row with no border at all, so once a harness's own startup banner scrolled out of the capture window the classifier read the composer as unknown forever. fm_backend_herdr_send_text_submit never confirmed "empty", so escalate_flush never cleared state/.subsuper-escalations, and the away-mode daemon retyped and resubmitted the same buffered digest every housekeeping cycle - reproduced live against a real herdr+claude pane (5+ identical deliveries in 40s). The classifier now recognizes an unbordered (bare) composer row led by a known prompt glyph alongside the existing bordered shape, keeping whichever match is bottom-most so a stale decorative box never outranks the live composer. * no-mistakes(review): Narrow herdr bare prompt matcher * no-mistakes(document): Sync herdr composer docs * fix(backends): confirm Herdr submits with native agent state (#323) * fix(herdr): confirm message submit via native agent-state, not composer text fm_backend_herdr_send_text_submit now confirms a landed submit by polling herdr's own agent-state (agent get) for the idle->working transition instead of reading composer content. Composer scraping remains, unchanged, for the away-mode daemon's pre-injection empty-box guard only. This fixes the practical effect of the codex idle-tip gap from the 2026-07-07 incident: codex's dynamic idle-composer hint text can no longer misread as pending and block/mis-confirm a send, since confirmation no longer looks at composer text at all. Verified empirically against real claude and codex agents (timing, swallowed-Enter, unreadable-target, and already-busy-target scenarios), and against the real away-mode daemon end-to-end after updating its synthetic supervisor-pane test fixture to register itself as a real herdr agent (herdr's own report-agent primitive) so it can still exercise the new confirmation path. * no-mistakes(review): Captain, harden herdr submit confirmation * no-mistakes(review): Captain, harden herdr submit confirmation * no-mistakes(document): Sync Herdr submit docs * no-mistakes: apply CI fixes * feat: add quota-balanced crew dispatch selection (#327) * Add quota-balanced dispatch selection * no-mistakes(document): Document dispatch selector guidance * fix(session-start): respawn dead secondmate agents conservatively * fix(session-start): deterministically respawn dead-shell secondmates A secondmate agent that exits leaves its backend pane alive as a bare shell. The session-start endpoint check only verified pane presence, so recovery and the watcher (which exempts secondmates from stale-pane detection) never noticed - evidence 2026-07-07: every secondmate in one fleet was found sitting at a dead zsh shell. Add fm_backend_agent_alive (bin/fm-backend.sh), a deeper per-backend liveness probe distinct from pane presence: fm_backend_tmux_agent_alive classifies the pane's live foreground process via tmux's own pane_current_command, and fm_backend_herdr_agent_alive reuses the already-verified pane_agent_state husk classifier. Both are conservative: anything ambiguous reports unknown, never a false dead. Wire this into a new session-start-only, locked-and-primary-only sweep in bin/fm-bootstrap.sh that kills and respawns only a confidently dead secondmate endpoint, leaving alive/unknown readings untouched - idempotent by construction, so repeated runs converge without duplicating agents. * no-mistakes(review): Guard raw secondmate liveness respawns * no-mistakes(review): Fix detect-only bootstrap test * no-mistakes(test): Pin liveness fixture harness * no-mistakes(document): Sync secondmate liveness docs * no-mistakes: apply CI fixes * fix: emit stable secondmate nudge selectors (#331) * Fix NUDGE_SECONDMATES to print stable fm-<id> selectors. Session-start secondmate sync used to accumulate raw backend window targets into NUDGE_SECONDMATES, but the liveness sweep in the same bootstrap run can respawn secondmates onto new endpoints. fm-send with those stale explicit targets bypasses meta resolution and fails, while fm-<id> resolves correctly. Accumulate fm-<id> in process_secondmate, update the bootstrap/update contracts and /updatefirstmate skill, and add a herdr respawn regression test. * no-mistakes(review): Captain, guard herdr regression jq dependency * no-mistakes(document): Document stable secondmate nudge selectors * no-mistakes(lint): Fix shell lint hints * feat: require bootstrap detection for AXI tools (#332) * Make tasks-axi and quota-axi required bootstrap tools Add both to the normal toolchain checks alongside lavish-axi, keep the tasks-axi 0.1.1+ compatibility gate, and report quota-axi through the standard MISSING install-consent flow. TASKS_AXI: available remains a backlog-backend capability signal only; manual opt-out no longer suppresses the missing-tool report. Update bootstrap tests and point docs/configuration.md at the canonical toolchain contract. * no-mistakes(review): Clarify manual backlog bootstrap reporting * no-mistakes(document): Document bootstrap AXI tools * bearings: delete today's report before recreating (#333) Replace overwrite-in-place wording with explicit delete-then-create instructions so agents do not modify an existing daily report file. * feat: guard primary turn ends across harnesses (#339) * Add primary turn-end guards for all harnesses * no-mistakes(review): Normalize Codex hook cwd resolution * no-mistakes(review): Fix OpenCode guard worktree anchoring * no-mistakes(review): Anchor Codex guard outside nested roots * no-mistakes(review): Anchor Codex guard to hook root * no-mistakes(review): Avoid Grok permission escalation * no-mistakes(document): Sync turn-end guard docs * fix: resolve backend selectors by exact task id first (#342) * fix backend selector task id resolution * no-mistakes(document): Document selector resolution behavior * fix: scale bootstrap fleet-sync timeout (#341) * fix bootstrap fleet sync timeout * no-mistakes(review): Fix bootstrap fleet-sync timeout regressions * no-mistakes(document): Sync bootstrap timeout docs * no-mistakes(lint): Clean ShellCheck directives * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat: add fleet snapshot and view commands (#343) * Add fleet snapshot and view * no-mistakes(review): Fix fleet snapshot parsing and overrides * no-mistakes(review): Fix secondmate fleet rendering * no-mistakes(review): Fix backlog title and completion parsing * no-mistakes(review): Include durable scout reports * no-mistakes(review): Fix fleet snapshot edge cases * no-mistakes(review): Captain: gate fleet hints on current state * no-mistakes(review): Captain: parse bracketed Done PR artifacts * no-mistakes(document): Sync fleet snapshot docs * fix(fm-send): fail loudly on unresolvable send targets (#254) * Make fm-send fail loudly on unresolved targets * no-mistakes(review): Document fm-send FM_HOME contract * Fix fm-send readiness docs and backend send path * Fix fm-send docs for cmux and X skill metadata * Make gotmp teardown test home-explicit * Scope watcher warning wording to fm-send * Fix fm-send review findings * Verify explicit tmux targets before sending * Isolate turnend guard test home * no-mistakes(document): Documented fm-send FM_HOME/backend guard additions missing from doc inventories --------- Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> * fix: deliver AFK escalations through herdr supervisors (#353) * fix afk codex ghost composer delivery * no-mistakes(review): Harden AFK startup flag writes * no-mistakes(review): Harden AFK daemon liveness checks * no-mistakes(document): Sync AFK herdr docs * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * feat: add harness-aware supervision (#367) * Add harness-aware supervision * no-mistakes(review): Captain, harden watcher supervision regressions * no-mistakes(review): Captain, harden watcher supervision cadence * no-mistakes(review): Harden watcher supervision ownership * no-mistakes(review): Captain, harden Pi extension marker * no-mistakes(review): Captain, harden Pi supervision restart checks * no-mistakes(review): Harden watcher ownership checks * no-mistakes(review): Captain, harden Pi supervision loading * no-mistakes(review): Captain, require Pi guard extension loading * no-mistakes(review): Captain, harden watcher supervision recovery * no-mistakes(test): Fix fm-send baseline log filtering * no-mistakes(document): Sync harness supervision docs * no-mistakes: apply CI fixes * fix: split X-mode replies by platform (#369) * fix: make x replies split by platform * no-mistakes(review): Captain: preserve Discord recovery relink context * no-mistakes(test): Captain: keep split markers outside fences * no-mistakes(document): Sync X-mode reply docs * fix: make stow memory writes inspect before update (#372) * docs: make stow inspect-then-update * no-mistakes(review): Remove unsupported archive-body guidance * no-mistakes(review): Clarify stow read-before-write exception * no-mistakes(test): Require archive-body for stow task notes * no-mistakes(document): Sync stow memory docs * no-mistakes(lint): Silence ShellCheck source warning * fix(watcher): wait when arm attaches to a healthy watcher (#375) * fix: attach-and-wait when arm finds a healthy watcher Grok and Claude re-arm after every turn with work in flight. When a watcher was already healthy, fm-watch-arm exited immediately with watcher: healthy, which completed the harness background task and injected an empty false wake. Attach to the live identity-matched holder instead, stay until that cycle ends, then exit 0 so notify fires for a real end-of-cycle. The peer-startup-race path uses the same contract. --restart and the started path are unchanged. * no-mistakes(review): Gate restart watcher peer attach * no-mistakes(document): Sync watcher arm docs * feat(pi): simplify primary session launch (#386) * docs(readme): reformat Quick Start and recommend Grok equally with Claude Code * no-mistakes(review): Captain: align harness launch guidance * no-mistakes(review): Captain, clarify Pi supervised launch * no-mistakes(review): Captain, document Pi first-launch bridge * feat(pi): track primary watcher extension for plain-pi launch Move Pi's primary watcher bridge from a generated state/ file to a tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end guard extension already works: self-hashing version, project-local auto-discovery after one-time Pi trust. This drops the state/-generation step and dual -e requirement from the happy path, so Pi's Quick Start launch becomes plain 'pi', the same friction class as 'claude' and 'grok --trust'. - bin/fm-pi-watch-extension.sh is removed; nothing generates the extension anymore since it is committed. - fm-session-start.sh and fm-supervision-instructions.sh resolve the watcher extension path from FM_ROOT instead of state/, and the session-start diagnostic now points at restarting plain pi after trust, with -e as a documented fallback. - fm-spawn.sh points Pi secondmate launches at the tracked extension path in the secondmate home instead of generating a state/ copy. - README Quick Start Pi block is now just 'pi' plus a trust note. - Tests, docs, and the harness-adapters skill updated to match. * fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016 * feat(supervision): prevent unsafe watcher-arm commands (#387) * feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that denies a primary shell command backgrounding, piping, or bundling the watcher arm/checkpoint, or force-killing the watcher process broadly - the exact shapes that silently took Grok's supervision down. Wires it into all five verified harnesses (grok, claude, codex, opencode, pi), each validated empirically against the real harness. Also fixes a grok 0.2.93 regression discovered during that validation: the existing turnend-guard Stop hook's bare root variable broke grok's own variable pre-substitution and silently no-op'd the hook. * no-mistakes(review): Harden watcher arm validation * no-mistakes(review): Harden arm guard metacharacter checks * no-mistakes(review): Harden nested shell arm guard * fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then A && B || C is not if-then-else; C can run when A is true. Replace both occurrences of the quote-state early-continue with an explicit if/then. * fix(pi): restore primary watcher supervision lifecycle (#397) * fix Pi primary supervision lifecycle * no-mistakes(document): Synchronize Pi primary extension documentation * fix: keep persistent secondmates out of the main backlog (#398) * fix secondmate backlog guidance * no-mistakes(review): Require reasons for captain backlog holds * no-mistakes(test): Document secondmate handoff skill requirement * fix secondmate teardown reminder * no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract * fix(backlog-handoff): move full item blocks including indented bodies (#401) * fix(backlog-handoff): move full item blocks including indented bodies fm-backlog-handoff only moved the checklist header line, so multi-line item bodies were left orphaned in the source backlog and never reached the secondmate. Move the full block (header plus indented body lines) atomically, treating body membership by indentation so lines like ## Intent stay with the item, and add regression coverage. * no-mistakes(review): Captain: preserve EOF handoff terminators * no-mistakes(review): treat blank lines inside item bodies as movable body * no-mistakes(document): sync backlog-handoff docs with full-block move behavior * feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402) * guard Herdr lab lifecycle in briefs * no-mistakes(review): Fix Herdr lab helper and provisioning safety * no-mistakes(review): Captain, harden Herdr lab lifecycle safety * no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range * no-mistakes(review): reject leading options in Herdr lab run guard * no-mistakes(review): strip leading non-alnum in Herdr lab name generator * no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag * no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab * no-mistakes: apply CI fixes * fix(watcher): classify arm-command seatbelt by execution position (#403) * fix watcher arm command policy * no-mistakes(review): Harden watcher command policy parsing * no-mistakes(review): Captain: harden watcher policy parsing * no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter * no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps * no-mistakes(review): fail closed on loop-wrapped broad watcher kills * no-mistakes(document): sync docs for watcher-arm command-position policy * fix: reconcile existing AGENTS.md safely (#405) * fix(agents-md): inject self-governance section into existing AGENTS.md fm-ensure-agents-md.sh only appended the canonical "## Maintaining this file" section on skeleton create or CLAUDE.md promotion, so an existing AGENTS.md that lacked it exited unchanged and forced hand-copying the wording during a rollout across existing projects. Call the already- idempotent ensure_maintenance_section on the existing-AGENTS.md paths and report whether the file changed; a re-run and an already-complete file stay byte-identical. Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase literal target dangles once the tree lands on a case-sensitive filesystem. Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and CLAUDE.md-promotion regressions still pass. Docs updated to match. * no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency * no-mistakes(review): Preserve CRLF during maintenance-section injection * no-mistakes(review): Captain: harden dangling-symlink regression coverage * no-mistakes(document): Document agent-memory injection outcomes * feat: support project-less secondmate homes (#409) * feat(secondmate): support project-less homes via --no-projects fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit --no-projects signal to scaffold, seed, and register a secondmate home whose subject is the firstmate repo itself (no clones). The signal is mutually exclusive with a project list; omitting both still fails loudly so an accidental omission is never a silent project-less seed. The registry line renders an empty projects: field, which spawn and the snapshot already tolerate. Docs updated in the secondmate-provisioning skill and both script headers. * no-mistakes(review): Captain: document project-less secondmate flow * no-mistakes(review): Captain: refuse project-less reseeding of populated homes * fix(seed): fail closed on unreadable project data * no-mistakes(review): Captain: reject stale projectful charters * no-mistakes(review): Captain: fail closed on unsafe project paths * no-mistakes(review): Captain: validate project-less charter clone sections * no-mistakes(document): Document project-less secondmate seeding * fix: delegate backlog handoffs to tasks-axi (#411) * wip(handoff): record verified delegation design + tasks-axi mv blocker No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically move a blocked-by-linked item set across backlogs (deadlocks both orders, no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked pending a tasks-axi connected-set mv enhancement; note captures the verified design, semantics, test/CI/doc changes, and resume checklist. * refactor(handoff): delegate the item move to tasks-axi mv fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog format and the source of the PR #401 body-orphaning drift. Delete it and delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic multi-id), the single owner of the format: a connected set (blocker plus dependents) moves together with blocked-by preserved, item blocks stay byte-exact, and destination section placement holds. The helper keeps only the fleet-level validation tasks-axi cannot know - secondmate-home resolution, the seeded-home safety checks, the In-flight refusal, and idempotent per-key reporting - and is atomic: on any move failure nothing moves. Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now exercises the delegated path and skips cleanly when tasks-axi is absent; the two whole-file fixtures move to tasks-axi's canonical whitespace. The lifecycle-e2e and safety move-cases gain the same skip guard. CI installs tasks-axi so the delegated path is exercised. Docs state that config/backlog-backend=manual governs firstmate's own hand-editing, not this validated helper, which delegates fleet-wide because bootstrap requires tasks-axi on PATH. Remove the now-redundant WIP design note. * no-mistakes(review): Captain: harden atomic backlog handoffs * no-mistakes(review): Captain: enforce queued-only backlog handoffs * no-mistakes(review): Captain: harden handoff section parsing * no-mistakes(document): Document delegated backlog handoffs * no-mistakes(lint): Silence ShellCheck source diagnostics * fix: ignore secondmate home marker during sync (#417) * fix: gitignore the secondmate home marker bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into every seeded secondmate home. A secondmate home is a worktree of the firstmate repo, so any plain `git status --porcelain` dirtiness check counted the untracked marker and the home read as dirty forever: fleet-sync reported it STUCK and the local fast-forward convergence sweeps risked leaving it stale on firstmate updates. Add .fm-secondmate-home to the tracked .gitignore so the marker is invisible to every dirtiness check uniformly, without weakening fleet-sync's deliberate untracked-counting for project clones. Convergence chicken-and-egg: existing homes predate the fix and it only arrives by fast-forward. The already-present marker-tolerant ff-skip (ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate, and spawn pre-launch) advances such a home past the fix commit, after which .gitignore takes over - no hand intervention. Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home reading clean, an existing marker-only home converging then reading clean, and a genuinely dirty home still skipping. * no-mistakes(review): Captain: document standalone-clone update path * no-mistakes(document): Document secondmate marker migration * fix(composer): prevent dead-shell message injection (#416) * fix(composer): stop reading dead-shell prompts as empty agent composers Consolidate composer empty/pending/unknown classification into one shared owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This replaces four drifting copies of the glyph decision. Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now classified unknown (a dead shell, unsafe for injection), not empty. It is only empty inside a bordered composer box (the harness's own prompt). Agent glyphs ❯ (claude) and › (codex) read empty either way. The away-mode injector (inject_msg) now requires an affirmatively-empty composer, deferring on pending or unknown, so an escalation can never be typed into (or executed by) a pane whose agent exited to its login shell. Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner; per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay green. * no-mistakes(review): Captain: harden composer safety checks * no-mistakes(test): Stabilize Herdr prune safety setup * no-mistakes(document): Document composer injection safety * no-mistakes(lint): Clean composer safety lint * no-mistakes: apply CI fixes * feat(watcher): add paused external-wait supervision (#421) * feat(watcher): add paused/awaiting-external crew state A crew (or firstmate steering it) can declare a deliberate wait on a known external dependency with a paused: <reason> status. Both the always-on watcher and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh vocabulary instead of tripping the possible-wedge stale escalation, and re-surface it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused distinctly. A crew that goes idle without declaring a pause classifies exactly as before. Docs and brief scaffold state lists updated; tests colocated. * no-mistakes(review): Captain: fix paused-state transitions * init * no-mistakes(review): Captain: fix paused-state supervision transitions * no-mistakes(review): Captain: fix paused supervision handoffs * no-mistakes(review): Reconcile paused supervision markers * no-mistakes(review): Captain: prioritize paused states over captain relevance * no-mistakes(review): Captain: preserve paused-working wedge timer * no-mistakes(review): Captain: honor configured pause verb in briefs * no-mistakes(test): Captain: fix AFK paused watcher handoff * no-mistakes(document): Document declared external waits * no-mistakes(lint): Clean paused-state lint --------- Co-authored-by: fmtest <fmtest@example.invalid> * fix: preserve X-mode follow-up platform limits (#425) * fix(x-mode): make follow-up platform splitting immune to link ordering A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280 chars because fm-x-link only learned the platform from the inbox payload, and the fmx-respond ack path can drain that inbox file before the task is linked. A link recorded after cleanup silently lost the platform and the splitter defaulted to the X 280-char budget. Make platform resolution ordering-proof: - fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a new fmx_request_relay_context helper (POST /connector/request-context) when neither the inbox payload nor carry flags carry it. The request_id survives the inbox drain, so a post-cleanup link still learns the right split budget. Best-effort: no token/curl or a non-2xx relay degrades to the loud warning below rather than a silent X default. - fm-x-link warns loudly when no platform source resolves, so the loss is never silent. - The fmx-respond procedure now orders link-before-inbox-cleanup so the fast local path stays correct without a relay round-trip. Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE message even when linked after inbox cleanup, and an unresolvable platform warns loudly instead of splitting silently. docs/configuration.md documents the request-context lookup. The relay endpoint is the companion durable change (see done status); until it ships, the link-before-cleanup reorder keeps the normal path correct. * no-mistakes(document): Document X-mode platform recovery * fix(composer): handle ANSI ghost text safely (#429) * fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor Away-mode injection wedged all night on the primary claude-on-herdr pane: the herdr composer classifier never stripped generic dim ghost text (only a narrow codex bold-wrapped byte-pattern check), so claude's rotating prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI pane read preserves - read as real pending input and every escalation deferred (6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s). Consolidate ghost extraction into one fleet-wide ANSI-aware owner, fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX, default 128, dark-theme assumption). Both ANSI-capable backends route through it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a dark box border dropped with the ghost does not lose the composer shape. This also closes the documented grok TRUECOLOR placeholder gap by the same mechanism (harness-adapters skill note updated). Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle) and the incident write-up are in docs/herdr-backend.md; deterministic regressions feed the exact captured bytes through the real classifiers (tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text (never a realistic real-input color) are corrected to a bright 38;2;224;222;244, preserving the truecolor payload-skip parser intent. * no-mistakes(review): Preserve dark shell prompt safety * no-mistakes(review): Harden erased shell prompt classification * no-mistakes(document): Document shared composer ghost extraction * no-mistakes(lint): Normalize tmux comment punctuation * fix(spawn): make tmux window handling robust under non-default config (#134) * test: isolate session-start suite from ambient harness markers (#432) * fix(session-start): isolate harness env markers in suite runner Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in run_session_start so ambient interactive shells cannot override the suite's fake ps harness (local-vs-CI split on the pi supervision case). * no-mistakes(document): Correct Pi marker documentation * fix(teardown): retry transient index locks during worktree return (#435) * fix(teardown): retry treehouse return on transient index.lock Killed crew git ops can leave a short-lived worktree index.lock that makes treehouse return fail. Retry on that error signature with a bounded wait (env-overridable), never force-delete a live lock, and only then fall back to the existing provably-stale cleanup path. * no-mistakes(review): Harden teardown retry configuration * no-mistakes(document): Document teardown index-lock retry behavior * no-mistakes(lint): Fix empty shell variable assignments * fix: complete brief help and consolidate documentation (#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage * fix: detect Git and centralize backend configuration (#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation * feat(daemon): add backend-independent wedge alerts (#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior * docs: centralize firstmate operating contracts (#447) * docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR #438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue #363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation * fix(cmux): close last workspace during teardown (#449) * fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior * fix: recover orphaned packed-refs locks during fleet sync (#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to…
fongryan
added a commit
to fongryan/firstmate
that referenced
this pull request
Jul 12, 2026
* fix: complete brief help and consolidate documentation (kunchenguid#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage * fix: detect Git and centralize backend configuration (kunchenguid#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation * feat(daemon): add backend-independent wedge alerts (kunchenguid#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior * docs: centralize firstmate operating contracts (kunchenguid#447) * docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR kunchenguid#438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue kunchenguid#363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation * fix(cmux): close last workspace during teardown (kunchenguid#449) * fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior * fix: recover orphaned packed-refs locks during fleet sync (kunchenguid#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to create '...packed-refs.lock': File exists", leaving the clone unsynced. On that signature only, fm-fleet-sync.sh now retries the fetch with a bounded wait (transient locks self-clear), then removes the lock and retries once more ONLY when it is provably stale: still present, mtime age past a threshold, and no lsof holder of the lock file or of the clone worktree itself (a live git keeps that as its cwd even in the window after it closes the lock and before it exits). A live lock, a missing lsof, any failed check, or any other fetch failure keeps today's behavior. Every wait/retry/removal prints to stderr, and a successful recovery also prints one "recovered:" summary to stdout so a session-start refresh - which discards fleet-sync stderr and relays only stdout - still surfaces it. The shared "is this git lock provably abandoned?" proof is extracted into bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/. * no-mistakes(review): Captain, remove obsolete teardown wake dependency * no-mistakes(document): Document packed-refs lock recovery architecture * feat(herdr): escalate blocked panes immediately (kunchenguid#472) * feat(herdr): immediate blocked-state escalation via native events.subscribe push Fold herdr's native pane.agent_status_changed stream into the single watcher so a crew entering blocked wakes its supervisor sub-second (measured 0.129s) instead of after the ~240s stale-pane wedge timer. - bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape plus the single-owner status->action policy table (blocked=actionable, working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling). - bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe subscriber over one connection for all this home's herdr panes, subscribing to ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe marker and a reconnect level-reconcile. Version/schema capability gate. - bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so the watcher stays backend-agnostic and the shape+policy are reusable. - bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal wait primitive (replacing the blind sleep POLL for push-capable homes), behind a source guard so the splice is unit-testable; secondmate/paused exemptions; map pane->window->task and enqueue a stale wake. No second watcher process; the single-cycle invariant and every guard/beacon/turn-end mechanism are unchanged. - Polling stays the permanent fail-closed backstop: below-capability, subscribe failure, and repeated runtime failures all degrade to sleep. - Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/ fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events) plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries the dated evidence and retires the old gap note. * no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs * no-mistakes(review): Captain, commit markers after wake and reuse capability cache * no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs * no-mistakes(review): Captain, subscribe before Herdr reconciliation * no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe * no-mistakes(test): Captain: include lock library in teardown fixture * no-mistakes(document): Captain: document Herdr immediate blocked escalation * fix: clarify shellcheck conditionals * docs(readme): reposition firstmate as an agent distro (kunchenguid#473) * feat: add deterministic bounded bearings snapshots (kunchenguid#475) * feat(bearings): deterministic bearings snapshot + durable decision model Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR discovery and checks happen only under --include-prs, which fails soft. Every dropped surface is marked in omitted[] with the flag that reveals it, and the prs: line states when checks were not requested, so absence is never silent. Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib gains status_open_decisions, the one authoritative keyed open/resolved fold over the whole status stream: needs-decision/blocked opens a keyed entry, only an explicit keyed resolution (or, for run-backed tasks, run-step advancement) closes it, so a later unrelated done/paused can no longer mask a still-open captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives pending_decision/blocked_event from it; the canonical schema stays complete. Point the /bearings skill at the one command; add the resolved: writer line to ship, scout, and secondmate briefs. Register the script and add regression tests for the output bound, TOON/JSON parity, local-only default, opt-in PR fetch, partial-failure degradation, decision durability, and report pointers. * fix(bearings): completed scout report is a pointer, not a pending decision A completed scout that raised a needs-decision and then finished (done) without a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103 case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh cleared a stale decision only for a live run-step/pane activity read, so a terminal task whose current state is read from the status log (a scout or ship that reached done/failed) never cleared its stale, never-keyed-resolved needs-decision, and it lingered as pending. The open-decision set is still derived purely from the keyed fold - never from a report body or decision-like prose - and reconciled against the crew lifecycle. Extend that reconciliation so a terminal done/failed state on a single-owner task (scout or ship), whose deliverable is its report or PR, also clears the set; a completed scout now surfaces only as a report pointer. Secondmates are excluded from the terminal clear (persistent, multiplexed stream), which keeps the unrelated-event masking fix intact. Add regression tests: a completed scout with decision-like report prose is a pointer not pending (canonical + end-to-end), and a scout still parked at a decision stays pending so the terminal clear never over-fires. * no-mistakes(review): Captain, preserve keyed decisions across shared status parsing * no-mistakes(review): Captain, close blockers and harden keyed decision parsing * no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout * no-mistakes(review): Captain, bound bearings sections and fail closed * no-mistakes(review): Captain, disclose capped per-repository PR results * no-mistakes(document): Refresh bearings documentation and status contracts * fix(bearings): avoid ambiguous worktree guard * fix: enforce deterministic ShellCheck parity (kunchenguid#481) * fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity Firstmate PRs passed local no-mistakes validation but failed CI's "Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034). Two divergences caused it: 1. The no-mistakes gate had no commands.lint, so its lint step never ran the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that CI runs. Confirmed from state.sqlite: the lint step_result recorded findings:null with no lint agent invocation. 2. CI's shellcheck floated with the runner image while local ran a newer build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck rejected an SC2015 that the newer local one no longer emits. Establish bin/fm-lint.sh as the single owner of the lint definition: the file set, the config, and the pinned shellcheck version (0.11.0, printed via --required-version). Both CI (.github/workflows/ci.yml) and the no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs the exact version it names and logs the resolved version, and fm-lint.sh refuses to lint under any other version. This is not a CI relaxation: it adopts shellcheck 0.11.0's rule set consistently, dropping only the upstream-retired, false-positive-prone SC2015; default severity and every still-supported finding stay enforced (no severity downgrade, no excludes). tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs and logs the pinned version, that the owner refuses a non-pinned shellcheck, and that it rejects a real lint defect the old no-op gate passed. * no-mistakes(review): Captain, harden deterministic ShellCheck parity * no-mistakes(review): Captain, neutralize ambient ShellCheck overrides * feat: guard primary shells from persistent cd commands (kunchenguid#483) * feat: add cd-guard PreToolUse seatbelt for the primary shell A stray persistent top-level `cd projects/<clone>` in the primary firstmate shell relocates the shell, so a later firstmate-owned command (a backlog write, an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the home. The cd-guard denies exactly that command shape before it runs, across all five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt. - bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate lexer; that file's CLI now runs only when invoked directly). - bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness output rendering, and primary-checkout scoping - fires in a secondmate's own primary session, inert in crew/scout child worktrees and non-firstmate repos. - Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents; per-harness hooks only call the owner. - Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd, and command cd). Allows git -C, subshell / bash -c / env -C / make -C / find -execdir, pipeline and background forms, and cd-as-data. Fails open on malformed input; agent-mistake threat model. - tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix, end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring. - docs/cd-guard.md: full contract plus live validation (claude, codex, opencode, pi blocked end-to-end; grok live run blocked by an API balance limit, with mechanism parity and deterministic coverage recorded). * no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage * no-mistakes(review): Captain, allow path-qualified command wrappers * no-mistakes(review): Captain, allow non-executing command queries * no-mistakes(test): Captain, clarify cd-guard safe-path remediation * docs: clarify cd guard guidance * no-mistakes(document): Clarify cd-guard safe target guidance * brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (kunchenguid#267) Crews must never stop, restart, or update the shared no-mistakes daemon since one instance serves every firstmate lane/home; a restart kills other lanes' in-flight pipeline runs and forces expensive re-runs. Encodes this as a new numbered rule in both the ship-task and scout-task brief scaffolds. Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> * feat: add atomic activation intake * fix: harden activation intake transaction * fix: gate activation intake test hooks * fix: block activation test-root symlink escapes * fix: make activation lock takeover ownership-safe --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: mielyemitchell <mielyemitchell@gmail.com> Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> Co-authored-by: Ryan Fong <noreply@anthropic.com>
blazingbunny
added a commit
to blazingbunny/firstmate
that referenced
this pull request
Jul 13, 2026
…sessions (#1) * feat(pi): simplify primary session launch (kunchenguid#386) * docs(readme): reformat Quick Start and recommend Grok equally with Claude Code * no-mistakes(review): Captain: align harness launch guidance * no-mistakes(review): Captain, clarify Pi supervised launch * no-mistakes(review): Captain, document Pi first-launch bridge * feat(pi): track primary watcher extension for plain-pi launch Move Pi's primary watcher bridge from a generated state/ file to a tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end guard extension already works: self-hashing version, project-local auto-discovery after one-time Pi trust. This drops the state/-generation step and dual -e requirement from the happy path, so Pi's Quick Start launch becomes plain 'pi', the same friction class as 'claude' and 'grok --trust'. - bin/fm-pi-watch-extension.sh is removed; nothing generates the extension anymore since it is committed. - fm-session-start.sh and fm-supervision-instructions.sh resolve the watcher extension path from FM_ROOT instead of state/, and the session-start diagnostic now points at restarting plain pi after trust, with -e as a documented fallback. - fm-spawn.sh points Pi secondmate launches at the tracked extension path in the secondmate home instead of generating a state/ copy. - README Quick Start Pi block is now just 'pi' plus a trust note. - Tests, docs, and the harness-adapters skill updated to match. * fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016 * feat(supervision): prevent unsafe watcher-arm commands (kunchenguid#387) * feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that denies a primary shell command backgrounding, piping, or bundling the watcher arm/checkpoint, or force-killing the watcher process broadly - the exact shapes that silently took Grok's supervision down. Wires it into all five verified harnesses (grok, claude, codex, opencode, pi), each validated empirically against the real harness. Also fixes a grok 0.2.93 regression discovered during that validation: the existing turnend-guard Stop hook's bare root variable broke grok's own variable pre-substitution and silently no-op'd the hook. * no-mistakes(review): Harden watcher arm validation * no-mistakes(review): Harden arm guard metacharacter checks * no-mistakes(review): Harden nested shell arm guard * fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then A && B || C is not if-then-else; C can run when A is true. Replace both occurrences of the quote-state early-continue with an explicit if/then. * fix(pi): restore primary watcher supervision lifecycle (kunchenguid#397) * fix Pi primary supervision lifecycle * no-mistakes(document): Synchronize Pi primary extension documentation * fix: keep persistent secondmates out of the main backlog (kunchenguid#398) * fix secondmate backlog guidance * no-mistakes(review): Require reasons for captain backlog holds * no-mistakes(test): Document secondmate handoff skill requirement * fix secondmate teardown reminder * no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract * fix(backlog-handoff): move full item blocks including indented bodies (kunchenguid#401) * fix(backlog-handoff): move full item blocks including indented bodies fm-backlog-handoff only moved the checklist header line, so multi-line item bodies were left orphaned in the source backlog and never reached the secondmate. Move the full block (header plus indented body lines) atomically, treating body membership by indentation so lines like ## Intent stay with the item, and add regression coverage. * no-mistakes(review): Captain: preserve EOF handoff terminators * no-mistakes(review): treat blank lines inside item bodies as movable body * no-mistakes(document): sync backlog-handoff docs with full-block move behavior * feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (kunchenguid#402) * guard Herdr lab lifecycle in briefs * no-mistakes(review): Fix Herdr lab helper and provisioning safety * no-mistakes(review): Captain, harden Herdr lab lifecycle safety * no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range * no-mistakes(review): reject leading options in Herdr lab run guard * no-mistakes(review): strip leading non-alnum in Herdr lab name generator * no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag * no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab * no-mistakes: apply CI fixes * fix(watcher): classify arm-command seatbelt by execution position (kunchenguid#403) * fix watcher arm command policy * no-mistakes(review): Harden watcher command policy parsing * no-mistakes(review): Captain: harden watcher policy parsing * no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter * no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps * no-mistakes(review): fail closed on loop-wrapped broad watcher kills * no-mistakes(document): sync docs for watcher-arm command-position policy * fix: reconcile existing AGENTS.md safely (kunchenguid#405) * fix(agents-md): inject self-governance section into existing AGENTS.md fm-ensure-agents-md.sh only appended the canonical "## Maintaining this file" section on skeleton create or CLAUDE.md promotion, so an existing AGENTS.md that lacked it exited unchanged and forced hand-copying the wording during a rollout across existing projects. Call the already- idempotent ensure_maintenance_section on the existing-AGENTS.md paths and report whether the file changed; a re-run and an already-complete file stay byte-identical. Also fixes kunchenguid#389: refuse a case-variant real memory file (e.g. a lowercase agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase literal target dangles once the tree lands on a case-sensitive filesystem. Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and CLAUDE.md-promotion regressions still pass. Docs updated to match. * no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency * no-mistakes(review): Preserve CRLF during maintenance-section injection * no-mistakes(review): Captain: harden dangling-symlink regression coverage * no-mistakes(document): Document agent-memory injection outcomes * feat: support project-less secondmate homes (kunchenguid#409) * feat(secondmate): support project-less homes via --no-projects fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit --no-projects signal to scaffold, seed, and register a secondmate home whose subject is the firstmate repo itself (no clones). The signal is mutually exclusive with a project list; omitting both still fails loudly so an accidental omission is never a silent project-less seed. The registry line renders an empty projects: field, which spawn and the snapshot already tolerate. Docs updated in the secondmate-provisioning skill and both script headers. * no-mistakes(review): Captain: document project-less secondmate flow * no-mistakes(review): Captain: refuse project-less reseeding of populated homes * fix(seed): fail closed on unreadable project data * no-mistakes(review): Captain: reject stale projectful charters * no-mistakes(review): Captain: fail closed on unsafe project paths * no-mistakes(review): Captain: validate project-less charter clone sections * no-mistakes(document): Document project-less secondmate seeding * fix: delegate backlog handoffs to tasks-axi (kunchenguid#411) * wip(handoff): record verified delegation design + tasks-axi mv blocker No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically move a blocked-by-linked item set across backlogs (deadlocks both orders, no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked pending a tasks-axi connected-set mv enhancement; note captures the verified design, semantics, test/CI/doc changes, and resume checklist. * refactor(handoff): delegate the item move to tasks-axi mv fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog format and the source of the PR kunchenguid#401 body-orphaning drift. Delete it and delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic multi-id), the single owner of the format: a connected set (blocker plus dependents) moves together with blocked-by preserved, item blocks stay byte-exact, and destination section placement holds. The helper keeps only the fleet-level validation tasks-axi cannot know - secondmate-home resolution, the seeded-home safety checks, the In-flight refusal, and idempotent per-key reporting - and is atomic: on any move failure nothing moves. Tests: fm-backlog-handoff.test.sh keeps PR kunchenguid#401's regression matrix but now exercises the delegated path and skips cleanly when tasks-axi is absent; the two whole-file fixtures move to tasks-axi's canonical whitespace. The lifecycle-e2e and safety move-cases gain the same skip guard. CI installs tasks-axi so the delegated path is exercised. Docs state that config/backlog-backend=manual governs firstmate's own hand-editing, not this validated helper, which delegates fleet-wide because bootstrap requires tasks-axi on PATH. Remove the now-redundant WIP design note. * no-mistakes(review): Captain: harden atomic backlog handoffs * no-mistakes(review): Captain: enforce queued-only backlog handoffs * no-mistakes(review): Captain: harden handoff section parsing * no-mistakes(document): Document delegated backlog handoffs * no-mistakes(lint): Silence ShellCheck source diagnostics * fix: ignore secondmate home marker during sync (kunchenguid#417) * fix: gitignore the secondmate home marker bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into every seeded secondmate home. A secondmate home is a worktree of the firstmate repo, so any plain `git status --porcelain` dirtiness check counted the untracked marker and the home read as dirty forever: fleet-sync reported it STUCK and the local fast-forward convergence sweeps risked leaving it stale on firstmate updates. Add .fm-secondmate-home to the tracked .gitignore so the marker is invisible to every dirtiness check uniformly, without weakening fleet-sync's deliberate untracked-counting for project clones. Convergence chicken-and-egg: existing homes predate the fix and it only arrives by fast-forward. The already-present marker-tolerant ff-skip (ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate, and spawn pre-launch) advances such a home past the fix commit, after which .gitignore takes over - no hand intervention. Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home reading clean, an existing marker-only home converging then reading clean, and a genuinely dirty home still skipping. * no-mistakes(review): Captain: document standalone-clone update path * no-mistakes(document): Document secondmate marker migration * fix(composer): prevent dead-shell message injection (kunchenguid#416) * fix(composer): stop reading dead-shell prompts as empty agent composers Consolidate composer empty/pending/unknown classification into one shared owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This replaces four drifting copies of the glyph decision. Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now classified unknown (a dead shell, unsafe for injection), not empty. It is only empty inside a bordered composer box (the harness's own prompt). Agent glyphs ❯ (claude) and › (codex) read empty either way. The away-mode injector (inject_msg) now requires an affirmatively-empty composer, deferring on pending or unknown, so an escalation can never be typed into (or executed by) a pane whose agent exited to its login shell. Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner; per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay green. * no-mistakes(review): Captain: harden composer safety checks * no-mistakes(test): Stabilize Herdr prune safety setup * no-mistakes(document): Document composer injection safety * no-mistakes(lint): Clean composer safety lint * no-mistakes: apply CI fixes * feat(watcher): add paused external-wait supervision (kunchenguid#421) * feat(watcher): add paused/awaiting-external crew state A crew (or firstmate steering it) can declare a deliberate wait on a known external dependency with a paused: <reason> status. Both the always-on watcher and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh vocabulary instead of tripping the possible-wedge stale escalation, and re-surface it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused distinctly. A crew that goes idle without declaring a pause classifies exactly as before. Docs and brief scaffold state lists updated; tests colocated. * no-mistakes(review): Captain: fix paused-state transitions * init * no-mistakes(review): Captain: fix paused-state supervision transitions * no-mistakes(review): Captain: fix paused supervision handoffs * no-mistakes(review): Reconcile paused supervision markers * no-mistakes(review): Captain: prioritize paused states over captain relevance * no-mistakes(review): Captain: preserve paused-working wedge timer * no-mistakes(review): Captain: honor configured pause verb in briefs * no-mistakes(test): Captain: fix AFK paused watcher handoff * no-mistakes(document): Document declared external waits * no-mistakes(lint): Clean paused-state lint --------- Co-authored-by: fmtest <fmtest@example.invalid> * fix: preserve X-mode follow-up platform limits (kunchenguid#425) * fix(x-mode): make follow-up platform splitting immune to link ordering A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280 chars because fm-x-link only learned the platform from the inbox payload, and the fmx-respond ack path can drain that inbox file before the task is linked. A link recorded after cleanup silently lost the platform and the splitter defaulted to the X 280-char budget. Make platform resolution ordering-proof: - fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a new fmx_request_relay_context helper (POST /connector/request-context) when neither the inbox payload nor carry flags carry it. The request_id survives the inbox drain, so a post-cleanup link still learns the right split budget. Best-effort: no token/curl or a non-2xx relay degrades to the loud warning below rather than a silent X default. - fm-x-link warns loudly when no platform source resolves, so the loss is never silent. - The fmx-respond procedure now orders link-before-inbox-cleanup so the fast local path stays correct without a relay round-trip. Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE message even when linked after inbox cleanup, and an unresolvable platform warns loudly instead of splitting silently. docs/configuration.md documents the request-context lookup. The relay endpoint is the companion durable change (see done status); until it ships, the link-before-cleanup reorder keeps the normal path correct. * no-mistakes(document): Document X-mode platform recovery * fix(composer): handle ANSI ghost text safely (kunchenguid#429) * fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor Away-mode injection wedged all night on the primary claude-on-herdr pane: the herdr composer classifier never stripped generic dim ghost text (only a narrow codex bold-wrapped byte-pattern check), so claude's rotating prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI pane read preserves - read as real pending input and every escalation deferred (6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s). Consolidate ghost extraction into one fleet-wide ANSI-aware owner, fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX, default 128, dark-theme assumption). Both ANSI-capable backends route through it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a dark box border dropped with the ghost does not lose the composer shape. This also closes the documented grok TRUECOLOR placeholder gap by the same mechanism (harness-adapters skill note updated). Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle) and the incident write-up are in docs/herdr-backend.md; deterministic regressions feed the exact captured bytes through the real classifiers (tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text (never a realistic real-input color) are corrected to a bright 38;2;224;222;244, preserving the truecolor payload-skip parser intent. * no-mistakes(review): Preserve dark shell prompt safety * no-mistakes(review): Harden erased shell prompt classification * no-mistakes(document): Document shared composer ghost extraction * no-mistakes(lint): Normalize tmux comment punctuation * fix(spawn): make tmux window handling robust under non-default config (kunchenguid#134) * test: isolate session-start suite from ambient harness markers (kunchenguid#432) * fix(session-start): isolate harness env markers in suite runner Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in run_session_start so ambient interactive shells cannot override the suite's fake ps harness (local-vs-CI split on the pi supervision case). * no-mistakes(document): Correct Pi marker documentation * fix(teardown): retry transient index locks during worktree return (kunchenguid#435) * fix(teardown): retry treehouse return on transient index.lock Killed crew git ops can leave a short-lived worktree index.lock that makes treehouse return fail. Retry on that error signature with a bounded wait (env-overridable), never force-delete a live lock, and only then fall back to the existing provably-stale cleanup path. * no-mistakes(review): Harden teardown retry configuration * no-mistakes(document): Document teardown index-lock retry behavior * no-mistakes(lint): Fix empty shell variable assignments * fix: complete brief help and consolidate documentation (kunchenguid#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage * fix: detect Git and centralize backend configuration (kunchenguid#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation * feat(daemon): add backend-independent wedge alerts (kunchenguid#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior * docs: centralize firstmate operating contracts (kunchenguid#447) * docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR kunchenguid#438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue kunchenguid#363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation * fix(cmux): close last workspace during teardown (kunchenguid#449) * fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior * fix: recover orphaned packed-refs locks during fleet sync (kunchenguid#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to create '...packed-refs.lock': File exists", leaving the clone unsynced. On that signature only, fm-fleet-sync.sh now retries the fetch with a bounded wait (transient locks self-clear), then removes the lock and retries once more ONLY when it is provably stale: still present, mtime age past a threshold, and no lsof holder of the lock file or of the clone worktree itself (a live git keeps that as its cwd even in the window after it closes the lock and before it exits). A live lock, a missing lsof, any failed check, or any other fetch failure keeps today's behavior. Every wait/retry/removal prints to stderr, and a successful recovery also prints one "recovered:" summary to stdout so a session-start refresh - which discards fleet-sync stderr and relays only stdout - still surfaces it. The shared "is this git lock provably abandoned?" proof is extracted into bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/. * no-mistakes(review): Captain, remove obsolete teardown wake dependency * no-mistakes(document): Document packed-refs lock recovery architecture * feat(herdr): escalate blocked panes immediately (kunchenguid#472) * feat(herdr): immediate blocked-state escalation via native events.subscribe push Fold herdr's native pane.agent_status_changed stream into the single watcher so a crew entering blocked wakes its supervisor sub-second (measured 0.129s) instead of after the ~240s stale-pane wedge timer. - bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape plus the single-owner status->action policy table (blocked=actionable, working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling). - bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe subscriber over one connection for all this home's herdr panes, subscribing to ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe marker and a reconnect level-reconcile. Version/schema capability gate. - bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so the watcher stays backend-agnostic and the shape+policy are reusable. - bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal wait primitive (replacing the blind sleep POLL for push-capable homes), behind a source guard so the splice is unit-testable; secondmate/paused exemptions; map pane->window->task and enqueue a stale wake. No second watcher process; the single-cycle invariant and every guard/beacon/turn-end mechanism are unchanged. - Polling stays the permanent fail-closed backstop: below-capability, subscribe failure, and repeated runtime failures all degrade to sleep. - Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/ fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events) plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries the dated evidence and retires the old gap note. * no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs * no-mistakes(review): Captain, commit markers after wake and reuse capability cache * no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs * no-mistakes(review): Captain, subscribe before Herdr reconciliation * no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe * no-mistakes(test): Captain: include lock library in teardown fixture * no-mistakes(document): Captain: document Herdr immediate blocked escalation * fix: clarify shellcheck conditionals * docs(readme): reposition firstmate as an agent distro (kunchenguid#473) * feat: add deterministic bounded bearings snapshots (kunchenguid#475) * feat(bearings): deterministic bearings snapshot + durable decision model Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR discovery and checks happen only under --include-prs, which fails soft. Every dropped surface is marked in omitted[] with the flag that reveals it, and the prs: line states when checks were not requested, so absence is never silent. Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib gains status_open_decisions, the one authoritative keyed open/resolved fold over the whole status stream: needs-decision/blocked opens a keyed entry, only an explicit keyed resolution (or, for run-backed tasks, run-step advancement) closes it, so a later unrelated done/paused can no longer mask a still-open captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives pending_decision/blocked_event from it; the canonical schema stays complete. Point the /bearings skill at the one command; add the resolved: writer line to ship, scout, and secondmate briefs. Register the script and add regression tests for the output bound, TOON/JSON parity, local-only default, opt-in PR fetch, partial-failure degradation, decision durability, and report pointers. * fix(bearings): completed scout report is a pointer, not a pending decision A completed scout that raised a needs-decision and then finished (done) without a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103 case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh cleared a stale decision only for a live run-step/pane activity read, so a terminal task whose current state is read from the status log (a scout or ship that reached done/failed) never cleared its stale, never-keyed-resolved needs-decision, and it lingered as pending. The open-decision set is still derived purely from the keyed fold - never from a report body or decision-like prose - and reconciled against the crew lifecycle. Extend that reconciliation so a terminal done/failed state on a single-owner task (scout or ship), whose deliverable is its report or PR, also clears the set; a completed scout now surfaces only as a report pointer. Secondmates are excluded from the terminal clear (persistent, multiplexed stream), which keeps the unrelated-event masking fix intact. Add regression tests: a completed scout with decision-like report prose is a pointer not pending (canonical + end-to-end), and a scout still parked at a decision stays pending so the terminal clear never over-fires. * no-mistakes(review): Captain, preserve keyed decisions across shared status parsing * no-mistakes(review): Captain, close blockers and harden keyed decision parsing * no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout * no-mistakes(review): Captain, bound bearings sections and fail closed * no-mistakes(review): Captain, disclose capped per-repository PR results * no-mistakes(document): Refresh bearings documentation and status contracts * fix(bearings): avoid ambiguous worktree guard * fix: enforce deterministic ShellCheck parity (kunchenguid#481) * fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity Firstmate PRs passed local no-mistakes validation but failed CI's "Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034). Two divergences caused it: 1. The no-mistakes gate had no commands.lint, so its lint step never ran the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that CI runs. Confirmed from state.sqlite: the lint step_result recorded findings:null with no lint agent invocation. 2. CI's shellcheck floated with the runner image while local ran a newer build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck rejected an SC2015 that the newer local one no longer emits. Establish bin/fm-lint.sh as the single owner of the lint definition: the file set, the config, and the pinned shellcheck version (0.11.0, printed via --required-version). Both CI (.github/workflows/ci.yml) and the no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs the exact version it names and logs the resolved version, and fm-lint.sh refuses to lint under any other version. This is not a CI relaxation: it adopts shellcheck 0.11.0's rule set consistently, dropping only the upstream-retired, false-positive-prone SC2015; default severity and every still-supported finding stay enforced (no severity downgrade, no excludes). tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs and logs the pinned version, that the owner refuses a non-pinned shellcheck, and that it rejects a real lint defect the old no-op gate passed. * no-mistakes(review): Captain, harden deterministic ShellCheck parity * no-mistakes(review): Captain, neutralize ambient ShellCheck overrides * feat: guard primary shells from persistent cd commands (kunchenguid#483) * feat: add cd-guard PreToolUse seatbelt for the primary shell A stray persistent top-level `cd projects/<clone>` in the primary firstmate shell relocates the shell, so a later firstmate-owned command (a backlog write, an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the home. The cd-guard denies exactly that command shape before it runs, across all five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt. - bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate lexer; that file's CLI now runs only when invoked directly). - bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness output rendering, and primary-checkout scoping - fires in a secondmate's own primary session, inert in crew/scout child worktrees and non-firstmate repos. - Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents; per-harness hooks only call the owner. - Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd, and command cd). Allows git -C, subshell / bash -c / env -C / make -C / find -execdir, pipeline and background forms, and cd-as-data. Fails open on malformed input; agent-mistake threat model. - tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix, end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring. - docs/cd-guard.md: full contract plus live validation (claude, codex, opencode, pi blocked end-to-end; grok live run blocked by an API balance limit, with mechanism parity and deterministic coverage recorded). * no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage * no-mistakes(review): Captain, allow path-qualified command wrappers * no-mistakes(review): Captain, allow non-executing command queries * no-mistakes(test): Captain, clarify cd-guard safe-path remediation * docs: clarify cd guard guidance * no-mistakes(document): Clarify cd-guard safe target guidance * brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (kunchenguid#267) Crews must never stop, restart, or update the shared no-mistakes daemon since one instance serves every firstmate lane/home; a restart kills other lanes' in-flight pipeline runs and forces expensive re-runs. Encodes this as a new numbered rule in both the ship-task and scout-task brief scaffolds. Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> * feat: make bearings concise and accurate (kunchenguid#485) * feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render /bearings skill (one owner of the chat-response format): mandate the four always-present chat sections - Captain's Call, Recently Landed, Underway, Charted Next - each with an explicit empty-state sentence, no At Anchor, materially shorter than and linking to the report file. Resolves the ambiguous Check first / Decisions pending split into one strict captain-action section. fm-crew-state: the log fallback derives current state only from a real run-state verb, so a trailing decision-closing resolved: event no longer renders a healthy idle crew (typically a secondmate) as unknown with the resolution prose as its detail. The keyed-decision contract in fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner. fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done records from registered secondmate homes, reusing the single backlog parser and the one secondmate-home enumerator (meta home= with data/secondmates.md fallback); no network, per-home capped. fm-bearings-snapshot: landed now merges main-home Done with the secondmate roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure (also fixing the previously-silent landed truncation); --all-landed reveals the full set. tests: resolved-event state render, secondmate landed aggregation with caps and omitted[] disclosure, Captain's Call anti-leak, and the four-section contract. * no-mistakes(review): Captain, ensure bearings reveals all landed work * no-mistakes(document): Document bearings accuracy contracts * fix: harden away-mode daemon lifecycle (kunchenguid#490) * fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle Away-mode entry left "make the daemon a tracked background terminal" to the operator; on a pi/herdr primary that meant splitting the captain's active pane, which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches the daemon in a non-visible tracked terminal per backend (herdr dedicated --no-focus workspace, detached tmux session), never a split, pins the captain pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal id, and tears it down or reconciles a leaked one by that id. No shell &. Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared with the daemon (one owner). Fix the stale subsuper-artifact leak: clear the prior away session's delivery cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon before clearing state/.afk so its shutdown flush runs instead of being a no-op. Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md, docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md. * no-mistakes(review): Captain, serialize AFK launcher lifecycle safely * no-mistakes(review): Captain, harden AFK launcher lifecycle races * no-mistakes(review): Captain, ensure AFK daemon launch readiness * no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown * no-mistakes(review): Captain, preserve AFK reconciliation records uniformly * no-mistakes(review): Captain, harden AFK recovery state durability * no-mistakes(review): Captain, harden AFK tmux ownership checks * no-mistakes(review): Captain, simplify AFK lifecycle failure handling * no-mistakes(review): Captain, require confirmed AFK daemon shutdown * no-mistakes(review): Captain, confirm AFK exit by process identity * no-mistakes(document): Align AFK launcher lifecycle documentation * no-mistakes: apply CI fixes * Add fm-jules-check.sh — watcher poll arm for Jules coding-agent sessions bin/fm-jules-check.sh mirrors fm-pr-check.sh's structure: appends jules_session= to state/<id>.meta idempotently, then writes state/<id>.check.sh that polls jules-axi watch for state transitions. The generated check.sh is thin (4 lines) — jules-axi watch already owns the no-change-is-silent contract. tests/fm-jules-check.test.sh covers: first-arm meta append, idempotent re-arm, check.sh invocation shape, silence-on-no-output, transition echo, and graceful exit when jules-axi is absent. * no-mistakes(review): fix: clear stale jules watermark on session re-arm, register script in docs * no-mistakes(document): Fixed 3 stale doc references to check-poll wake mechanics that omitted the new Jules session poll; no gaps remain unresolved. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: fmtest <fmtest@example.invalid> Co-authored-by: Pierre Marais <pierremarais67@gmail.com> Co-authored-by: mielyemitchell <mielyemitchell@gmail.com> Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
This was referenced Jul 14, 2026
fongryan
added a commit
to fongryan/firstmate
that referenced
this pull request
Jul 16, 2026
…n current (#3) * fix: complete brief help and consolidate documentation (kunchenguid#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage * fix: detect Git and centralize backend configuration (kunchenguid#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation * feat(daemon): add backend-independent wedge alerts (kunchenguid#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior * docs: centralize firstmate operating contracts (kunchenguid#447) * docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR kunchenguid#438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue kunchenguid#363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation * fix(cmux): close last workspace during teardown (kunchenguid#449) * fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior * fix: recover orphaned packed-refs locks during fleet sync (kunchenguid#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to create '...packed-refs.lock': File exists", leaving the clone unsynced. On that signature only, fm-fleet-sync.sh now retries the fetch with a bounded wait (transient locks self-clear), then removes the lock and retries once more ONLY when it is provably stale: still present, mtime age past a threshold, and no lsof holder of the lock file or of the clone worktree itself (a live git keeps that as its cwd even in the window after it closes the lock and before it exits). A live lock, a missing lsof, any failed check, or any other fetch failure keeps today's behavior. Every wait/retry/removal prints to stderr, and a successful recovery also prints one "recovered:" summary to stdout so a session-start refresh - which discards fleet-sync stderr and relays only stdout - still surfaces it. The shared "is this git lock provably abandoned?" proof is extracted into bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/. * no-mistakes(review): Captain, remove obsolete teardown wake dependency * no-mistakes(document): Document packed-refs lock recovery architecture * feat(herdr): escalate blocked panes immediately (kunchenguid#472) * feat(herdr): immediate blocked-state escalation via native events.subscribe push Fold herdr's native pane.agent_status_changed stream into the single watcher so a crew entering blocked wakes its supervisor sub-second (measured 0.129s) instead of after the ~240s stale-pane wedge timer. - bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape plus the single-owner status->action policy table (blocked=actionable, working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling). - bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe subscriber over one connection for all this home's herdr panes, subscribing to ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe marker and a reconnect level-reconcile. Version/schema capability gate. - bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so the watcher stays backend-agnostic and the shape+policy are reusable. - bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal wait primitive (replacing the blind sleep POLL for push-capable homes), behind a source guard so the splice is unit-testable; secondmate/paused exemptions; map pane->window->task and enqueue a stale wake. No second watcher process; the single-cycle invariant and every guard/beacon/turn-end mechanism are unchanged. - Polling stays the permanent fail-closed backstop: below-capability, subscribe failure, and repeated runtime failures all degrade to sleep. - Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/ fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events) plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries the dated evidence and retires the old gap note. * no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs * no-mistakes(review): Captain, commit markers after wake and reuse capability cache * no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs * no-mistakes(review): Captain, subscribe before Herdr reconciliation * no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe * no-mistakes(test): Captain: include lock library in teardown fixture * no-mistakes(document): Captain: document Herdr immediate blocked escalation * fix: clarify shellcheck conditionals * docs(readme): reposition firstmate as an agent distro (kunchenguid#473) * feat: add deterministic bounded bearings snapshots (kunchenguid#475) * feat(bearings): deterministic bearings snapshot + durable decision model Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR discovery and checks happen only under --include-prs, which fails soft. Every dropped surface is marked in omitted[] with the flag that reveals it, and the prs: line states when checks were not requested, so absence is never silent. Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib gains status_open_decisions, the one authoritative keyed open/resolved fold over the whole status stream: needs-decision/blocked opens a keyed entry, only an explicit keyed resolution (or, for run-backed tasks, run-step advancement) closes it, so a later unrelated done/paused can no longer mask a still-open captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives pending_decision/blocked_event from it; the canonical schema stays complete. Point the /bearings skill at the one command; add the resolved: writer line to ship, scout, and secondmate briefs. Register the script and add regression tests for the output bound, TOON/JSON parity, local-only default, opt-in PR fetch, partial-failure degradation, decision durability, and report pointers. * fix(bearings): completed scout report is a pointer, not a pending decision A completed scout that raised a needs-decision and then finished (done) without a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103 case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh cleared a stale decision only for a live run-step/pane activity read, so a terminal task whose current state is read from the status log (a scout or ship that reached done/failed) never cleared its stale, never-keyed-resolved needs-decision, and it lingered as pending. The open-decision set is still derived purely from the keyed fold - never from a report body or decision-like prose - and reconciled against the crew lifecycle. Extend that reconciliation so a terminal done/failed state on a single-owner task (scout or ship), whose deliverable is its report or PR, also clears the set; a completed scout now surfaces only as a report pointer. Secondmates are excluded from the terminal clear (persistent, multiplexed stream), which keeps the unrelated-event masking fix intact. Add regression tests: a completed scout with decision-like report prose is a pointer not pending (canonical + end-to-end), and a scout still parked at a decision stays pending so the terminal clear never over-fires. * no-mistakes(review): Captain, preserve keyed decisions across shared status parsing * no-mistakes(review): Captain, close blockers and harden keyed decision parsing * no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout * no-mistakes(review): Captain, bound bearings sections and fail closed * no-mistakes(review): Captain, disclose capped per-repository PR results * no-mistakes(document): Refresh bearings documentation and status contracts * fix(bearings): avoid ambiguous worktree guard * fix: enforce deterministic ShellCheck parity (kunchenguid#481) * fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity Firstmate PRs passed local no-mistakes validation but failed CI's "Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034). Two divergences caused it: 1. The no-mistakes gate had no commands.lint, so its lint step never ran the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that CI runs. Confirmed from state.sqlite: the lint step_result recorded findings:null with no lint agent invocation. 2. CI's shellcheck floated with the runner image while local ran a newer build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck rejected an SC2015 that the newer local one no longer emits. Establish bin/fm-lint.sh as the single owner of the lint definition: the file set, the config, and the pinned shellcheck version (0.11.0, printed via --required-version). Both CI (.github/workflows/ci.yml) and the no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs the exact version it names and logs the resolved version, and fm-lint.sh refuses to lint under any other version. This is not a CI relaxation: it adopts shellcheck 0.11.0's rule set consistently, dropping only the upstream-retired, false-positive-prone SC2015; default severity and every still-supported finding stay enforced (no severity downgrade, no excludes). tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs and logs the pinned version, that the owner refuses a non-pinned shellcheck, and that it rejects a real lint defect the old no-op gate passed. * no-mistakes(review): Captain, harden deterministic ShellCheck parity * no-mistakes(review): Captain, neutralize ambient ShellCheck overrides * feat: guard primary shells from persistent cd commands (kunchenguid#483) * feat: add cd-guard PreToolUse seatbelt for the primary shell A stray persistent top-level `cd projects/<clone>` in the primary firstmate shell relocates the shell, so a later firstmate-owned command (a backlog write, an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the home. The cd-guard denies exactly that command shape before it runs, across all five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt. - bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate lexer; that file's CLI now runs only when invoked directly). - bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness output rendering, and primary-checkout scoping - fires in a secondmate's own primary session, inert in crew/scout child worktrees and non-firstmate repos. - Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents; per-harness hooks only call the owner. - Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd, and command cd). Allows git -C, subshell / bash -c / env -C / make -C / find -execdir, pipeline and background forms, and cd-as-data. Fails open on malformed input; agent-mistake threat model. - tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix, end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring. - docs/cd-guard.md: full contract plus live validation (claude, codex, opencode, pi blocked end-to-end; grok live run blocked by an API balance limit, with mechanism parity and deterministic coverage recorded). * no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage * no-mistakes(review): Captain, allow path-qualified command wrappers * no-mistakes(review): Captain, allow non-executing command queries * no-mistakes(test): Captain, clarify cd-guard safe-path remediation * docs: clarify cd guard guidance * no-mistakes(document): Clarify cd-guard safe target guidance * brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (kunchenguid#267) Crews must never stop, restart, or update the shared no-mistakes daemon since one instance serves every firstmate lane/home; a restart kills other lanes' in-flight pipeline runs and forces expensive re-runs. Encodes this as a new numbered rule in both the ship-task and scout-task brief scaffolds. Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> * feat: add closed-loop task lifecycle supervision * fix: close lifecycle supervision gaps * feat: import legacy tasks into lifecycle loop * feat: make bearings concise and accurate (kunchenguid#485) * feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render /bearings skill (one owner of the chat-response format): mandate the four always-present chat sections - Captain's Call, Recently Landed, Underway, Charted Next - each with an explicit empty-state sentence, no At Anchor, materially shorter than and linking to the report file. Resolves the ambiguous Check first / Decisions pending split into one strict captain-action section. fm-crew-state: the log fallback derives current state only from a real run-state verb, so a trailing decision-closing resolved: event no longer renders a healthy idle crew (typically a secondmate) as unknown with the resolution prose as its detail. The keyed-decision contract in fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner. fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done records from registered secondmate homes, reusing the single backlog parser and the one secondmate-home enumerator (meta home= with data/secondmates.md fallback); no network, per-home capped. fm-bearings-snapshot: landed now merges main-home Done with the secondmate roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure (also fixing the previously-silent landed truncation); --all-landed reveals the full set. tests: resolved-event state render, secondmate landed aggregation with caps and omitted[] disclosure, Captain's Call anti-leak, and the four-section contract. * no-mistakes(review): Captain, ensure bearings reveals all landed work * no-mistakes(document): Document bearings accuracy contracts * fix: harden away-mode daemon lifecycle (kunchenguid#490) * fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle Away-mode entry left "make the daemon a tracked background terminal" to the operator; on a pi/herdr primary that meant splitting the captain's active pane, which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches the daemon in a non-visible tracked terminal per backend (herdr dedicated --no-focus workspace, detached tmux session), never a split, pins the captain pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal id, and tears it down or reconciles a leaked one by that id. No shell &. Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared with the daemon (one owner). Fix the stale subsuper-artifact leak: clear the prior away session's delivery cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon before clearing state/.afk so its shutdown flush runs instead of being a no-op. Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md, docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md. * no-mistakes(review): Captain, serialize AFK launcher lifecycle safely * no-mistakes(review): Captain, harden AFK launcher lifecycle races * no-mistakes(review): Captain, ensure AFK daemon launch readiness * no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown * no-mistakes(review): Captain, preserve AFK reconciliation records uniformly * no-mistakes(review): Captain, harden AFK recovery state durability * no-mistakes(review): Captain, harden AFK tmux ownership checks * no-mistakes(review): Captain, simplify AFK lifecycle failure handling * no-mistakes(review): Captain, require confirmed AFK daemon shutdown * no-mistakes(review): Captain, confirm AFK exit by process identity * no-mistakes(document): Align AFK launcher lifecycle documentation * no-mistakes: apply CI fixes * fix: prevent no-mistakes gate agents from driving the fleet (kunchenguid#518) * feat: contain no-mistakes gate agents from driving the fleet Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/ fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE is set, and via an unspoofable git-common-dir backstop when invoked from a no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker unset. A normal firstmate session has neither signal and is unaffected. Set disable_project_settings: true in the tracked .no-mistakes.yaml so the installed pipeline neutralizes gate agents' project instructions for this repo (trusted-only, honored from the default branch). firstmate's own suite runs from a gate worktree during validation, so the shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated tests/fm-gate-refuse.test.sh strips it to verify real refusal. * no-mistakes(review): Captain, refuse empty no-mistakes gate markers * no-mistakes(document): Document no-mistakes gate authority boundary * fix: guard secondmate primary sessions from blind turn ends (kunchenguid#505) * fix: guard secondmate own-home turn ends Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the 'no turn ends blind' backstop fires in a secondmate's own primary session, matching the cd-guard's scope: the own home is guarded, child crew/scout worktrees stay exempt via the retained git-dir/git-common-dir test. This was pure scoping from the guard's primary-only origin and guarded against no secondmate-specific hazard. Add secondmate regression tests (blind-turn block, idle-by-default, stop_hook_active loop guard, deferred-death recovery loop, child-worktree exemption) and record the autonomous background-notify re-invoke measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md. * no-mistakes(document): Correct secondmate guard documentation, captain * fix: force-include marked secondmate homes in turn-end guard The prior remove-only form (just deleting the .fm-secondmate-home check) left the DEFAULT secondmate topology unguarded: a treehouse-leased home is a linked git worktree (git-dir != git-common-dir), which the retained git-dir exemption still skipped, so its own primary session could still end a turn blind. Invert the marker: a genuinely-marked home is force-included as a guarded primary (treehouse-leased linked OR git-cloned plain), and the git-dir exemption applies only to UNMARKED child worktrees. Marker validation (regular non-symlink file, non-empty id-token content) blocks a stray or empty marker from spoofing inclusion. Add real linked-worktree regression tests: a treehouse-leased LINKED secondmate home is guarded, a stray/empty marker stays exempt, and the unmarked child worktree stays exempt - the topology the plain git-init fixtures masked. Predicates, in-flight gate, and loop guard untouched. * fix: force ASCII collation in secondmate marker validation Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the [A-Za-z0-9._-] id allowlist matches under C collation, not the ambient locale - a locale-crafted non-ASCII marker id can no longer slip through the range match and spoof force-inclusion of a linked child worktree. Add a regression test proving a non-ASCII marker id is rejected and the linked worktree stays exempt. * no-mistakes(test): fix backend baseline gate-refusal dependency * no-mistakes(document): Correct secondmate turn-end guard documentation * fix: make bootstrap diagnostics backend-aware (kunchenguid#519) * fix: make bootstrap required-tool detection backend-aware Bootstrap demanded tmux and treehouse for every backend except orca, so a herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux. Required tools now follow the resolved backend via the single-owner fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux), and treehouse for session-provider-only backends (orca owns its worktree). The treehouse lease-support check is gated to backends that use treehouse. Adds install hints for herdr/zellij/cmux, regression tests for the full backend dependency matrix (herdr-without-tmux repro plus each boundary), and updates the authoritative Toolchain docs. * no-mistakes(review): Captain, prevent executing Herdr install guidance * no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics * no-mistakes(review): Captain, separate manual dependency remediation * no-mistakes(review): Captain, align bootstrap diagnostic consumers * no-mistakes(document): Align backend adapter dependency comments * fix: preserve follow-up platform context after inbox cleanup (kunchenguid#520) * fix: recover X/Discord follow-up platform after inbox cleanup A milestone follow-up posted directly by request_id after the inbox was drained - and with no task link, because one persistent secondmate's single x_request slot collides across concurrent requests - resolved platform only from the local inbox, so a >280 Discord reply silently defaulted to the X 280-char budget and threaded as (1/2). - fm-x-poll records a durable per-request reply context (state/x-context/<rid>.json) at stash time, keyed by request_id so concurrent requests never overwrite each other; it survives inbox cleanup and restart. - fm-x-reply resolves platform/budget through registry -> inbox -> relay (the relay lookup confined to a live follow-up), recovering the original platform independent of task-link availability. - Fail-safe: a follow-up whose platform/budget cannot be authoritatively resolved and that would split is refused (exit 8) and held for retry, never wrongly split; fm-x-followup keeps the link on that exit. - fm-x-dismiss clears the durable context for a dismissed mention. Refactors reply-context extraction into a single owner and adds regression coverage for all four cases. * no-mistakes(review): Captain, fail closed on incomplete follow-up context * no-mistakes(review): Captain, bound X context registry retention * no-mistakes(review): Captain, align context retention with answer binding * no-mistakes(document): Align X follow-up context documentation * no-mistakes(document): Align durable X follow-up documentation * fix: preserve secondmate routing markers in terminal sends (kunchenguid#533) * fix: preserve secondmate routing markers * no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends * no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift * no-mistakes(document): Refresh Herdr marker documentation * fix: bound firstmate secondmate recovery during bootstrap * fix: align Grok effort handling with 0.2.99 (kunchenguid#527) * fix: align grok effort docs and spawn with 0.2.99 ceiling grok 0.2.99 accepts only low|medium|high for --reasoning-effort and rejects both xhigh and max. Omit unsupported values on spawn, flag them in crew-dispatch validation, and update harness-adapters. * no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies * no-mistakes(document): Clarify Grok effort documentation ownership * fix: bound and rotate secondmate recovery * fix: preserve bounded recovery failures * feat: replace Treehouse with Git worktrees * test: isolate dispatch profile fixtures * test: isolate backend worktree root * fix: register detached lifecycle worktrees * fix: harden lifecycle startup and spawn cleanup * docs: retire Treehouse provider references * fix(firstmate): keep supervision keeper alive * docs(firstmate): document supervision keeper * fix: reject shared codex app-server as firstmate holder * chore: make no-mistakes optional * fix: resolve residual conflict markers after merge The auto-merge left several orphan <<<<<< / >>>>> markers and a few nested blocks where git recorded a content conflict but the regex-based script-resolution left lines like >>>>>>> origin/main stranded between clean sections. This commit removes them, takes the union of HEAD-vs-origin/main local-variable declarations in three of the affected files, and keeps both sides' content where they were complementary additions. No behavior change; the affected scripts all pass 'bash -n'. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: mielyemitchell <mielyemitchell@gmail.com> Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> Co-authored-by: Ryan Fong <noreply@anthropic.com> Co-authored-by: Korallis <lee.barry84@gmail.com>
Marl0nL
added a commit
to Marl0nL/firstmate
that referenced
this pull request
Jul 16, 2026
* fix: ignore secondmate home marker during sync (kunchenguid#417) * fix: gitignore the secondmate home marker bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into every seeded secondmate home. A secondmate home is a worktree of the firstmate repo, so any plain `git status --porcelain` dirtiness check counted the untracked marker and the home read as dirty forever: fleet-sync reported it STUCK and the local fast-forward convergence sweeps risked leaving it stale on firstmate updates. Add .fm-secondmate-home to the tracked .gitignore so the marker is invisible to every dirtiness check uniformly, without weakening fleet-sync's deliberate untracked-counting for project clones. Convergence chicken-and-egg: existing homes predate the fix and it only arrives by fast-forward. The already-present marker-tolerant ff-skip (ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate, and spawn pre-launch) advances such a home past the fix commit, after which .gitignore takes over - no hand intervention. Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home reading clean, an existing marker-only home converging then reading clean, and a genuinely dirty home still skipping. * no-mistakes(review): Captain: document standalone-clone update path * no-mistakes(document): Document secondmate marker migration * fix(composer): prevent dead-shell message injection (kunchenguid#416) * fix(composer): stop reading dead-shell prompts as empty agent composers Consolidate composer empty/pending/unknown classification into one shared owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This replaces four drifting copies of the glyph decision. Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now classified unknown (a dead shell, unsafe for injection), not empty. It is only empty inside a bordered composer box (the harness's own prompt). Agent glyphs ❯ (claude) and › (codex) read empty either way. The away-mode injector (inject_msg) now requires an affirmatively-empty composer, deferring on pending or unknown, so an escalation can never be typed into (or executed by) a pane whose agent exited to its login shell. Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner; per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay green. * no-mistakes(review): Captain: harden composer safety checks * no-mistakes(test): Stabilize Herdr prune safety setup * no-mistakes(document): Document composer injection safety * no-mistakes(lint): Clean composer safety lint * no-mistakes: apply CI fixes * feat(watcher): add paused external-wait supervision (kunchenguid#421) * feat(watcher): add paused/awaiting-external crew state A crew (or firstmate steering it) can declare a deliberate wait on a known external dependency with a paused: <reason> status. Both the always-on watcher and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh vocabulary instead of tripping the possible-wedge stale escalation, and re-surface it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused distinctly. A crew that goes idle without declaring a pause classifies exactly as before. Docs and brief scaffold state lists updated; tests colocated. * no-mistakes(review): Captain: fix paused-state transitions * init * no-mistakes(review): Captain: fix paused-state supervision transitions * no-mistakes(review): Captain: fix paused supervision handoffs * no-mistakes(review): Reconcile paused supervision markers * no-mistakes(review): Captain: prioritize paused states over captain relevance * no-mistakes(review): Captain: preserve paused-working wedge timer * no-mistakes(review): Captain: honor configured pause verb in briefs * no-mistakes(test): Captain: fix AFK paused watcher handoff * no-mistakes(document): Document declared external waits * no-mistakes(lint): Clean paused-state lint --------- Co-authored-by: fmtest <fmtest@example.invalid> * fix: preserve X-mode follow-up platform limits (kunchenguid#425) * fix(x-mode): make follow-up platform splitting immune to link ordering A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280 chars because fm-x-link only learned the platform from the inbox payload, and the fmx-respond ack path can drain that inbox file before the task is linked. A link recorded after cleanup silently lost the platform and the splitter defaulted to the X 280-char budget. Make platform resolution ordering-proof: - fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a new fmx_request_relay_context helper (POST /connector/request-context) when neither the inbox payload nor carry flags carry it. The request_id survives the inbox drain, so a post-cleanup link still learns the right split budget. Best-effort: no token/curl or a non-2xx relay degrades to the loud warning below rather than a silent X default. - fm-x-link warns loudly when no platform source resolves, so the loss is never silent. - The fmx-respond procedure now orders link-before-inbox-cleanup so the fast local path stays correct without a relay round-trip. Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE message even when linked after inbox cleanup, and an unresolvable platform warns loudly instead of splitting silently. docs/configuration.md documents the request-context lookup. The relay endpoint is the companion durable change (see done status); until it ships, the link-before-cleanup reorder keeps the normal path correct. * no-mistakes(document): Document X-mode platform recovery * fix(composer): handle ANSI ghost text safely (kunchenguid#429) * fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor Away-mode injection wedged all night on the primary claude-on-herdr pane: the herdr composer classifier never stripped generic dim ghost text (only a narrow codex bold-wrapped byte-pattern check), so claude's rotating prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI pane read preserves - read as real pending input and every escalation deferred (6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s). Consolidate ghost extraction into one fleet-wide ANSI-aware owner, fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX, default 128, dark-theme assumption). Both ANSI-capable backends route through it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a dark box border dropped with the ghost does not lose the composer shape. This also closes the documented grok TRUECOLOR placeholder gap by the same mechanism (harness-adapters skill note updated). Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle) and the incident write-up are in docs/herdr-backend.md; deterministic regressions feed the exact captured bytes through the real classifiers (tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text (never a realistic real-input color) are corrected to a bright 38;2;224;222;244, preserving the truecolor payload-skip parser intent. * no-mistakes(review): Preserve dark shell prompt safety * no-mistakes(review): Harden erased shell prompt classification * no-mistakes(document): Document shared composer ghost extraction * no-mistakes(lint): Normalize tmux comment punctuation * fix(spawn): make tmux window handling robust under non-default config (kunchenguid#134) * test: isolate session-start suite from ambient harness markers (kunchenguid#432) * fix(session-start): isolate harness env markers in suite runner Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in run_session_start so ambient interactive shells cannot override the suite's fake ps harness (local-vs-CI split on the pi supervision case). * no-mistakes(document): Correct Pi marker documentation * fix(teardown): retry transient index locks during worktree return (kunchenguid#435) * fix(teardown): retry treehouse return on transient index.lock Killed crew git ops can leave a short-lived worktree index.lock that makes treehouse return fail. Retry on that error signature with a bounded wait (env-overridable), never force-delete a live lock, and only then fall back to the existing provably-stale cleanup path. * no-mistakes(review): Harden teardown retry configuration * no-mistakes(document): Document teardown index-lock retry behavior * no-mistakes(lint): Fix empty shell variable assignments * fix: complete brief help and consolidate documentation (kunchenguid#438) * docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage * fix: detect Git and centralize backend configuration (kunchenguid#445) * docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation * feat(daemon): add backend-independent wedge alerts (kunchenguid#444) * feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior * docs: centralize firstmate operating contracts (kunchenguid#447) * docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR kunchenguid#438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue kunchenguid#363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation * fix(cmux): close last workspace during teardown (kunchenguid#449) * fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior * fix: recover orphaned packed-refs locks during fleet sync (kunchenguid#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to create '...packed-refs.lock': File exists", leaving the clone unsynced. On that signature only, fm-fleet-sync.sh now retries the fetch with a bounded wait (transient locks self-clear), then removes the lock and retries once more ONLY when it is provably stale: still present, mtime age past a threshold, and no lsof holder of the lock file or of the clone worktree itself (a live git keeps that as its cwd even in the window after it closes the lock and before it exits). A live lock, a missing lsof, any failed check, or any other fetch failure keeps today's behavior. Every wait/retry/removal prints to stderr, and a successful recovery also prints one "recovered:" summary to stdout so a session-start refresh - which discards fleet-sync stderr and relays only stdout - still surfaces it. The shared "is this git lock provably abandoned?" proof is extracted into bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/. * no-mistakes(review): Captain, remove obsolete teardown wake dependency * no-mistakes(document): Document packed-refs lock recovery architecture * feat(herdr): escalate blocked panes immediately (kunchenguid#472) * feat(herdr): immediate blocked-state escalation via native events.subscribe push Fold herdr's native pane.agent_status_changed stream into the single watcher so a crew entering blocked wakes its supervisor sub-second (measured 0.129s) instead of after the ~240s stale-pane wedge timer. - bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape plus the single-owner status->action policy table (blocked=actionable, working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling). - bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe subscriber over one connection for all this home's herdr panes, subscribing to ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe marker and a reconnect level-reconcile. Version/schema capability gate. - bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so the watcher stays backend-agnostic and the shape+policy are reusable. - bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal wait primitive (replacing the blind sleep POLL for push-capable homes), behind a source guard so the splice is unit-testable; secondmate/paused exemptions; map pane->window->task and enqueue a stale wake. No second watcher process; the single-cycle invariant and every guard/beacon/turn-end mechanism are unchanged. - Polling stays the permanent fail-closed backstop: below-capability, subscribe failure, and repeated runtime failures all degrade to sleep. - Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/ fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events) plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries the dated evidence and retires the old gap note. * no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs * no-mistakes(review): Captain, commit markers after wake and reuse capability cache * no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs * no-mistakes(review): Captain, subscribe before Herdr reconciliation * no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe * no-mistakes(test): Captain: include lock library in teardown fixture * no-mistakes(document): Captain: document Herdr immediate blocked escalation * fix: clarify shellcheck conditionals * docs(readme): reposition firstmate as an agent distro (kunchenguid#473) * feat: add deterministic bounded bearings snapshots (kunchenguid#475) * feat(bearings): deterministic bearings snapshot + durable decision model Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR discovery and checks happen only under --include-prs, which fails soft. Every dropped surface is marked in omitted[] with the flag that reveals it, and the prs: line states when checks were not requested, so absence is never silent. Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib gains status_open_decisions, the one authoritative keyed open/resolved fold over the whole status stream: needs-decision/blocked opens a keyed entry, only an explicit keyed resolution (or, for run-backed tasks, run-step advancement) closes it, so a later unrelated done/paused can no longer mask a still-open captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives pending_decision/blocked_event from it; the canonical schema stays complete. Point the /bearings skill at the one command; add the resolved: writer line to ship, scout, and secondmate briefs. Register the script and add regression tests for the output bound, TOON/JSON parity, local-only default, opt-in PR fetch, partial-failure degradation, decision durability, and report pointers. * fix(bearings): completed scout report is a pointer, not a pending decision A completed scout that raised a needs-decision and then finished (done) without a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103 case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh cleared a stale decision only for a live run-step/pane activity read, so a terminal task whose current state is read from the status log (a scout or ship that reached done/failed) never cleared its stale, never-keyed-resolved needs-decision, and it lingered as pending. The open-decision set is still derived purely from the keyed fold - never from a report body or decision-like prose - and reconciled against the crew lifecycle. Extend that reconciliation so a terminal done/failed state on a single-owner task (scout or ship), whose deliverable is its report or PR, also clears the set; a completed scout now surfaces only as a report pointer. Secondmates are excluded from the terminal clear (persistent, multiplexed stream), which keeps the unrelated-event masking fix intact. Add regression tests: a completed scout with decision-like report prose is a pointer not pending (canonical + end-to-end), and a scout still parked at a decision stays pending so the terminal clear never over-fires. * no-mistakes(review): Captain, preserve keyed decisions across shared status parsing * no-mistakes(review): Captain, close blockers and harden keyed decision parsing * no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout * no-mistakes(review): Captain, bound bearings sections and fail closed * no-mistakes(review): Captain, disclose capped per-repository PR results * no-mistakes(document): Refresh bearings documentation and status contracts * fix(bearings): avoid ambiguous worktree guard * fix: enforce deterministic ShellCheck parity (kunchenguid#481) * fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity Firstmate PRs passed local no-mistakes validation but failed CI's "Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034). Two divergences caused it: 1. The no-mistakes gate had no commands.lint, so its lint step never ran the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that CI runs. Confirmed from state.sqlite: the lint step_result recorded findings:null with no lint agent invocation. 2. CI's shellcheck floated with the runner image while local ran a newer build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck rejected an SC2015 that the newer local one no longer emits. Establish bin/fm-lint.sh as the single owner of the lint definition: the file set, the config, and the pinned shellcheck version (0.11.0, printed via --required-version). Both CI (.github/workflows/ci.yml) and the no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs the exact version it names and logs the resolved version, and fm-lint.sh refuses to lint under any other version. This is not a CI relaxation: it adopts shellcheck 0.11.0's rule set consistently, dropping only the upstream-retired, false-positive-prone SC2015; default severity and every still-supported finding stay enforced (no severity downgrade, no excludes). tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs and logs the pinned version, that the owner refuses a non-pinned shellcheck, and that it rejects a real lint defect the old no-op gate passed. * no-mistakes(review): Captain, harden deterministic ShellCheck parity * no-mistakes(review): Captain, neutralize ambient ShellCheck overrides * feat: guard primary shells from persistent cd commands (kunchenguid#483) * feat: add cd-guard PreToolUse seatbelt for the primary shell A stray persistent top-level `cd projects/<clone>` in the primary firstmate shell relocates the shell, so a later firstmate-owned command (a backlog write, an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the home. The cd-guard denies exactly that command shape before it runs, across all five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt. - bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate lexer; that file's CLI now runs only when invoked directly). - bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness output rendering, and primary-checkout scoping - fires in a secondmate's own primary session, inert in crew/scout child worktrees and non-firstmate repos. - Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents; per-harness hooks only call the owner. - Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd, and command cd). Allows git -C, subshell / bash -c / env -C / make -C / find -execdir, pipeline and background forms, and cd-as-data. Fails open on malformed input; agent-mistake threat model. - tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix, end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring. - docs/cd-guard.md: full contract plus live validation (claude, codex, opencode, pi blocked end-to-end; grok live run blocked by an API balance limit, with mechanism parity and deterministic coverage recorded). * no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage * no-mistakes(review): Captain, allow path-qualified command wrappers * no-mistakes(review): Captain, allow non-executing command queries * no-mistakes(test): Captain, clarify cd-guard safe-path remediation * docs: clarify cd guard guidance * no-mistakes(document): Clarify cd-guard safe target guidance * brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (kunchenguid#267) Crews must never stop, restart, or update the shared no-mistakes daemon since one instance serves every firstmate lane/home; a restart kills other lanes' in-flight pipeline runs and forces expensive re-runs. Encodes this as a new numbered rule in both the ship-task and scout-task brief scaffolds. Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> * feat: make bearings concise and accurate (kunchenguid#485) * feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render /bearings skill (one owner of the chat-response format): mandate the four always-present chat sections - Captain's Call, Recently Landed, Underway, Charted Next - each with an explicit empty-state sentence, no At Anchor, materially shorter than and linking to the report file. Resolves the ambiguous Check first / Decisions pending split into one strict captain-action section. fm-crew-state: the log fallback derives current state only from a real run-state verb, so a trailing decision-closing resolved: event no longer renders a healthy idle crew (typically a secondmate) as unknown with the resolution prose as its detail. The keyed-decision contract in fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner. fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done records from registered secondmate homes, reusing the single backlog parser and the one secondmate-home enumerator (meta home= with data/secondmates.md fallback); no network, per-home capped. fm-bearings-snapshot: landed now merges main-home Done with the secondmate roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure (also fixing the previously-silent landed truncation); --all-landed reveals the full set. tests: resolved-event state render, secondmate landed aggregation with caps and omitted[] disclosure, Captain's Call anti-leak, and the four-section contract. * no-mistakes(review): Captain, ensure bearings reveals all landed work * no-mistakes(document): Document bearings accuracy contracts * fix: harden away-mode daemon lifecycle (kunchenguid#490) * fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle Away-mode entry left "make the daemon a tracked background terminal" to the operator; on a pi/herdr primary that meant splitting the captain's active pane, which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches the daemon in a non-visible tracked terminal per backend (herdr dedicated --no-focus workspace, detached tmux session), never a split, pins the captain pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal id, and tears it down or reconciles a leaked one by that id. No shell &. Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared with the daemon (one owner). Fix the stale subsuper-artifact leak: clear the prior away session's delivery cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon before clearing state/.afk so its shutdown flush runs instead of being a no-op. Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md, docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md. * no-mistakes(review): Captain, serialize AFK launcher lifecycle safely * no-mistakes(review): Captain, harden AFK launcher lifecycle races * no-mistakes(review): Captain, ensure AFK daemon launch readiness * no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown * no-mistakes(review): Captain, preserve AFK reconciliation records uniformly * no-mistakes(review): Captain, harden AFK recovery state durability * no-mistakes(review): Captain, harden AFK tmux ownership checks * no-mistakes(review): Captain, simplify AFK lifecycle failure handling * no-mistakes(review): Captain, require confirmed AFK daemon shutdown * no-mistakes(review): Captain, confirm AFK exit by process identity * no-mistakes(document): Align AFK launcher lifecycle documentation * no-mistakes: apply CI fixes * fix: prevent no-mistakes gate agents from driving the fleet (kunchenguid#518) * feat: contain no-mistakes gate agents from driving the fleet Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/ fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE is set, and via an unspoofable git-common-dir backstop when invoked from a no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker unset. A normal firstmate session has neither signal and is unaffected. Set disable_project_settings: true in the tracked .no-mistakes.yaml so the installed pipeline neutralizes gate agents' project instructions for this repo (trusted-only, honored from the default branch). firstmate's own suite runs from a gate worktree during validation, so the shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated tests/fm-gate-refuse.test.sh strips it to verify real refusal. * no-mistakes(review): Captain, refuse empty no-mistakes gate markers * no-mistakes(document): Document no-mistakes gate authority boundary * fix: guard secondmate primary sessions from blind turn ends (kunchenguid#505) * fix: guard secondmate own-home turn ends Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the 'no turn ends blind' backstop fires in a secondmate's own primary session, matching the cd-guard's scope: the own home is guarded, child crew/scout worktrees stay exempt via the retained git-dir/git-common-dir test. This was pure scoping from the guard's primary-only origin and guarded against no secondmate-specific hazard. Add secondmate regression tests (blind-turn block, idle-by-default, stop_hook_active loop guard, deferred-death recovery loop, child-worktree exemption) and record the autonomous background-notify re-invoke measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md. * no-mistakes(document): Correct secondmate guard documentation, captain * fix: force-include marked secondmate homes in turn-end guard The prior remove-only form (just deleting the .fm-secondmate-home check) left the DEFAULT secondmate topology unguarded: a treehouse-leased home is a linked git worktree (git-dir != git-common-dir), which the retained git-dir exemption still skipped, so its own primary session could still end a turn blind. Invert the marker: a genuinely-marked home is force-included as a guarded primary (treehouse-leased linked OR git-cloned plain), and the git-dir exemption applies only to UNMARKED child worktrees. Marker validation (regular non-symlink file, non-empty id-token content) blocks a stray or empty marker from spoofing inclusion. Add real linked-worktree regression tests: a treehouse-leased LINKED secondmate home is guarded, a stray/empty marker stays exempt, and the unmarked child worktree stays exempt - the topology the plain git-init fixtures masked. Predicates, in-flight gate, and loop guard untouched. * fix: force ASCII collation in secondmate marker validation Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the [A-Za-z0-9._-] id allowlist matches under C collation, not the ambient locale - a locale-crafted non-ASCII marker id can no longer slip through the range match and spoof force-inclusion of a linked child worktree. Add a regression test proving a non-ASCII marker id is rejected and the linked worktree stays exempt. * no-mistakes(test): fix backend baseline gate-refusal dependency * no-mistakes(document): Correct secondmate turn-end guard documentation * fix: make bootstrap diagnostics backend-aware (kunchenguid#519) * fix: make bootstrap required-tool detection backend-aware Bootstrap demanded tmux and treehouse for every backend except orca, so a herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux. Required tools now follow the resolved backend via the single-owner fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux), and treehouse for session-provider-only backends (orca owns its worktree). The treehouse lease-support check is gated to backends that use treehouse. Adds install hints for herdr/zellij/cmux, regression tests for the full backend dependency matrix (herdr-without-tmux repro plus each boundary), and updates the authoritative Toolchain docs. * no-mistakes(review): Captain, prevent executing Herdr install guidance * no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics * no-mistakes(review): Captain, separate manual dependency remediation * no-mistakes(review): Captain, align bootstrap diagnostic consumers * no-mistakes(document): Align backend adapter dependency comments * fix: preserve follow-up platform context after inbox cleanup (kunchenguid#520) * fix: recover X/Discord follow-up platform after inbox cleanup A milestone follow-up posted directly by request_id after the inbox was drained - and with no task link, because one persistent secondmate's single x_request slot collides across concurrent requests - resolved platform only from the local inbox, so a >280 Discord reply silently defaulted to the X 280-char budget and threaded as (1/2). - fm-x-poll records a durable per-request reply context (state/x-context/<rid>.json) at stash time, keyed by request_id so concurrent requests never overwrite each other; it survives inbox cleanup and restart. - fm-x-reply resolves platform/budget through registry -> inbox -> relay (the relay lookup confined to a live follow-up), recovering the original platform independent of task-link availability. - Fail-safe: a follow-up whose platform/budget cannot be authoritatively resolved and that would split is refused (exit 8) and held for retry, never wrongly split; fm-x-followup keeps the link on that exit. - fm-x-dismiss clears the durable context for a dismissed mention. Refactors reply-context extraction into a single owner and adds regression coverage for all four cases. * no-mistakes(review): Captain, fail closed on incomplete follow-up context * no-mistakes(review): Captain, bound X context registry retention * no-mistakes(review): Captain, align context retention with answer binding * no-mistakes(document): Align X follow-up context documentation * no-mistakes(document): Align durable X follow-up documentation * fix: preserve secondmate routing markers in terminal sends (kunchenguid#533) * fix: preserve secondmate routing markers * no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends * no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift * no-mistakes(document): Refresh Herdr marker documentation * fix: align Grok effort handling with 0.2.99 (kunchenguid#527) * fix: align grok effort docs and spawn with 0.2.99 ceiling grok 0.2.99 accepts only low|medium|high for --reasoning-effort and rejects both xhigh and max. Omit unsupported values on spawn, flag them in crew-dispatch validation, and update harness-adapters. * no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies * no-mistakes(document): Clarify Grok effort documentation ownership * fix: derive bearings from authoritative secondmate state (kunchenguid#555) * fix: make bearings use secondmate home state * test: anonymize bearings fixtures * no-mistakes(review): Bound parent activity evidence scans, captain * no-mistakes(review): Preserve structured secondmate authority and bounds, captain * no-mistakes(review): Preserve registry completeness and child inventory, captain * no-mistakes(review): Reconcile parent evidence by verb and key, captain * no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain * no-mistakes(document): Document bearings local snapshot and PR opt-in * no-mistakes(lint): Fix fleet snapshot ShellCheck findings * no-mistakes: apply CI fixes * fix: restore fleet snapshots on stock macOS Bash (kunchenguid#578) * fix: restore stock macOS snapshot parsing * no-mistakes(document): Clarify Linux gate and macOS CI coverage * fix(afk): make Pi escalation and return catch-up reliable (kunchenguid#587) * fix: close away-mode blocker supervision gap * no-mistakes(review): Gate teardown retries and verify U+2063 dedupe * no-mistakes(test): Fail closed on incomplete Pi composer separators * no-mistakes(document): Document Pi composer recognition and return gating * feat: support Pi max reasoning profiles (kunchenguid#537) * support Pi max thinking profiles * no-mistakes(review): Captain, allow Pi max dispatch profiles * chore: no-mistakes(document): Clarify yolo response ownership (kunchenguid#595) * Clarify validation response ownership * no-mistakes(document): Clarify yolo response ownership * feat: establish instruction ownership foundation (kunchenguid#619) * Add instruction owners foundation * no-mistakes(document): Refresh project-management owner pointers * fix: compress Firstmate contract and enforce delivery rigor ownership (kunchenguid#626) * docs: compress firstmate operating contract * docs: make delivery rigor single-owner PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion. * no-mistakes(review): Honor configured merge authority across faster delivery paths * no-mistakes(document): Align docs with compressed operating contract --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: fmtest <fmtest@example.invalid> Co-authored-by: Pierre Marais <pierremarais67@gmail.com> Co-authored-by: mielyemitchell <mielyemitchell@gmail.com> Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com> Co-authored-by: Korallis <lee.barry84@gmail.com> Co-authored-by: Marlon <marlonleicester@gmail.com>
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Jul 17, 2026
* feat(backends): add experimental cmux runtime backend (#246)
* feat(backends): add cmux runtime backend (experimental)
Session-provider-only adapter for cmux (bin/backends/cmux.sh), mirroring
zellij/herdr structurally, wired into fm-backend.sh and fm-spawn.sh with
--secondmate refused for now. Verified against the real cmux 0.64.17 app:
send does not auto-submit, cwd is creation-time-frozen (zellij-shape,
pwd-marker-probe workaround), close-surface refuses on a workspace's last
surface (falls back to close-workspace), workspace ids do not survive a
relaunch, and the control socket defaults to cmuxOnly access (requires a
one-time password-mode setup, documented in docs/cmux-backend.md). Also
found and fixed a live bug during development: read-screen fails on a
surface that has never been written to, so liveness now uses list-panes
instead. Fake-CLI unit suite (40 tests), a real-binary smoke test, and a
full spawn/steer/peek/done/merge/teardown E2E pass against a real claude
crewmate all pass, including the popup/second-Enter regression class.
* no-mistakes(review): Harden cmux recovery and password parsing
* no-mistakes(review): Harden cmux capture failure handling
* no-mistakes(review): Mark cmux test scripts executable
* no-mistakes(review): Scope cmux workspaces and teardown
* no-mistakes(review): Captain, honor cmux password config override
* no-mistakes(review): Captain, hash cmux home labels
* no-mistakes(document): Sync cmux backend docs
* feat(agents): add firstmate coding guidelines skill (#248)
* Add firstmate-coding-guidelines skill (AGENTS.md diet PR 0)
Encodes the knowledge-placement decision tree, one-owner rule, and
inline-stub pattern from the diet analysis so future contributions stop
adding conditional detail inline. AGENTS.md gets one section-13 trigger
line; fm-brief.sh's REPO argument has no reliable signal for "this is
firstmate's own repo", so the load instruction goes in CONTRIBUTING.md's
Development section instead of the scaffold.
* no-mistakes(review): Captain, align tracked-material trigger scope
* no-mistakes(document): Sync coding guidelines docs
* no-mistakes(lint): Fix Markdown style issues
* fix: add turn-end supervision guard (#249)
* feat: structural Stop-hook backstop for primary turn-end supervision
fm-guard.sh is pull-based: it only warns when some other supervision
script happens to run, so a primary session that ends a turn without
re-arming the watcher and then runs no further fleet-touching command
can sit blind for hours (the 2026-07-04 incident this fixes).
Add bin/fm-turnend-guard.sh, a Claude Code Stop hook registered in the
tracked .claude/settings.json, that fires on every primary turn end and
blocks (exit 2, verified empirically to force continuation) when work
is in flight with no fresh watcher beacon. It never blocks more than
once per turn, using Claude Code's own stop_hook_active loop-guard
field, and scopes itself to the actual primary checkout only (inert in
crewmate/scout worktrees and secondmate homes).
Factor the shared "in-flight but no live watcher" predicate out of
fm-guard.sh into bin/fm-supervision-lib.sh so the pull-based banner and
the push-based hook can never drift on what "unhealthy" means.
Document the verified Stop-hook mechanism and scoping in
docs/turnend-guard.md, add a harness-adapters note, and cover the
predicate and hook with tests/fm-turnend-guard.test.sh.
* no-mistakes(review): Respect active home in turnend guard
* no-mistakes(review): Require live watcher for turn-end guard
* no-mistakes(review): Captain: portable turn-end timing
* no-mistakes(document): Sync turn-end guard documentation
* feat(backends): auto-detect cmux runtime (#250)
* feat(backends): auto-detect cmux runtime from CMUX_WORKSPACE_ID
Wires cmux into fm_backend_detect the same way herdr already is: a
firstmate process running inside a cmux-spawned terminal now spawns
new tasks into cmux by default, no config needed. Verified from cmux's
own shipped source that CMUX_WORKSPACE_ID/CMUX_SURFACE_ID/CMUX_SOCKET_PATH
are unconditionally, non-overridably injected into every terminal
surface it spawns, and that cmux's own CLI treats CMUX_WORKSPACE_ID as
its own ambient-target fallback - the same role $TMUX/HERDR_ENV play for
their backends. CMUX_WORKSPACE_ID is checked last (after $TMUX and
HERDR_ENV=1) since cmux is a terminal application, not a nestable
multiplexer. Socket auth (config/cmux-socket-password) stays required
regardless of how the backend was selected; the existing spawn refusal
now also names the config/backend=tmux / --backend tmux opt-out for a
caller who never explicitly chose cmux.
A live env dump inside a real cmux terminal was not obtained safely on
the shared dev machine (documented in docs/cmux-backend.md); this rests
on the source read instead, mirroring this doc's existing
verified-from-source precedent.
* no-mistakes(review): Fix cmux autodetect docs and tests
* no-mistakes(document): Document cmux auto-detection
* fix(afk): support herdr away-mode injection (#251)
* fix(afk): make the away-mode daemon backend-aware for herdr
bin/fm-supervise-daemon.sh discovered its supervisor pane and injected
via raw tmux calls only, so /afk failed outright on a herdr-based
fleet (TMUX_PANE unset, firstmate:0 fallback unresolvable).
Discovery now resolves backend (tmux|herdr) and target independently,
mirroring fm-backend.sh's own runtime auto-detection, with an explicit
FM_SUPERVISOR_BACKEND override alongside the existing FM_SUPERVISOR_TARGET.
zellij/orca refuse loudly at startup instead of misapplying tmux
primitives. Injection (pane-exists probe, busy-guard, composer-guard,
verified submit) now dispatches through bin/fm-backend.sh's generic
primitives, adding a new fm_backend_composer_state dispatcher; the
tmux path is byte-identical to before. Also fixes a pre-existing bug
in fm_backend_target_exists's herdr arm (missing --session, so it
silently misrouted once more than one herdr server was running) found
while verifying this end to end against a real isolated herdr session.
Classification, batching, max-defer, the marker contract, locks, and
wake-queue handling are unchanged - this is a transport-layer fix.
* no-mistakes(review): Corroborate Herdr idle busy state
* no-mistakes(review): Stabilize Herdr daemon startup wait
* no-mistakes(review): Captain, route cmux composer and update AFK docs
* no-mistakes(document): Document AFK supervisor backend support
* docs(agents): move X-mode procedures out of AGENTS (#253)
* docs(agents): collapse X-mode section 14 into fmx-respond/docs pointers
AGENTS.md diet PR 1 of 3 (agentsmd-diet-s2 report, move-plan items 1-2).
Replaces section 14's "Answering"/"Completion follow-up"/"Conversations"/
"Length and threads"/"Preview / dry-run" blocks (54 lines) and the
"Mechanism" narrative (6 lines) with two short pointers: fmx-respond
(section 13) for the procedure, docs/configuration.md "X mode (.env)"
for the wire protocol. Net -55 lines in AGENTS.md.
Destination edits landed first, deletions second (q4 discipline):
- docs/configuration.md: added the "purely additive, watcher untouched"
guarantee that AGENTS.md's Mechanism block stated but configuration.md
did not.
- fmx-respond/SKILL.md: added the x-mode-error wake boundary (report as
a blocker, do not load this skill), the --image flag for replies and
follow-ups, the "images are for real artifacts, not prose" rule, and
the dry-run compact-image-marker behavior - none of these were
previously in the skill even though AGENTS.md described them, so they
were genuine gaps, not pre-existing duplication. Also made the skill's
own "Completion follow-up" section the sole, full owner of that
procedure instead of deferring to AGENTS.md section 14 for substance
that no longer lives there (two internal cross-references updated to
point at section 8's terminal-wake trigger and the skill's own section
instead).
Mechanical line-by-line audit of every removed AGENTS.md line:
Mechanism block (6 lines removed):
- bootstrap artifact-writing description -> already owned by
docs/configuration.md "X mode (.env)" (locked-bootstrap paragraph)
- check-shim/poll mechanism description -> already owned by
docs/configuration.md same section
- missing-deps/x-mode-error diagnostic description -> already owned by
docs/configuration.md ("Relay auth or config problems...") plus
bin/fm-x-poll.sh's own header comment for the missing-curl/jq mechanics
- opt-out artifact removal description -> already owned by
docs/configuration.md same section
- "purely additive, no edit to fm-watch.sh/fm-watch-arm.sh/fm-wake-lib.sh/
afk daemon" guarantee -> MOVED to docs/configuration.md (added in this
PR; this fact had no other home before)
Answering/Completion follow-up/Conversations/Length and threads/
Preview-dry-run blocks (54 lines removed):
- x-mention wake -> load fmx-respond: already owned by section 13's
existing trigger line (unchanged) and restated in the new pointer
- x-mode-error wake -> report as blocker, don't load fmx-respond: MOVED
to fmx-respond/SKILL.md (added in this PR)
- inbox-draining, classification, acting, reply composition, submission,
cleanup-on-success/failure: already owned by fmx-respond/SKILL.md
"Procedure" section (unchanged, pre-existing)
- owner-only routing / captain-as-asker framing: already owned by
fmx-respond/SKILL.md "The asker is your own captain" section
- standing X-mode authorization / autonomous posting / dry-run as only
non-posting path: already owned by fmx-respond/SKILL.md same section
- acknowledge-first -> act -> follow-up shape, three-case classification:
already owned by fmx-respond/SKILL.md "A request to act on" section
- destructive/irreversible/security-sensitive escalation guardrail:
already owned by fmx-respond/SKILL.md "Public channel..." section and
Procedure step 2c
- dismiss-instead-of-reply for pure acknowledgments, relay re-offer
prevention, dry-run honoring: already owned by fmx-respond/SKILL.md
Procedure steps 2b/2c/2e-skip and docs/configuration.md
- public-safety bar (no task ids/internals/captain-private/secrets):
already owned by fmx-respond/SKILL.md "The reply is public" section
- never-inline-into-shell-command / --text-file or stdin: already owned
by fmx-respond/SKILL.md Procedure step 2e and Notes
- --image flag for replies (formats, base64, no-inline guarantee): MOVED
to fmx-respond/SKILL.md Procedure step 2e (added in this PR - this was
not previously in the skill)
- fm-x-link field names (x_request=, x_request_ts=, x_followups=):
already owned by AGENTS.md section 2's state/<id>.meta field list
(untouched, out of scope for this PR) and fmx-respond/SKILL.md
- carry-count/carry-ts relink behavior, three-follow-up budget, milestone
sparingness, --check/--text-file posting, connector/followup wire
detail, --final clearing, cap/window graceful degradation: already
owned by fmx-respond/SKILL.md "Completion follow-up" section (now sole
owner) and docs/configuration.md wire-protocol paragraphs
- --image flag for follow-ups: MOVED to fmx-respond/SKILL.md "Completion
follow-up" section (added in this PR - genuine gap)
- "failed task still gets an honest final follow-up": already owned by
fmx-respond/SKILL.md "Completion follow-up" section
- FMX_DRY_RUN whole-loop previewability: already owned by
fmx-respond/SKILL.md "Dry-run / preview mode" section
- in_reply_to conversation continuity, untrusted-thread handling,
follow-up worthiness judgment, relay-owned self-reply guard/cap:
already owned by fmx-respond/SKILL.md "The direct ask is the captain's"
section and Notes (one bullet is a verbatim match)
- concise-by-default / no hand-numbered threads: already owned by
fmx-respond/SKILL.md "Voice" section
- auto-split behavior, char/tweet caps, premium-independence, wire shape
({text}/{text,texts}): behavior already owned by fmx-respond/SKILL.md
Voice section; exact defaults and wire shape already owned by
docs/configuration.md; "premium-independent" mechanics already owned
by bin/fm-x-reply.sh's own header comment
- "images are for real artifacts, not prose": MOVED to fmx-respond/
SKILL.md "Voice" section (added in this PR - genuine gap)
- image-on-thread wire behavior: already owned by docs/configuration.md;
reinforced in fmx-respond/SKILL.md's new --image note
- dry-run POST-body shape, endpoint marker, truthy-value definition,
jq-only dependency, end-to-end testability, x-outbox inspection:
already owned by fmx-respond/SKILL.md "Dry-run / preview mode" section
(several near-verbatim matches) and docs/configuration.md wire detail
- dry-run compact image marker: MOVED to fmx-respond/SKILL.md "Dry-run /
preview mode" section (added in this PR - genuine gap)
Section 8's terminal-wake completion-follow-up trigger (the one fact
required to survive inline) is untouched and already present; the new
section 14 pointer references it instead of restating it.
Nothing outside section 14 (plus the two destination files) is touched.
Full test suite green, including all 74 fm-x-mode.test.sh checks.
* no-mistakes(review): Preserve X-linked follow-up triggers
* no-mistakes(review): Fix x-mode error trigger
* no-mistakes(document): Docs cross-reference synchronized
* no-mistakes(lint): Clean Markdown lint pass
* docs: trim duplicated harness guidance (#255)
* docs(agents): trim section 4 harness/secondmate duplication
AGENTS.md diet PR 2 of 3 (data/agentsmd-diet-s2/report.md, move-plan
items 3-4; redundancy item 2 folded into item 3).
Removed the five claude/codex/grok/pi/opencode model/effort-flag
bullets from section 4 - byte-for-byte duplicated by
harness-adapters' "Launch profile axes" table (which is already a
superset: it carries verified CLI versions per adapter that the
AGENTS.md bullets lacked). Replaced with a one-line pointer; the
skill is already loaded before every spawn per section 4's own
closing trigger, so no new trigger was needed.
Moved the config/secondmate-harness model/effort pin-format detail
(the `<harness> [<model>] [<effort>]` line format, the
secondmate-model/secondmate-effort accessors, back-compat, and the
durability-across-respawn behavior) into secondmate-provisioning,
which is already a mandatory load at every secondmate lifecycle
touchpoint. Added the destination content to the skill first, then
replaced the AGENTS.md paragraph with a 3-line pointer.
Mechanical audit - every removed line's new home:
- 5 harness bullets (claude/codex/grok/pi/opencode model+effort
flags, per-harness max-omission rationale) -> already present in
harness-adapters SKILL.md's "Launch profile axes" table (lines
53-59), confirmed fact-by-fact before deleting.
- "config/secondmate-harness may also pin..." paragraph (pin format,
bare-harness back-compat, secondmate-model/secondmate-effort
accessors, per-spawn override precedence, respawn durability,
secondmate-only scope) -> secondmate-provisioning SKILL.md's
"Charter and seed" section, added verbatim before this trim.
- The following paragraph (inheritable config: crew-dispatch.json,
crew-harness, backlog-backend) is untouched - out of scope for
this PR, still inline.
- The bootstrap CREW_DISPATCH effort-mismatch diagnostic sentence is
untouched - not part of the five-bullet duplication, stays inline.
No script changes. Section 4 shrinks from 104 to 87 lines
(958 -> 901 total AGENTS.md lines) with zero facts lost: every fact
is reachable through harness-adapters or secondmate-provisioning,
both already mandatory loads at the relevant lifecycle points.
* no-mistakes(document): Align secondmate skill triggers
* no-mistakes(lint): Markdown style clean
* fix: anchor turn-end Stop hook to project root (#256)
* Fix turn-end Stop hook to use CLAUDE_PROJECT_DIR path
Claude Code runs hook commands via /bin/sh from the session cwd, so the
bare relative bin/fm-turnend-guard.sh path fails when cwd is not the repo
root. Anchor the command with "$CLAUDE_PROJECT_DIR"/bin/fm-turnend-guard.sh
instead; verified CLAUDE_PROJECT_DIR is set on Stop hooks in Claude Code
2.1.201. Document the cwd caveat and add a settings.json regression test.
* no-mistakes(document): Document Stop hook path anchoring
* docs: trim firstmate agent guidance duplication (#258)
* docs: trim AGENTS.md redundancy (diet PR 3/3)
Consolidates five duplicated passages to a single owner each, per
data/agentsmd-diet-s2/report.md redundancy items c3-c7:
- Inheritable-config propagation mechanism: owned by section 3 (where
the sweep runs); sections 4 and 7 keep compact references. Section 4
retains its one genuinely unique fact (crew-harness inherit-vs-fallback
semantics), just no longer restates the propagation mechanism itself.
- Landed-work definition: owned by section 7's ship-teardown detail
(PR-containment mechanics, pr= discovery fallback); section 1's hard
rule #3 keeps the rule plus a three-case summary and a pointer.
- Backend meta-field enumeration: owned by docs/configuration.md
("Runtime backend", already comprehensive including cmux) and each
backend's own doc; AGENTS.md keeps only the fields common to every
task plus a pointer.
- Dropped one redundant restatement of "silence is correct while
waiting" in section 8.
- Worktree-tangle guard explanation: owned by section 8 (already the
fuller, cross-referenced version); section 3's TANGLE bullet keeps
the remediation action and points at section 8 for the why.
Also adds two captain-requested single-sentence rules: invoke bin/
scripts by absolute $FM_ROOT path after any cd away from the home, and
a backend spawn refusal must be surfaced to the captain rather than
silently worked around by switching backends.
AGENTS.md: 901 -> 889 lines, 112355 -> 108560 bytes.
* no-mistakes(review): Clarify post-cd bin invocation guidance
* no-mistakes(document): Sync AGENTS trim docs
* no-mistakes(lint): Fix Markdown line style
* feat(backends): improve cmux detection and socket-mode guidance (#259)
* feat(backends): cmux detection fallbacks and socket-mode matrix
Workstream A: cmux's bundled claude wrapper strips every CMUX_* env var on
its passthrough path (reproduced live 2026-07-04, cmux 0.64.17), so a
claude-harness firstmate inside a cmux tab has no CMUX_WORKSPACE_ID.
fm_backend_detect now falls back - macOS-only, only when the primary marker
is absent - to __CFBundleIdentifier=com.cmuxterm.app and then a process
ancestry walk resolved by bundle id (lsappinfo) plus a bundle-shaped ps comm
match. Innermost-first ordering is unchanged and absorbs the
tmux-inside-cmux bundle-id false positive; the auto-detect NOTICE names the
winning fallback signal.
Workstream B: the five socketControlMode values were traced through cmux
source (commit 9c91710e3f58): off/cmuxOnly can never admit an external CLI,
automation admits same-user clients with no secret (0600 socket only),
password needs the auth handshake, allowAll opens the socket to every local
user (0666). Automation mode is now the documented recommendation; the
adapter's refusals name every viable mode, classify Invalid password as
unauth, and the launch-timeout message names the off-mode possibility.
Docs carry the wrapper-strip empirical record, the fallback contract and
authority split, and the full mode matrix with rationale; tests cover the
new detection paths, the nested false positive, and the refusal wording.
* no-mistakes(review): Document cmux fallback detection
* no-mistakes(review): Update cmux architecture docs
* no-mistakes(document): Align cmux backend docs
* fix(backends): scope zellij tabs by firstmate home (#252)
* fix(backends): home-scope zellij tab titles to close cross-home collision gap
Zellij's one shared "firstmate" session has no per-home split and enforces
no tab-name uniqueness, so two firstmate homes with colliding task ids could
send/peek/close each other's tabs - the same gap a no-mistakes review gate
caught for cmux (docs/cmux-backend.md). Ports that fix: every new tab is
created with a home-scoped title (fm-<home-label>-<id>), and every
list/find/recover/kill path scopes matches to this home's own tag. A tab
spawned before this change still matches via its old untagged bare title,
but only when unambiguous - two live tabs sharing a bare title refuse rather
than guessing which one is ours.
Factors the home-label/hash derivation shared with cmux into
bin/fm-backend-hometag-lib.sh so the two adapters can't drift.
* no-mistakes(review): Fix zellij child teardown home tag
* no-mistakes(review): Fix zellij teardown and selector scoping
* no-mistakes(document): Sync zellij home-scope docs
* fix: sync project clones after merged PR wakes (#293)
* fix(fleet-sync): auto-sync on merged-PR wake, accept project name
fm-fleet-sync.sh's single-project form failed on a bare project name
("not a directory"), forcing hand-typed full paths (4 manual runs in
one incident). It now resolves a bare name or projects/<name> against
the home's projects dir.
AGENTS.md now encodes the trigger: a wake whose status reports a
merged PR for a project cloned in this home runs fleet-sync for that
project as part of handling the wake, so a secondmate-reported merge
does not leave the primary's clone stale until the next session start
or teardown.
* no-mistakes(review): Fix fleet-sync project name shadowing
* no-mistakes(document): sync fleet-sync docs
* fix: canonicalize spawn worktree path checks (#294)
* fix(spawn): canonicalize worktree-isolation guard against symlinked project prefixes
fm-spawn.sh compared a logical PROJ_ABS against the physically-resolved
pane cwd every backend reports, so a project reached through a symlinked
prefix (e.g. macOS's /tmp -> /private/tmp) could trip the isolation
guard's false refusal before treehouse ever moved the pane. Canonicalize
once into PROJ_ABS_REAL and compare against that everywhere instead.
* no-mistakes(review): Canonicalize spawn cwd comparisons
* no-mistakes(document): Refresh symlinked spawn docs
* docs: add Orca operator skill (#276)
* docs: add Orca operator skill
* no-mistakes(document): Document Orca checklist
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* fix: surface green PRs during CI monitoring (#297)
* fix(crew-state): detect green-PR CI monitoring, escalate repeat wedges
fm-crew-state.sh's ci step never distinguishes "still waiting on checks"
from "checks green, waiting on merge" via axi status alone, since a repo
that defers merge to the captain keeps the ci step at status=running for
the whole monitor phase. Read the ci step's own log tail (axi logs) for
the checks-passed marker and surface done instead of a false "validating
(running)" - verified against the real PR #252 run's ci.log.
The watcher's wedge timer can re-escalate the same stale pane forever
without ever signaling that it is a repeat; track a per-pane consecutive
escalation count and add a demand-deep-inspection marker to the wake
payload once it crosses a threshold, so the supervisor can no longer
dismiss each one as an isolated, still-validating pane.
Also clarify the ship-brief's checks-green line: it is owed at the
CI-ready return point, not after the background monitor-until-merge
loop finishes.
* no-mistakes(review): Captain, distinguish pending no-checks CI marker
* no-mistakes(review): Harden CI relapse handling
* no-mistakes(review): Block stale done during fixing
* no-mistakes(review): Captain, tighten CI status gating
* no-mistakes(review): Captain, harden stale CI green handling
* no-mistakes(review): Captain, recognize ranged CI rearm markers
* no-mistakes(document): Sync crew-state supervision docs
* fix(teardown): recover provably stale git index locks (#296)
* fix(teardown): recover from a stale worktree git index.lock
A crew process killed mid-git-operation can leave a stale
.git/worktrees/<wt>/index.lock behind, making fm-teardown.sh's
`treehouse return --force` fail closed. On that failure, retry once
after a short wait (the owning process may be exiting), then remove
the lock and retry once more only when it is provably stale: old
enough by mtime and lsof shows no live holder on the lock or the
worktree itself. A lock that isn't provably stale is left in place and
the original failure still surfaces.
* no-mistakes(review): Harden teardown lock refusal paths
* no-mistakes(review): Harden stale-lock teardown safety rechecks
* no-mistakes(review): Harden stale teardown lock checks
* no-mistakes(document): Document teardown lock recovery
* feat(bin): encode project AGENTS.md authoring bar with canonical self-governance section (#307)
* Encode project AGENTS authoring bar
* no-mistakes(review): Captain, centralize CLAUDE promotion governance
* no-mistakes(review): make ensure_maintenance_section idempotent-success, drop || true guards
* no-mistakes(review): separate appended maintenance section on newline-less CLAUDE.md promotion
* no-mistakes(review): assert maintenance heading present before separator check in test
* no-mistakes(document): sync docs with AGENTS.md authoring bar and self-governance
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* feat(skills): add captain-invocable bearings status-report skill (#300)
* Add captain-invocable bearings skill
Generates a pick-up-where-I-left-off status report from live fleet
state to data/status-report-<YYYY-MM-DD>.md plus a concise chat
summary. Read-mostly procedure: reads backlog, per-task crew state
via bin/fm-crew-state.sh, open PRs via gh-axi, scout reports,
pending decisions, and date-gated queued work; composes the
exemplar's sections (TL;DR, Check first, Landed, In flight, Plans,
Decisions pending, Date-gated/queued); never tears down, merges, or
mutates task state as a side effect.
* no-mistakes(document): docs: list new /bearings skill in README built-in skills table
* fix(watcher): make PID identity locale-invariant (#285)
* fix(watcher): pin LC_ALL=C in fm_pid_identity for locale-invariant identity
ps's lstart date format follows the caller's LC_TIME/LC_ALL. The watcher records
its process identity under one locale, but arm/guard/turn-end re-read it under the
machine's ambient locale. On a non-C locale (e.g. ko_KR) the two strings differ
only in the date portion, so fm_watcher_lock_matches_pid / fm_watcher_healthy
reject a genuinely live watcher - breaking fm-watch-arm.sh, fm-guard.sh, and
fm-turnend-guard.sh on every non-C-locale machine.
Pin LC_ALL=C on that one ps call so the write and read sides agree regardless of
machine locale, matching the LC_ALL=C determinism the file already uses elsewhere.
Add a colocated regression test asserting fm_pid_identity is locale-invariant
across exported LC_ALL/LC_TIME.
* no-mistakes(document): Document watcher PID identity coverage
* docs: document codex app backend contract (#222)
* docs: reconcile Codex App backend contract
* no-mistakes(document): Sync backend docs
* docs: clarify Codex Desktop bridge blocker
* no-mistakes(document): Align Codex App backend docs
* no-mistakes(test): Captain, stabilize watcher self-eviction test cadence
* no-mistakes(document): Document Codex App backend contract
* no-mistakes(document): Captain, document blocked codex-app coverage
* docs: make Codex App contract doc authoritative
* no-mistakes(document): Align Codex App backend docs
* docs: redact local Codex App smoke paths
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* docs: add Codex Desktop coordination skill (#275)
* docs: add Codex App coordination skill
* no-mistakes(review): Captain, mark Codex App skill agent-only
* no-mistakes(document): Document Codex App backend boundary
* no-mistakes(document): Captain, document Codex Desktop backend boundary
* no-mistakes(lint): Captain, lint clean
* no-mistakes(document): Document Codex Desktop boundaries
* docs: narrow Codex App skill playbook
* fix(afk): stop herdr escalation redelivery loop (#317)
* fix(afk): recognize unbordered herdr composer rows to stop escalation redelivery loop
fm_backend_herdr_composer_state only recognized bordered composer rows
(the grok shape). Real claude and codex render their live input row
with no border at all, so once a harness's own startup banner scrolled
out of the capture window the classifier read the composer as unknown
forever. fm_backend_herdr_send_text_submit never confirmed "empty", so
escalate_flush never cleared state/.subsuper-escalations, and the
away-mode daemon retyped and resubmitted the same buffered digest every
housekeeping cycle - reproduced live against a real herdr+claude pane
(5+ identical deliveries in 40s).
The classifier now recognizes an unbordered (bare) composer row led by
a known prompt glyph alongside the existing bordered shape, keeping
whichever match is bottom-most so a stale decorative box never
outranks the live composer.
* no-mistakes(review): Narrow herdr bare prompt matcher
* no-mistakes(document): Sync herdr composer docs
* fix(backends): confirm Herdr submits with native agent state (#323)
* fix(herdr): confirm message submit via native agent-state, not composer text
fm_backend_herdr_send_text_submit now confirms a landed submit by polling
herdr's own agent-state (agent get) for the idle->working transition instead
of reading composer content. Composer scraping remains, unchanged, for the
away-mode daemon's pre-injection empty-box guard only.
This fixes the practical effect of the codex idle-tip gap from the
2026-07-07 incident: codex's dynamic idle-composer hint text can no longer
misread as pending and block/mis-confirm a send, since confirmation no
longer looks at composer text at all. Verified empirically against real
claude and codex agents (timing, swallowed-Enter, unreadable-target, and
already-busy-target scenarios), and against the real away-mode daemon
end-to-end after updating its synthetic supervisor-pane test fixture to
register itself as a real herdr agent (herdr's own report-agent primitive)
so it can still exercise the new confirmation path.
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(document): Sync Herdr submit docs
* no-mistakes: apply CI fixes
* feat: add quota-balanced crew dispatch selection (#327)
* Add quota-balanced dispatch selection
* no-mistakes(document): Document dispatch selector guidance
* fix(session-start): respawn dead secondmate agents conservatively
* fix(session-start): deterministically respawn dead-shell secondmates
A secondmate agent that exits leaves its backend pane alive as a bare
shell. The session-start endpoint check only verified pane presence, so
recovery and the watcher (which exempts secondmates from stale-pane
detection) never noticed - evidence 2026-07-07: every secondmate in one
fleet was found sitting at a dead zsh shell.
Add fm_backend_agent_alive (bin/fm-backend.sh), a deeper per-backend
liveness probe distinct from pane presence: fm_backend_tmux_agent_alive
classifies the pane's live foreground process via tmux's own
pane_current_command, and fm_backend_herdr_agent_alive reuses the
already-verified pane_agent_state husk classifier. Both are conservative:
anything ambiguous reports unknown, never a false dead.
Wire this into a new session-start-only, locked-and-primary-only sweep in
bin/fm-bootstrap.sh that kills and respawns only a confidently dead
secondmate endpoint, leaving alive/unknown readings untouched - idempotent
by construction, so repeated runs converge without duplicating agents.
* no-mistakes(review): Guard raw secondmate liveness respawns
* no-mistakes(review): Fix detect-only bootstrap test
* no-mistakes(test): Pin liveness fixture harness
* no-mistakes(document): Sync secondmate liveness docs
* no-mistakes: apply CI fixes
* fix: emit stable secondmate nudge selectors (#331)
* Fix NUDGE_SECONDMATES to print stable fm-<id> selectors.
Session-start secondmate sync used to accumulate raw backend window targets
into NUDGE_SECONDMATES, but the liveness sweep in the same bootstrap run can
respawn secondmates onto new endpoints. fm-send with those stale explicit
targets bypasses meta resolution and fails, while fm-<id> resolves correctly.
Accumulate fm-<id> in process_secondmate, update the bootstrap/update contracts
and /updatefirstmate skill, and add a herdr respawn regression test.
* no-mistakes(review): Captain, guard herdr regression jq dependency
* no-mistakes(document): Document stable secondmate nudge selectors
* no-mistakes(lint): Fix shell lint hints
* feat: require bootstrap detection for AXI tools (#332)
* Make tasks-axi and quota-axi required bootstrap tools
Add both to the normal toolchain checks alongside lavish-axi, keep the
tasks-axi 0.1.1+ compatibility gate, and report quota-axi through the
standard MISSING install-consent flow. TASKS_AXI: available remains a
backlog-backend capability signal only; manual opt-out no longer suppresses
the missing-tool report.
Update bootstrap tests and point docs/configuration.md at the canonical
toolchain contract.
* no-mistakes(review): Clarify manual backlog bootstrap reporting
* no-mistakes(document): Document bootstrap AXI tools
* bearings: delete today's report before recreating (#333)
Replace overwrite-in-place wording with explicit delete-then-create
instructions so agents do not modify an existing daily report file.
* feat: guard primary turn ends across harnesses (#339)
* Add primary turn-end guards for all harnesses
* no-mistakes(review): Normalize Codex hook cwd resolution
* no-mistakes(review): Fix OpenCode guard worktree anchoring
* no-mistakes(review): Anchor Codex guard outside nested roots
* no-mistakes(review): Anchor Codex guard to hook root
* no-mistakes(review): Avoid Grok permission escalation
* no-mistakes(document): Sync turn-end guard docs
* fix: resolve backend selectors by exact task id first (#342)
* fix backend selector task id resolution
* no-mistakes(document): Document selector resolution behavior
* fix: scale bootstrap fleet-sync timeout (#341)
* fix bootstrap fleet sync timeout
* no-mistakes(review): Fix bootstrap fleet-sync timeout regressions
* no-mistakes(document): Sync bootstrap timeout docs
* no-mistakes(lint): Clean ShellCheck directives
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add fleet snapshot and view commands (#343)
* Add fleet snapshot and view
* no-mistakes(review): Fix fleet snapshot parsing and overrides
* no-mistakes(review): Fix secondmate fleet rendering
* no-mistakes(review): Fix backlog title and completion parsing
* no-mistakes(review): Include durable scout reports
* no-mistakes(review): Fix fleet snapshot edge cases
* no-mistakes(review): Captain: gate fleet hints on current state
* no-mistakes(review): Captain: parse bracketed Done PR artifacts
* no-mistakes(document): Sync fleet snapshot docs
* fix(fm-send): fail loudly on unresolvable send targets (#254)
* Make fm-send fail loudly on unresolved targets
* no-mistakes(review): Document fm-send FM_HOME contract
* Fix fm-send readiness docs and backend send path
* Fix fm-send docs for cmux and X skill metadata
* Make gotmp teardown test home-explicit
* Scope watcher warning wording to fm-send
* Fix fm-send review findings
* Verify explicit tmux targets before sending
* Isolate turnend guard test home
* no-mistakes(document): Documented fm-send FM_HOME/backend guard additions missing from doc inventories
---------
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* fix: deliver AFK escalations through herdr supervisors (#353)
* fix afk codex ghost composer delivery
* no-mistakes(review): Harden AFK startup flag writes
* no-mistakes(review): Harden AFK daemon liveness checks
* no-mistakes(document): Sync AFK herdr docs
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add harness-aware supervision (#367)
* Add harness-aware supervision
* no-mistakes(review): Captain, harden watcher supervision regressions
* no-mistakes(review): Captain, harden watcher supervision cadence
* no-mistakes(review): Harden watcher supervision ownership
* no-mistakes(review): Captain, harden Pi extension marker
* no-mistakes(review): Captain, harden Pi supervision restart checks
* no-mistakes(review): Harden watcher ownership checks
* no-mistakes(review): Captain, harden Pi supervision loading
* no-mistakes(review): Captain, require Pi guard extension loading
* no-mistakes(review): Captain, harden watcher supervision recovery
* no-mistakes(test): Fix fm-send baseline log filtering
* no-mistakes(document): Sync harness supervision docs
* no-mistakes: apply CI fixes
* fix: split X-mode replies by platform (#369)
* fix: make x replies split by platform
* no-mistakes(review): Captain: preserve Discord recovery relink context
* no-mistakes(test): Captain: keep split markers outside fences
* no-mistakes(document): Sync X-mode reply docs
* fix: make stow memory writes inspect before update (#372)
* docs: make stow inspect-then-update
* no-mistakes(review): Remove unsupported archive-body guidance
* no-mistakes(review): Clarify stow read-before-write exception
* no-mistakes(test): Require archive-body for stow task notes
* no-mistakes(document): Sync stow memory docs
* no-mistakes(lint): Silence ShellCheck source warning
* fix(watcher): wait when arm attaches to a healthy watcher (#375)
* fix: attach-and-wait when arm finds a healthy watcher
Grok and Claude re-arm after every turn with work in flight. When a
watcher was already healthy, fm-watch-arm exited immediately with
watcher: healthy, which completed the harness background task and
injected an empty false wake.
Attach to the live identity-matched holder instead, stay until that
cycle ends, then exit 0 so notify fires for a real end-of-cycle. The
peer-startup-race path uses the same contract. --restart and the
started path are unchanged.
* no-mistakes(review): Gate restart watcher peer attach
* no-mistakes(document): Sync watcher arm docs
* feat(pi): simplify primary session launch (#386)
* docs(readme): reformat Quick Start and recommend Grok equally with Claude Code
* no-mistakes(review): Captain: align harness launch guidance
* no-mistakes(review): Captain, clarify Pi supervised launch
* no-mistakes(review): Captain, document Pi first-launch bridge
* feat(pi): track primary watcher extension for plain-pi launch
Move Pi's primary watcher bridge from a generated state/ file to a
tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end
guard extension already works: self-hashing version, project-local
auto-discovery after one-time Pi trust. This drops the state/-generation
step and dual -e requirement from the happy path, so Pi's Quick Start
launch becomes plain 'pi', the same friction class as 'claude' and
'grok --trust'.
- bin/fm-pi-watch-extension.sh is removed; nothing generates the
extension anymore since it is committed.
- fm-session-start.sh and fm-supervision-instructions.sh resolve the
watcher extension path from FM_ROOT instead of state/, and the
session-start diagnostic now points at restarting plain pi after
trust, with -e as a documented fallback.
- fm-spawn.sh points Pi secondmate launches at the tracked extension
path in the secondmate home instead of generating a state/ copy.
- README Quick Start Pi block is now just 'pi' plus a trust note.
- Tests, docs, and the harness-adapters skill updated to match.
* fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016
* feat(supervision): prevent unsafe watcher-arm commands (#387)
* feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns
Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that
denies a primary shell command backgrounding, piping, or bundling the
watcher arm/checkpoint, or force-killing the watcher process broadly -
the exact shapes that silently took Grok's supervision down. Wires it
into all five verified harnesses (grok, claude, codex, opencode, pi),
each validated empirically against the real harness.
Also fixes a grok 0.2.93 regression discovered during that validation: the
existing turnend-guard Stop hook's bare root variable broke grok's own
variable pre-substitution and silently no-op'd the hook.
* no-mistakes(review): Harden watcher arm validation
* no-mistakes(review): Harden arm guard metacharacter checks
* no-mistakes(review): Harden nested shell arm guard
* fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then
A && B || C is not if-then-else; C can run when A is true. Replace both
occurrences of the quote-state early-continue with an explicit if/then.
* fix(pi): restore primary watcher supervision lifecycle (#397)
* fix Pi primary supervision lifecycle
* no-mistakes(document): Synchronize Pi primary extension documentation
* fix: keep persistent secondmates out of the main backlog (#398)
* fix secondmate backlog guidance
* no-mistakes(review): Require reasons for captain backlog holds
* no-mistakes(test): Document secondmate handoff skill requirement
* fix secondmate teardown reminder
* no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract
* fix(backlog-handoff): move full item blocks including indented bodies (#401)
* fix(backlog-handoff): move full item blocks including indented bodies
fm-backlog-handoff only moved the checklist header line, so multi-line
item bodies were left orphaned in the source backlog and never reached
the secondmate. Move the full block (header plus indented body lines)
atomically, treating body membership by indentation so lines like
## Intent stay with the item, and add regression coverage.
* no-mistakes(review): Captain: preserve EOF handoff terminators
* no-mistakes(review): treat blank lines inside item bodies as movable body
* no-mistakes(document): sync backlog-handoff docs with full-block move behavior
* feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402)
* guard Herdr lab lifecycle in briefs
* no-mistakes(review): Fix Herdr lab helper and provisioning safety
* no-mistakes(review): Captain, harden Herdr lab lifecycle safety
* no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range
* no-mistakes(review): reject leading options in Herdr lab run guard
* no-mistakes(review): strip leading non-alnum in Herdr lab name generator
* no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag
* no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab
* no-mistakes: apply CI fixes
* fix(watcher): classify arm-command seatbelt by execution position (#403)
* fix watcher arm command policy
* no-mistakes(review): Harden watcher command policy parsing
* no-mistakes(review): Captain: harden watcher policy parsing
* no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter
* no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps
* no-mistakes(review): fail closed on loop-wrapped broad watcher kills
* no-mistakes(document): sync docs for watcher-arm command-position policy
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed…
vitalNohj
added a commit
to vitalNohj/firstmate
that referenced
this pull request
Jul 24, 2026
* feat(pi): simplify primary session launch (#386)
* docs(readme): reformat Quick Start and recommend Grok equally with Claude Code
* no-mistakes(review): Captain: align harness launch guidance
* no-mistakes(review): Captain, clarify Pi supervised launch
* no-mistakes(review): Captain, document Pi first-launch bridge
* feat(pi): track primary watcher extension for plain-pi launch
Move Pi's primary watcher bridge from a generated state/ file to a
tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end
guard extension already works: self-hashing version, project-local
auto-discovery after one-time Pi trust. This drops the state/-generation
step and dual -e requirement from the happy path, so Pi's Quick Start
launch becomes plain 'pi', the same friction class as 'claude' and
'grok --trust'.
- bin/fm-pi-watch-extension.sh is removed; nothing generates the
extension anymore since it is committed.
- fm-session-start.sh and fm-supervision-instructions.sh resolve the
watcher extension path from FM_ROOT instead of state/, and the
session-start diagnostic now points at restarting plain pi after
trust, with -e as a documented fallback.
- fm-spawn.sh points Pi secondmate launches at the tracked extension
path in the secondmate home instead of generating a state/ copy.
- README Quick Start Pi block is now just 'pi' plus a trust note.
- Tests, docs, and the harness-adapters skill updated to match.
* fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016
* feat(supervision): prevent unsafe watcher-arm commands (#387)
* feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns
Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that
denies a primary shell command backgrounding, piping, or bundling the
watcher arm/checkpoint, or force-killing the watcher process broadly -
the exact shapes that silently took Grok's supervision down. Wires it
into all five verified harnesses (grok, claude, codex, opencode, pi),
each validated empirically against the real harness.
Also fixes a grok 0.2.93 regression discovered during that validation: the
existing turnend-guard Stop hook's bare root variable broke grok's own
variable pre-substitution and silently no-op'd the hook.
* no-mistakes(review): Harden watcher arm validation
* no-mistakes(review): Harden arm guard metacharacter checks
* no-mistakes(review): Harden nested shell arm guard
* fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then
A && B || C is not if-then-else; C can run when A is true. Replace both
occurrences of the quote-state early-continue with an explicit if/then.
* fix(pi): restore primary watcher supervision lifecycle (#397)
* fix Pi primary supervision lifecycle
* no-mistakes(document): Synchronize Pi primary extension documentation
* fix: keep persistent secondmates out of the main backlog (#398)
* fix secondmate backlog guidance
* no-mistakes(review): Require reasons for captain backlog holds
* no-mistakes(test): Document secondmate handoff skill requirement
* fix secondmate teardown reminder
* no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract
* fix(backlog-handoff): move full item blocks including indented bodies (#401)
* fix(backlog-handoff): move full item blocks including indented bodies
fm-backlog-handoff only moved the checklist header line, so multi-line
item bodies were left orphaned in the source backlog and never reached
the secondmate. Move the full block (header plus indented body lines)
atomically, treating body membership by indentation so lines like
## Intent stay with the item, and add regression coverage.
* no-mistakes(review): Captain: preserve EOF handoff terminators
* no-mistakes(review): treat blank lines inside item bodies as movable body
* no-mistakes(document): sync backlog-handoff docs with full-block move behavior
* feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402)
* guard Herdr lab lifecycle in briefs
* no-mistakes(review): Fix Herdr lab helper and provisioning safety
* no-mistakes(review): Captain, harden Herdr lab lifecycle safety
* no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range
* no-mistakes(review): reject leading options in Herdr lab run guard
* no-mistakes(review): strip leading non-alnum in Herdr lab name generator
* no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag
* no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab
* no-mistakes: apply CI fixes
* fix(watcher): classify arm-command seatbelt by execution position (#403)
* fix watcher arm command policy
* no-mistakes(review): Harden watcher command policy parsing
* no-mistakes(review): Captain: harden watcher policy parsing
* no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter
* no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps
* no-mistakes(review): fail closed on loop-wrapped broad watcher kills
* no-mistakes(document): sync docs for watcher-arm command-position policy
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", leaving the clone unsynced.
On that signature only, fm-fleet-sync.sh now retries the fetch with a
bounded wait (transient locks self-clear), then removes the lock and
retries once more ONLY when it is provably stale: still present, mtime
age past a threshold, and no lsof holder of the lock file or of the clone
worktree itself (a live git keeps that as its cwd even in the window after
it closes the lock and before it exits). A live lock, a missing lsof, any
failed check, or any other fetch failure keeps today's behavior. Every
wait/retry/removal prints to stderr, and a successful recovery also prints
one "recovered:" summary to stdout so a session-start refresh - which
discards fleet-sync stderr and relays only stdout - still surfaces it.
The shared "is this git lock provably abandoned?" proof is extracted into
bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and
fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh
gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/.
* no-mistakes(review): Captain, remove obsolete teardown wake dependency
* no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): escalate blocked panes immediately (#472)
* feat(herdr): immediate blocked-state escalation via native events.subscribe push
Fold herdr's native pane.agent_status_changed stream into the single watcher so
a crew entering blocked wakes its supervisor sub-second (measured 0.129s)
instead of after the ~240s stale-pane wedge timer.
- bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape
plus the single-owner status->action policy table (blocked=actionable,
working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling).
- bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe
subscriber over one connection for all this home's herdr panes, subscribing to
ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe
marker and a reconnect level-reconcile. Version/schema capability gate.
- bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so
the watcher stays backend-agnostic and the shape+policy are reusable.
- bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal
wait primitive (replacing the blind sleep POLL for push-capable homes),
behind a source guard so the splice is unit-testable; secondmate/paused
exemptions; map pane->window->task and enqueue a stale wake. No second
watcher process; the single-cycle invariant and every guard/beacon/turn-end
mechanism are unchanged.
- Polling stays the permanent fail-closed backstop: below-capability, subscribe
failure, and repeated runtime failures all degrade to sleep.
- Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/
fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events)
plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries
the dated evidence and retires the old gap note.
* no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs
* no-mistakes(review): Captain, commit markers after wake and reuse capability cache
* no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs
* no-mistakes(review): Captain, subscribe before Herdr reconciliation
* no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe
* no-mistakes(test): Captain: include lock library in teardown fixture
* no-mistakes(document): Captain: document Herdr immediate blocked escalation
* fix: clarify shellcheck conditionals
* docs(readme): reposition firstmate as an agent distro (#473)
* feat: add deterministic bounded bearings snapshots (#475)
* feat(bearings): deterministic bearings snapshot + durable decision model
Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the
canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR
discovery and checks happen only under --include-prs, which fails soft. Every
dropped surface is marked in omitted[] with the flag that reveals it, and the
prs: line states when checks were not requested, so absence is never silent.
Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib
gains status_open_decisions, the one authoritative keyed open/resolved fold over
the whole status stream: needs-decision/blocked opens a keyed entry, only an
explicit keyed resolution (or, for run-backed tasks, run-step advancement)
closes it, so a later unrelated done/paused can no longer mask a still-open
captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives
pending_decision/blocked_event from it; the canonical schema stays complete.
Point the /bearings skill at the one command; add the resolved: writer line to
ship, scout, and secondmate briefs. Register the script and add regression
tests for the output bound, TOON/JSON parity, local-only default, opt-in PR
fetch, partial-failure degradation, decision durability, and report pointers.
* fix(bearings): completed scout report is a pointer, not a pending decision
A completed scout that raised a needs-decision and then finished (done) without
a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103
case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh
cleared a stale decision only for a live run-step/pane activity read, so a
terminal task whose current state is read from the status log (a scout or ship
that reached done/failed) never cleared its stale, never-keyed-resolved
needs-decision, and it lingered as pending.
The open-decision set is still derived purely from the keyed fold - never from a
report body or decision-like prose - and reconciled against the crew lifecycle.
Extend that reconciliation so a terminal done/failed state on a single-owner
task (scout or ship), whose deliverable is its report or PR, also clears the set;
a completed scout now surfaces only as a report pointer. Secondmates are excluded
from the terminal clear (persistent, multiplexed stream), which keeps the
unrelated-event masking fix intact. Add regression tests: a completed scout with
decision-like report prose is a pointer not pending (canonical + end-to-end), and
a scout still parked at a decision stays pending so the terminal clear never
over-fires.
* no-mistakes(review): Captain, preserve keyed decisions across shared status parsing
* no-mistakes(review): Captain, close blockers and harden keyed decision parsing
* no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout
* no-mistakes(review): Captain, bound bearings sections and fail closed
* no-mistakes(review): Captain, disclose capped per-repository PR results
* no-mistakes(document): Refresh bearings documentation and status contracts
* fix(bearings): avoid ambiguous worktree guard
* fix: enforce deterministic ShellCheck parity (#481)
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity
Firstmate PRs passed local no-mistakes validation but failed CI's
"Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034).
Two divergences caused it:
1. The no-mistakes gate had no commands.lint, so its lint step never ran
the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that
CI runs. Confirmed from state.sqlite: the lint step_result recorded
findings:null with no lint agent invocation.
2. CI's shellcheck floated with the runner image while local ran a newer
build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck
rejected an SC2015 that the newer local one no longer emits.
Establish bin/fm-lint.sh as the single owner of the lint definition: the
file set, the config, and the pinned shellcheck version (0.11.0, printed
via --required-version). Both CI (.github/workflows/ci.yml) and the
no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs
the exact version it names and logs the resolved version, and fm-lint.sh
refuses to lint under any other version. This is not a CI relaxation: it
adopts shellcheck 0.11.0's rule set consistently, dropping only the
upstream-retired, false-positive-prone SC2015; default severity and every
still-supported finding stay enforced (no severity downgrade, no excludes).
tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs
and logs the pinned version, that the owner refuses a non-pinned shellcheck,
and that it rejects a real lint defect the old no-op gate passed.
* no-mistakes(review): Captain, harden deterministic ShellCheck parity
* no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: guard primary shells from persistent cd commands (#483)
* feat: add cd-guard PreToolUse seatbelt for the primary shell
A stray persistent top-level `cd projects/<clone>` in the primary firstmate
shell relocates the shell, so a later firstmate-owned command (a backlog write,
an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the
home. The cd-guard denies exactly that command shape before it runs, across all
five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt.
- bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the
shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate
lexer; that file's CLI now runs only when invoked directly).
- bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness
output rendering, and primary-checkout scoping - fires in a secondmate's own
primary session, inert in crew/scout child worktrees and non-firstmate repos.
- Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents;
per-harness hooks only call the owner.
- Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd,
and command cd). Allows git -C, subshell / bash -c / env -C / make -C /
find -execdir, pipeline and background forms, and cd-as-data. Fails open on
malformed input; agent-mistake threat model.
- tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix,
end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring.
- docs/cd-guard.md: full contract plus live validation (claude, codex,
opencode, pi blocked end-to-end; grok live run blocked by an API balance
limit, with mechanism parity and deterministic coverage recorded).
* no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage
* no-mistakes(review): Captain, allow path-qualified command wrappers
* no-mistakes(review): Captain, allow non-executing command queries
* no-mistakes(test): Captain, clarify cd-guard safe-path remediation
* docs: clarify cd guard guidance
* no-mistakes(document): Clarify cd-guard safe target guidance
* brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (#267)
Crews must never stop, restart, or update the shared no-mistakes
daemon since one instance serves every firstmate lane/home; a restart
kills other lanes' in-flight pipeline runs and forces expensive
re-runs. Encodes this as a new numbered rule in both the ship-task and
scout-task brief scaffolds.
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat: make bearings concise and accurate (#485)
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render
/bearings skill (one owner of the chat-response format): mandate the four
always-present chat sections - Captain's Call, Recently Landed, Underway,
Charted Next - each with an explicit empty-state sentence, no At Anchor,
materially shorter than and linking to the report file. Resolves the ambiguous
Check first / Decisions pending split into one strict captain-action section.
fm-crew-state: the log fallback derives current state only from a real
run-state verb, so a trailing decision-closing resolved: event no longer
renders a healthy idle crew (typically a secondmate) as unknown with the
resolution prose as its detail. The keyed-decision contract in
fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner.
fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done
records from registered secondmate homes, reusing the single backlog parser and
the one secondmate-home enumerator (meta home= with data/secondmates.md
fallback); no network, per-home capped.
fm-bearings-snapshot: landed now merges main-home Done with the secondmate
roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure
(also fixing the previously-silent landed truncation); --all-landed reveals the
full set.
tests: resolved-event state render, secondmate landed aggregation with caps and
omitted[] disclosure, Captain's Call anti-leak, and the four-section contract.
* no-mistakes(review): Captain, ensure bearings reveals all landed work
* no-mistakes(document): Document bearings accuracy contracts
* fix: harden away-mode daemon lifecycle (#490)
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle
Away-mode entry left "make the daemon a tracked background terminal" to the
operator; on a pi/herdr primary that meant splitting the captain's active pane,
which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches
the daemon in a non-visible tracked terminal per backend (herdr dedicated
--no-focus workspace, detached tmux session), never a split, pins the captain
pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal
id, and tears it down or reconciles a leaked one by that id. No shell &.
Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared
with the daemon (one owner).
Fix the stale subsuper-artifact leak: clear the prior away session's delivery
cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon
before clearing state/.afk so its shutdown flush runs instead of being a no-op.
Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab
session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md,
docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md.
* no-mistakes(review): Captain, serialize AFK launcher lifecycle safely
* no-mistakes(review): Captain, harden AFK launcher lifecycle races
* no-mistakes(review): Captain, ensure AFK daemon launch readiness
* no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown
* no-mistakes(review): Captain, preserve AFK reconciliation records uniformly
* no-mistakes(review): Captain, harden AFK recovery state durability
* no-mistakes(review): Captain, harden AFK tmux ownership checks
* no-mistakes(review): Captain, simplify AFK lifecycle failure handling
* no-mistakes(review): Captain, require confirmed AFK daemon shutdown
* no-mistakes(review): Captain, confirm AFK exit by process identity
* no-mistakes(document): Align AFK launcher lifecycle documentation
* no-mistakes: apply CI fixes
* fix: prevent no-mistakes gate agents from driving the fleet (#518)
* feat: contain no-mistakes gate agents from driving the fleet
Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.
Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).
firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.
* no-mistakes(review): Captain, refuse empty no-mistakes gate markers
* no-mistakes(document): Document no-mistakes gate authority boundary
* fix: guard secondmate primary sessions from blind turn ends (#505)
* fix: guard secondmate own-home turn ends
Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the
'no turn ends blind' backstop fires in a secondmate's own primary session,
matching the cd-guard's scope: the own home is guarded, child crew/scout
worktrees stay exempt via the retained git-dir/git-common-dir test. This
was pure scoping from the guard's primary-only origin and guarded against
no secondmate-specific hazard.
Add secondmate regression tests (blind-turn block, idle-by-default,
stop_hook_active loop guard, deferred-death recovery loop, child-worktree
exemption) and record the autonomous background-notify re-invoke
measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md.
* no-mistakes(document): Correct secondmate guard documentation, captain
* fix: force-include marked secondmate homes in turn-end guard
The prior remove-only form (just deleting the .fm-secondmate-home check)
left the DEFAULT secondmate topology unguarded: a treehouse-leased home is
a linked git worktree (git-dir != git-common-dir), which the retained
git-dir exemption still skipped, so its own primary session could still end
a turn blind. Invert the marker: a genuinely-marked home is force-included
as a guarded primary (treehouse-leased linked OR git-cloned plain), and the
git-dir exemption applies only to UNMARKED child worktrees. Marker
validation (regular non-symlink file, non-empty id-token content) blocks a
stray or empty marker from spoofing inclusion.
Add real linked-worktree regression tests: a treehouse-leased LINKED
secondmate home is guarded, a stray/empty marker stays exempt, and the
unmarked child worktree stays exempt - the topology the plain git-init
fixtures masked. Predicates, in-flight gate, and loop guard untouched.
* fix: force ASCII collation in secondmate marker validation
Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the
[A-Za-z0-9._-] id allowlist matches under C collation, not the ambient
locale - a locale-crafted non-ASCII marker id can no longer slip through
the range match and spoof force-inclusion of a linked child worktree.
Add a regression test proving a non-ASCII marker id is rejected and the
linked worktree stays exempt.
* no-mistakes(test): fix backend baseline gate-refusal dependency
* no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap diagnostics backend-aware (#519)
* fix: make bootstrap required-tool detection backend-aware
Bootstrap demanded tmux and treehouse for every backend except orca, so a
herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux.
Required tools now follow the resolved backend via the single-owner
fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own
session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux),
and treehouse for session-provider-only backends (orca owns its worktree).
The treehouse lease-support check is gated to backends that use treehouse.
Adds install hints for herdr/zellij/cmux, regression tests for the full
backend dependency matrix (herdr-without-tmux repro plus each boundary),
and updates the authoritative Toolchain docs.
* no-mistakes(review): Captain, prevent executing Herdr install guidance
* no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics
* no-mistakes(review): Captain, separate manual dependency remediation
* no-mistakes(review): Captain, align bootstrap diagnostic consumers
* no-mistakes(document): Align backend adapter dependency comments
* fix: preserve follow-up platform context after inbox cleanup (#520)
* fix: recover X/Discord follow-up platform after inbox cleanup
A milestone follow-up posted directly by request_id after the inbox was
drained - and with no task link, because one persistent secondmate's single
x_request slot collides across concurrent requests - resolved platform only
from the local inbox, so a >280 Discord reply silently defaulted to the X
280-char budget and threaded as (1/2).
- fm-x-poll records a durable per-request reply context
(state/x-context/<rid>.json) at stash time, keyed by request_id so
concurrent requests never overwrite each other; it survives inbox cleanup
and restart.
- fm-x-reply resolves platform/budget through registry -> inbox -> relay
(the relay lookup confined to a live follow-up), recovering the original
platform independent of task-link availability.
- Fail-safe: a follow-up whose platform/budget cannot be authoritatively
resolved and that would split is refused (exit 8) and held for retry,
never wrongly split; fm-x-followup keeps the link on that exit.
- fm-x-dismiss clears the durable context for a dismissed mention.
Refactors reply-context extraction into a single owner and adds regression
coverage for all four cases.
* no-mistakes(review): Captain, fail closed on incomplete follow-up context
* no-mistakes(review): Captain, bound X context registry retention
* no-mistakes(review): Captain, align context retention with answer binding
* no-mistakes(document): Align X follow-up context documentation
* no-mistakes(document): Align durable X follow-up documentation
* no-mistakes(review): make cursor stop-hook install idempotent; align checkpoint fallback chain
* fix: preserve secondmate routing markers in terminal sends (#533)
* fix: preserve secondmate routing markers
* no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends
* no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift
* no-mistakes(document): Refresh Herdr marker documentation
* fix: align Grok effort handling with 0.2.99 (#527)
* fix: align grok effort docs and spawn with 0.2.99 ceiling
grok 0.2.99 accepts only low|medium|high for --reasoning-effort and
rejects both xhigh and max. Omit unsupported values on spawn, flag them
in crew-dispatch validation, and update harness-adapters.
* no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies
* no-mistakes(document): Clarify Grok effort documentation ownership
* fix: derive bearings from authoritative secondmate state (#555)
* fix: make bearings use secondmate home state
* test: anonymize bearings fixtures
* no-mistakes(review): Bound parent activity evidence scans, captain
* no-mistakes(review): Preserve structured secondmate authority and bounds, captain
* no-mistakes(review): Preserve registry completeness and child inventory, captain
* no-mistakes(review): Reconcile parent evidence by verb and key, captain
* no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain
* no-mistakes(document): Document bearings local snapshot and PR opt-in
* no-mistakes(lint): Fix fleet snapshot ShellCheck findings
* no-mistakes: apply CI fixes
* fix: restore fleet snapshots on stock macOS Bash (#578)
* fix: restore stock macOS snapshot parsing
* no-mistakes(document): Clarify Linux gate and macOS CI coverage
* fix(afk): make Pi escalation and return catch-up reliable (#587)
* fix: close away-mode blocker supervision gap
* no-mistakes(review): Gate teardown retries and verify U+2063 dedupe
* no-mistakes(test): Fail closed on incomplete Pi composer separators
* no-mistakes(document): Document Pi composer recognition and return gating
* feat: support Pi max reasoning profiles (#537)
* support Pi max thinking profiles
* no-mistakes(review): Captain, allow Pi max dispatch profiles
* chore: no-mistakes(document): Clarify yolo response ownership (#595)
* Clarify validation response ownership
* no-mistakes(document): Clarify yolo response ownership
* feat: establish instruction ownership foundation (#619)
* Add instruction owners foundation
* no-mistakes(document): Refresh project-management owner pointers
* fix: compress Firstmate contract and enforce delivery rigor ownership (#626)
* docs: compress firstmate operating contract
* docs: make delivery rigor single-owner
PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion.
* no-mistakes(review): Honor configured merge authority across faster delivery paths
* no-mistakes(document): Align docs with compressed operating contract
* feat: add durable captain decision holds (#593)
* Add durable captain decision holds
* no-mistakes(review): Validate decision hold retries and origin paths
* no-mistakes(review): Enforce durable decision lifecycle boundaries
* no-mistakes(review): Harden decision display and partial retry recovery
* no-mistakes(test): Update scout teardown fixtures for decision inventory
* no-mistakes(document): Align decision lifecycle and scout teardown documentation
* no-mistakes: apply CI fixes
* no-mistakes(review): Reconcile terminal decision holds
* no-mistakes(document): Align captain decision-hold documentation
* fix(bin): harden PR check artifacts (#556)
* fix: harden PR check artifacts
* fix: close PR check migration gaps
* fix: close PR check retry gaps
* fix: clarify migration outcomes and ESM boundary
* fix: keep failed migrations authoritative
* test: use inert PR validation fixtures
* no-mistakes(review): Reserve noncanonical PR quarantine namespace
* no-mistakes(review): Prevalidate final PR-check teardown artifacts
* no-mistakes(review): Preserve X metadata and validate teardown IDs
* no-mistakes(review): Initialize migration state before watcher exclusion
* no-mistakes(review): Isolate failed poll migrations from bootstrap recovery
* no-mistakes(review): Allow safe polling during incomplete private repairs
* no-mistakes(review): Authenticate watcher checks at execution time
* no-mistakes(review): Preserve custom checks with hash-bound registration
* no-mistakes(review): Clean custom check snapshots on watcher signals
* no-mistakes(review): Stop watcher checks promptly on signals
* no-mistakes(review): Terminate watcher check groups before cleanup
* no-mistakes(document): Correct stale X-mode watcher documentation
* fix: drain returned watcher check groups
* no-mistakes(review): Harden quarantine links and recover validated replacement polls
* no-mistakes(review): Preserve X mode across shim version transitions
* no-mistakes(review): Refresh legacy X shims before marker short-circuits
* no-mistakes(document): Correct persisted PR-check artifact documentation
* no-mistakes(document): Correct stale PR-check documentation
* fix: bind PR poll repair provenance
* no-mistakes(review): Enforce single-link ownership for custom check artifacts
* no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs
* no-mistakes(review): Separate task creation and legacy teardown validation
* no-mistakes(review): Restore safe legacy operations and teardown validation
* no-mistakes(review): Disambiguate migration obligations and preserve legacy retries
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits
* no-mistakes(document): Document private poll artifact safety contracts
* no-mistakes(lint): Suppress intentional literal-dollar lint finding
* fix: migrate historical X poll identity
* fix: harden PR check artifacts
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* fix: migrate historical X poll identity
* no-mistakes(review): Harden X-mode artifact publication against symlink corruption
* no-mistakes(review): Guard X artifact publication
* no-mistakes(review): Enforce private X artifact reads
* no-mistakes(test): Fix backend compatibility fixture dependencies
* no-mistakes(document): Refresh PR-check documentation
* no-mistakes(lint): Remove unused x-mode test locals
* no-mistakes: apply CI fixes
* fix(bin): compact session-start backlog digest (#636)
* fix: compact session-start backlog digest
* no-mistakes(test): Fix legacy backend fixture helper
* no-mistakes(test): Fix watcher exit wait helper
* no-mistakes(document): Document compact backlog digest
* fix: dedupe stale watcher guard banners (#637)
* fix: dedupe stale watcher guard banner
* no-mistakes(review): Keep read-only guard state nonmutating
* no-mistakes(document): Clarify stale watcher docs
* fix(bin): balance bearings landed baseline (#640)
* fix: balance bearings landed defaults
* no-mistakes(document): Document balanced landed baseline
* no-mistakes: apply CI fixes
* fix: clarify captain-facing translation contract (#644)
* docs: clarify captain-facing translation contract
* no-mistakes(review): Restore runtime fallback mandate
* no-mistakes(document): Align Bearings translation wording
* fix(bin): make bootstrap output and nudges deterministic (#646)
* fix: make bootstrap nudges deterministic
* no-mistakes(review): Honor state override for bootstrap nudges
* no-mistakes(review): Update benign bootstrap documentation labels
* no-mistakes(review): Validate bootstrap nudge retry markers
* no-mistakes(document): Align bootstrap nudge documentation
* no-mistakes: apply CI fixes
* docs(secondmate-provisioning): clarify concise registry ownership (#649)
* Clarify concise secondmate registry contract
* no-mistakes(review): Expand secondmate registry boilerplate coverage
* no-mistakes(document): Point route docs to owner
* fix(bin): strip quoted blocked_by values during decision hold resolve (#654)
* fix(bin): strip quotes on blocked_by in decision-hold resolve
tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary
membership test only matched middle elements. Strip surrounding quotes
before matching so first and last hold ids resolve correctly.
* no-mistakes(document): Refresh decision-hold regression evidence
* feat(secondmate): inherit shared captain preferences (#656)
* feat(secondmate): inherit shared captain preferences
* no-mistakes(review): Honor shared captain data overrides
* no-mistakes(review): Honor bootstrap data override registry
* no-mistakes(document): Refresh shared inheritance docs
* no-mistakes(document): Clarify inherited local-material docs
* feat: gate local agent secret injection (#658)
* feat(spawn): gate local agent secret injection
* fix(spawn): align final Keychain slot
* no-mistakes: apply CI fixes
* test: isolate Herdr autodetect smoke sessions (#662)
* test: isolate herdr autodetect smoke session
* no-mistakes(review): Restored autodetect smoke gate bypass
* no-mistakes(test): Harden Herdr lab provisioning
* no-mistakes(document): Refresh Herdr lab docs
* docs: adopt under way for active work (#666)
* Revert "feat: gate local agent secret injection (#658)" (#668)
This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef.
* fix(pi): distinguish stale locks when arming watcher (#681)
* fix(pi): distinguish stale locks when arming watcher
* no-mistakes(test): Stabilize watcher extension async waits
* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate house vocabulary (#685)
* fix: accept secondmate as house vocabulary
* no-mistakes(test): Update captain vocabulary contract test
* no-mistakes(document): Align secondmate documentation vocabulary
* fix(bin): parse handoff homes after registry parentheticals (#686)
* fix: parse secondmate home after pre-field parentheses
Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.
* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges (#687)
* feat: add native session-start nudges
* no-mistakes(document): Document nudge script inventory
* docs: call built-in defaults the firstmate repo, not template (#688)
Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
* fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205)
* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn
Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(brief): harden fm-brief regression coverage for parse and scaffolds
Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the
Co-authored-by: Cursor <cursoragent@cursor.com>
#166 apostrophe regression cannot return unnoticed.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): keep watcher supervision continuous across child cycles (#693)
* fix: make watcher supervision continuous
* no-mistakes(review): Bound watcher retries and log attached signals
* no-mistakes(review): Add bounded successor-recovery wake fallbacks
* no-mistakes(review): Prevent overlapping successor-arm retries
* no-mistakes(review): Resume supervision after late arm closes
* no-mistakes(review): Bind OpenCode recovery to attempted arm
* no-mistakes(test): Synchronize peer beacon regression fixture
* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures
* no-mistakes(document): Captain: document watcher successor protocol behavior
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: fetch current PR head for review diffs (#722)
* fix: always fetch PR head for review diffs
Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.
* no-mistakes(test): Isolate session-start nudge tests from gate state
* no-mistakes(document): Correct review-diff documentation
* docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736)
* docs: resolve five contract contradictions
* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): correct Grok exit guidance (#742)
* docs(harness): reverify grok exit command
* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for parked crew (#743)
* fix(watcher): bound stale wakes for exited paused crew
* no-mistakes(review): Gate pause suppression on confirmed agent death
* no-mistakes(test): Fixed stale pause cadence
* no-mistakes(document): Document dead-agent hold cadence
* fix(supervision): distinguish ordinary wakes from recovery (#744)
* fix(supervision): distinguish ordinary wakes from repair
* no-mistakes(review): Make passive guard follow-ups recovery-only
* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes (#745)
* fix(x-mode): dedupe pending mention wakes
* no-mistakes(review): fix x-poll claim error deduplication
* no-mistakes(review): separate claim diagnostics from relay recovery
* no-mistakes(document): Document X-mode once-only mention wakes
* feat(wake): enrich drained signals with bounded status context (#747)
* feat(wake): enrich drained signal context
* no-mistakes(review): Bound wake enrichment reads
* no-mistakes(document): Document wake-drain annotations
* docs(wake): explain at-least-once drain boundary
* no-mistakes(review): Prevent symlink races in wake annotations
* no-mistakes(review): Exercise wake symlink race regression
* test: document intentional AFK marker subprocesses
* fix(wake): isolate annotation marker state
* docs: default future planning to Markdown
* feat(herdr): add optional presentation spaces (#784)
* feat(herdr): add optional presentation spaces
* no-mistakes(review): Harden Herdr projection creation and spawn serialization
* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission
* no-mistakes(test): Correct stale Orca metadata failure fixture
* no-mistakes(document): Document Herdr presentation projection accurately
* fix(send): treat opencode busy-queued composer state as submitted (#775)
* fix(send): treat opencode busy-queued composer state as submitted
When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row. The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.
Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted). On an idle pane, keep returning "pending"
(genuine swallow detection preserved).
Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pa…
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Jul 26, 2026
* feat(backends): add experimental cmux runtime backend (#246)
* feat(backends): add cmux runtime backend (experimental)
Session-provider-only adapter for cmux (bin/backends/cmux.sh), mirroring
zellij/herdr structurally, wired into fm-backend.sh and fm-spawn.sh with
--secondmate refused for now. Verified against the real cmux 0.64.17 app:
send does not auto-submit, cwd is creation-time-frozen (zellij-shape,
pwd-marker-probe workaround), close-surface refuses on a workspace's last
surface (falls back to close-workspace), workspace ids do not survive a
relaunch, and the control socket defaults to cmuxOnly access (requires a
one-time password-mode setup, documented in docs/cmux-backend.md). Also
found and fixed a live bug during development: read-screen fails on a
surface that has never been written to, so liveness now uses list-panes
instead. Fake-CLI unit suite (40 tests), a real-binary smoke test, and a
full spawn/steer/peek/done/merge/teardown E2E pass against a real claude
crewmate all pass, including the popup/second-Enter regression class.
* no-mistakes(review): Harden cmux recovery and password parsing
* no-mistakes(review): Harden cmux capture failure handling
* no-mistakes(review): Mark cmux test scripts executable
* no-mistakes(review): Scope cmux workspaces and teardown
* no-mistakes(review): Captain, honor cmux password config override
* no-mistakes(review): Captain, hash cmux home labels
* no-mistakes(document): Sync cmux backend docs
* feat(agents): add firstmate coding guidelines skill (#248)
* Add firstmate-coding-guidelines skill (AGENTS.md diet PR 0)
Encodes the knowledge-placement decision tree, one-owner rule, and
inline-stub pattern from the diet analysis so future contributions stop
adding conditional detail inline. AGENTS.md gets one section-13 trigger
line; fm-brief.sh's REPO argument has no reliable signal for "this is
firstmate's own repo", so the load instruction goes in CONTRIBUTING.md's
Development section instead of the scaffold.
* no-mistakes(review): Captain, align tracked-material trigger scope
* no-mistakes(document): Sync coding guidelines docs
* no-mistakes(lint): Fix Markdown style issues
* fix: add turn-end supervision guard (#249)
* feat: structural Stop-hook backstop for primary turn-end supervision
fm-guard.sh is pull-based: it only warns when some other supervision
script happens to run, so a primary session that ends a turn without
re-arming the watcher and then runs no further fleet-touching command
can sit blind for hours (the 2026-07-04 incident this fixes).
Add bin/fm-turnend-guard.sh, a Claude Code Stop hook registered in the
tracked .claude/settings.json, that fires on every primary turn end and
blocks (exit 2, verified empirically to force continuation) when work
is in flight with no fresh watcher beacon. It never blocks more than
once per turn, using Claude Code's own stop_hook_active loop-guard
field, and scopes itself to the actual primary checkout only (inert in
crewmate/scout worktrees and secondmate homes).
Factor the shared "in-flight but no live watcher" predicate out of
fm-guard.sh into bin/fm-supervision-lib.sh so the pull-based banner and
the push-based hook can never drift on what "unhealthy" means.
Document the verified Stop-hook mechanism and scoping in
docs/turnend-guard.md, add a harness-adapters note, and cover the
predicate and hook with tests/fm-turnend-guard.test.sh.
* no-mistakes(review): Respect active home in turnend guard
* no-mistakes(review): Require live watcher for turn-end guard
* no-mistakes(review): Captain: portable turn-end timing
* no-mistakes(document): Sync turn-end guard documentation
* feat(backends): auto-detect cmux runtime (#250)
* feat(backends): auto-detect cmux runtime from CMUX_WORKSPACE_ID
Wires cmux into fm_backend_detect the same way herdr already is: a
firstmate process running inside a cmux-spawned terminal now spawns
new tasks into cmux by default, no config needed. Verified from cmux's
own shipped source that CMUX_WORKSPACE_ID/CMUX_SURFACE_ID/CMUX_SOCKET_PATH
are unconditionally, non-overridably injected into every terminal
surface it spawns, and that cmux's own CLI treats CMUX_WORKSPACE_ID as
its own ambient-target fallback - the same role $TMUX/HERDR_ENV play for
their backends. CMUX_WORKSPACE_ID is checked last (after $TMUX and
HERDR_ENV=1) since cmux is a terminal application, not a nestable
multiplexer. Socket auth (config/cmux-socket-password) stays required
regardless of how the backend was selected; the existing spawn refusal
now also names the config/backend=tmux / --backend tmux opt-out for a
caller who never explicitly chose cmux.
A live env dump inside a real cmux terminal was not obtained safely on
the shared dev machine (documented in docs/cmux-backend.md); this rests
on the source read instead, mirroring this doc's existing
verified-from-source precedent.
* no-mistakes(review): Fix cmux autodetect docs and tests
* no-mistakes(document): Document cmux auto-detection
* fix(afk): support herdr away-mode injection (#251)
* fix(afk): make the away-mode daemon backend-aware for herdr
bin/fm-supervise-daemon.sh discovered its supervisor pane and injected
via raw tmux calls only, so /afk failed outright on a herdr-based
fleet (TMUX_PANE unset, firstmate:0 fallback unresolvable).
Discovery now resolves backend (tmux|herdr) and target independently,
mirroring fm-backend.sh's own runtime auto-detection, with an explicit
FM_SUPERVISOR_BACKEND override alongside the existing FM_SUPERVISOR_TARGET.
zellij/orca refuse loudly at startup instead of misapplying tmux
primitives. Injection (pane-exists probe, busy-guard, composer-guard,
verified submit) now dispatches through bin/fm-backend.sh's generic
primitives, adding a new fm_backend_composer_state dispatcher; the
tmux path is byte-identical to before. Also fixes a pre-existing bug
in fm_backend_target_exists's herdr arm (missing --session, so it
silently misrouted once more than one herdr server was running) found
while verifying this end to end against a real isolated herdr session.
Classification, batching, max-defer, the marker contract, locks, and
wake-queue handling are unchanged - this is a transport-layer fix.
* no-mistakes(review): Corroborate Herdr idle busy state
* no-mistakes(review): Stabilize Herdr daemon startup wait
* no-mistakes(review): Captain, route cmux composer and update AFK docs
* no-mistakes(document): Document AFK supervisor backend support
* docs(agents): move X-mode procedures out of AGENTS (#253)
* docs(agents): collapse X-mode section 14 into fmx-respond/docs pointers
AGENTS.md diet PR 1 of 3 (agentsmd-diet-s2 report, move-plan items 1-2).
Replaces section 14's "Answering"/"Completion follow-up"/"Conversations"/
"Length and threads"/"Preview / dry-run" blocks (54 lines) and the
"Mechanism" narrative (6 lines) with two short pointers: fmx-respond
(section 13) for the procedure, docs/configuration.md "X mode (.env)"
for the wire protocol. Net -55 lines in AGENTS.md.
Destination edits landed first, deletions second (q4 discipline):
- docs/configuration.md: added the "purely additive, watcher untouched"
guarantee that AGENTS.md's Mechanism block stated but configuration.md
did not.
- fmx-respond/SKILL.md: added the x-mode-error wake boundary (report as
a blocker, do not load this skill), the --image flag for replies and
follow-ups, the "images are for real artifacts, not prose" rule, and
the dry-run compact-image-marker behavior - none of these were
previously in the skill even though AGENTS.md described them, so they
were genuine gaps, not pre-existing duplication. Also made the skill's
own "Completion follow-up" section the sole, full owner of that
procedure instead of deferring to AGENTS.md section 14 for substance
that no longer lives there (two internal cross-references updated to
point at section 8's terminal-wake trigger and the skill's own section
instead).
Mechanical line-by-line audit of every removed AGENTS.md line:
Mechanism block (6 lines removed):
- bootstrap artifact-writing description -> already owned by
docs/configuration.md "X mode (.env)" (locked-bootstrap paragraph)
- check-shim/poll mechanism description -> already owned by
docs/configuration.md same section
- missing-deps/x-mode-error diagnostic description -> already owned by
docs/configuration.md ("Relay auth or config problems...") plus
bin/fm-x-poll.sh's own header comment for the missing-curl/jq mechanics
- opt-out artifact removal description -> already owned by
docs/configuration.md same section
- "purely additive, no edit to fm-watch.sh/fm-watch-arm.sh/fm-wake-lib.sh/
afk daemon" guarantee -> MOVED to docs/configuration.md (added in this
PR; this fact had no other home before)
Answering/Completion follow-up/Conversations/Length and threads/
Preview-dry-run blocks (54 lines removed):
- x-mention wake -> load fmx-respond: already owned by section 13's
existing trigger line (unchanged) and restated in the new pointer
- x-mode-error wake -> report as blocker, don't load fmx-respond: MOVED
to fmx-respond/SKILL.md (added in this PR)
- inbox-draining, classification, acting, reply composition, submission,
cleanup-on-success/failure: already owned by fmx-respond/SKILL.md
"Procedure" section (unchanged, pre-existing)
- owner-only routing / captain-as-asker framing: already owned by
fmx-respond/SKILL.md "The asker is your own captain" section
- standing X-mode authorization / autonomous posting / dry-run as only
non-posting path: already owned by fmx-respond/SKILL.md same section
- acknowledge-first -> act -> follow-up shape, three-case classification:
already owned by fmx-respond/SKILL.md "A request to act on" section
- destructive/irreversible/security-sensitive escalation guardrail:
already owned by fmx-respond/SKILL.md "Public channel..." section and
Procedure step 2c
- dismiss-instead-of-reply for pure acknowledgments, relay re-offer
prevention, dry-run honoring: already owned by fmx-respond/SKILL.md
Procedure steps 2b/2c/2e-skip and docs/configuration.md
- public-safety bar (no task ids/internals/captain-private/secrets):
already owned by fmx-respond/SKILL.md "The reply is public" section
- never-inline-into-shell-command / --text-file or stdin: already owned
by fmx-respond/SKILL.md Procedure step 2e and Notes
- --image flag for replies (formats, base64, no-inline guarantee): MOVED
to fmx-respond/SKILL.md Procedure step 2e (added in this PR - this was
not previously in the skill)
- fm-x-link field names (x_request=, x_request_ts=, x_followups=):
already owned by AGENTS.md section 2's state/<id>.meta field list
(untouched, out of scope for this PR) and fmx-respond/SKILL.md
- carry-count/carry-ts relink behavior, three-follow-up budget, milestone
sparingness, --check/--text-file posting, connector/followup wire
detail, --final clearing, cap/window graceful degradation: already
owned by fmx-respond/SKILL.md "Completion follow-up" section (now sole
owner) and docs/configuration.md wire-protocol paragraphs
- --image flag for follow-ups: MOVED to fmx-respond/SKILL.md "Completion
follow-up" section (added in this PR - genuine gap)
- "failed task still gets an honest final follow-up": already owned by
fmx-respond/SKILL.md "Completion follow-up" section
- FMX_DRY_RUN whole-loop previewability: already owned by
fmx-respond/SKILL.md "Dry-run / preview mode" section
- in_reply_to conversation continuity, untrusted-thread handling,
follow-up worthiness judgment, relay-owned self-reply guard/cap:
already owned by fmx-respond/SKILL.md "The direct ask is the captain's"
section and Notes (one bullet is a verbatim match)
- concise-by-default / no hand-numbered threads: already owned by
fmx-respond/SKILL.md "Voice" section
- auto-split behavior, char/tweet caps, premium-independence, wire shape
({text}/{text,texts}): behavior already owned by fmx-respond/SKILL.md
Voice section; exact defaults and wire shape already owned by
docs/configuration.md; "premium-independent" mechanics already owned
by bin/fm-x-reply.sh's own header comment
- "images are for real artifacts, not prose": MOVED to fmx-respond/
SKILL.md "Voice" section (added in this PR - genuine gap)
- image-on-thread wire behavior: already owned by docs/configuration.md;
reinforced in fmx-respond/SKILL.md's new --image note
- dry-run POST-body shape, endpoint marker, truthy-value definition,
jq-only dependency, end-to-end testability, x-outbox inspection:
already owned by fmx-respond/SKILL.md "Dry-run / preview mode" section
(several near-verbatim matches) and docs/configuration.md wire detail
- dry-run compact image marker: MOVED to fmx-respond/SKILL.md "Dry-run /
preview mode" section (added in this PR - genuine gap)
Section 8's terminal-wake completion-follow-up trigger (the one fact
required to survive inline) is untouched and already present; the new
section 14 pointer references it instead of restating it.
Nothing outside section 14 (plus the two destination files) is touched.
Full test suite green, including all 74 fm-x-mode.test.sh checks.
* no-mistakes(review): Preserve X-linked follow-up triggers
* no-mistakes(review): Fix x-mode error trigger
* no-mistakes(document): Docs cross-reference synchronized
* no-mistakes(lint): Clean Markdown lint pass
* docs: trim duplicated harness guidance (#255)
* docs(agents): trim section 4 harness/secondmate duplication
AGENTS.md diet PR 2 of 3 (data/agentsmd-diet-s2/report.md, move-plan
items 3-4; redundancy item 2 folded into item 3).
Removed the five claude/codex/grok/pi/opencode model/effort-flag
bullets from section 4 - byte-for-byte duplicated by
harness-adapters' "Launch profile axes" table (which is already a
superset: it carries verified CLI versions per adapter that the
AGENTS.md bullets lacked). Replaced with a one-line pointer; the
skill is already loaded before every spawn per section 4's own
closing trigger, so no new trigger was needed.
Moved the config/secondmate-harness model/effort pin-format detail
(the `<harness> [<model>] [<effort>]` line format, the
secondmate-model/secondmate-effort accessors, back-compat, and the
durability-across-respawn behavior) into secondmate-provisioning,
which is already a mandatory load at every secondmate lifecycle
touchpoint. Added the destination content to the skill first, then
replaced the AGENTS.md paragraph with a 3-line pointer.
Mechanical audit - every removed line's new home:
- 5 harness bullets (claude/codex/grok/pi/opencode model+effort
flags, per-harness max-omission rationale) -> already present in
harness-adapters SKILL.md's "Launch profile axes" table (lines
53-59), confirmed fact-by-fact before deleting.
- "config/secondmate-harness may also pin..." paragraph (pin format,
bare-harness back-compat, secondmate-model/secondmate-effort
accessors, per-spawn override precedence, respawn durability,
secondmate-only scope) -> secondmate-provisioning SKILL.md's
"Charter and seed" section, added verbatim before this trim.
- The following paragraph (inheritable config: crew-dispatch.json,
crew-harness, backlog-backend) is untouched - out of scope for
this PR, still inline.
- The bootstrap CREW_DISPATCH effort-mismatch diagnostic sentence is
untouched - not part of the five-bullet duplication, stays inline.
No script changes. Section 4 shrinks from 104 to 87 lines
(958 -> 901 total AGENTS.md lines) with zero facts lost: every fact
is reachable through harness-adapters or secondmate-provisioning,
both already mandatory loads at the relevant lifecycle points.
* no-mistakes(document): Align secondmate skill triggers
* no-mistakes(lint): Markdown style clean
* fix: anchor turn-end Stop hook to project root (#256)
* Fix turn-end Stop hook to use CLAUDE_PROJECT_DIR path
Claude Code runs hook commands via /bin/sh from the session cwd, so the
bare relative bin/fm-turnend-guard.sh path fails when cwd is not the repo
root. Anchor the command with "$CLAUDE_PROJECT_DIR"/bin/fm-turnend-guard.sh
instead; verified CLAUDE_PROJECT_DIR is set on Stop hooks in Claude Code
2.1.201. Document the cwd caveat and add a settings.json regression test.
* no-mistakes(document): Document Stop hook path anchoring
* docs: trim firstmate agent guidance duplication (#258)
* docs: trim AGENTS.md redundancy (diet PR 3/3)
Consolidates five duplicated passages to a single owner each, per
data/agentsmd-diet-s2/report.md redundancy items c3-c7:
- Inheritable-config propagation mechanism: owned by section 3 (where
the sweep runs); sections 4 and 7 keep compact references. Section 4
retains its one genuinely unique fact (crew-harness inherit-vs-fallback
semantics), just no longer restates the propagation mechanism itself.
- Landed-work definition: owned by section 7's ship-teardown detail
(PR-containment mechanics, pr= discovery fallback); section 1's hard
rule #3 keeps the rule plus a three-case summary and a pointer.
- Backend meta-field enumeration: owned by docs/configuration.md
("Runtime backend", already comprehensive including cmux) and each
backend's own doc; AGENTS.md keeps only the fields common to every
task plus a pointer.
- Dropped one redundant restatement of "silence is correct while
waiting" in section 8.
- Worktree-tangle guard explanation: owned by section 8 (already the
fuller, cross-referenced version); section 3's TANGLE bullet keeps
the remediation action and points at section 8 for the why.
Also adds two captain-requested single-sentence rules: invoke bin/
scripts by absolute $FM_ROOT path after any cd away from the home, and
a backend spawn refusal must be surfaced to the captain rather than
silently worked around by switching backends.
AGENTS.md: 901 -> 889 lines, 112355 -> 108560 bytes.
* no-mistakes(review): Clarify post-cd bin invocation guidance
* no-mistakes(document): Sync AGENTS trim docs
* no-mistakes(lint): Fix Markdown line style
* feat(backends): improve cmux detection and socket-mode guidance (#259)
* feat(backends): cmux detection fallbacks and socket-mode matrix
Workstream A: cmux's bundled claude wrapper strips every CMUX_* env var on
its passthrough path (reproduced live 2026-07-04, cmux 0.64.17), so a
claude-harness firstmate inside a cmux tab has no CMUX_WORKSPACE_ID.
fm_backend_detect now falls back - macOS-only, only when the primary marker
is absent - to __CFBundleIdentifier=com.cmuxterm.app and then a process
ancestry walk resolved by bundle id (lsappinfo) plus a bundle-shaped ps comm
match. Innermost-first ordering is unchanged and absorbs the
tmux-inside-cmux bundle-id false positive; the auto-detect NOTICE names the
winning fallback signal.
Workstream B: the five socketControlMode values were traced through cmux
source (commit 9c91710e3f58): off/cmuxOnly can never admit an external CLI,
automation admits same-user clients with no secret (0600 socket only),
password needs the auth handshake, allowAll opens the socket to every local
user (0666). Automation mode is now the documented recommendation; the
adapter's refusals name every viable mode, classify Invalid password as
unauth, and the launch-timeout message names the off-mode possibility.
Docs carry the wrapper-strip empirical record, the fallback contract and
authority split, and the full mode matrix with rationale; tests cover the
new detection paths, the nested false positive, and the refusal wording.
* no-mistakes(review): Document cmux fallback detection
* no-mistakes(review): Update cmux architecture docs
* no-mistakes(document): Align cmux backend docs
* fix(backends): scope zellij tabs by firstmate home (#252)
* fix(backends): home-scope zellij tab titles to close cross-home collision gap
Zellij's one shared "firstmate" session has no per-home split and enforces
no tab-name uniqueness, so two firstmate homes with colliding task ids could
send/peek/close each other's tabs - the same gap a no-mistakes review gate
caught for cmux (docs/cmux-backend.md). Ports that fix: every new tab is
created with a home-scoped title (fm-<home-label>-<id>), and every
list/find/recover/kill path scopes matches to this home's own tag. A tab
spawned before this change still matches via its old untagged bare title,
but only when unambiguous - two live tabs sharing a bare title refuse rather
than guessing which one is ours.
Factors the home-label/hash derivation shared with cmux into
bin/fm-backend-hometag-lib.sh so the two adapters can't drift.
* no-mistakes(review): Fix zellij child teardown home tag
* no-mistakes(review): Fix zellij teardown and selector scoping
* no-mistakes(document): Sync zellij home-scope docs
* fix: sync project clones after merged PR wakes (#293)
* fix(fleet-sync): auto-sync on merged-PR wake, accept project name
fm-fleet-sync.sh's single-project form failed on a bare project name
("not a directory"), forcing hand-typed full paths (4 manual runs in
one incident). It now resolves a bare name or projects/<name> against
the home's projects dir.
AGENTS.md now encodes the trigger: a wake whose status reports a
merged PR for a project cloned in this home runs fleet-sync for that
project as part of handling the wake, so a secondmate-reported merge
does not leave the primary's clone stale until the next session start
or teardown.
* no-mistakes(review): Fix fleet-sync project name shadowing
* no-mistakes(document): sync fleet-sync docs
* fix: canonicalize spawn worktree path checks (#294)
* fix(spawn): canonicalize worktree-isolation guard against symlinked project prefixes
fm-spawn.sh compared a logical PROJ_ABS against the physically-resolved
pane cwd every backend reports, so a project reached through a symlinked
prefix (e.g. macOS's /tmp -> /private/tmp) could trip the isolation
guard's false refusal before treehouse ever moved the pane. Canonicalize
once into PROJ_ABS_REAL and compare against that everywhere instead.
* no-mistakes(review): Canonicalize spawn cwd comparisons
* no-mistakes(document): Refresh symlinked spawn docs
* docs: add Orca operator skill (#276)
* docs: add Orca operator skill
* no-mistakes(document): Document Orca checklist
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* fix: surface green PRs during CI monitoring (#297)
* fix(crew-state): detect green-PR CI monitoring, escalate repeat wedges
fm-crew-state.sh's ci step never distinguishes "still waiting on checks"
from "checks green, waiting on merge" via axi status alone, since a repo
that defers merge to the captain keeps the ci step at status=running for
the whole monitor phase. Read the ci step's own log tail (axi logs) for
the checks-passed marker and surface done instead of a false "validating
(running)" - verified against the real PR #252 run's ci.log.
The watcher's wedge timer can re-escalate the same stale pane forever
without ever signaling that it is a repeat; track a per-pane consecutive
escalation count and add a demand-deep-inspection marker to the wake
payload once it crosses a threshold, so the supervisor can no longer
dismiss each one as an isolated, still-validating pane.
Also clarify the ship-brief's checks-green line: it is owed at the
CI-ready return point, not after the background monitor-until-merge
loop finishes.
* no-mistakes(review): Captain, distinguish pending no-checks CI marker
* no-mistakes(review): Harden CI relapse handling
* no-mistakes(review): Block stale done during fixing
* no-mistakes(review): Captain, tighten CI status gating
* no-mistakes(review): Captain, harden stale CI green handling
* no-mistakes(review): Captain, recognize ranged CI rearm markers
* no-mistakes(document): Sync crew-state supervision docs
* fix(teardown): recover provably stale git index locks (#296)
* fix(teardown): recover from a stale worktree git index.lock
A crew process killed mid-git-operation can leave a stale
.git/worktrees/<wt>/index.lock behind, making fm-teardown.sh's
`treehouse return --force` fail closed. On that failure, retry once
after a short wait (the owning process may be exiting), then remove
the lock and retry once more only when it is provably stale: old
enough by mtime and lsof shows no live holder on the lock or the
worktree itself. A lock that isn't provably stale is left in place and
the original failure still surfaces.
* no-mistakes(review): Harden teardown lock refusal paths
* no-mistakes(review): Harden stale-lock teardown safety rechecks
* no-mistakes(review): Harden stale teardown lock checks
* no-mistakes(document): Document teardown lock recovery
* feat(bin): encode project AGENTS.md authoring bar with canonical self-governance section (#307)
* Encode project AGENTS authoring bar
* no-mistakes(review): Captain, centralize CLAUDE promotion governance
* no-mistakes(review): make ensure_maintenance_section idempotent-success, drop || true guards
* no-mistakes(review): separate appended maintenance section on newline-less CLAUDE.md promotion
* no-mistakes(review): assert maintenance heading present before separator check in test
* no-mistakes(document): sync docs with AGENTS.md authoring bar and self-governance
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* feat(skills): add captain-invocable bearings status-report skill (#300)
* Add captain-invocable bearings skill
Generates a pick-up-where-I-left-off status report from live fleet
state to data/status-report-<YYYY-MM-DD>.md plus a concise chat
summary. Read-mostly procedure: reads backlog, per-task crew state
via bin/fm-crew-state.sh, open PRs via gh-axi, scout reports,
pending decisions, and date-gated queued work; composes the
exemplar's sections (TL;DR, Check first, Landed, In flight, Plans,
Decisions pending, Date-gated/queued); never tears down, merges, or
mutates task state as a side effect.
* no-mistakes(document): docs: list new /bearings skill in README built-in skills table
* fix(watcher): make PID identity locale-invariant (#285)
* fix(watcher): pin LC_ALL=C in fm_pid_identity for locale-invariant identity
ps's lstart date format follows the caller's LC_TIME/LC_ALL. The watcher records
its process identity under one locale, but arm/guard/turn-end re-read it under the
machine's ambient locale. On a non-C locale (e.g. ko_KR) the two strings differ
only in the date portion, so fm_watcher_lock_matches_pid / fm_watcher_healthy
reject a genuinely live watcher - breaking fm-watch-arm.sh, fm-guard.sh, and
fm-turnend-guard.sh on every non-C-locale machine.
Pin LC_ALL=C on that one ps call so the write and read sides agree regardless of
machine locale, matching the LC_ALL=C determinism the file already uses elsewhere.
Add a colocated regression test asserting fm_pid_identity is locale-invariant
across exported LC_ALL/LC_TIME.
* no-mistakes(document): Document watcher PID identity coverage
* docs: document codex app backend contract (#222)
* docs: reconcile Codex App backend contract
* no-mistakes(document): Sync backend docs
* docs: clarify Codex Desktop bridge blocker
* no-mistakes(document): Align Codex App backend docs
* no-mistakes(test): Captain, stabilize watcher self-eviction test cadence
* no-mistakes(document): Document Codex App backend contract
* no-mistakes(document): Captain, document blocked codex-app coverage
* docs: make Codex App contract doc authoritative
* no-mistakes(document): Align Codex App backend docs
* docs: redact local Codex App smoke paths
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* docs: add Codex Desktop coordination skill (#275)
* docs: add Codex App coordination skill
* no-mistakes(review): Captain, mark Codex App skill agent-only
* no-mistakes(document): Document Codex App backend boundary
* no-mistakes(document): Captain, document Codex Desktop backend boundary
* no-mistakes(lint): Captain, lint clean
* no-mistakes(document): Document Codex Desktop boundaries
* docs: narrow Codex App skill playbook
* fix(afk): stop herdr escalation redelivery loop (#317)
* fix(afk): recognize unbordered herdr composer rows to stop escalation redelivery loop
fm_backend_herdr_composer_state only recognized bordered composer rows
(the grok shape). Real claude and codex render their live input row
with no border at all, so once a harness's own startup banner scrolled
out of the capture window the classifier read the composer as unknown
forever. fm_backend_herdr_send_text_submit never confirmed "empty", so
escalate_flush never cleared state/.subsuper-escalations, and the
away-mode daemon retyped and resubmitted the same buffered digest every
housekeeping cycle - reproduced live against a real herdr+claude pane
(5+ identical deliveries in 40s).
The classifier now recognizes an unbordered (bare) composer row led by
a known prompt glyph alongside the existing bordered shape, keeping
whichever match is bottom-most so a stale decorative box never
outranks the live composer.
* no-mistakes(review): Narrow herdr bare prompt matcher
* no-mistakes(document): Sync herdr composer docs
* fix(backends): confirm Herdr submits with native agent state (#323)
* fix(herdr): confirm message submit via native agent-state, not composer text
fm_backend_herdr_send_text_submit now confirms a landed submit by polling
herdr's own agent-state (agent get) for the idle->working transition instead
of reading composer content. Composer scraping remains, unchanged, for the
away-mode daemon's pre-injection empty-box guard only.
This fixes the practical effect of the codex idle-tip gap from the
2026-07-07 incident: codex's dynamic idle-composer hint text can no longer
misread as pending and block/mis-confirm a send, since confirmation no
longer looks at composer text at all. Verified empirically against real
claude and codex agents (timing, swallowed-Enter, unreadable-target, and
already-busy-target scenarios), and against the real away-mode daemon
end-to-end after updating its synthetic supervisor-pane test fixture to
register itself as a real herdr agent (herdr's own report-agent primitive)
so it can still exercise the new confirmation path.
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(document): Sync Herdr submit docs
* no-mistakes: apply CI fixes
* feat: add quota-balanced crew dispatch selection (#327)
* Add quota-balanced dispatch selection
* no-mistakes(document): Document dispatch selector guidance
* fix(session-start): respawn dead secondmate agents conservatively
* fix(session-start): deterministically respawn dead-shell secondmates
A secondmate agent that exits leaves its backend pane alive as a bare
shell. The session-start endpoint check only verified pane presence, so
recovery and the watcher (which exempts secondmates from stale-pane
detection) never noticed - evidence 2026-07-07: every secondmate in one
fleet was found sitting at a dead zsh shell.
Add fm_backend_agent_alive (bin/fm-backend.sh), a deeper per-backend
liveness probe distinct from pane presence: fm_backend_tmux_agent_alive
classifies the pane's live foreground process via tmux's own
pane_current_command, and fm_backend_herdr_agent_alive reuses the
already-verified pane_agent_state husk classifier. Both are conservative:
anything ambiguous reports unknown, never a false dead.
Wire this into a new session-start-only, locked-and-primary-only sweep in
bin/fm-bootstrap.sh that kills and respawns only a confidently dead
secondmate endpoint, leaving alive/unknown readings untouched - idempotent
by construction, so repeated runs converge without duplicating agents.
* no-mistakes(review): Guard raw secondmate liveness respawns
* no-mistakes(review): Fix detect-only bootstrap test
* no-mistakes(test): Pin liveness fixture harness
* no-mistakes(document): Sync secondmate liveness docs
* no-mistakes: apply CI fixes
* fix: emit stable secondmate nudge selectors (#331)
* Fix NUDGE_SECONDMATES to print stable fm-<id> selectors.
Session-start secondmate sync used to accumulate raw backend window targets
into NUDGE_SECONDMATES, but the liveness sweep in the same bootstrap run can
respawn secondmates onto new endpoints. fm-send with those stale explicit
targets bypasses meta resolution and fails, while fm-<id> resolves correctly.
Accumulate fm-<id> in process_secondmate, update the bootstrap/update contracts
and /updatefirstmate skill, and add a herdr respawn regression test.
* no-mistakes(review): Captain, guard herdr regression jq dependency
* no-mistakes(document): Document stable secondmate nudge selectors
* no-mistakes(lint): Fix shell lint hints
* feat: require bootstrap detection for AXI tools (#332)
* Make tasks-axi and quota-axi required bootstrap tools
Add both to the normal toolchain checks alongside lavish-axi, keep the
tasks-axi 0.1.1+ compatibility gate, and report quota-axi through the
standard MISSING install-consent flow. TASKS_AXI: available remains a
backlog-backend capability signal only; manual opt-out no longer suppresses
the missing-tool report.
Update bootstrap tests and point docs/configuration.md at the canonical
toolchain contract.
* no-mistakes(review): Clarify manual backlog bootstrap reporting
* no-mistakes(document): Document bootstrap AXI tools
* bearings: delete today's report before recreating (#333)
Replace overwrite-in-place wording with explicit delete-then-create
instructions so agents do not modify an existing daily report file.
* feat: guard primary turn ends across harnesses (#339)
* Add primary turn-end guards for all harnesses
* no-mistakes(review): Normalize Codex hook cwd resolution
* no-mistakes(review): Fix OpenCode guard worktree anchoring
* no-mistakes(review): Anchor Codex guard outside nested roots
* no-mistakes(review): Anchor Codex guard to hook root
* no-mistakes(review): Avoid Grok permission escalation
* no-mistakes(document): Sync turn-end guard docs
* fix: resolve backend selectors by exact task id first (#342)
* fix backend selector task id resolution
* no-mistakes(document): Document selector resolution behavior
* fix: scale bootstrap fleet-sync timeout (#341)
* fix bootstrap fleet sync timeout
* no-mistakes(review): Fix bootstrap fleet-sync timeout regressions
* no-mistakes(document): Sync bootstrap timeout docs
* no-mistakes(lint): Clean ShellCheck directives
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add fleet snapshot and view commands (#343)
* Add fleet snapshot and view
* no-mistakes(review): Fix fleet snapshot parsing and overrides
* no-mistakes(review): Fix secondmate fleet rendering
* no-mistakes(review): Fix backlog title and completion parsing
* no-mistakes(review): Include durable scout reports
* no-mistakes(review): Fix fleet snapshot edge cases
* no-mistakes(review): Captain: gate fleet hints on current state
* no-mistakes(review): Captain: parse bracketed Done PR artifacts
* no-mistakes(document): Sync fleet snapshot docs
* fix(fm-send): fail loudly on unresolvable send targets (#254)
* Make fm-send fail loudly on unresolved targets
* no-mistakes(review): Document fm-send FM_HOME contract
* Fix fm-send readiness docs and backend send path
* Fix fm-send docs for cmux and X skill metadata
* Make gotmp teardown test home-explicit
* Scope watcher warning wording to fm-send
* Fix fm-send review findings
* Verify explicit tmux targets before sending
* Isolate turnend guard test home
* no-mistakes(document): Documented fm-send FM_HOME/backend guard additions missing from doc inventories
---------
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* fix: deliver AFK escalations through herdr supervisors (#353)
* fix afk codex ghost composer delivery
* no-mistakes(review): Harden AFK startup flag writes
* no-mistakes(review): Harden AFK daemon liveness checks
* no-mistakes(document): Sync AFK herdr docs
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add harness-aware supervision (#367)
* Add harness-aware supervision
* no-mistakes(review): Captain, harden watcher supervision regressions
* no-mistakes(review): Captain, harden watcher supervision cadence
* no-mistakes(review): Harden watcher supervision ownership
* no-mistakes(review): Captain, harden Pi extension marker
* no-mistakes(review): Captain, harden Pi supervision restart checks
* no-mistakes(review): Harden watcher ownership checks
* no-mistakes(review): Captain, harden Pi supervision loading
* no-mistakes(review): Captain, require Pi guard extension loading
* no-mistakes(review): Captain, harden watcher supervision recovery
* no-mistakes(test): Fix fm-send baseline log filtering
* no-mistakes(document): Sync harness supervision docs
* no-mistakes: apply CI fixes
* fix: split X-mode replies by platform (#369)
* fix: make x replies split by platform
* no-mistakes(review): Captain: preserve Discord recovery relink context
* no-mistakes(test): Captain: keep split markers outside fences
* no-mistakes(document): Sync X-mode reply docs
* fix: make stow memory writes inspect before update (#372)
* docs: make stow inspect-then-update
* no-mistakes(review): Remove unsupported archive-body guidance
* no-mistakes(review): Clarify stow read-before-write exception
* no-mistakes(test): Require archive-body for stow task notes
* no-mistakes(document): Sync stow memory docs
* no-mistakes(lint): Silence ShellCheck source warning
* fix(watcher): wait when arm attaches to a healthy watcher (#375)
* fix: attach-and-wait when arm finds a healthy watcher
Grok and Claude re-arm after every turn with work in flight. When a
watcher was already healthy, fm-watch-arm exited immediately with
watcher: healthy, which completed the harness background task and
injected an empty false wake.
Attach to the live identity-matched holder instead, stay until that
cycle ends, then exit 0 so notify fires for a real end-of-cycle. The
peer-startup-race path uses the same contract. --restart and the
started path are unchanged.
* no-mistakes(review): Gate restart watcher peer attach
* no-mistakes(document): Sync watcher arm docs
* feat(pi): simplify primary session launch (#386)
* docs(readme): reformat Quick Start and recommend Grok equally with Claude Code
* no-mistakes(review): Captain: align harness launch guidance
* no-mistakes(review): Captain, clarify Pi supervised launch
* no-mistakes(review): Captain, document Pi first-launch bridge
* feat(pi): track primary watcher extension for plain-pi launch
Move Pi's primary watcher bridge from a generated state/ file to a
tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end
guard extension already works: self-hashing version, project-local
auto-discovery after one-time Pi trust. This drops the state/-generation
step and dual -e requirement from the happy path, so Pi's Quick Start
launch becomes plain 'pi', the same friction class as 'claude' and
'grok --trust'.
- bin/fm-pi-watch-extension.sh is removed; nothing generates the
extension anymore since it is committed.
- fm-session-start.sh and fm-supervision-instructions.sh resolve the
watcher extension path from FM_ROOT instead of state/, and the
session-start diagnostic now points at restarting plain pi after
trust, with -e as a documented fallback.
- fm-spawn.sh points Pi secondmate launches at the tracked extension
path in the secondmate home instead of generating a state/ copy.
- README Quick Start Pi block is now just 'pi' plus a trust note.
- Tests, docs, and the harness-adapters skill updated to match.
* fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016
* feat(supervision): prevent unsafe watcher-arm commands (#387)
* feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns
Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that
denies a primary shell command backgrounding, piping, or bundling the
watcher arm/checkpoint, or force-killing the watcher process broadly -
the exact shapes that silently took Grok's supervision down. Wires it
into all five verified harnesses (grok, claude, codex, opencode, pi),
each validated empirically against the real harness.
Also fixes a grok 0.2.93 regression discovered during that validation: the
existing turnend-guard Stop hook's bare root variable broke grok's own
variable pre-substitution and silently no-op'd the hook.
* no-mistakes(review): Harden watcher arm validation
* no-mistakes(review): Harden arm guard metacharacter checks
* no-mistakes(review): Harden nested shell arm guard
* fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then
A && B || C is not if-then-else; C can run when A is true. Replace both
occurrences of the quote-state early-continue with an explicit if/then.
* fix(pi): restore primary watcher supervision lifecycle (#397)
* fix Pi primary supervision lifecycle
* no-mistakes(document): Synchronize Pi primary extension documentation
* fix: keep persistent secondmates out of the main backlog (#398)
* fix secondmate backlog guidance
* no-mistakes(review): Require reasons for captain backlog holds
* no-mistakes(test): Document secondmate handoff skill requirement
* fix secondmate teardown reminder
* no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract
* fix(backlog-handoff): move full item blocks including indented bodies (#401)
* fix(backlog-handoff): move full item blocks including indented bodies
fm-backlog-handoff only moved the checklist header line, so multi-line
item bodies were left orphaned in the source backlog and never reached
the secondmate. Move the full block (header plus indented body lines)
atomically, treating body membership by indentation so lines like
## Intent stay with the item, and add regression coverage.
* no-mistakes(review): Captain: preserve EOF handoff terminators
* no-mistakes(review): treat blank lines inside item bodies as movable body
* no-mistakes(document): sync backlog-handoff docs with full-block move behavior
* feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402)
* guard Herdr lab lifecycle in briefs
* no-mistakes(review): Fix Herdr lab helper and provisioning safety
* no-mistakes(review): Captain, harden Herdr lab lifecycle safety
* no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range
* no-mistakes(review): reject leading options in Herdr lab run guard
* no-mistakes(review): strip leading non-alnum in Herdr lab name generator
* no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag
* no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab
* no-mistakes: apply CI fixes
* fix(watcher): classify arm-command seatbelt by execution position (#403)
* fix watcher arm command policy
* no-mistakes(review): Harden watcher command policy parsing
* no-mistakes(review): Captain: harden watcher policy parsing
* no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter
* no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps
* no-mistakes(review): fail closed on loop-wrapped broad watcher kills
* no-mistakes(document): sync docs for watcher-arm command-position policy
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", l…
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Aug 2, 2026
* feat(backends): add experimental cmux runtime backend (#246)
* feat(backends): add cmux runtime backend (experimental)
Session-provider-only adapter for cmux (bin/backends/cmux.sh), mirroring
zellij/herdr structurally, wired into fm-backend.sh and fm-spawn.sh with
--secondmate refused for now. Verified against the real cmux 0.64.17 app:
send does not auto-submit, cwd is creation-time-frozen (zellij-shape,
pwd-marker-probe workaround), close-surface refuses on a workspace's last
surface (falls back to close-workspace), workspace ids do not survive a
relaunch, and the control socket defaults to cmuxOnly access (requires a
one-time password-mode setup, documented in docs/cmux-backend.md). Also
found and fixed a live bug during development: read-screen fails on a
surface that has never been written to, so liveness now uses list-panes
instead. Fake-CLI unit suite (40 tests), a real-binary smoke test, and a
full spawn/steer/peek/done/merge/teardown E2E pass against a real claude
crewmate all pass, including the popup/second-Enter regression class.
* no-mistakes(review): Harden cmux recovery and password parsing
* no-mistakes(review): Harden cmux capture failure handling
* no-mistakes(review): Mark cmux test scripts executable
* no-mistakes(review): Scope cmux workspaces and teardown
* no-mistakes(review): Captain, honor cmux password config override
* no-mistakes(review): Captain, hash cmux home labels
* no-mistakes(document): Sync cmux backend docs
* feat(agents): add firstmate coding guidelines skill (#248)
* Add firstmate-coding-guidelines skill (AGENTS.md diet PR 0)
Encodes the knowledge-placement decision tree, one-owner rule, and
inline-stub pattern from the diet analysis so future contributions stop
adding conditional detail inline. AGENTS.md gets one section-13 trigger
line; fm-brief.sh's REPO argument has no reliable signal for "this is
firstmate's own repo", so the load instruction goes in CONTRIBUTING.md's
Development section instead of the scaffold.
* no-mistakes(review): Captain, align tracked-material trigger scope
* no-mistakes(document): Sync coding guidelines docs
* no-mistakes(lint): Fix Markdown style issues
* fix: add turn-end supervision guard (#249)
* feat: structural Stop-hook backstop for primary turn-end supervision
fm-guard.sh is pull-based: it only warns when some other supervision
script happens to run, so a primary session that ends a turn without
re-arming the watcher and then runs no further fleet-touching command
can sit blind for hours (the 2026-07-04 incident this fixes).
Add bin/fm-turnend-guard.sh, a Claude Code Stop hook registered in the
tracked .claude/settings.json, that fires on every primary turn end and
blocks (exit 2, verified empirically to force continuation) when work
is in flight with no fresh watcher beacon. It never blocks more than
once per turn, using Claude Code's own stop_hook_active loop-guard
field, and scopes itself to the actual primary checkout only (inert in
crewmate/scout worktrees and secondmate homes).
Factor the shared "in-flight but no live watcher" predicate out of
fm-guard.sh into bin/fm-supervision-lib.sh so the pull-based banner and
the push-based hook can never drift on what "unhealthy" means.
Document the verified Stop-hook mechanism and scoping in
docs/turnend-guard.md, add a harness-adapters note, and cover the
predicate and hook with tests/fm-turnend-guard.test.sh.
* no-mistakes(review): Respect active home in turnend guard
* no-mistakes(review): Require live watcher for turn-end guard
* no-mistakes(review): Captain: portable turn-end timing
* no-mistakes(document): Sync turn-end guard documentation
* feat(backends): auto-detect cmux runtime (#250)
* feat(backends): auto-detect cmux runtime from CMUX_WORKSPACE_ID
Wires cmux into fm_backend_detect the same way herdr already is: a
firstmate process running inside a cmux-spawned terminal now spawns
new tasks into cmux by default, no config needed. Verified from cmux's
own shipped source that CMUX_WORKSPACE_ID/CMUX_SURFACE_ID/CMUX_SOCKET_PATH
are unconditionally, non-overridably injected into every terminal
surface it spawns, and that cmux's own CLI treats CMUX_WORKSPACE_ID as
its own ambient-target fallback - the same role $TMUX/HERDR_ENV play for
their backends. CMUX_WORKSPACE_ID is checked last (after $TMUX and
HERDR_ENV=1) since cmux is a terminal application, not a nestable
multiplexer. Socket auth (config/cmux-socket-password) stays required
regardless of how the backend was selected; the existing spawn refusal
now also names the config/backend=tmux / --backend tmux opt-out for a
caller who never explicitly chose cmux.
A live env dump inside a real cmux terminal was not obtained safely on
the shared dev machine (documented in docs/cmux-backend.md); this rests
on the source read instead, mirroring this doc's existing
verified-from-source precedent.
* no-mistakes(review): Fix cmux autodetect docs and tests
* no-mistakes(document): Document cmux auto-detection
* fix(afk): support herdr away-mode injection (#251)
* fix(afk): make the away-mode daemon backend-aware for herdr
bin/fm-supervise-daemon.sh discovered its supervisor pane and injected
via raw tmux calls only, so /afk failed outright on a herdr-based
fleet (TMUX_PANE unset, firstmate:0 fallback unresolvable).
Discovery now resolves backend (tmux|herdr) and target independently,
mirroring fm-backend.sh's own runtime auto-detection, with an explicit
FM_SUPERVISOR_BACKEND override alongside the existing FM_SUPERVISOR_TARGET.
zellij/orca refuse loudly at startup instead of misapplying tmux
primitives. Injection (pane-exists probe, busy-guard, composer-guard,
verified submit) now dispatches through bin/fm-backend.sh's generic
primitives, adding a new fm_backend_composer_state dispatcher; the
tmux path is byte-identical to before. Also fixes a pre-existing bug
in fm_backend_target_exists's herdr arm (missing --session, so it
silently misrouted once more than one herdr server was running) found
while verifying this end to end against a real isolated herdr session.
Classification, batching, max-defer, the marker contract, locks, and
wake-queue handling are unchanged - this is a transport-layer fix.
* no-mistakes(review): Corroborate Herdr idle busy state
* no-mistakes(review): Stabilize Herdr daemon startup wait
* no-mistakes(review): Captain, route cmux composer and update AFK docs
* no-mistakes(document): Document AFK supervisor backend support
* docs(agents): move X-mode procedures out of AGENTS (#253)
* docs(agents): collapse X-mode section 14 into fmx-respond/docs pointers
AGENTS.md diet PR 1 of 3 (agentsmd-diet-s2 report, move-plan items 1-2).
Replaces section 14's "Answering"/"Completion follow-up"/"Conversations"/
"Length and threads"/"Preview / dry-run" blocks (54 lines) and the
"Mechanism" narrative (6 lines) with two short pointers: fmx-respond
(section 13) for the procedure, docs/configuration.md "X mode (.env)"
for the wire protocol. Net -55 lines in AGENTS.md.
Destination edits landed first, deletions second (q4 discipline):
- docs/configuration.md: added the "purely additive, watcher untouched"
guarantee that AGENTS.md's Mechanism block stated but configuration.md
did not.
- fmx-respond/SKILL.md: added the x-mode-error wake boundary (report as
a blocker, do not load this skill), the --image flag for replies and
follow-ups, the "images are for real artifacts, not prose" rule, and
the dry-run compact-image-marker behavior - none of these were
previously in the skill even though AGENTS.md described them, so they
were genuine gaps, not pre-existing duplication. Also made the skill's
own "Completion follow-up" section the sole, full owner of that
procedure instead of deferring to AGENTS.md section 14 for substance
that no longer lives there (two internal cross-references updated to
point at section 8's terminal-wake trigger and the skill's own section
instead).
Mechanical line-by-line audit of every removed AGENTS.md line:
Mechanism block (6 lines removed):
- bootstrap artifact-writing description -> already owned by
docs/configuration.md "X mode (.env)" (locked-bootstrap paragraph)
- check-shim/poll mechanism description -> already owned by
docs/configuration.md same section
- missing-deps/x-mode-error diagnostic description -> already owned by
docs/configuration.md ("Relay auth or config problems...") plus
bin/fm-x-poll.sh's own header comment for the missing-curl/jq mechanics
- opt-out artifact removal description -> already owned by
docs/configuration.md same section
- "purely additive, no edit to fm-watch.sh/fm-watch-arm.sh/fm-wake-lib.sh/
afk daemon" guarantee -> MOVED to docs/configuration.md (added in this
PR; this fact had no other home before)
Answering/Completion follow-up/Conversations/Length and threads/
Preview-dry-run blocks (54 lines removed):
- x-mention wake -> load fmx-respond: already owned by section 13's
existing trigger line (unchanged) and restated in the new pointer
- x-mode-error wake -> report as blocker, don't load fmx-respond: MOVED
to fmx-respond/SKILL.md (added in this PR)
- inbox-draining, classification, acting, reply composition, submission,
cleanup-on-success/failure: already owned by fmx-respond/SKILL.md
"Procedure" section (unchanged, pre-existing)
- owner-only routing / captain-as-asker framing: already owned by
fmx-respond/SKILL.md "The asker is your own captain" section
- standing X-mode authorization / autonomous posting / dry-run as only
non-posting path: already owned by fmx-respond/SKILL.md same section
- acknowledge-first -> act -> follow-up shape, three-case classification:
already owned by fmx-respond/SKILL.md "A request to act on" section
- destructive/irreversible/security-sensitive escalation guardrail:
already owned by fmx-respond/SKILL.md "Public channel..." section and
Procedure step 2c
- dismiss-instead-of-reply for pure acknowledgments, relay re-offer
prevention, dry-run honoring: already owned by fmx-respond/SKILL.md
Procedure steps 2b/2c/2e-skip and docs/configuration.md
- public-safety bar (no task ids/internals/captain-private/secrets):
already owned by fmx-respond/SKILL.md "The reply is public" section
- never-inline-into-shell-command / --text-file or stdin: already owned
by fmx-respond/SKILL.md Procedure step 2e and Notes
- --image flag for replies (formats, base64, no-inline guarantee): MOVED
to fmx-respond/SKILL.md Procedure step 2e (added in this PR - this was
not previously in the skill)
- fm-x-link field names (x_request=, x_request_ts=, x_followups=):
already owned by AGENTS.md section 2's state/<id>.meta field list
(untouched, out of scope for this PR) and fmx-respond/SKILL.md
- carry-count/carry-ts relink behavior, three-follow-up budget, milestone
sparingness, --check/--text-file posting, connector/followup wire
detail, --final clearing, cap/window graceful degradation: already
owned by fmx-respond/SKILL.md "Completion follow-up" section (now sole
owner) and docs/configuration.md wire-protocol paragraphs
- --image flag for follow-ups: MOVED to fmx-respond/SKILL.md "Completion
follow-up" section (added in this PR - genuine gap)
- "failed task still gets an honest final follow-up": already owned by
fmx-respond/SKILL.md "Completion follow-up" section
- FMX_DRY_RUN whole-loop previewability: already owned by
fmx-respond/SKILL.md "Dry-run / preview mode" section
- in_reply_to conversation continuity, untrusted-thread handling,
follow-up worthiness judgment, relay-owned self-reply guard/cap:
already owned by fmx-respond/SKILL.md "The direct ask is the captain's"
section and Notes (one bullet is a verbatim match)
- concise-by-default / no hand-numbered threads: already owned by
fmx-respond/SKILL.md "Voice" section
- auto-split behavior, char/tweet caps, premium-independence, wire shape
({text}/{text,texts}): behavior already owned by fmx-respond/SKILL.md
Voice section; exact defaults and wire shape already owned by
docs/configuration.md; "premium-independent" mechanics already owned
by bin/fm-x-reply.sh's own header comment
- "images are for real artifacts, not prose": MOVED to fmx-respond/
SKILL.md "Voice" section (added in this PR - genuine gap)
- image-on-thread wire behavior: already owned by docs/configuration.md;
reinforced in fmx-respond/SKILL.md's new --image note
- dry-run POST-body shape, endpoint marker, truthy-value definition,
jq-only dependency, end-to-end testability, x-outbox inspection:
already owned by fmx-respond/SKILL.md "Dry-run / preview mode" section
(several near-verbatim matches) and docs/configuration.md wire detail
- dry-run compact image marker: MOVED to fmx-respond/SKILL.md "Dry-run /
preview mode" section (added in this PR - genuine gap)
Section 8's terminal-wake completion-follow-up trigger (the one fact
required to survive inline) is untouched and already present; the new
section 14 pointer references it instead of restating it.
Nothing outside section 14 (plus the two destination files) is touched.
Full test suite green, including all 74 fm-x-mode.test.sh checks.
* no-mistakes(review): Preserve X-linked follow-up triggers
* no-mistakes(review): Fix x-mode error trigger
* no-mistakes(document): Docs cross-reference synchronized
* no-mistakes(lint): Clean Markdown lint pass
* docs: trim duplicated harness guidance (#255)
* docs(agents): trim section 4 harness/secondmate duplication
AGENTS.md diet PR 2 of 3 (data/agentsmd-diet-s2/report.md, move-plan
items 3-4; redundancy item 2 folded into item 3).
Removed the five claude/codex/grok/pi/opencode model/effort-flag
bullets from section 4 - byte-for-byte duplicated by
harness-adapters' "Launch profile axes" table (which is already a
superset: it carries verified CLI versions per adapter that the
AGENTS.md bullets lacked). Replaced with a one-line pointer; the
skill is already loaded before every spawn per section 4's own
closing trigger, so no new trigger was needed.
Moved the config/secondmate-harness model/effort pin-format detail
(the `<harness> [<model>] [<effort>]` line format, the
secondmate-model/secondmate-effort accessors, back-compat, and the
durability-across-respawn behavior) into secondmate-provisioning,
which is already a mandatory load at every secondmate lifecycle
touchpoint. Added the destination content to the skill first, then
replaced the AGENTS.md paragraph with a 3-line pointer.
Mechanical audit - every removed line's new home:
- 5 harness bullets (claude/codex/grok/pi/opencode model+effort
flags, per-harness max-omission rationale) -> already present in
harness-adapters SKILL.md's "Launch profile axes" table (lines
53-59), confirmed fact-by-fact before deleting.
- "config/secondmate-harness may also pin..." paragraph (pin format,
bare-harness back-compat, secondmate-model/secondmate-effort
accessors, per-spawn override precedence, respawn durability,
secondmate-only scope) -> secondmate-provisioning SKILL.md's
"Charter and seed" section, added verbatim before this trim.
- The following paragraph (inheritable config: crew-dispatch.json,
crew-harness, backlog-backend) is untouched - out of scope for
this PR, still inline.
- The bootstrap CREW_DISPATCH effort-mismatch diagnostic sentence is
untouched - not part of the five-bullet duplication, stays inline.
No script changes. Section 4 shrinks from 104 to 87 lines
(958 -> 901 total AGENTS.md lines) with zero facts lost: every fact
is reachable through harness-adapters or secondmate-provisioning,
both already mandatory loads at the relevant lifecycle points.
* no-mistakes(document): Align secondmate skill triggers
* no-mistakes(lint): Markdown style clean
* fix: anchor turn-end Stop hook to project root (#256)
* Fix turn-end Stop hook to use CLAUDE_PROJECT_DIR path
Claude Code runs hook commands via /bin/sh from the session cwd, so the
bare relative bin/fm-turnend-guard.sh path fails when cwd is not the repo
root. Anchor the command with "$CLAUDE_PROJECT_DIR"/bin/fm-turnend-guard.sh
instead; verified CLAUDE_PROJECT_DIR is set on Stop hooks in Claude Code
2.1.201. Document the cwd caveat and add a settings.json regression test.
* no-mistakes(document): Document Stop hook path anchoring
* docs: trim firstmate agent guidance duplication (#258)
* docs: trim AGENTS.md redundancy (diet PR 3/3)
Consolidates five duplicated passages to a single owner each, per
data/agentsmd-diet-s2/report.md redundancy items c3-c7:
- Inheritable-config propagation mechanism: owned by section 3 (where
the sweep runs); sections 4 and 7 keep compact references. Section 4
retains its one genuinely unique fact (crew-harness inherit-vs-fallback
semantics), just no longer restates the propagation mechanism itself.
- Landed-work definition: owned by section 7's ship-teardown detail
(PR-containment mechanics, pr= discovery fallback); section 1's hard
rule #3 keeps the rule plus a three-case summary and a pointer.
- Backend meta-field enumeration: owned by docs/configuration.md
("Runtime backend", already comprehensive including cmux) and each
backend's own doc; AGENTS.md keeps only the fields common to every
task plus a pointer.
- Dropped one redundant restatement of "silence is correct while
waiting" in section 8.
- Worktree-tangle guard explanation: owned by section 8 (already the
fuller, cross-referenced version); section 3's TANGLE bullet keeps
the remediation action and points at section 8 for the why.
Also adds two captain-requested single-sentence rules: invoke bin/
scripts by absolute $FM_ROOT path after any cd away from the home, and
a backend spawn refusal must be surfaced to the captain rather than
silently worked around by switching backends.
AGENTS.md: 901 -> 889 lines, 112355 -> 108560 bytes.
* no-mistakes(review): Clarify post-cd bin invocation guidance
* no-mistakes(document): Sync AGENTS trim docs
* no-mistakes(lint): Fix Markdown line style
* feat(backends): improve cmux detection and socket-mode guidance (#259)
* feat(backends): cmux detection fallbacks and socket-mode matrix
Workstream A: cmux's bundled claude wrapper strips every CMUX_* env var on
its passthrough path (reproduced live 2026-07-04, cmux 0.64.17), so a
claude-harness firstmate inside a cmux tab has no CMUX_WORKSPACE_ID.
fm_backend_detect now falls back - macOS-only, only when the primary marker
is absent - to __CFBundleIdentifier=com.cmuxterm.app and then a process
ancestry walk resolved by bundle id (lsappinfo) plus a bundle-shaped ps comm
match. Innermost-first ordering is unchanged and absorbs the
tmux-inside-cmux bundle-id false positive; the auto-detect NOTICE names the
winning fallback signal.
Workstream B: the five socketControlMode values were traced through cmux
source (commit 9c91710e3f58): off/cmuxOnly can never admit an external CLI,
automation admits same-user clients with no secret (0600 socket only),
password needs the auth handshake, allowAll opens the socket to every local
user (0666). Automation mode is now the documented recommendation; the
adapter's refusals name every viable mode, classify Invalid password as
unauth, and the launch-timeout message names the off-mode possibility.
Docs carry the wrapper-strip empirical record, the fallback contract and
authority split, and the full mode matrix with rationale; tests cover the
new detection paths, the nested false positive, and the refusal wording.
* no-mistakes(review): Document cmux fallback detection
* no-mistakes(review): Update cmux architecture docs
* no-mistakes(document): Align cmux backend docs
* fix(backends): scope zellij tabs by firstmate home (#252)
* fix(backends): home-scope zellij tab titles to close cross-home collision gap
Zellij's one shared "firstmate" session has no per-home split and enforces
no tab-name uniqueness, so two firstmate homes with colliding task ids could
send/peek/close each other's tabs - the same gap a no-mistakes review gate
caught for cmux (docs/cmux-backend.md). Ports that fix: every new tab is
created with a home-scoped title (fm-<home-label>-<id>), and every
list/find/recover/kill path scopes matches to this home's own tag. A tab
spawned before this change still matches via its old untagged bare title,
but only when unambiguous - two live tabs sharing a bare title refuse rather
than guessing which one is ours.
Factors the home-label/hash derivation shared with cmux into
bin/fm-backend-hometag-lib.sh so the two adapters can't drift.
* no-mistakes(review): Fix zellij child teardown home tag
* no-mistakes(review): Fix zellij teardown and selector scoping
* no-mistakes(document): Sync zellij home-scope docs
* fix: sync project clones after merged PR wakes (#293)
* fix(fleet-sync): auto-sync on merged-PR wake, accept project name
fm-fleet-sync.sh's single-project form failed on a bare project name
("not a directory"), forcing hand-typed full paths (4 manual runs in
one incident). It now resolves a bare name or projects/<name> against
the home's projects dir.
AGENTS.md now encodes the trigger: a wake whose status reports a
merged PR for a project cloned in this home runs fleet-sync for that
project as part of handling the wake, so a secondmate-reported merge
does not leave the primary's clone stale until the next session start
or teardown.
* no-mistakes(review): Fix fleet-sync project name shadowing
* no-mistakes(document): sync fleet-sync docs
* fix: canonicalize spawn worktree path checks (#294)
* fix(spawn): canonicalize worktree-isolation guard against symlinked project prefixes
fm-spawn.sh compared a logical PROJ_ABS against the physically-resolved
pane cwd every backend reports, so a project reached through a symlinked
prefix (e.g. macOS's /tmp -> /private/tmp) could trip the isolation
guard's false refusal before treehouse ever moved the pane. Canonicalize
once into PROJ_ABS_REAL and compare against that everywhere instead.
* no-mistakes(review): Canonicalize spawn cwd comparisons
* no-mistakes(document): Refresh symlinked spawn docs
* docs: add Orca operator skill (#276)
* docs: add Orca operator skill
* no-mistakes(document): Document Orca checklist
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* fix: surface green PRs during CI monitoring (#297)
* fix(crew-state): detect green-PR CI monitoring, escalate repeat wedges
fm-crew-state.sh's ci step never distinguishes "still waiting on checks"
from "checks green, waiting on merge" via axi status alone, since a repo
that defers merge to the captain keeps the ci step at status=running for
the whole monitor phase. Read the ci step's own log tail (axi logs) for
the checks-passed marker and surface done instead of a false "validating
(running)" - verified against the real PR #252 run's ci.log.
The watcher's wedge timer can re-escalate the same stale pane forever
without ever signaling that it is a repeat; track a per-pane consecutive
escalation count and add a demand-deep-inspection marker to the wake
payload once it crosses a threshold, so the supervisor can no longer
dismiss each one as an isolated, still-validating pane.
Also clarify the ship-brief's checks-green line: it is owed at the
CI-ready return point, not after the background monitor-until-merge
loop finishes.
* no-mistakes(review): Captain, distinguish pending no-checks CI marker
* no-mistakes(review): Harden CI relapse handling
* no-mistakes(review): Block stale done during fixing
* no-mistakes(review): Captain, tighten CI status gating
* no-mistakes(review): Captain, harden stale CI green handling
* no-mistakes(review): Captain, recognize ranged CI rearm markers
* no-mistakes(document): Sync crew-state supervision docs
* fix(teardown): recover provably stale git index locks (#296)
* fix(teardown): recover from a stale worktree git index.lock
A crew process killed mid-git-operation can leave a stale
.git/worktrees/<wt>/index.lock behind, making fm-teardown.sh's
`treehouse return --force` fail closed. On that failure, retry once
after a short wait (the owning process may be exiting), then remove
the lock and retry once more only when it is provably stale: old
enough by mtime and lsof shows no live holder on the lock or the
worktree itself. A lock that isn't provably stale is left in place and
the original failure still surfaces.
* no-mistakes(review): Harden teardown lock refusal paths
* no-mistakes(review): Harden stale-lock teardown safety rechecks
* no-mistakes(review): Harden stale teardown lock checks
* no-mistakes(document): Document teardown lock recovery
* feat(bin): encode project AGENTS.md authoring bar with canonical self-governance section (#307)
* Encode project AGENTS authoring bar
* no-mistakes(review): Captain, centralize CLAUDE promotion governance
* no-mistakes(review): make ensure_maintenance_section idempotent-success, drop || true guards
* no-mistakes(review): separate appended maintenance section on newline-less CLAUDE.md promotion
* no-mistakes(review): assert maintenance heading present before separator check in test
* no-mistakes(document): sync docs with AGENTS.md authoring bar and self-governance
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* feat(skills): add captain-invocable bearings status-report skill (#300)
* Add captain-invocable bearings skill
Generates a pick-up-where-I-left-off status report from live fleet
state to data/status-report-<YYYY-MM-DD>.md plus a concise chat
summary. Read-mostly procedure: reads backlog, per-task crew state
via bin/fm-crew-state.sh, open PRs via gh-axi, scout reports,
pending decisions, and date-gated queued work; composes the
exemplar's sections (TL;DR, Check first, Landed, In flight, Plans,
Decisions pending, Date-gated/queued); never tears down, merges, or
mutates task state as a side effect.
* no-mistakes(document): docs: list new /bearings skill in README built-in skills table
* fix(watcher): make PID identity locale-invariant (#285)
* fix(watcher): pin LC_ALL=C in fm_pid_identity for locale-invariant identity
ps's lstart date format follows the caller's LC_TIME/LC_ALL. The watcher records
its process identity under one locale, but arm/guard/turn-end re-read it under the
machine's ambient locale. On a non-C locale (e.g. ko_KR) the two strings differ
only in the date portion, so fm_watcher_lock_matches_pid / fm_watcher_healthy
reject a genuinely live watcher - breaking fm-watch-arm.sh, fm-guard.sh, and
fm-turnend-guard.sh on every non-C-locale machine.
Pin LC_ALL=C on that one ps call so the write and read sides agree regardless of
machine locale, matching the LC_ALL=C determinism the file already uses elsewhere.
Add a colocated regression test asserting fm_pid_identity is locale-invariant
across exported LC_ALL/LC_TIME.
* no-mistakes(document): Document watcher PID identity coverage
* docs: document codex app backend contract (#222)
* docs: reconcile Codex App backend contract
* no-mistakes(document): Sync backend docs
* docs: clarify Codex Desktop bridge blocker
* no-mistakes(document): Align Codex App backend docs
* no-mistakes(test): Captain, stabilize watcher self-eviction test cadence
* no-mistakes(document): Document Codex App backend contract
* no-mistakes(document): Captain, document blocked codex-app coverage
* docs: make Codex App contract doc authoritative
* no-mistakes(document): Align Codex App backend docs
* docs: redact local Codex App smoke paths
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* docs: add Codex Desktop coordination skill (#275)
* docs: add Codex App coordination skill
* no-mistakes(review): Captain, mark Codex App skill agent-only
* no-mistakes(document): Document Codex App backend boundary
* no-mistakes(document): Captain, document Codex Desktop backend boundary
* no-mistakes(lint): Captain, lint clean
* no-mistakes(document): Document Codex Desktop boundaries
* docs: narrow Codex App skill playbook
* fix(afk): stop herdr escalation redelivery loop (#317)
* fix(afk): recognize unbordered herdr composer rows to stop escalation redelivery loop
fm_backend_herdr_composer_state only recognized bordered composer rows
(the grok shape). Real claude and codex render their live input row
with no border at all, so once a harness's own startup banner scrolled
out of the capture window the classifier read the composer as unknown
forever. fm_backend_herdr_send_text_submit never confirmed "empty", so
escalate_flush never cleared state/.subsuper-escalations, and the
away-mode daemon retyped and resubmitted the same buffered digest every
housekeeping cycle - reproduced live against a real herdr+claude pane
(5+ identical deliveries in 40s).
The classifier now recognizes an unbordered (bare) composer row led by
a known prompt glyph alongside the existing bordered shape, keeping
whichever match is bottom-most so a stale decorative box never
outranks the live composer.
* no-mistakes(review): Narrow herdr bare prompt matcher
* no-mistakes(document): Sync herdr composer docs
* fix(backends): confirm Herdr submits with native agent state (#323)
* fix(herdr): confirm message submit via native agent-state, not composer text
fm_backend_herdr_send_text_submit now confirms a landed submit by polling
herdr's own agent-state (agent get) for the idle->working transition instead
of reading composer content. Composer scraping remains, unchanged, for the
away-mode daemon's pre-injection empty-box guard only.
This fixes the practical effect of the codex idle-tip gap from the
2026-07-07 incident: codex's dynamic idle-composer hint text can no longer
misread as pending and block/mis-confirm a send, since confirmation no
longer looks at composer text at all. Verified empirically against real
claude and codex agents (timing, swallowed-Enter, unreadable-target, and
already-busy-target scenarios), and against the real away-mode daemon
end-to-end after updating its synthetic supervisor-pane test fixture to
register itself as a real herdr agent (herdr's own report-agent primitive)
so it can still exercise the new confirmation path.
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(document): Sync Herdr submit docs
* no-mistakes: apply CI fixes
* feat: add quota-balanced crew dispatch selection (#327)
* Add quota-balanced dispatch selection
* no-mistakes(document): Document dispatch selector guidance
* fix(session-start): respawn dead secondmate agents conservatively
* fix(session-start): deterministically respawn dead-shell secondmates
A secondmate agent that exits leaves its backend pane alive as a bare
shell. The session-start endpoint check only verified pane presence, so
recovery and the watcher (which exempts secondmates from stale-pane
detection) never noticed - evidence 2026-07-07: every secondmate in one
fleet was found sitting at a dead zsh shell.
Add fm_backend_agent_alive (bin/fm-backend.sh), a deeper per-backend
liveness probe distinct from pane presence: fm_backend_tmux_agent_alive
classifies the pane's live foreground process via tmux's own
pane_current_command, and fm_backend_herdr_agent_alive reuses the
already-verified pane_agent_state husk classifier. Both are conservative:
anything ambiguous reports unknown, never a false dead.
Wire this into a new session-start-only, locked-and-primary-only sweep in
bin/fm-bootstrap.sh that kills and respawns only a confidently dead
secondmate endpoint, leaving alive/unknown readings untouched - idempotent
by construction, so repeated runs converge without duplicating agents.
* no-mistakes(review): Guard raw secondmate liveness respawns
* no-mistakes(review): Fix detect-only bootstrap test
* no-mistakes(test): Pin liveness fixture harness
* no-mistakes(document): Sync secondmate liveness docs
* no-mistakes: apply CI fixes
* fix: emit stable secondmate nudge selectors (#331)
* Fix NUDGE_SECONDMATES to print stable fm-<id> selectors.
Session-start secondmate sync used to accumulate raw backend window targets
into NUDGE_SECONDMATES, but the liveness sweep in the same bootstrap run can
respawn secondmates onto new endpoints. fm-send with those stale explicit
targets bypasses meta resolution and fails, while fm-<id> resolves correctly.
Accumulate fm-<id> in process_secondmate, update the bootstrap/update contracts
and /updatefirstmate skill, and add a herdr respawn regression test.
* no-mistakes(review): Captain, guard herdr regression jq dependency
* no-mistakes(document): Document stable secondmate nudge selectors
* no-mistakes(lint): Fix shell lint hints
* feat: require bootstrap detection for AXI tools (#332)
* Make tasks-axi and quota-axi required bootstrap tools
Add both to the normal toolchain checks alongside lavish-axi, keep the
tasks-axi 0.1.1+ compatibility gate, and report quota-axi through the
standard MISSING install-consent flow. TASKS_AXI: available remains a
backlog-backend capability signal only; manual opt-out no longer suppresses
the missing-tool report.
Update bootstrap tests and point docs/configuration.md at the canonical
toolchain contract.
* no-mistakes(review): Clarify manual backlog bootstrap reporting
* no-mistakes(document): Document bootstrap AXI tools
* bearings: delete today's report before recreating (#333)
Replace overwrite-in-place wording with explicit delete-then-create
instructions so agents do not modify an existing daily report file.
* feat: guard primary turn ends across harnesses (#339)
* Add primary turn-end guards for all harnesses
* no-mistakes(review): Normalize Codex hook cwd resolution
* no-mistakes(review): Fix OpenCode guard worktree anchoring
* no-mistakes(review): Anchor Codex guard outside nested roots
* no-mistakes(review): Anchor Codex guard to hook root
* no-mistakes(review): Avoid Grok permission escalation
* no-mistakes(document): Sync turn-end guard docs
* fix: resolve backend selectors by exact task id first (#342)
* fix backend selector task id resolution
* no-mistakes(document): Document selector resolution behavior
* fix: scale bootstrap fleet-sync timeout (#341)
* fix bootstrap fleet sync timeout
* no-mistakes(review): Fix bootstrap fleet-sync timeout regressions
* no-mistakes(document): Sync bootstrap timeout docs
* no-mistakes(lint): Clean ShellCheck directives
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add fleet snapshot and view commands (#343)
* Add fleet snapshot and view
* no-mistakes(review): Fix fleet snapshot parsing and overrides
* no-mistakes(review): Fix secondmate fleet rendering
* no-mistakes(review): Fix backlog title and completion parsing
* no-mistakes(review): Include durable scout reports
* no-mistakes(review): Fix fleet snapshot edge cases
* no-mistakes(review): Captain: gate fleet hints on current state
* no-mistakes(review): Captain: parse bracketed Done PR artifacts
* no-mistakes(document): Sync fleet snapshot docs
* fix(fm-send): fail loudly on unresolvable send targets (#254)
* Make fm-send fail loudly on unresolved targets
* no-mistakes(review): Document fm-send FM_HOME contract
* Fix fm-send readiness docs and backend send path
* Fix fm-send docs for cmux and X skill metadata
* Make gotmp teardown test home-explicit
* Scope watcher warning wording to fm-send
* Fix fm-send review findings
* Verify explicit tmux targets before sending
* Isolate turnend guard test home
* no-mistakes(document): Documented fm-send FM_HOME/backend guard additions missing from doc inventories
---------
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* fix: deliver AFK escalations through herdr supervisors (#353)
* fix afk codex ghost composer delivery
* no-mistakes(review): Harden AFK startup flag writes
* no-mistakes(review): Harden AFK daemon liveness checks
* no-mistakes(document): Sync AFK herdr docs
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add harness-aware supervision (#367)
* Add harness-aware supervision
* no-mistakes(review): Captain, harden watcher supervision regressions
* no-mistakes(review): Captain, harden watcher supervision cadence
* no-mistakes(review): Harden watcher supervision ownership
* no-mistakes(review): Captain, harden Pi extension marker
* no-mistakes(review): Captain, harden Pi supervision restart checks
* no-mistakes(review): Harden watcher ownership checks
* no-mistakes(review): Captain, harden Pi supervision loading
* no-mistakes(review): Captain, require Pi guard extension loading
* no-mistakes(review): Captain, harden watcher supervision recovery
* no-mistakes(test): Fix fm-send baseline log filtering
* no-mistakes(document): Sync harness supervision docs
* no-mistakes: apply CI fixes
* fix: split X-mode replies by platform (#369)
* fix: make x replies split by platform
* no-mistakes(review): Captain: preserve Discord recovery relink context
* no-mistakes(test): Captain: keep split markers outside fences
* no-mistakes(document): Sync X-mode reply docs
* fix: make stow memory writes inspect before update (#372)
* docs: make stow inspect-then-update
* no-mistakes(review): Remove unsupported archive-body guidance
* no-mistakes(review): Clarify stow read-before-write exception
* no-mistakes(test): Require archive-body for stow task notes
* no-mistakes(document): Sync stow memory docs
* no-mistakes(lint): Silence ShellCheck source warning
* fix(watcher): wait when arm attaches to a healthy watcher (#375)
* fix: attach-and-wait when arm finds a healthy watcher
Grok and Claude re-arm after every turn with work in flight. When a
watcher was already healthy, fm-watch-arm exited immediately with
watcher: healthy, which completed the harness background task and
injected an empty false wake.
Attach to the live identity-matched holder instead, stay until that
cycle ends, then exit 0 so notify fires for a real end-of-cycle. The
peer-startup-race path uses the same contract. --restart and the
started path are unchanged.
* no-mistakes(review): Gate restart watcher peer attach
* no-mistakes(document): Sync watcher arm docs
* feat(pi): simplify primary session launch (#386)
* docs(readme): reformat Quick Start and recommend Grok equally with Claude Code
* no-mistakes(review): Captain: align harness launch guidance
* no-mistakes(review): Captain, clarify Pi supervised launch
* no-mistakes(review): Captain, document Pi first-launch bridge
* feat(pi): track primary watcher extension for plain-pi launch
Move Pi's primary watcher bridge from a generated state/ file to a
tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end
guard extension already works: self-hashing version, project-local
auto-discovery after one-time Pi trust. This drops the state/-generation
step and dual -e requirement from the happy path, so Pi's Quick Start
launch becomes plain 'pi', the same friction class as 'claude' and
'grok --trust'.
- bin/fm-pi-watch-extension.sh is removed; nothing generates the
extension anymore since it is committed.
- fm-session-start.sh and fm-supervision-instructions.sh resolve the
watcher extension path from FM_ROOT instead of state/, and the
session-start diagnostic now points at restarting plain pi after
trust, with -e as a documented fallback.
- fm-spawn.sh points Pi secondmate launches at the tracked extension
path in the secondmate home instead of generating a state/ copy.
- README Quick Start Pi block is now just 'pi' plus a trust note.
- Tests, docs, and the harness-adapters skill updated to match.
* fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016
* feat(supervision): prevent unsafe watcher-arm commands (#387)
* feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns
Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that
denies a primary shell command backgrounding, piping, or bundling the
watcher arm/checkpoint, or force-killing the watcher process broadly -
the exact shapes that silently took Grok's supervision down. Wires it
into all five verified harnesses (grok, claude, codex, opencode, pi),
each validated empirically against the real harness.
Also fixes a grok 0.2.93 regression discovered during that validation: the
existing turnend-guard Stop hook's bare root variable broke grok's own
variable pre-substitution and silently no-op'd the hook.
* no-mistakes(review): Harden watcher arm validation
* no-mistakes(review): Harden arm guard metacharacter checks
* no-mistakes(review): Harden nested shell arm guard
* fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then
A && B || C is not if-then-else; C can run when A is true. Replace both
occurrences of the quote-state early-continue with an explicit if/then.
* fix(pi): restore primary watcher supervision lifecycle (#397)
* fix Pi primary supervision lifecycle
* no-mistakes(document): Synchronize Pi primary extension documentation
* fix: keep persistent secondmates out of the main backlog (#398)
* fix secondmate backlog guidance
* no-mistakes(review): Require reasons for captain backlog holds
* no-mistakes(test): Document secondmate handoff skill requirement
* fix secondmate teardown reminder
* no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract
* fix(backlog-handoff): move full item blocks including indented bodies (#401)
* fix(backlog-handoff): move full item blocks including indented bodies
fm-backlog-handoff only moved the checklist header line, so multi-line
item bodies were left orphaned in the source backlog and never reached
the secondmate. Move the full block (header plus indented body lines)
atomically, treating body membership by indentation so lines like
## Intent stay with the item, and add regression coverage.
* no-mistakes(review): Captain: preserve EOF handoff terminators
* no-mistakes(review): treat blank lines inside item bodies as movable body
* no-mistakes(document): sync backlog-handoff docs with full-block move behavior
* feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402)
* guard Herdr lab lifecycle in briefs
* no-mistakes(review): Fix Herdr lab helper and provisioning safety
* no-mistakes(review): Captain, harden Herdr lab lifecycle safety
* no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range
* no-mistakes(review): reject leading options in Herdr lab run guard
* no-mistakes(review): strip leading non-alnum in Herdr lab name generator
* no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag
* no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab
* no-mistakes: apply CI fixes
* fix(watcher): classify arm-command seatbelt by execution position (#403)
* fix watcher arm command policy
* no-mistakes(review): Harden watcher command policy parsing
* no-mistakes(review): Captain: harden watcher policy parsing
* no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter
* no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps
* no-mistakes(review): fail closed on loop-wrapped broad watcher kills
* no-mistakes(document): sync docs for watcher-arm command-position policy
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File ex…
This was referenced Aug 2, 2026
Merged
vipentti
pushed a commit
to vipentti/firstmate
that referenced
this pull request
Aug 5, 2026
* docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Aug 9, 2026
* docs: trim firstmate agent guidance duplication (#258)
* docs: trim AGENTS.md redundancy (diet PR 3/3)
Consolidates five duplicated passages to a single owner each, per
data/agentsmd-diet-s2/report.md redundancy items c3-c7:
- Inheritable-config propagation mechanism: owned by section 3 (where
the sweep runs); sections 4 and 7 keep compact references. Section 4
retains its one genuinely unique fact (crew-harness inherit-vs-fallback
semantics), just no longer restates the propagation mechanism itself.
- Landed-work definition: owned by section 7's ship-teardown detail
(PR-containment mechanics, pr= discovery fallback); section 1's hard
rule #3 keeps the rule plus a three-case summary and a pointer.
- Backend meta-field enumeration: owned by docs/configuration.md
("Runtime backend", already comprehensive including cmux) and each
backend's own doc; AGENTS.md keeps only the fields common to every
task plus a pointer.
- Dropped one redundant restatement of "silence is correct while
waiting" in section 8.
- Worktree-tangle guard explanation: owned by section 8 (already the
fuller, cross-referenced version); section 3's TANGLE bullet keeps
the remediation action and points at section 8 for the why.
Also adds two captain-requested single-sentence rules: invoke bin/
scripts by absolute $FM_ROOT path after any cd away from the home, and
a backend spawn refusal must be surfaced to the captain rather than
silently worked around by switching backends.
AGENTS.md: 901 -> 889 lines, 112355 -> 108560 bytes.
* no-mistakes(review): Clarify post-cd bin invocation guidance
* no-mistakes(document): Sync AGENTS trim docs
* no-mistakes(lint): Fix Markdown line style
* feat(backends): improve cmux detection and socket-mode guidance (#259)
* feat(backends): cmux detection fallbacks and socket-mode matrix
Workstream A: cmux's bundled claude wrapper strips every CMUX_* env var on
its passthrough path (reproduced live 2026-07-04, cmux 0.64.17), so a
claude-harness firstmate inside a cmux tab has no CMUX_WORKSPACE_ID.
fm_backend_detect now falls back - macOS-only, only when the primary marker
is absent - to __CFBundleIdentifier=com.cmuxterm.app and then a process
ancestry walk resolved by bundle id (lsappinfo) plus a bundle-shaped ps comm
match. Innermost-first ordering is unchanged and absorbs the
tmux-inside-cmux bundle-id false positive; the auto-detect NOTICE names the
winning fallback signal.
Workstream B: the five socketControlMode values were traced through cmux
source (commit 9c91710e3f58): off/cmuxOnly can never admit an external CLI,
automation admits same-user clients with no secret (0600 socket only),
password needs the auth handshake, allowAll opens the socket to every local
user (0666). Automation mode is now the documented recommendation; the
adapter's refusals name every viable mode, classify Invalid password as
unauth, and the launch-timeout message names the off-mode possibility.
Docs carry the wrapper-strip empirical record, the fallback contract and
authority split, and the full mode matrix with rationale; tests cover the
new detection paths, the nested false positive, and the refusal wording.
* no-mistakes(review): Document cmux fallback detection
* no-mistakes(review): Update cmux architecture docs
* no-mistakes(document): Align cmux backend docs
* fix(backends): scope zellij tabs by firstmate home (#252)
* fix(backends): home-scope zellij tab titles to close cross-home collision gap
Zellij's one shared "firstmate" session has no per-home split and enforces
no tab-name uniqueness, so two firstmate homes with colliding task ids could
send/peek/close each other's tabs - the same gap a no-mistakes review gate
caught for cmux (docs/cmux-backend.md). Ports that fix: every new tab is
created with a home-scoped title (fm-<home-label>-<id>), and every
list/find/recover/kill path scopes matches to this home's own tag. A tab
spawned before this change still matches via its old untagged bare title,
but only when unambiguous - two live tabs sharing a bare title refuse rather
than guessing which one is ours.
Factors the home-label/hash derivation shared with cmux into
bin/fm-backend-hometag-lib.sh so the two adapters can't drift.
* no-mistakes(review): Fix zellij child teardown home tag
* no-mistakes(review): Fix zellij teardown and selector scoping
* no-mistakes(document): Sync zellij home-scope docs
* fix: sync project clones after merged PR wakes (#293)
* fix(fleet-sync): auto-sync on merged-PR wake, accept project name
fm-fleet-sync.sh's single-project form failed on a bare project name
("not a directory"), forcing hand-typed full paths (4 manual runs in
one incident). It now resolves a bare name or projects/<name> against
the home's projects dir.
AGENTS.md now encodes the trigger: a wake whose status reports a
merged PR for a project cloned in this home runs fleet-sync for that
project as part of handling the wake, so a secondmate-reported merge
does not leave the primary's clone stale until the next session start
or teardown.
* no-mistakes(review): Fix fleet-sync project name shadowing
* no-mistakes(document): sync fleet-sync docs
* fix: canonicalize spawn worktree path checks (#294)
* fix(spawn): canonicalize worktree-isolation guard against symlinked project prefixes
fm-spawn.sh compared a logical PROJ_ABS against the physically-resolved
pane cwd every backend reports, so a project reached through a symlinked
prefix (e.g. macOS's /tmp -> /private/tmp) could trip the isolation
guard's false refusal before treehouse ever moved the pane. Canonicalize
once into PROJ_ABS_REAL and compare against that everywhere instead.
* no-mistakes(review): Canonicalize spawn cwd comparisons
* no-mistakes(document): Refresh symlinked spawn docs
* docs: add Orca operator skill (#276)
* docs: add Orca operator skill
* no-mistakes(document): Document Orca checklist
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* fix: surface green PRs during CI monitoring (#297)
* fix(crew-state): detect green-PR CI monitoring, escalate repeat wedges
fm-crew-state.sh's ci step never distinguishes "still waiting on checks"
from "checks green, waiting on merge" via axi status alone, since a repo
that defers merge to the captain keeps the ci step at status=running for
the whole monitor phase. Read the ci step's own log tail (axi logs) for
the checks-passed marker and surface done instead of a false "validating
(running)" - verified against the real PR #252 run's ci.log.
The watcher's wedge timer can re-escalate the same stale pane forever
without ever signaling that it is a repeat; track a per-pane consecutive
escalation count and add a demand-deep-inspection marker to the wake
payload once it crosses a threshold, so the supervisor can no longer
dismiss each one as an isolated, still-validating pane.
Also clarify the ship-brief's checks-green line: it is owed at the
CI-ready return point, not after the background monitor-until-merge
loop finishes.
* no-mistakes(review): Captain, distinguish pending no-checks CI marker
* no-mistakes(review): Harden CI relapse handling
* no-mistakes(review): Block stale done during fixing
* no-mistakes(review): Captain, tighten CI status gating
* no-mistakes(review): Captain, harden stale CI green handling
* no-mistakes(review): Captain, recognize ranged CI rearm markers
* no-mistakes(document): Sync crew-state supervision docs
* fix(teardown): recover provably stale git index locks (#296)
* fix(teardown): recover from a stale worktree git index.lock
A crew process killed mid-git-operation can leave a stale
.git/worktrees/<wt>/index.lock behind, making fm-teardown.sh's
`treehouse return --force` fail closed. On that failure, retry once
after a short wait (the owning process may be exiting), then remove
the lock and retry once more only when it is provably stale: old
enough by mtime and lsof shows no live holder on the lock or the
worktree itself. A lock that isn't provably stale is left in place and
the original failure still surfaces.
* no-mistakes(review): Harden teardown lock refusal paths
* no-mistakes(review): Harden stale-lock teardown safety rechecks
* no-mistakes(review): Harden stale teardown lock checks
* no-mistakes(document): Document teardown lock recovery
* feat(bin): encode project AGENTS.md authoring bar with canonical self-governance section (#307)
* Encode project AGENTS authoring bar
* no-mistakes(review): Captain, centralize CLAUDE promotion governance
* no-mistakes(review): make ensure_maintenance_section idempotent-success, drop || true guards
* no-mistakes(review): separate appended maintenance section on newline-less CLAUDE.md promotion
* no-mistakes(review): assert maintenance heading present before separator check in test
* no-mistakes(document): sync docs with AGENTS.md authoring bar and self-governance
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* feat(skills): add captain-invocable bearings status-report skill (#300)
* Add captain-invocable bearings skill
Generates a pick-up-where-I-left-off status report from live fleet
state to data/status-report-<YYYY-MM-DD>.md plus a concise chat
summary. Read-mostly procedure: reads backlog, per-task crew state
via bin/fm-crew-state.sh, open PRs via gh-axi, scout reports,
pending decisions, and date-gated queued work; composes the
exemplar's sections (TL;DR, Check first, Landed, In flight, Plans,
Decisions pending, Date-gated/queued); never tears down, merges, or
mutates task state as a side effect.
* no-mistakes(document): docs: list new /bearings skill in README built-in skills table
* fix(watcher): make PID identity locale-invariant (#285)
* fix(watcher): pin LC_ALL=C in fm_pid_identity for locale-invariant identity
ps's lstart date format follows the caller's LC_TIME/LC_ALL. The watcher records
its process identity under one locale, but arm/guard/turn-end re-read it under the
machine's ambient locale. On a non-C locale (e.g. ko_KR) the two strings differ
only in the date portion, so fm_watcher_lock_matches_pid / fm_watcher_healthy
reject a genuinely live watcher - breaking fm-watch-arm.sh, fm-guard.sh, and
fm-turnend-guard.sh on every non-C-locale machine.
Pin LC_ALL=C on that one ps call so the write and read sides agree regardless of
machine locale, matching the LC_ALL=C determinism the file already uses elsewhere.
Add a colocated regression test asserting fm_pid_identity is locale-invariant
across exported LC_ALL/LC_TIME.
* no-mistakes(document): Document watcher PID identity coverage
* docs: document codex app backend contract (#222)
* docs: reconcile Codex App backend contract
* no-mistakes(document): Sync backend docs
* docs: clarify Codex Desktop bridge blocker
* no-mistakes(document): Align Codex App backend docs
* no-mistakes(test): Captain, stabilize watcher self-eviction test cadence
* no-mistakes(document): Document Codex App backend contract
* no-mistakes(document): Captain, document blocked codex-app coverage
* docs: make Codex App contract doc authoritative
* no-mistakes(document): Align Codex App backend docs
* docs: redact local Codex App smoke paths
---------
Co-authored-by: Stephen Brouhard <vesta@stephens-macbook-air.tail2122af.ts.net>
* docs: add Codex Desktop coordination skill (#275)
* docs: add Codex App coordination skill
* no-mistakes(review): Captain, mark Codex App skill agent-only
* no-mistakes(document): Document Codex App backend boundary
* no-mistakes(document): Captain, document Codex Desktop backend boundary
* no-mistakes(lint): Captain, lint clean
* no-mistakes(document): Document Codex Desktop boundaries
* docs: narrow Codex App skill playbook
* fix(afk): stop herdr escalation redelivery loop (#317)
* fix(afk): recognize unbordered herdr composer rows to stop escalation redelivery loop
fm_backend_herdr_composer_state only recognized bordered composer rows
(the grok shape). Real claude and codex render their live input row
with no border at all, so once a harness's own startup banner scrolled
out of the capture window the classifier read the composer as unknown
forever. fm_backend_herdr_send_text_submit never confirmed "empty", so
escalate_flush never cleared state/.subsuper-escalations, and the
away-mode daemon retyped and resubmitted the same buffered digest every
housekeeping cycle - reproduced live against a real herdr+claude pane
(5+ identical deliveries in 40s).
The classifier now recognizes an unbordered (bare) composer row led by
a known prompt glyph alongside the existing bordered shape, keeping
whichever match is bottom-most so a stale decorative box never
outranks the live composer.
* no-mistakes(review): Narrow herdr bare prompt matcher
* no-mistakes(document): Sync herdr composer docs
* fix(backends): confirm Herdr submits with native agent state (#323)
* fix(herdr): confirm message submit via native agent-state, not composer text
fm_backend_herdr_send_text_submit now confirms a landed submit by polling
herdr's own agent-state (agent get) for the idle->working transition instead
of reading composer content. Composer scraping remains, unchanged, for the
away-mode daemon's pre-injection empty-box guard only.
This fixes the practical effect of the codex idle-tip gap from the
2026-07-07 incident: codex's dynamic idle-composer hint text can no longer
misread as pending and block/mis-confirm a send, since confirmation no
longer looks at composer text at all. Verified empirically against real
claude and codex agents (timing, swallowed-Enter, unreadable-target, and
already-busy-target scenarios), and against the real away-mode daemon
end-to-end after updating its synthetic supervisor-pane test fixture to
register itself as a real herdr agent (herdr's own report-agent primitive)
so it can still exercise the new confirmation path.
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(review): Captain, harden herdr submit confirmation
* no-mistakes(document): Sync Herdr submit docs
* no-mistakes: apply CI fixes
* feat: add quota-balanced crew dispatch selection (#327)
* Add quota-balanced dispatch selection
* no-mistakes(document): Document dispatch selector guidance
* fix(session-start): respawn dead secondmate agents conservatively
* fix(session-start): deterministically respawn dead-shell secondmates
A secondmate agent that exits leaves its backend pane alive as a bare
shell. The session-start endpoint check only verified pane presence, so
recovery and the watcher (which exempts secondmates from stale-pane
detection) never noticed - evidence 2026-07-07: every secondmate in one
fleet was found sitting at a dead zsh shell.
Add fm_backend_agent_alive (bin/fm-backend.sh), a deeper per-backend
liveness probe distinct from pane presence: fm_backend_tmux_agent_alive
classifies the pane's live foreground process via tmux's own
pane_current_command, and fm_backend_herdr_agent_alive reuses the
already-verified pane_agent_state husk classifier. Both are conservative:
anything ambiguous reports unknown, never a false dead.
Wire this into a new session-start-only, locked-and-primary-only sweep in
bin/fm-bootstrap.sh that kills and respawns only a confidently dead
secondmate endpoint, leaving alive/unknown readings untouched - idempotent
by construction, so repeated runs converge without duplicating agents.
* no-mistakes(review): Guard raw secondmate liveness respawns
* no-mistakes(review): Fix detect-only bootstrap test
* no-mistakes(test): Pin liveness fixture harness
* no-mistakes(document): Sync secondmate liveness docs
* no-mistakes: apply CI fixes
* fix: emit stable secondmate nudge selectors (#331)
* Fix NUDGE_SECONDMATES to print stable fm-<id> selectors.
Session-start secondmate sync used to accumulate raw backend window targets
into NUDGE_SECONDMATES, but the liveness sweep in the same bootstrap run can
respawn secondmates onto new endpoints. fm-send with those stale explicit
targets bypasses meta resolution and fails, while fm-<id> resolves correctly.
Accumulate fm-<id> in process_secondmate, update the bootstrap/update contracts
and /updatefirstmate skill, and add a herdr respawn regression test.
* no-mistakes(review): Captain, guard herdr regression jq dependency
* no-mistakes(document): Document stable secondmate nudge selectors
* no-mistakes(lint): Fix shell lint hints
* feat: require bootstrap detection for AXI tools (#332)
* Make tasks-axi and quota-axi required bootstrap tools
Add both to the normal toolchain checks alongside lavish-axi, keep the
tasks-axi 0.1.1+ compatibility gate, and report quota-axi through the
standard MISSING install-consent flow. TASKS_AXI: available remains a
backlog-backend capability signal only; manual opt-out no longer suppresses
the missing-tool report.
Update bootstrap tests and point docs/configuration.md at the canonical
toolchain contract.
* no-mistakes(review): Clarify manual backlog bootstrap reporting
* no-mistakes(document): Document bootstrap AXI tools
* bearings: delete today's report before recreating (#333)
Replace overwrite-in-place wording with explicit delete-then-create
instructions so agents do not modify an existing daily report file.
* feat: guard primary turn ends across harnesses (#339)
* Add primary turn-end guards for all harnesses
* no-mistakes(review): Normalize Codex hook cwd resolution
* no-mistakes(review): Fix OpenCode guard worktree anchoring
* no-mistakes(review): Anchor Codex guard outside nested roots
* no-mistakes(review): Anchor Codex guard to hook root
* no-mistakes(review): Avoid Grok permission escalation
* no-mistakes(document): Sync turn-end guard docs
* fix: resolve backend selectors by exact task id first (#342)
* fix backend selector task id resolution
* no-mistakes(document): Document selector resolution behavior
* fix: scale bootstrap fleet-sync timeout (#341)
* fix bootstrap fleet sync timeout
* no-mistakes(review): Fix bootstrap fleet-sync timeout regressions
* no-mistakes(document): Sync bootstrap timeout docs
* no-mistakes(lint): Clean ShellCheck directives
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add fleet snapshot and view commands (#343)
* Add fleet snapshot and view
* no-mistakes(review): Fix fleet snapshot parsing and overrides
* no-mistakes(review): Fix secondmate fleet rendering
* no-mistakes(review): Fix backlog title and completion parsing
* no-mistakes(review): Include durable scout reports
* no-mistakes(review): Fix fleet snapshot edge cases
* no-mistakes(review): Captain: gate fleet hints on current state
* no-mistakes(review): Captain: parse bracketed Done PR artifacts
* no-mistakes(document): Sync fleet snapshot docs
* fix(fm-send): fail loudly on unresolvable send targets (#254)
* Make fm-send fail loudly on unresolved targets
* no-mistakes(review): Document fm-send FM_HOME contract
* Fix fm-send readiness docs and backend send path
* Fix fm-send docs for cmux and X skill metadata
* Make gotmp teardown test home-explicit
* Scope watcher warning wording to fm-send
* Fix fm-send review findings
* Verify explicit tmux targets before sending
* Isolate turnend guard test home
* no-mistakes(document): Documented fm-send FM_HOME/backend guard additions missing from doc inventories
---------
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* fix: deliver AFK escalations through herdr supervisors (#353)
* fix afk codex ghost composer delivery
* no-mistakes(review): Harden AFK startup flag writes
* no-mistakes(review): Harden AFK daemon liveness checks
* no-mistakes(document): Sync AFK herdr docs
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* feat: add harness-aware supervision (#367)
* Add harness-aware supervision
* no-mistakes(review): Captain, harden watcher supervision regressions
* no-mistakes(review): Captain, harden watcher supervision cadence
* no-mistakes(review): Harden watcher supervision ownership
* no-mistakes(review): Captain, harden Pi extension marker
* no-mistakes(review): Captain, harden Pi supervision restart checks
* no-mistakes(review): Harden watcher ownership checks
* no-mistakes(review): Captain, harden Pi supervision loading
* no-mistakes(review): Captain, require Pi guard extension loading
* no-mistakes(review): Captain, harden watcher supervision recovery
* no-mistakes(test): Fix fm-send baseline log filtering
* no-mistakes(document): Sync harness supervision docs
* no-mistakes: apply CI fixes
* fix: split X-mode replies by platform (#369)
* fix: make x replies split by platform
* no-mistakes(review): Captain: preserve Discord recovery relink context
* no-mistakes(test): Captain: keep split markers outside fences
* no-mistakes(document): Sync X-mode reply docs
* fix: make stow memory writes inspect before update (#372)
* docs: make stow inspect-then-update
* no-mistakes(review): Remove unsupported archive-body guidance
* no-mistakes(review): Clarify stow read-before-write exception
* no-mistakes(test): Require archive-body for stow task notes
* no-mistakes(document): Sync stow memory docs
* no-mistakes(lint): Silence ShellCheck source warning
* fix(watcher): wait when arm attaches to a healthy watcher (#375)
* fix: attach-and-wait when arm finds a healthy watcher
Grok and Claude re-arm after every turn with work in flight. When a
watcher was already healthy, fm-watch-arm exited immediately with
watcher: healthy, which completed the harness background task and
injected an empty false wake.
Attach to the live identity-matched holder instead, stay until that
cycle ends, then exit 0 so notify fires for a real end-of-cycle. The
peer-startup-race path uses the same contract. --restart and the
started path are unchanged.
* no-mistakes(review): Gate restart watcher peer attach
* no-mistakes(document): Sync watcher arm docs
* feat(pi): simplify primary session launch (#386)
* docs(readme): reformat Quick Start and recommend Grok equally with Claude Code
* no-mistakes(review): Captain: align harness launch guidance
* no-mistakes(review): Captain, clarify Pi supervised launch
* no-mistakes(review): Captain, document Pi first-launch bridge
* feat(pi): track primary watcher extension for plain-pi launch
Move Pi's primary watcher bridge from a generated state/ file to a
tracked .pi/extensions/fm-primary-pi-watch.ts, matching how the turn-end
guard extension already works: self-hashing version, project-local
auto-discovery after one-time Pi trust. This drops the state/-generation
step and dual -e requirement from the happy path, so Pi's Quick Start
launch becomes plain 'pi', the same friction class as 'claude' and
'grok --trust'.
- bin/fm-pi-watch-extension.sh is removed; nothing generates the
extension anymore since it is committed.
- fm-session-start.sh and fm-supervision-instructions.sh resolve the
watcher extension path from FM_ROOT instead of state/, and the
session-start diagnostic now points at restarting plain pi after
trust, with -e as a documented fallback.
- fm-spawn.sh points Pi secondmate launches at the tracked extension
path in the secondmate home instead of generating a state/ copy.
- README Quick Start Pi block is now just 'pi' plus a trust note.
- Tests, docs, and the harness-adapters skill updated to match.
* fix(pi): drop backticks from session-start diagnostic to satisfy shellcheck SC2016
* feat(supervision): prevent unsafe watcher-arm commands (#387)
* feat(supervision): add PreToolUse seatbelt against watcher-arm anti-patterns
Adds bin/fm-arm-pretool-check.sh, a shared PreToolUse-style checker that
denies a primary shell command backgrounding, piping, or bundling the
watcher arm/checkpoint, or force-killing the watcher process broadly -
the exact shapes that silently took Grok's supervision down. Wires it
into all five verified harnesses (grok, claude, codex, opencode, pi),
each validated empirically against the real harness.
Also fixes a grok 0.2.93 regression discovered during that validation: the
existing turnend-guard Stop hook's bare root variable broke grok's own
variable pre-substitution and silently no-op'd the hook.
* no-mistakes(review): Harden watcher arm validation
* no-mistakes(review): Harden arm guard metacharacter checks
* no-mistakes(review): Harden nested shell arm guard
* fix(lint): rewrite SC2015 guards in fm-arm-pretool-check.sh as if/then
A && B || C is not if-then-else; C can run when A is true. Replace both
occurrences of the quote-state early-continue with an explicit if/then.
* fix(pi): restore primary watcher supervision lifecycle (#397)
* fix Pi primary supervision lifecycle
* no-mistakes(document): Synchronize Pi primary extension documentation
* fix: keep persistent secondmates out of the main backlog (#398)
* fix secondmate backlog guidance
* no-mistakes(review): Require reasons for captain backlog holds
* no-mistakes(test): Document secondmate handoff skill requirement
* fix secondmate teardown reminder
* no-mistakes(document): sync teardown reminder docs to work-items-only backlog contract
* fix(backlog-handoff): move full item blocks including indented bodies (#401)
* fix(backlog-handoff): move full item blocks including indented bodies
fm-backlog-handoff only moved the checklist header line, so multi-line
item bodies were left orphaned in the source backlog and never reached
the secondmate. Move the full block (header plus indented body lines)
atomically, treating body membership by indentation so lines like
## Intent stay with the item, and add regression coverage.
* no-mistakes(review): Captain: preserve EOF handoff terminators
* no-mistakes(review): treat blank lines inside item bodies as movable body
* no-mistakes(document): sync backlog-handoff docs with full-block move behavior
* feat(herdr): make Herdr lab lifecycle safety deterministic for briefs (#402)
* guard Herdr lab lifecycle in briefs
* no-mistakes(review): Fix Herdr lab helper and provisioning safety
* no-mistakes(review): Captain, harden Herdr lab lifecycle safety
* no-mistakes(review): fix Herdr lab test cleanup ordering and brief help range
* no-mistakes(review): reject leading options in Herdr lab run guard
* no-mistakes(review): strip leading non-alnum in Herdr lab name generator
* no-mistakes(document): document Herdr lab helper and --herdr-lab brief flag
* no-mistakes(lint): add shellcheck disable for deliberate SC2016 literals in fm-brief herdr-lab
* no-mistakes: apply CI fixes
* fix(watcher): classify arm-command seatbelt by execution position (#403)
* fix watcher arm command policy
* no-mistakes(review): Harden watcher command policy parsing
* no-mistakes(review): Captain: harden watcher policy parsing
* no-mistakes(review): harden watcher policy for expanded paths, direct-watch, and sound prefilter
* no-mistakes(review): close prefilter and classifier locale/ANSI-C watcher-path decode gaps
* no-mistakes(review): fail closed on loop-wrapped broad watcher kills
* no-mistakes(document): sync docs for watcher-arm command-position policy
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", leaving the clone unsynced.
On that signature only, fm-fleet-sync.sh now retries the fetch with a
bounded wait (transient locks self-clear), then removes the lock and
retries once more ONLY when it is provably stale: still present, mtime
age past a threshold, and no lsof holder of the lock file or of the clone
worktree itself (a live git keeps that as its cwd even in the window after
it closes the lock and before it exits). A live lock, a missing lsof, any
failed check, or any other fetch failure keeps today's behavior. Every
wait/retry/removal prints to stderr, and a successful recovery also prints
one "recovered:" summary to stdout so a session-start refresh - which
discards fleet-sync stderr and relays only stdout - still surfaces it.
The shared "is this git lock provably abandoned?" proof is extracted into
bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and
fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh
gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/.
* no-mistakes(review): Captain, remove obsolete teardown wake dependency
* no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): escalate blocked panes immediately (#472)
* feat(herdr): immediate blocked-state escalation via native events.subscribe push
Fold herdr's native pane.agent_status_changed stream into the single watcher so
a crew entering blocked wakes its supervisor sub-second (measured 0.129s)
instead of after the ~240s stale-pane wedge timer.
- bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape
plus the single-owner status->action policy table (blocked=actionable,
working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling).
- bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe
subscriber over one connection for all this home's herdr panes, subscribing to
ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe
marker and a reconnect level-reconcile. Version/schema capability gate.
- bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so
the watcher stays backend-agnostic and the shape+policy are reusable.
- bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal
wait primitive (replacing the blind sleep POLL for push-capable homes),
behind a source guard so the splice is unit-testable; secondmate/paused
exemptions; map pane->window->task and enqueue a stale wake. No second
watcher process; the single-cycle invariant and every guard/beacon/turn-end
mechanism are unchanged.
- Polling stays the permanent fail-closed backstop: below-capability, subscribe
failure, and repeated runtime failures all degrade to sleep.
- Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/
fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events)
plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries
the dated evidence and retires the old gap note.
* no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs
* no-mistakes(review): Captain, commit markers after wake and reuse capability cache
* no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs
* no-mistakes(review): Captain, subscribe before Herdr reconciliation
* no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe
* no-mistakes(test): Captain: include lock library in teardown fixture
* no-mistakes(document): Captain: document Herdr immediate blocked escalation
* fix: clarify shellcheck conditionals
* docs(readme): reposition firstmate as an agent distro (#473)
* feat: add deterministic bounded bearings snapshots (#475)
* feat(bearings): deterministic bearings snapshot + durable decision model
Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the
canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR
discovery and checks happen only under --include-prs, which fails soft. Every
dropped surface is marked in omitted[] with the flag that reveals it, and the
prs: line states when checks were not requested, so absence is never silent.
Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib
gains status_open_decisions, the one authoritative keyed open/resolved fold over
the whole status stream: needs-decision/blocked opens a keyed entry, only an
explicit keyed resolution (or, for run-backed tasks, run-step advancement)
closes it, so a later unrelated done/paused can no longer mask a still-open
captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives
pending_decision/blocked_event from it; the canonical schema stays complete.
Point the /bearings skill at the one command; add the resolved: writer line to
ship, scout, and secondmate briefs. Register the script and add regression
tests for the output bound, TOON/JSON parity, local-only default, opt-in PR
fetch, partial-failure degradation, decision durability, and report pointers.
* fix(bearings): completed scout report is a pointer, not a pending decision
A completed scout that raised a needs-decision and then finished (done) without
a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103
case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh
cleared a stale decision only for a live run-step/pane activity read, so a
terminal task whose current state is read from the status log (a scout or ship
that reached done/failed) never cleared its stale, never-keyed-resolved
needs-decision, and it lingered as pending.
The open-decision set is still derived purely from the keyed fold - never from a
report body or decision-like prose - and reconciled against the crew lifecycle.
Extend that reconciliation so a terminal done/failed state on a single-owner
task (scout or ship), whose deliverable is its report or PR, also clears the set;
a completed scout now surfaces only as a report pointer. Secondmates are excluded
from the terminal clear (persistent, multiplexed stream), which keeps the
unrelated-event masking fix intact. Add regression tests: a completed scout with
decision-like report prose is a pointer not pending (canonical + end-to-end), and
a scout still parked at a decision stays pending so the terminal clear never
over-fires.
* no-mistakes(review): Captain, preserve keyed decisions across shared status parsing
* no-mistakes(review): Captain, close blockers and harden keyed decision parsing
* no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout
* no-mistakes(review): Captain, bound bearings sections and fail closed
* no-mistakes(review): Captain, disclose capped per-repository PR results
* no-mistakes(document): Refresh bearings documentation and status contracts
* fix(bearings): avoid ambiguous worktree guard
* fix: enforce deterministic ShellCheck parity (#481)
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity
Firstmate PRs passed local no-mistakes validation but failed CI's
"Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034).
Two divergences caused it:
1. The no-mistakes gate had no commands.lint, so its lint step never ran
the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that
CI runs. Confirmed from state.sqlite: the lint step_result recorded
findings:null with no lint agent invocation.
2. CI's shellcheck floated with the runner image while local ran a newer
build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck
rejected an SC2015 that the newer local one no longer emits.
Establish bin/fm-lint.sh as the single owner of the lint definition: the
file set, the config, and the pinned shellcheck version (0.11.0, printed
via --required-version). Both CI (.github/workflows/ci.yml) and the
no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs
the exact version it names and logs the resolved version, and fm-lint.sh
refuses to lint under any other version. This is not a CI relaxation: it
adopts shellcheck 0.11.0's rule set consistently, dropping only the
upstream-retired, false-positive-prone SC2015; default severity and every
still-supported finding stay enforced (no severity downgrade, no excludes).
tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs
and logs the pinned version, that the owner refuses a non-pinned shellcheck,
and that it rejects a real lint defect the old no-op gate passed.
* no-mistakes(review): Captain, harden deterministic ShellCheck parity
* no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: guard primary shells from persistent cd commands (#483)
* feat: add cd-guard PreToolUse seatbelt for the primary shell
A stray persistent top-level `cd projects/<clone>` in the primary firstmate
shell relocates the shell, so a later firstmate-owned command (a backlog write,
an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the
home. The cd-guard denies exactly that command shape before it runs, across all
five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt.
- bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the
shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate
lexer; that file's CLI now runs only when invoked directly).
- bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness
output rendering, and primary-checkout scoping - fires in a secondmate's own
primary session, inert in crew/scout child worktrees and non-firstmate repos.
- Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents;
per-harness hooks only call the owner.
- Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd,
and command cd). Allows git -C, subshell / bash -c / env -C / make -C /
find -execdir, pipeline and background forms, and cd-as-data. Fails open on
malformed input; agent-mistake threat model.
- tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix,
end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring.
- docs/cd-guard.md: full contract plus live validation (claude, codex,
opencode, pi blocked end-to-end; grok live run blocked by an API balance
limit, with mechanism parity and deterministic coverage recorded).
* no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage
* no-mistakes(review): Captain, allow path-qualified command wrappers
* no-mistakes(review): Captain, allow non-executing command queries
* no-mistakes(test): Captain, clarify cd-guard safe-path remediation
* docs: clarify cd guard guidance
* no-mistakes(document): Clarify cd-guard safe target guidance
* brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (#267)
Crews must never stop, restart, or update the shared no-mistakes
daemon since one instance serves every firstmate lane/home; a restart
kills other lanes' in-flight pipeline runs and forces expensive
re-runs. Encodes this as a new numbered rule in both the ship-task and
scout-task brief scaffolds.
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat: make bearings concise and accurate (#485)
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render
/bearings skill (one owner of the chat-response format): mandate the four
always-present chat sections - Captain's Call, Recently Landed, Underway,
Charted Next - each with an explicit empty-state sentence, no At Anchor,
materially shorter than and linking to the report file. Resolves the ambiguous
Check first / Decisions pending split into one strict captain-action section.
fm-crew-state: the log fallback derives current state only from a real
run-state verb, so a trailing decision-closing resolved: event no longer
renders a healthy idle crew (typically a secondmate) as unknown with the
resolution prose as its detail. The keyed-decision contract in
fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner.
fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done
records from registered secondmate homes, reusing the single backlog parser and
the one secondmate-home enumerator (meta home= with data/secondmates.md
fallback); no network, per-home capped.
fm-bearings-snapshot: landed now merges main-home Done with the secondmate
roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure
(also fixing the previously-silent landed truncation); --all-landed reveals the
full set.
tests: resolved-event state render, secondmate landed aggregation with caps and
omitted[] disclosure, Captain's Call anti-leak, and the four-section contract.
* no-mistakes(review): Captain, ensure bearings reveals all landed work
* no-mistakes(document): Document bearings accuracy contracts
* fix: harden away-mode daemon lifecycle (#490)
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle
Away-mode entry left "make the daemon a tracked background terminal" to the
operator; on a pi/herdr primary that meant splitting the captain's active pane,
which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches
the daemon in a non-visible tracked terminal per backend (herdr dedicated
--no-focus workspace, detached tmux session), never a split, pins the captain
pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal
id, and tears it down or reconciles a leaked one by that id. No shell &.
Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared
with the daemon (one owner).
Fix the stale subsuper-artifact leak: clear the prior away session's delivery
cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon
before clearing state/.afk so its shutdown flush runs instead of being a no-op.
Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab
session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md,
docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md.
* no-mistakes(review): Captain, serialize AFK launcher lifecycle safely
* no-mistakes(review): Captain, harden AFK launcher lifecycle races
* no-mistakes(review): Captain, ensure AFK daemon launch readiness
* no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown
* no-mistakes(review): Captain, preserve AFK reconciliation records uniformly
* no-mistakes(review): Captain, harden AFK recovery state durability
* no-mistakes(review): Captain, harden AFK tmux ownership checks
* no-mistakes(review): Captain, simplify AFK lifecycle failure handling
* no-mistakes(review): Captain, require confirmed AFK daemon shutdown
* no-mistakes(review): Captain, confirm AFK exit by process identity
* no-mistakes(document): Align AFK launcher lifecycle documentation
* no-mistakes: apply CI fixes
* fix: prevent no-mistakes gate agents from driving the fleet (#518)
* feat: contain no-mistakes gate agents from driving the fleet
Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.
Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).
firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.
* no-mistakes(review): Captain, refuse empty no-mistakes gate markers
* no-mistakes(document): Document no-mistakes gate authority boundary
* fix: guard secondmate primary sessions from blind turn ends (#505)
* fix: guard secondmate own-home turn ends
Rem…
DereKk8
added a commit
to DereKk8/firstmate
that referenced
this pull request
Aug 15, 2026
* fix: reconcile existing AGENTS.md safely (#405)
* fix(agents-md): inject self-governance section into existing AGENTS.md
fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.
Also fixes #389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.
Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.
* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency
* no-mistakes(review): Preserve CRLF during maintenance-section injection
* no-mistakes(review): Captain: harden dangling-symlink regression coverage
* no-mistakes(document): Document agent-memory injection outcomes
* feat: support project-less secondmate homes (#409)
* feat(secondmate): support project-less homes via --no-projects
fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.
* no-mistakes(review): Captain: document project-less secondmate flow
* no-mistakes(review): Captain: refuse project-less reseeding of populated homes
* fix(seed): fail closed on unreadable project data
* no-mistakes(review): Captain: reject stale projectful charters
* no-mistakes(review): Captain: fail closed on unsafe project paths
* no-mistakes(review): Captain: validate project-less charter clone sections
* no-mistakes(document): Document project-less secondmate seeding
* fix: delegate backlog handoffs to tasks-axi (#411)
* wip(handoff): record verified delegation design + tasks-axi mv blocker
No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.
* refactor(handoff): delegate the item move to tasks-axi mv
fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR #401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.
Tests: fm-backlog-handoff.test.sh keeps PR #401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.
Remove the now-redundant WIP design note.
* no-mistakes(review): Captain: harden atomic backlog handoffs
* no-mistakes(review): Captain: enforce queued-only backlog handoffs
* no-mistakes(review): Captain: harden handoff section parsing
* no-mistakes(document): Document delegated backlog handoffs
* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: ignore secondmate home marker during sync (#417)
* fix: gitignore the secondmate home marker
bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.
Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.
Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.
Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.
* no-mistakes(review): Captain: document standalone-clone update path
* no-mistakes(document): Document secondmate marker migration
* fix(composer): prevent dead-shell message injection (#416)
* fix(composer): stop reading dead-shell prompts as empty agent composers
Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.
Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.
Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.
* no-mistakes(review): Captain: harden composer safety checks
* no-mistakes(test): Stabilize Herdr prune safety setup
* no-mistakes(document): Document composer injection safety
* no-mistakes(lint): Clean composer safety lint
* no-mistakes: apply CI fixes
* feat(watcher): add paused external-wait supervision (#421)
* feat(watcher): add paused/awaiting-external crew state
A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.
* no-mistakes(review): Captain: fix paused-state transitions
* init
* no-mistakes(review): Captain: fix paused-state supervision transitions
* no-mistakes(review): Captain: fix paused supervision handoffs
* no-mistakes(review): Reconcile paused supervision markers
* no-mistakes(review): Captain: prioritize paused states over captain relevance
* no-mistakes(review): Captain: preserve paused-working wedge timer
* no-mistakes(review): Captain: honor configured pause verb in briefs
* no-mistakes(test): Captain: fix AFK paused watcher handoff
* no-mistakes(document): Document declared external waits
* no-mistakes(lint): Clean paused-state lint
---------
Co-authored-by: fmtest <fmtest@example.invalid>
* fix: preserve X-mode follow-up platform limits (#425)
* fix(x-mode): make follow-up platform splitting immune to link ordering
A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.
Make platform resolution ordering-proof:
- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
new fmx_request_relay_context helper (POST /connector/request-context)
when neither the inbox payload nor carry flags carry it. The request_id
survives the inbox drain, so a post-cleanup link still learns the right
split budget. Best-effort: no token/curl or a non-2xx relay degrades to
the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
fast local path stays correct without a relay round-trip.
Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.
The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.
* no-mistakes(document): Document X-mode platform recovery
* fix(composer): handle ANSI ghost text safely (#429)
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor
Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).
Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.
This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).
Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.
* no-mistakes(review): Preserve dark shell prompt safety
* no-mistakes(review): Harden erased shell prompt classification
* no-mistakes(document): Document shared composer ghost extraction
* no-mistakes(lint): Normalize tmux comment punctuation
* fix(spawn): make tmux window handling robust under non-default config (#134)
* test: isolate session-start suite from ambient harness markers (#432)
* fix(session-start): isolate harness env markers in suite runner
Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).
* no-mistakes(document): Correct Pi marker documentation
* fix(teardown): retry transient index locks during worktree return (#435)
* fix(teardown): retry treehouse return on transient index.lock
Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.
* no-mistakes(review): Harden teardown retry configuration
* no-mistakes(document): Document teardown index-lock retry behavior
* no-mistakes(lint): Fix empty shell variable assignments
* fix: complete brief help and consolidate documentation (#438)
* docs: de-feature the scripts.md and CONTRIBUTING test inventories
Slice 1 of the documentation redundancy cleanup wave (firstmate scope).
docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.
CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).
Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.
No instruction-surface or behavior change; AGENTS.md untouched.
* no-mistakes(review): Captain, fix brief help and Grok test discovery
* no-mistakes(review): Captain: document Grok lock-holder test coverage
* fix: detect Git and centralize backend configuration (#445)
* docs: consolidate universal backend contracts into configuration.md
Slice 2 of the documentation redundancy cleanup wave (firstmate scope).
docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).
The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.
Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).
* docs: include git and GitHub auth in the toolchain owner list
The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.
* no-mistakes(review): Detect Git in bootstrap toolchain
* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): add backend-independent wedge alerts (#444)
* feat(daemon): backend-independent active alert for the wedge alarm
When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).
Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.
The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.
Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.
* no-mistakes(review): Bound wedge alarm notifier execution
* no-mistakes(review): Captain: harden wedge alarm notifier safety
* no-mistakes(review): Captain: harden wedge alarm test notifier isolation
* no-mistakes(review): Captain: harden wedge alarm throttling
* no-mistakes(review): Redact wedge alarm directive logs
* no-mistakes(review): Harden wedge alarm notifier safety
* no-mistakes(review): Track notifier process groups through cleanup
* no-mistakes(document): Document wedge-alarm active alert behavior
* docs: centralize firstmate operating contracts (#447)
* docs(agents): extract conditional AGENTS.md material to owned homes
Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.
Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
agent-only bootstrap-diagnostics skill, added to the section 13
trigger index; the detect-consent-install rule and the
do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
docs/configuration.md 'Crew dispatch profiles' (pointer direction
flipped); the intake procedure, precedence, backstop, and
never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
header (now the declared owner; usage() converted to the dynamic
header extraction pattern PR #438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
header (section 1's containment pointer retargeted); the fork benign
case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
'Event-driven supervision' (already the owner); every operative rule
(one live cycle, no turn ends blind, drain first, wake ladder,
never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
restatements -> secondmate-provisioning skill, now the declared
owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
'X mode (.env)', closing issue #363; activation semantics, the
fmx-respond trigger, and the terminal-wake final-follow-up duty
stay inline.
CLAUDE.md stays a symlink; no behavior or test change.
* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close last workspace during teardown (#449)
* fix(cmux): close the last/selected workspace in a window at teardown
cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).
Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.
Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.
* no-mistakes(review): Derive cmux count from membership snapshot
* no-mistakes(document): Document cmux last-workspace teardown behavior
* fix: recover orphaned packed-refs locks during fleet sync (#453)
* fix(fleet-sync): recover from an orphaned packed-refs.lock
A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", leaving the clone unsynced.
On that signature only, fm-fleet-sync.sh now retries the fetch with a
bounded wait (transient locks self-clear), then removes the lock and
retries once more ONLY when it is provably stale: still present, mtime
age past a threshold, and no lsof holder of the lock file or of the clone
worktree itself (a live git keeps that as its cwd even in the window after
it closes the lock and before it exits). A live lock, a missing lsof, any
failed check, or any other fetch failure keeps today's behavior. Every
wait/retry/removal prints to stderr, and a successful recovery also prints
one "recovered:" summary to stdout so a session-start refresh - which
discards fleet-sync stderr and relays only stdout - still surfaces it.
The shared "is this git lock provably abandoned?" proof is extracted into
bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and
fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh
gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/.
* no-mistakes(review): Captain, remove obsolete teardown wake dependency
* no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): escalate blocked panes immediately (#472)
* feat(herdr): immediate blocked-state escalation via native events.subscribe push
Fold herdr's native pane.agent_status_changed stream into the single watcher so
a crew entering blocked wakes its supervisor sub-second (measured 0.129s)
instead of after the ~240s stale-pane wedge timer.
- bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape
plus the single-owner status->action policy table (blocked=actionable,
working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling).
- bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe
subscriber over one connection for all this home's herdr panes, subscribing to
ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe
marker and a reconnect level-reconcile. Version/schema capability gate.
- bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so
the watcher stays backend-agnostic and the shape+policy are reusable.
- bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal
wait primitive (replacing the blind sleep POLL for push-capable homes),
behind a source guard so the splice is unit-testable; secondmate/paused
exemptions; map pane->window->task and enqueue a stale wake. No second
watcher process; the single-cycle invariant and every guard/beacon/turn-end
mechanism are unchanged.
- Polling stays the permanent fail-closed backstop: below-capability, subscribe
failure, and repeated runtime failures all degrade to sleep.
- Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/
fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events)
plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries
the dated evidence and retires the old gap note.
* no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs
* no-mistakes(review): Captain, commit markers after wake and reuse capability cache
* no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs
* no-mistakes(review): Captain, subscribe before Herdr reconciliation
* no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe
* no-mistakes(test): Captain: include lock library in teardown fixture
* no-mistakes(document): Captain: document Herdr immediate blocked escalation
* fix: clarify shellcheck conditionals
* docs(readme): reposition firstmate as an agent distro (#473)
* feat: add deterministic bounded bearings snapshots (#475)
* feat(bearings): deterministic bearings snapshot + durable decision model
Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the
canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR
discovery and checks happen only under --include-prs, which fails soft. Every
dropped surface is marked in omitted[] with the flag that reveals it, and the
prs: line states when checks were not requested, so absence is never silent.
Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib
gains status_open_decisions, the one authoritative keyed open/resolved fold over
the whole status stream: needs-decision/blocked opens a keyed entry, only an
explicit keyed resolution (or, for run-backed tasks, run-step advancement)
closes it, so a later unrelated done/paused can no longer mask a still-open
captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives
pending_decision/blocked_event from it; the canonical schema stays complete.
Point the /bearings skill at the one command; add the resolved: writer line to
ship, scout, and secondmate briefs. Register the script and add regression
tests for the output bound, TOON/JSON parity, local-only default, opt-in PR
fetch, partial-failure degradation, decision durability, and report pointers.
* fix(bearings): completed scout report is a pointer, not a pending decision
A completed scout that raised a needs-decision and then finished (done) without
a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103
case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh
cleared a stale decision only for a live run-step/pane activity read, so a
terminal task whose current state is read from the status log (a scout or ship
that reached done/failed) never cleared its stale, never-keyed-resolved
needs-decision, and it lingered as pending.
The open-decision set is still derived purely from the keyed fold - never from a
report body or decision-like prose - and reconciled against the crew lifecycle.
Extend that reconciliation so a terminal done/failed state on a single-owner
task (scout or ship), whose deliverable is its report or PR, also clears the set;
a completed scout now surfaces only as a report pointer. Secondmates are excluded
from the terminal clear (persistent, multiplexed stream), which keeps the
unrelated-event masking fix intact. Add regression tests: a completed scout with
decision-like report prose is a pointer not pending (canonical + end-to-end), and
a scout still parked at a decision stays pending so the terminal clear never
over-fires.
* no-mistakes(review): Captain, preserve keyed decisions across shared status parsing
* no-mistakes(review): Captain, close blockers and harden keyed decision parsing
* no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout
* no-mistakes(review): Captain, bound bearings sections and fail closed
* no-mistakes(review): Captain, disclose capped per-repository PR results
* no-mistakes(document): Refresh bearings documentation and status contracts
* fix(bearings): avoid ambiguous worktree guard
* fix: enforce deterministic ShellCheck parity (#481)
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity
Firstmate PRs passed local no-mistakes validation but failed CI's
"Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034).
Two divergences caused it:
1. The no-mistakes gate had no commands.lint, so its lint step never ran
the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that
CI runs. Confirmed from state.sqlite: the lint step_result recorded
findings:null with no lint agent invocation.
2. CI's shellcheck floated with the runner image while local ran a newer
build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck
rejected an SC2015 that the newer local one no longer emits.
Establish bin/fm-lint.sh as the single owner of the lint definition: the
file set, the config, and the pinned shellcheck version (0.11.0, printed
via --required-version). Both CI (.github/workflows/ci.yml) and the
no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs
the exact version it names and logs the resolved version, and fm-lint.sh
refuses to lint under any other version. This is not a CI relaxation: it
adopts shellcheck 0.11.0's rule set consistently, dropping only the
upstream-retired, false-positive-prone SC2015; default severity and every
still-supported finding stay enforced (no severity downgrade, no excludes).
tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs
and logs the pinned version, that the owner refuses a non-pinned shellcheck,
and that it rejects a real lint defect the old no-op gate passed.
* no-mistakes(review): Captain, harden deterministic ShellCheck parity
* no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: guard primary shells from persistent cd commands (#483)
* feat: add cd-guard PreToolUse seatbelt for the primary shell
A stray persistent top-level `cd projects/<clone>` in the primary firstmate
shell relocates the shell, so a later firstmate-owned command (a backlog write,
an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the
home. The cd-guard denies exactly that command shape before it runs, across all
five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt.
- bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the
shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate
lexer; that file's CLI now runs only when invoked directly).
- bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness
output rendering, and primary-checkout scoping - fires in a secondmate's own
primary session, inert in crew/scout child worktrees and non-firstmate repos.
- Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents;
per-harness hooks only call the owner.
- Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd,
and command cd). Allows git -C, subshell / bash -c / env -C / make -C /
find -execdir, pipeline and background forms, and cd-as-data. Fails open on
malformed input; agent-mistake threat model.
- tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix,
end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring.
- docs/cd-guard.md: full contract plus live validation (claude, codex,
opencode, pi blocked end-to-end; grok live run blocked by an API balance
limit, with mechanism parity and deterministic coverage recorded).
* no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage
* no-mistakes(review): Captain, allow path-qualified command wrappers
* no-mistakes(review): Captain, allow non-executing command queries
* no-mistakes(test): Captain, clarify cd-guard safe-path remediation
* docs: clarify cd guard guidance
* no-mistakes(document): Clarify cd-guard safe target guidance
* brief: add no-mistakes shared-daemon rule to ship and scout scaffolds (#267)
Crews must never stop, restart, or update the shared no-mistakes
daemon since one instance serves every firstmate lane/home; a restart
kills other lanes' in-flight pipeline runs and forces expensive
re-runs. Encodes this as a new numbered rule in both the ship-task and
scout-task brief scaffolds.
Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat: make bearings concise and accurate (#485)
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render
/bearings skill (one owner of the chat-response format): mandate the four
always-present chat sections - Captain's Call, Recently Landed, Underway,
Charted Next - each with an explicit empty-state sentence, no At Anchor,
materially shorter than and linking to the report file. Resolves the ambiguous
Check first / Decisions pending split into one strict captain-action section.
fm-crew-state: the log fallback derives current state only from a real
run-state verb, so a trailing decision-closing resolved: event no longer
renders a healthy idle crew (typically a secondmate) as unknown with the
resolution prose as its detail. The keyed-decision contract in
fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner.
fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done
records from registered secondmate homes, reusing the single backlog parser and
the one secondmate-home enumerator (meta home= with data/secondmates.md
fallback); no network, per-home capped.
fm-bearings-snapshot: landed now merges main-home Done with the secondmate
roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure
(also fixing the previously-silent landed truncation); --all-landed reveals the
full set.
tests: resolved-event state render, secondmate landed aggregation with caps and
omitted[] disclosure, Captain's Call anti-leak, and the four-section contract.
* no-mistakes(review): Captain, ensure bearings reveals all landed work
* no-mistakes(document): Document bearings accuracy contracts
* fix: harden away-mode daemon lifecycle (#490)
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle
Away-mode entry left "make the daemon a tracked background terminal" to the
operator; on a pi/herdr primary that meant splitting the captain's active pane,
which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches
the daemon in a non-visible tracked terminal per backend (herdr dedicated
--no-focus workspace, detached tmux session), never a split, pins the captain
pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal
id, and tears it down or reconciles a leaked one by that id. No shell &.
Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared
with the daemon (one owner).
Fix the stale subsuper-artifact leak: clear the prior away session's delivery
cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon
before clearing state/.afk so its shutdown flush runs instead of being a no-op.
Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab
session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md,
docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md.
* no-mistakes(review): Captain, serialize AFK launcher lifecycle safely
* no-mistakes(review): Captain, harden AFK launcher lifecycle races
* no-mistakes(review): Captain, ensure AFK daemon launch readiness
* no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown
* no-mistakes(review): Captain, preserve AFK reconciliation records uniformly
* no-mistakes(review): Captain, harden AFK recovery state durability
* no-mistakes(review): Captain, harden AFK tmux ownership checks
* no-mistakes(review): Captain, simplify AFK lifecycle failure handling
* no-mistakes(review): Captain, require confirmed AFK daemon shutdown
* no-mistakes(review): Captain, confirm AFK exit by process identity
* no-mistakes(document): Align AFK launcher lifecycle documentation
* no-mistakes: apply CI fixes
* fix: prevent no-mistakes gate agents from driving the fleet (#518)
* feat: contain no-mistakes gate agents from driving the fleet
Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.
Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).
firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.
* no-mistakes(review): Captain, refuse empty no-mistakes gate markers
* no-mistakes(document): Document no-mistakes gate authority boundary
* fix: guard secondmate primary sessions from blind turn ends (#505)
* fix: guard secondmate own-home turn ends
Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the
'no turn ends blind' backstop fires in a secondmate's own primary session,
matching the cd-guard's scope: the own home is guarded, child crew/scout
worktrees stay exempt via the retained git-dir/git-common-dir test. This
was pure scoping from the guard's primary-only origin and guarded against
no secondmate-specific hazard.
Add secondmate regression tests (blind-turn block, idle-by-default,
stop_hook_active loop guard, deferred-death recovery loop, child-worktree
exemption) and record the autonomous background-notify re-invoke
measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md.
* no-mistakes(document): Correct secondmate guard documentation, captain
* fix: force-include marked secondmate homes in turn-end guard
The prior remove-only form (just deleting the .fm-secondmate-home check)
left the DEFAULT secondmate topology unguarded: a treehouse-leased home is
a linked git worktree (git-dir != git-common-dir), which the retained
git-dir exemption still skipped, so its own primary session could still end
a turn blind. Invert the marker: a genuinely-marked home is force-included
as a guarded primary (treehouse-leased linked OR git-cloned plain), and the
git-dir exemption applies only to UNMARKED child worktrees. Marker
validation (regular non-symlink file, non-empty id-token content) blocks a
stray or empty marker from spoofing inclusion.
Add real linked-worktree regression tests: a treehouse-leased LINKED
secondmate home is guarded, a stray/empty marker stays exempt, and the
unmarked child worktree stays exempt - the topology the plain git-init
fixtures masked. Predicates, in-flight gate, and loop guard untouched.
* fix: force ASCII collation in secondmate marker validation
Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the
[A-Za-z0-9._-] id allowlist matches under C collation, not the ambient
locale - a locale-crafted non-ASCII marker id can no longer slip through
the range match and spoof force-inclusion of a linked child worktree.
Add a regression test proving a non-ASCII marker id is rejected and the
linked worktree stays exempt.
* no-mistakes(test): fix backend baseline gate-refusal dependency
* no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap diagnostics backend-aware (#519)
* fix: make bootstrap required-tool detection backend-aware
Bootstrap demanded tmux and treehouse for every backend except orca, so a
herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux.
Required tools now follow the resolved backend via the single-owner
fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own
session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux),
and treehouse for session-provider-only backends (orca owns its worktree).
The treehouse lease-support check is gated to backends that use treehouse.
Adds install hints for herdr/zellij/cmux, regression tests for the full
backend dependency matrix (herdr-without-tmux repro plus each boundary),
and updates the authoritative Toolchain docs.
* no-mistakes(review): Captain, prevent executing Herdr install guidance
* no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics
* no-mistakes(review): Captain, separate manual dependency remediation
* no-mistakes(review): Captain, align bootstrap diagnostic consumers
* no-mistakes(document): Align backend adapter dependency comments
* fix: preserve follow-up platform context after inbox cleanup (#520)
* fix: recover X/Discord follow-up platform after inbox cleanup
A milestone follow-up posted directly by request_id after the inbox was
drained - and with no task link, because one persistent secondmate's single
x_request slot collides across concurrent requests - resolved platform only
from the local inbox, so a >280 Discord reply silently defaulted to the X
280-char budget and threaded as (1/2).
- fm-x-poll records a durable per-request reply context
(state/x-context/<rid>.json) at stash time, keyed by request_id so
concurrent requests never overwrite each other; it survives inbox cleanup
and restart.
- fm-x-reply resolves platform/budget through registry -> inbox -> relay
(the relay lookup confined to a live follow-up), recovering the original
platform independent of task-link availability.
- Fail-safe: a follow-up whose platform/budget cannot be authoritatively
resolved and that would split is refused (exit 8) and held for retry,
never wrongly split; fm-x-followup keeps the link on that exit.
- fm-x-dismiss clears the durable context for a dismissed mention.
Refactors reply-context extraction into a single owner and adds regression
coverage for all four cases.
* no-mistakes(review): Captain, fail closed on incomplete follow-up context
* no-mistakes(review): Captain, bound X context registry retention
* no-mistakes(review): Captain, align context retention with answer binding
* no-mistakes(document): Align X follow-up context documentation
* no-mistakes(document): Align durable X follow-up documentation
* fix: preserve secondmate routing markers in terminal sends (#533)
* fix: preserve secondmate routing markers
* no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends
* no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift
* no-mistakes(document): Refresh Herdr marker documentation
* fix: align Grok effort handling with 0.2.99 (#527)
* fix: align grok effort docs and spawn with 0.2.99 ceiling
grok 0.2.99 accepts only low|medium|high for --reasoning-effort and
rejects both xhigh and max. Omit unsupported values on spawn, flag them
in crew-dispatch validation, and update harness-adapters.
* no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies
* no-mistakes(document): Clarify Grok effort documentation ownership
* fix: derive bearings from authoritative secondmate state (#555)
* fix: make bearings use secondmate home state
* test: anonymize bearings fixtures
* no-mistakes(review): Bound parent activity evidence scans, captain
* no-mistakes(review): Preserve structured secondmate authority and bounds, captain
* no-mistakes(review): Preserve registry completeness and child inventory, captain
* no-mistakes(review): Reconcile parent evidence by verb and key, captain
* no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain
* no-mistakes(document): Document bearings local snapshot and PR opt-in
* no-mistakes(lint): Fix fleet snapshot ShellCheck findings
* no-mistakes: apply CI fixes
* fix: restore fleet snapshots on stock macOS Bash (#578)
* fix: restore stock macOS snapshot parsing
* no-mistakes(document): Clarify Linux gate and macOS CI coverage
* fix(afk): make Pi escalation and return catch-up reliable (#587)
* fix: close away-mode blocker supervision gap
* no-mistakes(review): Gate teardown retries and verify U+2063 dedupe
* no-mistakes(test): Fail closed on incomplete Pi composer separators
* no-mistakes(document): Document Pi composer recognition and return gating
* feat: support Pi max reasoning profiles (#537)
* support Pi max thinking profiles
* no-mistakes(review): Captain, allow Pi max dispatch profiles
* chore: no-mistakes(document): Clarify yolo response ownership (#595)
* Clarify validation response ownership
* no-mistakes(document): Clarify yolo response ownership
* feat: establish instruction ownership foundation (#619)
* Add instruction owners foundation
* no-mistakes(document): Refresh project-management owner pointers
* fix: compress Firstmate contract and enforce delivery rigor ownership (#626)
* docs: compress firstmate operating contract
* docs: make delivery rigor single-owner
PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion.
* no-mistakes(review): Honor configured merge authority across faster delivery paths
* no-mistakes(document): Align docs with compressed operating contract
* feat: add durable captain decision holds (#593)
* Add durable captain decision holds
* no-mistakes(review): Validate decision hold retries and origin paths
* no-mistakes(review): Enforce durable decision lifecycle boundaries
* no-mistakes(review): Harden decision display and partial retry recovery
* no-mistakes(test): Update scout teardown fixtures for decision inventory
* no-mistakes(document): Align decision lifecycle and scout teardown documentation
* no-mistakes: apply CI fixes
* no-mistakes(review): Reconcile terminal decision holds
* no-mistakes(document): Align captain decision-hold documentation
* fix(bin): harden PR check artifacts (#556)
* fix: harden PR check artifacts
* fix: close PR check migration gaps
* fix: close PR check retry gaps
* fix: clarify migration outcomes and ESM boundary
* fix: keep failed migrations authoritative
* test: use inert PR validation fixtures
* no-mistakes(review): Reserve noncanonical PR quarantine namespace
* no-mistakes(review): Prevalidate final PR-check teardown artifacts
* no-mistakes(review): Preserve X metadata and validate teardown IDs
* no-mistakes(review): Initialize migration state before watcher exclusion
* no-mistakes(review): Isolate failed poll migrations from bootstrap recovery
* no-mistakes(review): Allow safe polling during incomplete private repairs
* no-mistakes(review): Authenticate watcher checks at execution time
* no-mistakes(review): Preserve custom checks with hash-bound registration
* no-mistakes(review): Clean custom check snapshots on watcher signals
* no-mistakes(review): Stop watcher checks promptly on signals
* no-mistakes(review): Terminate watcher check groups before cleanup
* no-mistakes(document): Correct stale X-mode watcher documentation
* fix: drain returned watcher check groups
* no-mistakes(review): Harden quarantine links and recover validated replacement polls
* no-mistakes(review): Preserve X mode across shim version transitions
* no-mistakes(review): Refresh legacy X shims before marker short-circuits
* no-mistakes(document): Correct persisted PR-check artifact documentation
* no-mistakes(document): Correct stale PR-check documentation
* fix: bind PR poll repair provenance
* no-mistakes(review): Enforce single-link ownership for custom check artifacts
* no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs
* no-mistakes(review): Separate task creation and legacy teardown validation
* no-mistakes(review): Restore safe legacy operations and teardown validation
* no-mistakes(review): Disambiguate migration obligations and preserve legacy retries
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits
* no-mistakes(document): Document private poll artifact safety contracts
* no-mistakes(lint): Suppress intentional literal-dollar lint finding
* fix: migrate historical X poll identity
* fix: harden PR check artifacts
* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence
* no-mistakes(review): Reconcile legacy migration retries and teardown collisions
* fix: migrate historical X poll identity
* no-mistakes(review): Harden X-mode artifact publication against symlink corruption
* no-mistakes(review): Guard X artifact publication
* no-mistakes(review): Enforce private X artifact reads
* no-mistakes(test): Fix backend compatibility fixture dependencies
* no-mistakes(document): Refresh PR-check documentation
* no-mistakes(lint): Remove unused x-mode test locals
* no-mistakes: apply CI fixes
* fix(bin): compact session-start backlog digest (#636)
* fix: compact session-start backlog digest
* no-mistakes(test): Fix legacy backend fixture helper
* no-mistakes(test): Fix watcher exit wait helper
* no-mistakes(document): Document compact backlog digest
* fix: dedupe stale watcher guard banners (#637)
* fix: dedupe stale watcher guard banner
* no-mistakes(review): Keep read-only guard state nonmutating
* no-mistakes(document): Clarify stale watcher docs
* fix(bin): balance bearings landed baseline (#640)
* fix: balance bearings landed defaults
* no-mistakes(document): Document balanced landed baseline
* no-mistakes: apply CI fixes
* fix: clarify captain-facing translation contract (#644)
* docs: clarify captain-facing translation contract
* no-mistakes(review): Restore runtime fallback mandate
* no-mistakes(document): Align Bearings translation wording
* fix(bin): make bootstrap output and nudges deterministic (#646)
* fix: make bootstrap nudges deterministic
* no-mistakes(review): Honor state override for bootstrap nudges
* no-mistakes(review): Update benign bootstrap documentation labels
* no-mistakes(review): Validate bootstrap nudge retry markers
* no-mistakes(document): Align bootstrap nudge documentation
* no-mistakes: apply CI fixes
* docs(secondmate-provisioning): clarify concise registry ownership (#649)
* Clarify concise secondmate registry contract
* no-mistakes(review): Expand secondmate registry boilerplate coverage
* no-mistakes(document): Point route docs to owner
* fix(bin): strip quoted blocked_by values during decision hold resolve (#654)
* fix(bin): strip quotes on blocked_by in decision-hold resolve
tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary
membership test only matched middle elements. Strip surrounding quotes
before matching so first and last hold ids resolve correctly.
* no-mistakes(document): Refresh decision-hold regression evidence
* feat(secondmate): inherit shared captain preferences (#656)
* feat(secondmate): inherit shared captain preferences
* no-mistakes(review): Honor shared captain data overrides
* no-mistakes(review): Honor bootstrap data override registry
* no-mistakes(document): Refresh shared inheritance docs
* no-mistakes(document): Clarify inherited local-material docs
* feat: gate local agent secret injection (#658)
* feat(spawn): gate local agent secret injection
* fix(spawn): align final Keychain slot
* no-mistakes: apply CI fixes
* test: isolate Herdr autodetect smoke sessions (#662)
* test: isolate herdr autodetect smoke session
* no-mistakes(review): Restored autodetect smoke gate bypass
* no-mistakes(test): Harden Herdr lab provisioning
* no-mistakes(document): Refresh Herdr lab docs
* docs: adopt under way for active work (#666)
* Revert "feat: gate local agent secret injection (#658)" (#668)
This reverts commit c27135cd9d35bc4c237d49b3b374da31fbd52eef.
* fix(pi): distinguish stale locks when arming watcher (#681)
* fix(pi): distinguish stale locks when arming watcher
* no-mistakes(test): Stabilize watcher extension async waits
* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate house vocabulary (#685)
* fix: accept secondmate as house vocabulary
* no-mistakes(test): Update captain vocabulary contract test
* no-mistakes(document): Align secondmate documentation vocabulary
* fix(bin): parse handoff homes after registry parentheticals (#686)
* fix: parse secondmate home after pre-field parentheses
Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.
* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges (#687)
* feat: add native session-start nudges
* no-mistakes(document): Document nudge script inventory
* docs: call built-in defaults the firstmate repo, not template (#688)
Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
* fix(bin): repair fm-brief.sh parse error and harden set -u array expansion (#205)
* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn
Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(brief): harden fm-brief regression coverage for parse and scaffolds
Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the
Co-authored-by: Cursor <cursoragent@cursor.com>
#166 apostrophe regression cannot return unnoticed.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(bin): keep watcher supervision continuous across child cycles (#693)
* fix: make watcher supervision continuous
* no-mistakes(review): Bound watcher retries and log attached signals
* no-mistakes(review): Add bounded successor-recovery wake fallbacks
* no-mistakes(review): Prevent overlapping successor-arm retries
* no-mistakes(review): Resume supervision after late arm closes
* no-mistakes(review): Bind OpenCode recovery to attempted arm
* no-mistakes(test): Synchronize peer beacon regression fixture
* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures
* no-mistakes(document): Captain: document watcher successor protocol behavior
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* no-mistakes: apply CI fixes
* fix: fetch current PR head for review diffs (#722)
* fix: always fetch PR head for review diffs
Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.
* no-mistakes(test): Isolate session-start nudge tests from gate state
* no-mistakes(document): Correct review-diff documentation
* docs: resolve five contract contradictions across AGENTS.md, README, and skills (#736)
* docs: resolve five contract contradictions
* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): correct Grok exit guidance (#742)
* docs(harness): reverify grok exit command
* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for parked crew (#743)
* fix(watcher): bound stale wakes for exited paused crew
* no-mistakes(review): Gate pause suppression on confirmed agent death
* no-mistakes(test): Fixed stale pause cadence
* no-mistakes(document): Document dead-agent hold cadence
* fix(supervision): distinguish ordinary wakes from recovery (#744)
* fix(supervision): distinguish ordinary wakes from repair
* no-mistakes(review): Make passive guard follow-ups recovery-only
* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes (#745)
* fix(x-mode): dedupe pending mention wakes
* no-mistakes(review): fix x-poll claim error deduplication
* no-mistakes(review): separate claim diagnostics from relay recovery
* no-mistakes(document): Document X-mode once-only mention wakes
* feat(wake): enrich drained signals with bounded status context (#747)
* feat(wake): enrich drained signal context
* no-mistakes(review): Bound wake enrichment reads
* no-mistakes(document): Document wake-drain annotations
* docs(wake): explain at-least-once drain boundary
* no-mistakes(review): Prevent symlink races in wake annotations
* no-mistakes(review): Exercise wake symlink race regression
* test: document intentional AFK marker subprocesses
* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces (#784)
* feat(herdr): add optional presentation spaces
* no-mistakes(review): Harden Herdr projection creation and spawn serialization
* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission
* no-mistakes(test): Correct stale Orca metadata failure fixture
* no-mistakes(document): Document Herdr presentation projection accurately
* fix(send): treat opencode busy-queued composer state as submitted (#775)
* fix(send): treat opencode busy-queued composer state as submitted
When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row. The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.
Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted). On an idle pane, keep returning "pending"
(genuine swallow detection preserved).
Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pane + composer clears on first Enter -> empty (existing path)
* docs: document busy-queued Enter exception across backend docs and skills
Add explanatory comments and backend documentation for the
busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn
but keeps typed text in composer until the turn ends):
- bin/fm-tmux-lib.sh: document the busy-aware fallback in the
file header and above fm_tmux_submit_enter_core
- .agents/skills/afk/SKILL.md: daemon-facing policy note
- .agents/skills/harness-adapters/SKILL.md: harness-specific fact
- docs/tmux-backend.md: submit-acknowledgement section with the
busy-queue exception
- docs/herdr-backend.md: record the known gap
- docs/architecture.md: cross-reference in the daemon section
* test(tmux): fix SC2181 and make busy-submit test executable
* fix(spawn): require two stable reads before accepting worktree path (#765)
* fix(spawn): require two stable reads before accepting worktree path
The treehouse-get worktree-detection loop in fm-spawn.sh accepted the
first pane_current_path read that differed from the project path, but
on some tmux/WSL setups a brand-new window transiently reports a
stale-but-real path before the pane actually settles into the
worktree. Since that stale path is itself a real, distinct git
checkout, it also passes validate_spawn_worktree's isolation check,
so the loop silently recorded the wrong worktree in state/<id>.meta
(and, for claude harness spawns, installed the turn-end hook there
too).
Require two consecutive polls to agree on the same non-project path
before accepting it, using the existing inter-poll sleep as the
confirmation gap so an already-settled pane isn't slowed down by an
extra cycle.
* fix(tests): drop unused CASE_DIR read in worktree-settle test
ShellCheck SC2034: CASE_DIR is split out of the case record but
never referenced; discard it with _ instead.
---------
Co-authored-by: Freudator86 <tim@allesknut.de>
* fix(bin): make watcher process identity immune to Linux wall-clock changes (#752)
* fix(watcher): stabilize Linux process identity
* no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix: prevent AFK idle stalls and stale run attribution (#758)
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state
Stop free-text tokens like "merged" from promoting nonterminal working: lines
to captain-relevant, so AFK no longer permanently suppresses idle recovery.
Defend wedge aging independently for nonterminal progress verbs, bind
no-mistakes current-state attribution to code identity (not branch alone),
and mark setup-complete as nonterminal in the ship brief scaffold.
* no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution
* no-mistakes(document): Document current-code-bound run attribution
* no-mistakes(test): Wait for stable Herdr shell readiness
* no-mistakes(test): Make Herdr and watcher readiness tests deterministic
* no-mistakes(test): Make tmux capture and watcher lifecycle deterministic
* no-mistakes(document): Document corrected supervision contracts
* fix(bin): allow safe teardown during watcher recovery (#750)
* fix: allow safe teardown during watcher recovery
* no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code
* no-mistakes(document): Sync continuity-gate docs to allow teardown recovery
* test: mark dynamic teardown fixture literal
* no-mistakes(document): docs: add teardown to continuity gate allow list
* feat(herdr): order presentation spaces while preserving focus (#790)
* feat(herdr): order presentation worker spaces
* fix(herdr): preserve focus during projected cleanup
* no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs
* no-mistakes(review): Serialize Herdr aborts with guarded focus regressions
* no-mistakes(review): Fall back flat when Herdr serialization is unavailable
* no-mistakes(test): Stabilize watcher startup and AFK handoff tests
* no-mistakes(document): Correct Herdr ordering and focus documentation
* fix(bin): send literal config reread nudges after pushes (#809)
* Send literal config reread after inherited config push
When declared inherited config changes under an already-running secondmate,
build a per-home instruction from validated destination post-write bytes and
deliver it on the routed secondmate path. Unchanged config sends nothing;
ABSENT represents removal; captain-shared is never inlined. Covers mid-session
config-push and the locked bootstrap convergence path without hardening spawn
against deliberate runtime choice.
* no-mistakes(review): Fix config reread framing, partial propagation, and respawn order
* no-mistakes(review): Send config rereads via durable single-line pointers
* no-mistakes(review): Make failed config rereads retryable
* no-mistakes(review): Make config reread retries generation-safe
* no-mistakes(review): Make config rereads durable and ordered
* no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode
* no-mistakes(review): Retain write retries and quarantine stale respawn generations
* no-mistakes(review): Preserve exact config reread retries and delivery order
* no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning
* no-mistakes(document):…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Slice 2 of the captain-ordered high-star documentation redundancy cleanup wave (firstmate scope), approved by the captain with docs/configuration.md as the declared shared owner: a docs-only consolidation establishing ONE owner for the three universal backend contracts the audit found copy-pasted across the five backend guides (the #372/#411/#332 five-doc re-edit class). (1) docs/configuration.md becomes the declared single owner, each with an explicit one-line ownership sentence so the Document step and humans know where the fact lives: the universal toolchain list (Toolchain section, now explicitly including git and gh with GitHub auth per the prior run's review finding, fixed between runs after that run failed terminally on the fix agent's usage limit), the task-selector vocabulary (Runtime backend section, the five selector-resolution sentences), and the tasks-axi compatibility definition (Backlog backend section). The Toolchain parenthetical defers tasks-axi version gates to the Backlog backend owner instead of restating them, and it absorbed the per-tool purpose clauses (treehouse pools worktrees, no-mistakes runs the pipeline, the axi tools' roles) that previously existed ONLY in the tmux guide's prerequisites - unique content deliberately moved into the owner rather than lost. (2) The five backend guides (tmux, herdr, zellij, orca, cmux) replace their verbatim universal-requirements parentheticals with a pointer to the owner plus only backend-specific items: herdr/zellij/cmux keep their jq and version-floor bullets, orca keeps its node-for-JSON-parsing bullet and the minus-tmux/treehouse exemption, tmux keeps its own install line. (3) Selector-vocabulary restatements in the zellij and cmux guides become pointers; architecture.md's and configuration.md's own trailing partial copies of the selector vocabulary are dropped. (4) CONTRIBUTING.md's 'Compatible means...' sentence becomes a pointer to the owner. (5) Two near-verbatim restatements of orca's bootstrap toolchain behavior (configuration.md carried it twice internally, the orca guide a third time) collapse into the single Toolchain copy. Deliberately untouched, per the wave contract: all empirical verification records in the backend guides, backend-specific setup/behavior/target-string sections, each guide's own backend-selection instructions (backend-specific user guidance, not a universal contract), and AGENTS.md plus the skills (those are the approved slice 3). Also deliberately deferred and documented rather than silently omitted: each guide's internal Setup-vs-Status selection-text duplication, which the audit did not name. PR body accounting (wave requirement): always-loaded AGENTS.md unchanged (slice 3 owns that reduction); full copies removed per contract: universal prereq list 6 copies -> 1 owner + 5 pointers, tasks-axi compatibility definition 7 copies -> 1 owner + pointers, selector vocabulary 4 prose copies -> 1 owner + 2 pointers, orca bootstrap-toolchain behavior 3 copies -> 1; files reduced to pointers: none removed, 7 files carry pointer lines instead of full copies; discoverability proof: every pointer names the owning file and section in the repo's established quoted-section link style, and the grep sweep shows zero remaining 'same universal requirements' phrasings or version-gate restatements outside the owner; checks run: full 59-script behavior suite green, repo-wide shellcheck green, CLAUDE.md and .claude/skills symlink assertions green, cross-reference grep sweep for all three contracts.
What Changed
tasks-axicompatibility contracts indocs/configuration.md, replacing duplicated backend-guide and contributor prose with pointers.Risk Assessment
✅ Low: Captain, the change is bounded documentation consolidation plus consistent Git prerequisite detection, with no material risk found.
Testing
The provided full-suite baseline was green; I reran the focused bootstrap behavior test and captured reader-facing documentation evidence showing
docs/configuration.mduniquely owns the requested contracts, with valid guide pointers and preserved backend-specific prerequisites.Evidence: Documentation contract validation
Reader-facing excerpts plus all ownership, pointer, duplicate-sweep, and relative-link checks passed.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 3 issues found → auto-fixed ✅
docs/configuration.md:171- The new toolchain owner says bootstrap detects and can installgit, butbin/fm-bootstrap.shneither probes nor installs it. Add matching bootstrap support or describe Git as an external prerequisite.docs/configuration.md:18- The new single-owner claim for tasks-axi compatibility is broader than the repository state: AGENTS.md still restates its compatibility gates. Either consolidate those copies or scope the ownership claim to this documentation subset.docs/configuration.md:53- The task-selector ownership claim is broader than the repository state: docs/herdr-backend.md and AGENTS.md still describe exact-id, legacy-label, and explicit-target resolution. Either replace them with pointers or narrow this ownership assertion.🔧 Fix: Detect Git in bootstrap toolchain
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"Baseline supplied as passed:command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"bash tests/fm-bootstrap.test.shManual reader-facing documentation owner, pointer, duplicate-sweep, and relative-link verification recorded in the evidence transcript.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.