Merge upstream firstmate up to 71db4be (excludes #593) - #6
Merged
Merged
Conversation
* fix: gitignore the secondmate home marker bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into every seeded secondmate home. A secondmate home is a worktree of the firstmate repo, so any plain `git status --porcelain` dirtiness check counted the untracked marker and the home read as dirty forever: fleet-sync reported it STUCK and the local fast-forward convergence sweeps risked leaving it stale on firstmate updates. Add .fm-secondmate-home to the tracked .gitignore so the marker is invisible to every dirtiness check uniformly, without weakening fleet-sync's deliberate untracked-counting for project clones. Convergence chicken-and-egg: existing homes predate the fix and it only arrives by fast-forward. The already-present marker-tolerant ff-skip (ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate, and spawn pre-launch) advances such a home past the fix commit, after which .gitignore takes over - no hand intervention. Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home reading clean, an existing marker-only home converging then reading clean, and a genuinely dirty home still skipping. * no-mistakes(review): Captain: document standalone-clone update path * no-mistakes(document): Document secondmate marker migration
* fix(composer): stop reading dead-shell prompts as empty agent composers Consolidate composer empty/pending/unknown classification into one shared owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This replaces four drifting copies of the glyph decision. Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now classified unknown (a dead shell, unsafe for injection), not empty. It is only empty inside a bordered composer box (the harness's own prompt). Agent glyphs ❯ (claude) and › (codex) read empty either way. The away-mode injector (inject_msg) now requires an affirmatively-empty composer, deferring on pending or unknown, so an escalation can never be typed into (or executed by) a pane whose agent exited to its login shell. Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner; per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay green. * no-mistakes(review): Captain: harden composer safety checks * no-mistakes(test): Stabilize Herdr prune safety setup * no-mistakes(document): Document composer injection safety * no-mistakes(lint): Clean composer safety lint * no-mistakes: apply CI fixes
* feat(watcher): add paused/awaiting-external crew state A crew (or firstmate steering it) can declare a deliberate wait on a known external dependency with a paused: <reason> status. Both the always-on watcher and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh vocabulary instead of tripping the possible-wedge stale escalation, and re-surface it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused distinctly. A crew that goes idle without declaring a pause classifies exactly as before. Docs and brief scaffold state lists updated; tests colocated. * no-mistakes(review): Captain: fix paused-state transitions * init * no-mistakes(review): Captain: fix paused-state supervision transitions * no-mistakes(review): Captain: fix paused supervision handoffs * no-mistakes(review): Reconcile paused supervision markers * no-mistakes(review): Captain: prioritize paused states over captain relevance * no-mistakes(review): Captain: preserve paused-working wedge timer * no-mistakes(review): Captain: honor configured pause verb in briefs * no-mistakes(test): Captain: fix AFK paused watcher handoff * no-mistakes(document): Document declared external waits * no-mistakes(lint): Clean paused-state lint --------- Co-authored-by: fmtest <fmtest@example.invalid>
* fix(x-mode): make follow-up platform splitting immune to link ordering A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280 chars because fm-x-link only learned the platform from the inbox payload, and the fmx-respond ack path can drain that inbox file before the task is linked. A link recorded after cleanup silently lost the platform and the splitter defaulted to the X 280-char budget. Make platform resolution ordering-proof: - fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a new fmx_request_relay_context helper (POST /connector/request-context) when neither the inbox payload nor carry flags carry it. The request_id survives the inbox drain, so a post-cleanup link still learns the right split budget. Best-effort: no token/curl or a non-2xx relay degrades to the loud warning below rather than a silent X default. - fm-x-link warns loudly when no platform source resolves, so the loss is never silent. - The fmx-respond procedure now orders link-before-inbox-cleanup so the fast local path stays correct without a relay round-trip. Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE message even when linked after inbox cleanup, and an unresolvable platform warns loudly instead of splitting silently. docs/configuration.md documents the request-context lookup. The relay endpoint is the companion durable change (see done status); until it ships, the link-before-cleanup reorder keeps the normal path correct. * no-mistakes(document): Document X-mode platform recovery
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor Away-mode injection wedged all night on the primary claude-on-herdr pane: the herdr composer classifier never stripped generic dim ghost text (only a narrow codex bold-wrapped byte-pattern check), so claude's rotating prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI pane read preserves - read as real pending input and every escalation deferred (6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s). Consolidate ghost extraction into one fleet-wide ANSI-aware owner, fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX, default 128, dark-theme assumption). Both ANSI-capable backends route through it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a dark box border dropped with the ghost does not lose the composer shape. This also closes the documented grok TRUECOLOR placeholder gap by the same mechanism (harness-adapters skill note updated). Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle) and the incident write-up are in docs/herdr-backend.md; deterministic regressions feed the exact captured bytes through the real classifiers (tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text (never a realistic real-input color) are corrected to a bright 38;2;224;222;244, preserving the truecolor payload-skip parser intent. * no-mistakes(review): Preserve dark shell prompt safety * no-mistakes(review): Harden erased shell prompt classification * no-mistakes(document): Document shared composer ghost extraction * no-mistakes(lint): Normalize tmux comment punctuation
…enguid#432) * fix(session-start): isolate harness env markers in suite runner Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in run_session_start so ambient interactive shells cannot override the suite's fake ps harness (local-vs-CI split on the pi supervision case). * no-mistakes(document): Correct Pi marker documentation
…nchenguid#435) * fix(teardown): retry treehouse return on transient index.lock Killed crew git ops can leave a short-lived worktree index.lock that makes treehouse return fail. Retry on that error signature with a bounded wait (env-overridable), never force-delete a live lock, and only then fall back to the existing provably-stale cleanup path. * no-mistakes(review): Harden teardown retry configuration * no-mistakes(document): Document teardown index-lock retry behavior * no-mistakes(lint): Fix empty shell variable assignments
* docs: de-feature the scripts.md and CONTRIBUTING test inventories Slice 1 of the documentation redundancy cleanup wave (firstmate scope). docs/scripts.md: every row is now one purpose clause; script headers are the declared owner of behavior, flags, and contracts. Coverage stays 61/61 scripts; bytes drop 19,922 -> 7,958. CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors discover tests by listing tests/*.test.sh and reading each script's own header, and gated tests print their own skip gates. The run commands, symlink assertions, and watcher smoke line are unchanged. Lines drop 135 -> 84 (18,797 -> 7,831 bytes). Two facts that existed only as inventory rows moved into their owners' headers first: fm-brief.sh's paused-vs-blocked scaffold distinction and fm-session-start.sh's Pi extension-loaded check. No instruction-surface or behavior change; AGENTS.md untouched. * no-mistakes(review): Captain, fix brief help and Grok test discovery * no-mistakes(review): Captain: document Grok lock-holder test coverage
* docs: consolidate universal backend contracts into configuration.md Slice 2 of the documentation redundancy cleanup wave (firstmate scope). docs/configuration.md is now the declared single owner of three universal contracts, each with an explicit ownership sentence: - the universal toolchain list (Toolchain), now also carrying the per-tool purpose clauses that previously lived only in the tmux guide; - the task-selector vocabulary (Runtime backend); - the tasks-axi compatibility definition (Backlog backend). The five backend guides' prerequisites replace their verbatim universal-requirements parentheticals (5 full copies) with a pointer plus only backend-specific items; zellij/cmux selector restatements and architecture.md's partial copy become pointers or are dropped; CONTRIBUTING's compatibility sentence becomes a pointer; two near-verbatim orca-bootstrap restatements (configuration.md Runtime backend, orca guide) collapse into the Toolchain owner copy. Backend-specific setup, behavior, target-string shapes, and every empirical verification record are untouched. AGENTS.md untouched (slice 3). * docs: include git and GitHub auth in the toolchain owner list The review flagged that the new universal-toolchain owner omitted git and GitHub authentication while every backend guide now defers its prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real universal requirements. * no-mistakes(review): Detect Git in bootstrap toolchain * no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): backend-independent active alert for the wedge alarm When away-mode injection wedges past max-defer, inject_wedge_alarm only actively signalled via the tmux status-line, which is skipped on non-tmux backends. A wedged claude-on-herdr primary left only the passive state/.subsuper-inject-wedged marker (2026-07-10 overnight incident). Add a config-gated active alert (config/wedge-alarm, local/gitignored; FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and its status-line is unreadable: an OS-level macOS notification (osascript), a herdr notification, or a captain-supplied command. Default-on (auto) so the alarm is never silent; each channel best-effort, degrading to the next and never crashing the daemon loop. The tmux flash and durable marker stay. The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the daemon is sourced (only tests do this; production execs it) the seam defaults to "discard", and tests/wake-helpers.sh points it at a recorder, so it is structurally impossible for any test to post a real notification. Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see docs/wedge-alarm.md. * no-mistakes(review): Bound wedge alarm notifier execution * no-mistakes(review): Captain: harden wedge alarm notifier safety * no-mistakes(review): Captain: harden wedge alarm test notifier isolation * no-mistakes(review): Captain: harden wedge alarm throttling * no-mistakes(review): Redact wedge alarm directive logs * no-mistakes(review): Harden wedge alarm notifier safety * no-mistakes(review): Track notifier process groups through cleanup * no-mistakes(document): Document wedge-alarm active alert behavior
* docs(agents): extract conditional AGENTS.md material to owned homes Slice 3 of the documentation redundancy cleanup wave (firstmate scope): the always-loaded instruction surface drops from 941 lines / 116,733 bytes (~29k tokens per session per fleet member) to 785 / 91,353 (~22.8k tokens), moving only audit-identified conditional and situational material while preserving every load-bearing invariant at its trigger point via the inline-stub pattern. Moves, each to one declared owner plus an inline stub: - section 3's bootstrap output-line handbook (~44 lines) -> new agent-only bootstrap-diagnostics skill, added to the section 13 trigger index; the detect-consent-install rule and the do-not-dispatch gate stay inline as safety-critical. - section 4's crew-dispatch JSON schema and field semantics -> docs/configuration.md 'Crew dispatch profiles' (pointer direction flipped); the intake procedure, precedence, backstop, and never-select-unverified rules stay inline. - section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh header (now the declared owner; usage() converted to the dynamic header extraction pattern PR kunchenguid#438 established for fm-brief.sh). - section 7's spawn resolution narrative and example sprawl -> bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is- a-blocker rule, and post-spawn duties stay inline. - section 7's teardown landed-work mechanics -> bin/fm-teardown.sh header (section 1's containment pointer retargeted); the fork benign case and never-force rule stay inline. - section 8's watcher classification narrative -> docs/architecture.md 'Event-driven supervision' (already the owner); every operative rule (one live cycle, no turn ends blind, drain first, wake ladder, never-pkill, guard responses) stays inline. - sections 3/4/6/7 secondmate sync, propagation, schema, and handoff restatements -> secondmate-provisioning skill, now the declared owner including the literal-file inheritance nuance. - section 14's X-mode cadence mechanism -> docs/configuration.md 'X mode (.env)', closing issue kunchenguid#363; activation semantics, the fmx-respond trigger, and the terminal-wake final-follow-up duty stay inline. CLAUDE.md stays a symlink; no behavior or test change. * no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close the last/selected workspace in a window at teardown cmux keeps every window at >=1 workspace, so close-workspace on the only workspace in a window silently no-ops (returns OK, workspace stays), and a window holding a live session cannot be closed over the control socket. That left a selected task workspace open at teardown (the last workspace in a window is always the selected one). Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill create a throwaway default sibling in the target's window before closing when the target is the last workspace there, so the close lands; the window keeps a fresh default workspace (cmux's own "closed the last tab" outcome). Non-last teardown closes directly, as before. Cover both kill branches plus the helper with fake-CLI unit tests, add a real-cmux window/count detection smoke assertion, and record the empirical evidence in docs/cmux-backend.md. * no-mistakes(review): Derive cmux count from membership snapshot * no-mistakes(document): Document cmux last-workspace teardown behavior
…d#453) * fix(fleet-sync): recover from an orphaned packed-refs.lock A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's timed-out fleet-sync kill or teardown's process kills - leaves a lock that makes the next sync's fetch fail with "Unable to create '...packed-refs.lock': File exists", leaving the clone unsynced. On that signature only, fm-fleet-sync.sh now retries the fetch with a bounded wait (transient locks self-clear), then removes the lock and retries once more ONLY when it is provably stale: still present, mtime age past a threshold, and no lsof holder of the lock file or of the clone worktree itself (a live git keeps that as its cwd even in the window after it closes the lock and before it exits). A live lock, a missing lsof, any failed check, or any other fetch failure keeps today's behavior. Every wait/retry/removal prints to stderr, and a successful recovery also prints one "recovered:" summary to stdout so a session-start refresh - which discards fleet-sync stderr and relays only stdout - still surfaces it. The shared "is this git lock provably abandoned?" proof is extracted into bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/. * no-mistakes(review): Captain, remove obsolete teardown wake dependency * no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): immediate blocked-state escalation via native events.subscribe push Fold herdr's native pane.agent_status_changed stream into the single watcher so a crew entering blocked wakes its supervisor sub-second (measured 0.129s) instead of after the ~240s stale-pane wedge timer. - bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape plus the single-owner status->action policy table (blocked=actionable, working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling). - bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe subscriber over one connection for all this home's herdr panes, subscribing to ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe marker and a reconnect level-reconcile. Version/schema capability gate. - bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so the watcher stays backend-agnostic and the shape+policy are reusable. - bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal wait primitive (replacing the blind sleep POLL for push-capable homes), behind a source guard so the splice is unit-testable; secondmate/paused exemptions; map pane->window->task and enqueue a stale wake. No second watcher process; the single-cycle invariant and every guard/beacon/turn-end mechanism are unchanged. - Polling stays the permanent fail-closed backstop: below-capability, subscribe failure, and repeated runtime failures all degrade to sleep. - Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/ fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events) plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries the dated evidence and retires the old gap note. * no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs * no-mistakes(review): Captain, commit markers after wake and reuse capability cache * no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs * no-mistakes(review): Captain, subscribe before Herdr reconciliation * no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe * no-mistakes(test): Captain: include lock library in teardown fixture * no-mistakes(document): Captain: document Herdr immediate blocked escalation * fix: clarify shellcheck conditionals
* feat(bearings): deterministic bearings snapshot + durable decision model Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR discovery and checks happen only under --include-prs, which fails soft. Every dropped surface is marked in omitted[] with the flag that reveals it, and the prs: line states when checks were not requested, so absence is never silent. Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib gains status_open_decisions, the one authoritative keyed open/resolved fold over the whole status stream: needs-decision/blocked opens a keyed entry, only an explicit keyed resolution (or, for run-backed tasks, run-step advancement) closes it, so a later unrelated done/paused can no longer mask a still-open captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives pending_decision/blocked_event from it; the canonical schema stays complete. Point the /bearings skill at the one command; add the resolved: writer line to ship, scout, and secondmate briefs. Register the script and add regression tests for the output bound, TOON/JSON parity, local-only default, opt-in PR fetch, partial-failure degradation, decision durability, and report pointers. * fix(bearings): completed scout report is a pointer, not a pending decision A completed scout that raised a needs-decision and then finished (done) without a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103 case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh cleared a stale decision only for a live run-step/pane activity read, so a terminal task whose current state is read from the status log (a scout or ship that reached done/failed) never cleared its stale, never-keyed-resolved needs-decision, and it lingered as pending. The open-decision set is still derived purely from the keyed fold - never from a report body or decision-like prose - and reconciled against the crew lifecycle. Extend that reconciliation so a terminal done/failed state on a single-owner task (scout or ship), whose deliverable is its report or PR, also clears the set; a completed scout now surfaces only as a report pointer. Secondmates are excluded from the terminal clear (persistent, multiplexed stream), which keeps the unrelated-event masking fix intact. Add regression tests: a completed scout with decision-like report prose is a pointer not pending (canonical + end-to-end), and a scout still parked at a decision stays pending so the terminal clear never over-fires. * no-mistakes(review): Captain, preserve keyed decisions across shared status parsing * no-mistakes(review): Captain, close blockers and harden keyed decision parsing * no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout * no-mistakes(review): Captain, bound bearings sections and fail closed * no-mistakes(review): Captain, disclose capped per-repository PR results * no-mistakes(document): Refresh bearings documentation and status contracts * fix(bearings): avoid ambiguous worktree guard
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity Firstmate PRs passed local no-mistakes validation but failed CI's "Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034). Two divergences caused it: 1. The no-mistakes gate had no commands.lint, so its lint step never ran the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that CI runs. Confirmed from state.sqlite: the lint step_result recorded findings:null with no lint agent invocation. 2. CI's shellcheck floated with the runner image while local ran a newer build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck rejected an SC2015 that the newer local one no longer emits. Establish bin/fm-lint.sh as the single owner of the lint definition: the file set, the config, and the pinned shellcheck version (0.11.0, printed via --required-version). Both CI (.github/workflows/ci.yml) and the no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs the exact version it names and logs the resolved version, and fm-lint.sh refuses to lint under any other version. This is not a CI relaxation: it adopts shellcheck 0.11.0's rule set consistently, dropping only the upstream-retired, false-positive-prone SC2015; default severity and every still-supported finding stay enforced (no severity downgrade, no excludes). tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs and logs the pinned version, that the owner refuses a non-pinned shellcheck, and that it rejects a real lint defect the old no-op gate passed. * no-mistakes(review): Captain, harden deterministic ShellCheck parity * no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: add cd-guard PreToolUse seatbelt for the primary shell A stray persistent top-level `cd projects/<clone>` in the primary firstmate shell relocates the shell, so a later firstmate-owned command (a backlog write, an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the home. The cd-guard denies exactly that command shape before it runs, across all five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt. - bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate lexer; that file's CLI now runs only when invoked directly). - bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness output rendering, and primary-checkout scoping - fires in a secondmate's own primary session, inert in crew/scout child worktrees and non-firstmate repos. - Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents; per-harness hooks only call the owner. - Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd, and command cd). Allows git -C, subshell / bash -c / env -C / make -C / find -execdir, pipeline and background forms, and cd-as-data. Fails open on malformed input; agent-mistake threat model. - tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix, end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring. - docs/cd-guard.md: full contract plus live validation (claude, codex, opencode, pi blocked end-to-end; grok live run blocked by an API balance limit, with mechanism parity and deterministic coverage recorded). * no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage * no-mistakes(review): Captain, allow path-qualified command wrappers * no-mistakes(review): Captain, allow non-executing command queries * no-mistakes(test): Captain, clarify cd-guard safe-path remediation * docs: clarify cd guard guidance * no-mistakes(document): Clarify cd-guard safe target guidance
…kunchenguid#267) Crews must never stop, restart, or update the shared no-mistakes daemon since one instance serves every firstmate lane/home; a restart kills other lanes' in-flight pipeline runs and forces expensive re-runs. Encodes this as a new numbered rule in both the ship-task and scout-task brief scaffolds. Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render /bearings skill (one owner of the chat-response format): mandate the four always-present chat sections - Captain's Call, Recently Landed, Underway, Charted Next - each with an explicit empty-state sentence, no At Anchor, materially shorter than and linking to the report file. Resolves the ambiguous Check first / Decisions pending split into one strict captain-action section. fm-crew-state: the log fallback derives current state only from a real run-state verb, so a trailing decision-closing resolved: event no longer renders a healthy idle crew (typically a secondmate) as unknown with the resolution prose as its detail. The keyed-decision contract in fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner. fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done records from registered secondmate homes, reusing the single backlog parser and the one secondmate-home enumerator (meta home= with data/secondmates.md fallback); no network, per-home capped. fm-bearings-snapshot: landed now merges main-home Done with the secondmate roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure (also fixing the previously-silent landed truncation); --all-landed reveals the full set. tests: resolved-event state render, secondmate landed aggregation with caps and omitted[] disclosure, Captain's Call anti-leak, and the four-section contract. * no-mistakes(review): Captain, ensure bearings reveals all landed work * no-mistakes(document): Document bearings accuracy contracts
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle Away-mode entry left "make the daemon a tracked background terminal" to the operator; on a pi/herdr primary that meant splitting the captain's active pane, which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches the daemon in a non-visible tracked terminal per backend (herdr dedicated --no-focus workspace, detached tmux session), never a split, pins the captain pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal id, and tears it down or reconciles a leaked one by that id. No shell &. Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared with the daemon (one owner). Fix the stale subsuper-artifact leak: clear the prior away session's delivery cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon before clearing state/.afk so its shutdown flush runs instead of being a no-op. Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md, docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md. * no-mistakes(review): Captain, serialize AFK launcher lifecycle safely * no-mistakes(review): Captain, harden AFK launcher lifecycle races * no-mistakes(review): Captain, ensure AFK daemon launch readiness * no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown * no-mistakes(review): Captain, preserve AFK reconciliation records uniformly * no-mistakes(review): Captain, harden AFK recovery state durability * no-mistakes(review): Captain, harden AFK tmux ownership checks * no-mistakes(review): Captain, simplify AFK lifecycle failure handling * no-mistakes(review): Captain, require confirmed AFK daemon shutdown * no-mistakes(review): Captain, confirm AFK exit by process identity * no-mistakes(document): Align AFK launcher lifecycle documentation * no-mistakes: apply CI fixes
…uid#518) * feat: contain no-mistakes gate agents from driving the fleet Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/ fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE is set, and via an unspoofable git-common-dir backstop when invoked from a no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker unset. A normal firstmate session has neither signal and is unaffected. Set disable_project_settings: true in the tracked .no-mistakes.yaml so the installed pipeline neutralizes gate agents' project instructions for this repo (trusted-only, honored from the default branch). firstmate's own suite runs from a gate worktree during validation, so the shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated tests/fm-gate-refuse.test.sh strips it to verify real refusal. * no-mistakes(review): Captain, refuse empty no-mistakes gate markers * no-mistakes(document): Document no-mistakes gate authority boundary
…uid#505) * fix: guard secondmate own-home turn ends Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the 'no turn ends blind' backstop fires in a secondmate's own primary session, matching the cd-guard's scope: the own home is guarded, child crew/scout worktrees stay exempt via the retained git-dir/git-common-dir test. This was pure scoping from the guard's primary-only origin and guarded against no secondmate-specific hazard. Add secondmate regression tests (blind-turn block, idle-by-default, stop_hook_active loop guard, deferred-death recovery loop, child-worktree exemption) and record the autonomous background-notify re-invoke measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md. * no-mistakes(document): Correct secondmate guard documentation, captain * fix: force-include marked secondmate homes in turn-end guard The prior remove-only form (just deleting the .fm-secondmate-home check) left the DEFAULT secondmate topology unguarded: a treehouse-leased home is a linked git worktree (git-dir != git-common-dir), which the retained git-dir exemption still skipped, so its own primary session could still end a turn blind. Invert the marker: a genuinely-marked home is force-included as a guarded primary (treehouse-leased linked OR git-cloned plain), and the git-dir exemption applies only to UNMARKED child worktrees. Marker validation (regular non-symlink file, non-empty id-token content) blocks a stray or empty marker from spoofing inclusion. Add real linked-worktree regression tests: a treehouse-leased LINKED secondmate home is guarded, a stray/empty marker stays exempt, and the unmarked child worktree stays exempt - the topology the plain git-init fixtures masked. Predicates, in-flight gate, and loop guard untouched. * fix: force ASCII collation in secondmate marker validation Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the [A-Za-z0-9._-] id allowlist matches under C collation, not the ambient locale - a locale-crafted non-ASCII marker id can no longer slip through the range match and spoof force-inclusion of a linked child worktree. Add a regression test proving a non-ASCII marker id is rejected and the linked worktree stays exempt. * no-mistakes(test): fix backend baseline gate-refusal dependency * no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap required-tool detection backend-aware Bootstrap demanded tmux and treehouse for every backend except orca, so a herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux. Required tools now follow the resolved backend via the single-owner fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux), and treehouse for session-provider-only backends (orca owns its worktree). The treehouse lease-support check is gated to backends that use treehouse. Adds install hints for herdr/zellij/cmux, regression tests for the full backend dependency matrix (herdr-without-tmux repro plus each boundary), and updates the authoritative Toolchain docs. * no-mistakes(review): Captain, prevent executing Herdr install guidance * no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics * no-mistakes(review): Captain, separate manual dependency remediation * no-mistakes(review): Captain, align bootstrap diagnostic consumers * no-mistakes(document): Align backend adapter dependency comments
…guid#520) * fix: recover X/Discord follow-up platform after inbox cleanup A milestone follow-up posted directly by request_id after the inbox was drained - and with no task link, because one persistent secondmate's single x_request slot collides across concurrent requests - resolved platform only from the local inbox, so a >280 Discord reply silently defaulted to the X 280-char budget and threaded as (1/2). - fm-x-poll records a durable per-request reply context (state/x-context/<rid>.json) at stash time, keyed by request_id so concurrent requests never overwrite each other; it survives inbox cleanup and restart. - fm-x-reply resolves platform/budget through registry -> inbox -> relay (the relay lookup confined to a live follow-up), recovering the original platform independent of task-link availability. - Fail-safe: a follow-up whose platform/budget cannot be authoritatively resolved and that would split is refused (exit 8) and held for retry, never wrongly split; fm-x-followup keeps the link on that exit. - fm-x-dismiss clears the durable context for a dismissed mention. Refactors reply-context extraction into a single owner and adds regression coverage for all four cases. * no-mistakes(review): Captain, fail closed on incomplete follow-up context * no-mistakes(review): Captain, bound X context registry retention * no-mistakes(review): Captain, align context retention with answer binding * no-mistakes(document): Align X follow-up context documentation * no-mistakes(document): Align durable X follow-up documentation
…id#533) * fix: preserve secondmate routing markers * no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends * no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift * no-mistakes(document): Refresh Herdr marker documentation
* fix: align grok effort docs and spawn with 0.2.99 ceiling grok 0.2.99 accepts only low|medium|high for --reasoning-effort and rejects both xhigh and max. Omit unsupported values on spawn, flag them in crew-dispatch validation, and update harness-adapters. * no-mistakes(test): Captain: refresh gotmp teardown fixture dependencies * no-mistakes(document): Clarify Grok effort documentation ownership
…#555) * fix: make bearings use secondmate home state * test: anonymize bearings fixtures * no-mistakes(review): Bound parent activity evidence scans, captain * no-mistakes(review): Preserve structured secondmate authority and bounds, captain * no-mistakes(review): Preserve registry completeness and child inventory, captain * no-mistakes(review): Reconcile parent evidence by verb and key, captain * no-mistakes(review): Treat unkeyed parent evidence as inconclusive, captain * no-mistakes(document): Document bearings local snapshot and PR opt-in * no-mistakes(lint): Fix fleet snapshot ShellCheck findings * no-mistakes: apply CI fixes
* fix: restore stock macOS snapshot parsing * no-mistakes(document): Clarify Linux gate and macOS CI coverage
…d#587) * fix: close away-mode blocker supervision gap * no-mistakes(review): Gate teardown retries and verify U+2063 dedupe * no-mistakes(test): Fail closed on incomplete Pi composer separators * no-mistakes(document): Document Pi composer recognition and return gating
* support Pi max thinking profiles * no-mistakes(review): Captain, allow Pi max dispatch profiles
…nguid#595) * Clarify validation response ownership * no-mistakes(document): Clarify yolo response ownership
* Add instruction owners foundation * no-mistakes(document): Refresh project-management owner pointers
…kunchenguid#626) * docs: compress firstmate operating contract * docs: make delivery rigor single-owner PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion. * no-mistakes(review): Honor configured merge authority across faster delivery paths * no-mistakes(document): Align docs with compressed operating contract
Takes 35 upstream commits (all CI-green), stopping deliberately at 71db4be rather than upstream/main. The one commit left behind, cd218f2 (kunchenguid#593, durable captain decision holds), is excluded for two independent reasons: its CI is red (the old-vs-new teardown conformance harness copies a fixed file list to a temp dir and omits the bin/fm-decision-hold.sh its own new teardown requires), and it adds an unresolved-decision gate to teardown while the teardown lifecycle is under validation here. Revisit later. Conflicts resolved on their merits: bin/fm-teardown.sh - the real one. Upstream 38086ea rewrote teardown_treehouse_return (bounded index.lock retries, output capture, error signature matching) in the same region as this fork's path fix (#5). Kept upstream's retry structure and threaded return_path through all three `treehouse return --force` call sites; the lock and git checks stay on $dir, which they resolve for themselves. Per docs/treehouse-path-contract.md: canonicalize firstmate-vs-firstmate comparisons, never canonicalize a path handed to treehouse. Upstream did not touch bin/fm-wake-lib.sh, so the internal-comparison half merged clean. AGENTS.md, CONTRIBUTING.md, docs/scripts.md - upstream restructured heavily (kunchenguid#619, kunchenguid#626, 8cd90fe), so took upstream's structure and re-applied this fork's content into it rather than defending the old layout. AGENTS.md keeps upstream's compressed contract plus the Crowsnest section, the fmc-respond and usage-monitor skill triggers, the quota-held intake classification, and the chat-mention check wake. Upstream retired AGENTS.md's layout tree and CONTRIBUTING.md's per-test listing in favor of the owning doc and each test's own header, so the fork lines that only existed inside those lists retire with the convention; docs/crowsnest.md and tests/fm-crowsnest.test.sh already own that detail. .gitignore, README.md - unions; deduped __pycache__/*.pyc, which upstream had added independently. Fork work preserved and verified: the treehouse path fix survives in both halves, .github/workflows/no-mistakes-required.yml stays deleted (upstream never touches it - permanent fork divergence), and the crowsnest and usage-monitor lines are intact. Verification: all 77 test scripts pass and bin/fm-lint.sh is clean on the pinned ShellCheck 0.11.0 that CI now installs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges upstream
kunchenguid/firstmateinto this fork, stopping deliberately at 71db4be, notupstream/main. That takes 35 upstream commits, all CI-green.Why not upstream/main
The one commit left behind is
cd218f2(kunchenguid#593, durable captain decision holds), which is 71db4be's direct child. Excluded for two independent reasons:bin/fm-decision-hold.shthat its new teardown requires.Worth revisiting once both settle.
The conflict that mattered:
bin/fm-teardown.shUpstream
38086earewroteteardown_treehouse_return— boundedindex.lockretries, stdout/stderr capture, error-signature matching — in the same function as this fork's path fix (#5).Resolved by keeping upstream's retry structure and threading
return_paththrough all threetreehouse return --forcecall sites. The lock and git checks deliberately stay on$dir, which git resolves for itself.This preserves the contract in
docs/treehouse-path-contract.md, which cuts both ways and is easy to "simplify" into a bug:fm_same_path)treehouse_recorded_path)No upstream change canonicalized a treehouse-bound path, so the two sides composed cleanly rather than fighting. Upstream does not touch
bin/fm-wake-lib.sh, so the internal-comparison half merged clean (verified, not assumed).Docs and contract files
Upstream restructured heavily (kunchenguid#619 instruction ownership, kunchenguid#626 contract compression,
8cd90fecentralize operating contracts). Per the task, I took upstream's structure and re-applied fork content into it rather than defending the old layout.fmc-respond/usage-monitorskill triggers, the quota-held intake classification, and thechat-mentioncheck wake — each rewritten in upstream's terse style.docs/crowsnest.mdandtests/fm-crowsnest.test.sh's header already own that detail, so nothing is lost.__pycache__/+*.pyc, which upstream had added independently at the top.Fork work preserved (verified, not assumed)
treehouse_recorded_pathin teardown,fm_same_pathin wake-lib. Both its tests pass.no-mistakes-required.ymlstays deleted. Confirmed upstream touches it nowhere between the merge base and 71db4be. Permanent fork divergence.fm-session-start.shshows no crowsnest refs, which is correct and not a regression — that content was always owned byfm-bootstrap.sh, which retains it.Verification
bin/fm-lint.shclean on the pinned ShellCheck 0.11.0. Note CI changed under this merge: lint is nowbin/fm-lint.shwith a pinned shellcheck (not a raw command), plus a new macOS stock-bash job.tests/fm-teardown.test.shandtests/fm-turnend-guard.test.shpass.kunchenguid#432 confirmed fixing a real bug we hit
tests/fm-session-start.test.shunder ambientCLAUDECODE=1, which is genuinely set in a Claude session:fm-send composer bug: NOT fixed (reported, not fixed here)
Tested against a real live Claude pane in an isolated throwaway tmux session, so the fleet was never touched. Upstream's composer rewrite (
0eaf293kunchenguid#429,a955a05kunchenguid#416) does not fix it:/help/helpvisibly executedSo
fm-send-queued-t7still reproduces after this merge, exactly as reported. Left alone per the task.