fix(wardroom): stop the preflight gate asserting a property of the machine - #3506
Closed
stoneevenson-biz wants to merge 56 commits into
Closed
stoneevenson-biz wants to merge 56 commits into
stoneevenson-biz wants to merge 56 commits into
Conversation
Ship and scout crewmate briefs now instruct delegating to subagents when useful: fan out for thinking, write in the single worktree yourself, escalate parallel-write work to firstmate.
gates/ (ledger.json + verify.sh shim to the global ledger CLI + rendered LEDGER.md), docs/ (context-watchdog spec + ADR-0001 on handoff-presence as the rehydrate trigger), and the gate-ledger definition-of-done clause in fm-brief.sh so crewmate briefs treat a green ledger as done for verifiable builds. Ledger currently reports g1/g2 red (stale fixtures) — fixed in a later commit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…OLE) The MEASURE→WATCH→REHYDRATE cycle that recycles a near-full session: - fm-ctx-statusline.sh / fm-ctx-stop-hook.sh write ctx-<key>.json sentinels. - fm-context-watch.sh polls them and, when a managed pane crosses its role threshold and is idle, runs the checkpoint→handoff-wait→/clear cycle (fm_ctx_fire_once); pure decision fns are sourceable + unit-tested. - fm-ctx-lib.sh holds the shared thresholds, role classification, window key, managed-scope, and inject cap. - fm-captain-bootstrap.sh rehydrates a /cleared pane from its handoff doc. Includes the FIRSTMATE_ROLE override in fm-ctx-lib.sh (fm_ctx_role) and fm-captain-bootstrap.sh: precedence FIRSTMATE_ROLE > FM_CTX_ROLE > cwd, only the validated captain|crew values force, anything else falls through unchanged. The Hermes captain wrapper depends on this signal. Gate tests g1..g5 included; g1/g2 fixtures are stale (planted sentinels omit managed:true, which the managed-scope gate now requires) and fail — fixed in the next commit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ession) The g1/g2 fixtures predate the managed-scope gate: their planted sentinels omitted "managed":true, so _ctx_eligible -> fm_ctx_sentinel_managed (a fail-closed gate that only fires on firstmate-owned panes) rejected them and fm_ctx_select returned empty. Add managed:true to both, matching the real measure-path sentinels and the g3/g4/g5 fixtures' intent. g2 additionally exports FM_TMUX_SESSION=$SESS so the fire-time _ctx_target_in_session re-confirmation accepts the disposable scratch session; without it the busy-guard was never the deciding factor and the LEDGER_MUTATE idle-pane mutation could not differentiate. Now g2's mutation correctly fires and trips the must-not-fire assertion. Fixture-only: no change to fm_ctx_select / _ctx_eligible / fm_ctx_can_fire, so daemon behavior is unchanged. ledger verify: green:5 red:0 wip:0. All five mutations still caught (non-vacuous). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…emon cycle) bin/fm-compact-crewmate.sh runs the watchdog's checkpoint -> handoff-wait -> /clear -> cooldown cycle NOW against one crewmate, for a secondmate (or the main firstmate) to invoke. It SOURCES fm-context-watch.sh and calls the same fm_ctx_fire_once the daemon's main loop calls — one implementation, no forked cycle (G6 asserts this statically + e2e on a scratch pane). Resolves the target via the existing fm_ctx_target_for; refuses if no live sentinel/target; guards an in-flight compact with a per-id mkdir lock (the daemon's lock primitive, scoped per key); honors the existing cooldown so repeat calls are idempotent no-ops. --resume frontier|restart (default frontier) writes a resume-<key>.directive sentinel that fm-captain-bootstrap.sh's rehydrate reads to inject a different, mode-correct directive after /clear (restart = re-read the compacted brief and start over; frontier = pick the Frontier back up), then consumes it so it fires once. The unwritten-directive path is byte-for-byte the prior frontier behavior, so the daemon-driven rehydrate is unchanged (G3/G4 still green). Gates: G6 (on-demand cycle + non-duplication + idempotency) and G7 (resume directive differentiation), both with caught mutations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
fm-context-watch.sh gains --scope/--home <home>: it re-points FM_HOME so _ctx_state_root resolves to <home>/state, scoping the poll set, cooldown markers, AND the singleton lock (state/.context-watch.lock) to that home. No scope -> unchanged global behavior (FM_STATE_OVERRIDE still wins), so the existing global daemon and all g1..g5 tests are unaffected. fm-spawn.sh, on a kind=secondmate launch, auto-starts a context-watch scoped to the secondmate's own home as a presence-gated background child. The scoped watch self-singletons on the home's lock, so a duplicate spawn or a recovery respawn no-ops instead of stacking daemons; it is idle-safe (only polls this home's ctx-*.json, fires the compact cycle for its own crewmates). The start is overridable via FM_CTX_WATCH_START_CMD (mirrors _ctx_send's FM_CTX_SEND_CMD) so tests record the intent instead of detaching a real loop, and FM_SECONDMATE_NO_WATCH=1 opts out. Gate G8 proves scope isolation: a watch scoped to home A selects only A's crewmate and never B's (and vice versa), while a global watch still sees both (no regression); mutation (scope ignored) leaks B into A and fails. fm-secondmate.test.sh now asserts a secondmate spawn starts a scoped watch for its own home (proven non-vacuous: disabling the auto-start fails the test) and stubs the start so no real daemon leaks. Also folds in fm-compact-crewmate's shellcheck SC1091 directive + explicit lock release (no RETURN-trap scoping under set -u), and the g6/g7 executable bit + mutation refinements so LEDGER_MUTATE flips each gate red. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…watch Document fm-compact-crewmate (on-demand single-crewmate compaction reusing the daemon's fire-once cycle, with the frontier/restart resume directive) and the per-secondmate scoped watch (FM_HOME-scoped poll/lock, auto-started on secondmate boot), plus gates G6-G8. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ledger verify re-confirmed g1..g5 green on this branch; only last_verified timestamps change (statuses unchanged: green:5 red:0). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…g home Two review findings from the gate pipeline: 1. Stale-handoff immediate /clear (fm-context-watch.sh): fm_ctx_fire_once sent the checkpoint then waited on `[ ! -e $handoff ]`, so a handoff-<key>.md left over from a prior cycle (never archived) let the loop fall straight through and /clear fired immediately — wiping the crewmate's CURRENT turn and rehydrating from the stale doc. Now baseline the existing handoff mtime before the checkpoint and only accept a handoff strictly NEWER than that baseline; a stale handoff no longer triggers /clear. Protects both the daemon and fm-compact-crewmate. New gate G9 (mutation-proven: a fresh post-checkpoint handoff still fires). 2. Scoped watch couldn't see its crewmates (fm-spawn.sh): the per-secondmate scoped watch polls $home/state, but crew launches set no FM_HOME, so the global statusLine wrote crew sentinels into the MAIN state dir — the secondmate's watch never saw its own crewmates. Now pin every crew/scout launch's FM_HOME to the SPAWNING home: a secondmate's crew route into the secondmate state dir; the main firstmate's crew keep landing in main state. Role is unaffected (fm_ctx_role is cwd-based; a crewmate's cwd is its worktree, never the home, so it stays crew — asserted). New gate in fm-secondmate.test.sh proving the FM_HOME pin. Full suite green (18 tests), shellcheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Post-rebase reconciliation after upstream restructured the test suite (kunchenguid#80) and the brief contract (kunchenguid#73): - bin/fm-brief.sh: collapse the gate-check definition-of-done clause, which had landed once per delivery-mode heredoc (3x), into a single GATE_CHECK variable appended once after the mode DOD. Identical brief output, no duplication. - tests/fm-secondmate-lifecycle-e2e.test.sh: re-land the scoped-watch coverage on the relocated spawn phase — stub the watch start via FM_CTX_WATCH_START_CMD (no leaked daemon) and assert a secondmate spawn starts a context-watch scoped to its own home; stub the recovery respawn's watch start too. - tests/fm-tangle-guard.test.sh: assert the crew-launch FM_HOME pin on the existing genuine-isolated-worktree success path (the post-kunchenguid#83 harness that can actually reach a successful ship launch), proving a crew launch pins FM_HOME to the spawning home for context-sentinel routing; teach its fake tmux to log send-keys so the launch env is observable. Full suite green (24 files), shellcheck clean, zero leaked daemons. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…se 1) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… inert) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…file Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…+ gate-q4 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…calate Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…mistakes mode) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… phase 2 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…b.sh Behavior-identical for fm-verify (gates q4/q6 stay green); ready for fm-intake. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sweep: FM_INTAKE_OVERRIDE for tangle-guard (spawn machinery suite); also fixes the pre-existing Phase-1 escape in fm-secondmate-safety (fm-pr-check assertion predates the quarterdeck verdict gate - FM_VERIFY_OVERRIDE, loudly). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…inding) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…honest budget - observability anchors to dirname($STATE) so suites draining temp state dirs never leak churn into the repo (gate-l1 now pins it + the secondmate skip) - .fm-secondmate-home homes skip logging (fleet ff-sync stays clean) - loop-budget 48/day relabeled as a chosen alarm; real cadence documented - loop-verifier: explicit do-not-invoke-directly; LOOP.md documents the loop-audit global dependency Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rook's definition now lives at ~/stone-skills/agents/loop-verifier/AGENT.md; this path becomes a symlink so fm-verify.sh keeps working unchanged and the agent has one canonical source. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Runtime churn (STATE.md, gates/, loop-run-log.md) plus two unlanded plan docs, captured during the 2026-08-26 config audit.
Without the colon, tmux can resolve $SES as a window index rather than a session name. Salvaged from an uncommitted edit in the hermes-jarvis-sm secondmate home before syncing that worktree.
Projection artifact of stone-skills#3: install.mjs now writes relative links, and loop-verifier installs into this repo via its .meta.json install_to.
Follows stone-skills' install.mjs: agent links are absolute so they resolve when mac-config mirrors them (its G11 gate). loop-verifier installs into this repo via its .meta.json install_to, so it moves with them.
Two registered secondmates, cellarsky-sm and hermes-jarvis-sm, were found on 2026-08-26 as bare zsh prompts rather than running agents. Their launch line had been typed into a shell that was not at a prompt, so the whole `claude --dangerously-skip-permissions "<charter>"` string was consumed as raw text and died on a zsh parse error. firstmate had already recorded a meta and registered them as live supervisors. Nothing noticed for weeks. Root cause: fm-spawn inferred shell readiness from pane_current_path changing after `treehouse get`. That proves a chdir happened; it does not prove the shell returned to a prompt, and tmux send-keys has no acknowledgment channel. The 'sleep 0.3' between the literal text and Enter was a timing guess with no backing signal. Adds two guards in fm-tmux-lib.sh: fm_tmux_wait_shell_ready - a bounded echo round-trip. Send a unique marker through the shell, wait to see it rendered back on its own line. Seeing it is positive proof the shell read a command line, ran it and printed a result. A probe landing in a busy shell is a harmless short printf, unlike a multi-KB launch string. fm_tmux_launch_failed - reads the pane after launch and fails loudly on a shell error rather than recording a meta for a pane holding nothing. Also fixes tests/lib.sh: fm_test_cleanup ended on a falsy test and returned 1, so any suite following the header's advice to call it from its own EXIT trap failed with every test passing. Verified red before green: with the readiness guard stubbed to return 0 (the pre-fix behavior) the deaf-shell case fails; restored, all six pass. Full suite 38/39 - fm-loop-l2 fails identically at 24f6891, before any of this.
…-digest) (#2) SessionStart captain block now carries a read-only snapshot of the section-5 recovery signals: lock-free wake-queue depth + last records (torn-tail rule), verbatim fm-watch-arm --status and fm-lock status relays, in-flight metas with last status line, afk flag, and the mandatory run-fm-wake-drain disclaimer, capped at ~2000 chars with explicit +N-more truncation. Adds a side-effect-free --status mode to fm-watch-arm.sh reusing healthy_watcher(). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
#3) firstmate drives panes through ~100 direct tmux calls across 15 scripts. That couples the fleet to one multiplexer and, worse, to that multiplexer's weakest property: send-keys is blind keystroke injection with no acknowledgment channel. Every supervision primitive built on it is therefore a heuristic - busy/idle is a regex over rendered text, readiness is a cwd poll, and delivery is a fixed sleep between the text and the Enter. Two secondmates died of exactly that on 2026-08-26. bin/fm-mux-lib.sh is the seam that lets a driver with real acknowledgment replace it. FM_MUX selects: tmux (default, byte-identical to today) or herdr. Six verbs, one contract: new_window, send, read, is_busy, wait_ready, close. The target is OPAQUE - session:window under tmux, a pane/agent id under herdr. Callers pass it back verbatim and never parse it; that rule is what keeps the drivers swappable. Where they differ is the point: tmux send is type-sleep-Enter with no confirmation, herdr send is a single acknowledged 'agent prompt --wait' that REFUSES a blocked agent (distinct rc 3) instead of typing into an approval dialog. tmux is_busy pattern-matches a footer; herdr is_busy reads real state. Verified falsifiable: removing the agent_blocked branch turns the blocked-agent case red; removing the FM_MUX override turns the rollback-switch case red. 9/9 seam tests, full suite 38 pass 1 fail (fm-loop-l2, failing at 24f6891 and unrelated), shellcheck clean. Implements W1 of docs/plans/cmux-herdr-surface-split.md. No call sites are migrated yet - the seam lands first, callers move behind it one at a time.
#4) * feat(herdr): one workspace per project, agent panes named for the work Captain's standing order, 2026-08-27: when the crew runs on herdr, every project gets its own workspace and every agent pane is named for the work it is doing rather than for a task id. That is not tidiness. herdr addresses agents BY NAME - 'herdr agent prompt <name>' - so the name is the address. 'afs-resources-r7' tells you nothing at a glance; 'afs/resource-registry' tells you what the pane is for without attaching to it, and makes the whole fleet readable from the workspace list. bin/fm-herdr-workspaces.sh reconciles workspaces against data/projects.md: (no args) print a plan and change nothing --apply create only the workspaces that are absent --name apply the convention to a live pane, normalising and length-capping Plan mode is the default because a tool that mutates on invocation is a tool people stop running. A registry line whose project directory does not exist is reported as NO-DIR rather than silently created. Refusals are actionable: a stopped server says to run 'herdr', a missing registry says which path it wanted. Verified falsifiable - removing the length cap turns the overlong-name case red. 8/8 suite cases, full suite 38 pass 1 fail (fm-loop-l2, failing at 24f6891 and unrelated), shellcheck clean. * fix(herdr): align the pane-name cap with its own documentation The header said keep names under ~28 chars; the code enforced 40. A doc and a guard disagreeing means one of them is lying to whoever reads it - and here the doc was right: herdr's sidebar_width is 30, so a 40-char name is truncated exactly where it stops being readable, which defeats the point of naming panes for their work. Code now enforces 28 and the comment cites the sidebar width as the reason, so the number is traceable rather than arbitrary. Found by a fresh-context review of the multi-project working model, not by me.
#5) * docs: promote the N-concurrent-firstmates design to a tracked spec The design lived in gitignored data/, so gates could not cite it as spec_ref and it was unreadable from any other clone. Promoted verbatim, with an appended Erratum section reconciling section 7's mutually unsatisfiable helper timeout (2s per helper x 2 helpers) against its own 1.5s total wall-clock ceiling, which gate m4 asserts. The total is authoritative; per-helper drops to 0.6s and a shared deadline bounds the whole helper phase. * test: freeze the boot-context gates red, before any implementation Five gate tests, none of which can pass on this commit. Red-before evidence, captured by running each exactly as the ledger harness runs it: m0 normal EXIT=1 gates/ledger.json 'gates' is a dict, must be a JSON array m1 normal EXIT=1 fm-boot-context.sh is not a registered SessionStart hook m2 normal EXIT=1 bin/fm-boot-context.sh must exist m4 normal EXIT=1 bin/fm-boot-context.sh must exist m5 normal EXIT=1 bin/fm-boot-context.sh must exist m0 asserts the real ledger's shape by pure read, and proves loud failure on a broken ledger against isolated fixture dirs under a temp root - it never invokes ledger against gates/, so it cannot recurse or rewrite tracked ledger state. m2 has two arms: writable (full-tree manifest identity, the real zero-writes proof) and read-only (chmod a-w on home, state, data and the fleet dir, proving the emitter still emits rather than silently dying). fs_usage needs root and macOS has no bind mounts, so chmod is the root-free observation harness. m1 is expected to stay red: its other half is a declared change in mac-config, which owns the rendered settings file. It asserts on the real registration, not on direct invocation - a direct-invocation test goes green while a real session still boots blind. Every test honours LEDGER_MUTATE. * fix(gates): restore ledger.json to the array shape the harness requires gates/verify.sh aborted on every invocation: Error: invalid ledger: gates must be an array at validateLedger (.../dist/ledger/schema.js:47:15) EXIT=1 It fails loudly and honestly - exec'd, so there is no swallowed exit code anywhere in the shim. The defect is the ledger file, not the harness. Archaeology settles which shape is canonical rather than guessing: every commit from 0dcc0f6 (2026-06-25) through 57bc27e stored 'gates' as an array. Commit fa04182 alone reshaped it to an object keyed by id while adding g-boot-digest, and nothing has run ledger since. So this is reverting a one-commit regression, not migrating to a new shape. The repair is a pure reshape - every key already equalled its entry's own id, order preserved, all 20 entries intact. Baseline captured immediately after the reshape, before adding anything: green:19 red:1 wip:1 (red = gate-l2-loop-audit-level, the accepted baseline) Registers five gates. m0 is green and frozen; the other four are red with first_observed_red stamped by the harness itself on a real red run: m0-ledger-shape frozen mutation-verified, not vacuous m1-hook-registered red blocked on the mac-config task m2-boot-emitter-is-read-only red m4-boot-budget-hostile red m5-digest-never-silent red m0's first_observed_red is the one timestamp written by hand rather than stamped by the harness, because its red condition was the harness being unable to load at all. It is the author date of adca03e, the commit at which the assertion was observed failing by the same command the harness runs: bash tests/fm-boot-m0.test.sh -> EXIT=1, 'gates must be a JSON array (got dict)' Note: ledger verify demotes frozen gates to green by design (freeze is re-earned via ledger freeze). g-boot-digest was frozen before this run and is re-frozen at the end of this branch. * feat(boot): a strictly read-only boot-context emitter with a tiered digest bin/fm-captain-bootstrap.sh was registered as a global SessionStart hook and was removed with cause: it mutates durable files on every session start, moving the pending handoff and deleting the resume directive on behalf of a subsystem whose writer is itself unregistered. The objection is correct on its facts, so this answers it rather than arguing with it. bin/fm-boot-context.sh reads. It writes nothing, moves nothing, deletes nothing, creates nothing, takes no lock. The rehydrate machinery stays out of the boot path until its writer is revived as its own task. That is what makes registering this script legal on the merits rather than by waiver. Read-only is a gate, not a promise (m2): a full boot is run with a recursive manifest of the home - path, type, size, mtime, ctime, inode, mode - before and after, and the manifest must be identical; then the same boot runs with the home, state/, data/ and the fleet dir held read-only via chmod a-w and must still emit valid output. fs_usage needs root and macOS has no bind mounts, so chmod is the root-free observation harness, and paired with the manifest it is stronger than syscall tracing: it proves both that no write happened and that none was needed. Two writes had to be closed to make that true unconditionally, both in relayed helpers rather than in the emitter: - fm-lock.sh created the state dir before dispatching, so its status mode created the directory it was reporting on. Moved to the acquire path. - fm-wake-lib.sh creates it at source time for seven consumers, six of which genuinely need it. Added FM_WAKE_LIB_READONLY=1, set only by this emitter. m2 now asserts a boot leaves a bare home completely empty. The tiered digest. The old block was 10,188 chars, already over the 10,000-char additionalContext cap, and paid in full by every session that activated it. Tier 1, universal: identity plus one line per fleet instance. A PEER IS NEVER ELIDED - answering 'what is the fleet doing' with zero tool calls is the whole point, and dropping a peer breaks it. Per-task detail may be elided. Tier 2, only for the session actually steering: spawn lifecycle, projects as generated one-liners rather than the 4,108-char verbatim dump, secondmates, backlog, reconciliation digest. Capped, with explicit (+N more) markers. Steering is read off the session lock and needs no new signal: free or stale means this session is about to take it; held means steering only if the holder is our own ancestor (resume, compact, /clear). If the ancestry probe cannot run the answer is 'not steering', so a failure costs a steering session one tool call rather than charging every incidental session the full block. Measured against the live primary home: 555 chars / ~138 tokens in 0.23s, down from 10,188 chars / ~2,550 tokens. A full steering block at 5 instances is 2,831. The budget (m4). A hook that overruns its declared timeout injects NOTHING, so the hazard is not slowness but a boot that silently lost all context. The design set a 1.5s ceiling and a 2s-per-helper timeout for two helpers, which is 4s and cannot satisfy its own gate; the spec's Erratum reconciles it. The total is authoritative at 1.5s (6x headroom under the timeout-10 convention) and per-helper drops to 0.6s, still 8.5x the slowest measured helper. Helpers also run under a shared deadline - each gets min(cap, budget left) and one whose budget is spent is skipped with a marker - so adding a helper can never breach the ceiling. The peer path costs zero execs by construction. Failure is never silent (m5). Every section builds inside its own guard; one that raises becomes an explicit UNAVAILABLE marker naming the section and the reason, and the other sections still render. The predecessor wrapped the whole digest in a bare 'except: pass', so a boot that lost all its fleet context was indistinguishable from a healthy one. Role resolution is deliberately unchanged. The activation-contract inversion is separate later work; the tiering here is what makes it affordable when it lands. * docs: the exact mac-config declaration that registers the boot hook Content only. Nothing here has been applied: no edit to ~/.claude/settings.json, no installer, no change to mac-config from this worktree. Registration cannot happen in firstmate. The settings file is rendered from mac-config's desired state and a render deletes undeclared keys, so a hand-written entry works until the next apply and then silently reverts - a boot path that is intermittently blind, worse than one reliably blind. Four edits, quoted at byte level against the live files: the registry.yaml hooks block, a policy.yaml hook_allowlist entry, the per-agent hook budget raised from 1 to 2, and the retirement of the spent fm-captain-bootstrap forbidden-substring rule plus its restatement in mac-config's own test suite. Two traps recorded so the applying task does not hit them: V15 matches the exact 4-tuple (agent, event, command, timeout), and an omitted timeout renders as null and will not match an allowlist entry carrying 10; and the settings file has two claimants, with mac-config the effective owner. On edit 4 the declaration states the design's case for deleting the rule and also names its weakness - deleting removes the one thing that would stop fm-captain-bootstrap being re-registered later - and records that edits 1-3 are sufficient on their own, since the new command does not contain the banned substring. The design is authoritative and calls for the deletion, so that is what is specified; the alternative is called out so the choice is deliberate. * docs: record the gate-ledger mechanics this branch had to rediscover Durable, project-intrinsic knowledge, none of it derivable from the code: the array-shape requirement and that it is validated on load so a bad shape kills every subcommand; that ledger verify executes test_refs and rewrites tracked ledger state, which is why a meta-gate needs a throwaway fixture dir and must never re-enter gates/verify.sh; that verify demotes frozen to green so freeze must be re-run; that freeze runs each test twice and demands the mutated run fail, so a test ignoring LEDGER_MUTATE freezes vacuous and only reveals it at the end; that born-green gates are refused, so a gate should be registered while genuinely red and let the harness stamp first_observed_red; and that spec_ref must be tracked, because data/ is gitignored. Also lists the five new boot gates, flagging m1 as expected-red. * chore(gates): sweep and freeze - WIP is exactly the two declared reds Full sweep: green:23 frozen:0 red:2 wip:2, then freeze on every gate that earned it. Freeze runs each test twice and demands the LEDGER_MUTATE=1 run fail, so all five frozen gates are mutation-verified and none is vacuous. m0-ledger-shape frozen m2-boot-emitter-is-read-only frozen m4-boot-budget-hostile frozen m5-digest-never-silent frozen g-boot-digest frozen restored; the sweep had demoted it Drain list is exactly: gate-l2-loop-audit-level the accepted red baseline, untouched by this work m1-hook-registered blocked on the mac-config declaration No gate that was green before this branch is red after it, and there is no new red gate other than m1. Baseline for that comparison was taken immediately after the ledger reshape and before anything was added: green:19 red:1. * fix(boot): close three defects an independent verifier found An adversarial re-proof of the gate claims refuted two of them. All three are real, and each gate is strengthened so the defect class cannot come back. 1. The 1.5s ceiling was unachievable, and m4 was a 28%-flaky gate. The budget clock started inside python, so it never counted bash startup, reading stdin, or starting the interpreter - about 0.3s. The helper phase deliberately spends 2 x 0.6s, so 1.2 + 0.3 landed exactly on the ceiling and breached it in 7 of 9 measured runs while the gate passed most of the time. A gate that is 72% green is worse than a red one: it launders a defect as flake. The start now comes from the wrapper via bash's EPOCHREALTIME builtin (free; one interpreter start as a bash 3.2 fallback), so the shared deadline sees real wall clock and shortens the second helper's slice on its own. The render reserve widens 0.2 -> 0.3 to absorb kill-and-wait overhead after a timeout. Worst of 15 wedged runs: 1.378s, margin 0.122s. m4 now times five runs and judges the WORST, so it is deterministic rather than sampled. 2. One peer could blow the injection cap on its own. peer_line interpolated id and watcher unbounded; a 20,000-char id produced a 20,882-char block. Tier 1 stays uncapped so no peer is ever elided, but the FIELDS are now bounded (40 and 24 chars). Same fleet: 2,603 chars. m4 gates it. 3. A malformed budget produced exit 1 and zero stdout - a silently blind boot. The three numeric settings were parsed at import time, outside the guard around main(), so FM_BOOT_TOTAL_BUDGET='' (an ordinary thing for a hook config to produce) took the whole hook down with a traceback and no output: exactly the failure this file was written to prevent. They now degrade to their defaults, and a non-numeric value leaves a visible marker naming the setting. Empty stays silent, being ordinary absence. m5 gates all three. Also, where 'a peer is never elided' and the injection cap genuinely collide - past ~145 peers, far outside the design's 12 - the peer rule wins and the block now says so, rather than quietly overrunning. A session that cannot see a peer has lost the thing this block exists to give it; an oversized block is a degraded read, not a blind one. m4 asserts all 200 peers plus the marker. Verified unchanged by the same pass: zero writes on a writable tree, no writes on a bare home (and that FM_WAKE_LIB_READONLY is load-bearing, not decorative), peers never elided at 12/40/150/200, the force-fail seam, mutation sensitivity, no scope leak, and the ledger reshape preserving all 20 entries field-identically. * chore(gates): re-sweep and re-freeze after the verifier fixes green:23 red:2 wip:2, unchanged. The four new frozen gates re-earned their mutation check against the corrected emitter, and g-boot-digest is restored to frozen after the sweep demoted it again. Drain list is still exactly the two declared reds: gate-l2-loop-audit-level (accepted baseline) and m1-hook-registered (blocked on mac-config). * test: freeze the declared-red skip gate red, before the runner exists Red-before evidence, by the same command the ledger harness runs: bash tests/fm-ci-declared-red.test.sh -> EXIT=1, 'tests/run-all.sh must exist' The gate's real subject is not that skipping works - it is that skipping is NARROW. A runner that skipped any currently-red gate would mask a regression the moment a working gate goes red, so a skip requires two independent conditions: the gate is red in the ledger AND its id is declared in gates/accepted-red.md. The test asserts an undeclared red gate still runs, still fails, and still fails the suite; that a declared gate which is green still runs; that a missing or unparseable ledger skips nothing; and that every skip is announced by name. * ci: skip a declared-red gate's test, out loud, and only that CI ran every tests/*.test.sh, so a gate that is deliberately red failed the job forever. Two are: gate-l2-loop-audit-level (accepted baseline) and m1-hook-registered (blocked on a mac-config change). Making either test pass would be the false green the gate ledger exists to prevent, so the suite skips them and the ledger stays the honest record of what is red. The skip is deliberately narrow. Two independent conditions must BOTH hold: the gate is red in gates/ledger.json, AND its id is declared with a reason in the new gates/accepted-red.md. Skipping on status alone would mask a real regression the moment a working gate went red - the failure would vanish from CI exactly when it mattered most. Skipping on the declaration alone would let a stale entry silence a test that had since been fixed. No skip is silent: each prints the gate, the test and the reason, plus a ::notice:: annotation under Actions, and the summary reports the counts. If the ledger or the declaration file is missing or unparseable, nothing is skipped at all and the runner says so - fail closed, never open. Gated by gate-ci-declared-red, frozen and mutation-verified, so the narrowness is machine-checked rather than trusted. Its test asserts an undeclared red gate still runs and still fails the suite, a declared gate that is green still runs, a missing or unparseable ledger skips nothing, and every skip is announced. Real suite through the runner: 46 ran, 2 skipped, 0 failed, exit 0. * chore(gates): re-sweep and re-freeze with gate-ci-declared-red green:24 red:2 wip:2. Six frozen gates, all mutation-verified, none vacuous. Drain list unchanged: the two declared reds, both now listed in gates/accepted-red.md with a reason and a route back to green. * style: shellcheck source directives so the new tests add no lint findings Under CI's own invocation (shellcheck bin/*.sh tests/*.sh) the six new test files now produce zero findings, so this branch leaves the lint job exactly where it found it: 20 findings, identical to main. Comments only, no behavior change. * chore(gates): re-freeze after the shellcheck directive edits * fix(boot): unreadable is not empty - a home we cannot read never reads as idle QUARTERDECK REJECT, and the verifier is right. Repro: FM_HOME=/tmp/nonexistent bash bin/fm-boot-context.sh Against a home that did not exist, the emitter printed a fully confident block with no marker anywhere: '0 in flight', 'Wake queue: empty', 'In-flight tasks: none'. Byte-identical to a genuinely idle, healthy home. Same for chmod 000 on state/ or on .wake-queue alone. This is the single most consequential lie this block can tell. Recovery keys off exactly those lines - they say there is nothing to reconcile - and it was reachable through an ordinary FM_HOME misconfiguration, which the N-concurrent-firstmates design makes routine. The file's own header claimed the opposite in as many words. m5 missed it because its observable was scoped to section builders that RAISE, and none of these paths raise: read_or, wake_queue and own_tasks each caught the error and returned the empty value. Every read now returns its problem alongside its text, and a problem always becomes a visible UNAVAILABLE marker. Two kinds of not-knowing, kept apart: STRUCTURAL - the home or state/ cannot be listed, so the COUNT is unknown. Renders '[UNAVAILABLE: ...]' and the digest refuses to make any 'nothing here' claim at all. DETAIL - state/ lists fine, one status file will not read. The count is real and survives as '3 in flight (+1 unreadable)'; only that task's line is marked. A detail we lost must not discard a fact we have. Absence is still absence: no .wake-queue genuinely means an empty queue and says so plainly. m5 asserts that too - a fix that marked everything would make the marker meaningless in the other direction. Also fixed, from the same review round: - P2, the backlog read took the first 1,200 chars silently, cutting records mid-line while every other capped path counts what it dropped. Now truncates on a line boundary with an explicit '(+N chars more)' marker. - A real fail-open in tests/run-all.sh, in the one place it promises to fail closed: the BADLEDGER sentinel was appended to the same stream the parse wrote rows to, so a ledger that emitted rows and THEN failed produced 'rows + BADLEDGER', the exact-match handling missed the sentinel, and skipping proceeded on a half-read ledger. The parse is now captured whole and used only if python exited cleanly, and it buffers rows so a mid-loop raise emits none. gate-ci-declared-red gains case F to freeze it. - m1 accepted any SessionStart command merely CONTAINING the emitter's filename, so a no-op like would have satisfied it - the same false-green class m1 refuses direct-invocation testing to avoid. It is red today so this was latent, but it would have bitten exactly when the dependent mac-config task landed. It now requires the command to invoke the script, and probes that the invocation actually emits additionalContext. Verified both ways: the no-op form is rejected, a real 'bash "<path>"' passes. - Single-file reads are bounded at 256KB. A 512MB status file measured at 3.6s against a 1.5s ceiling; oversized input is a defect or an attack either way. Gates: m0/m2/m4/m5 and gate-ci-declared-red green and mutation-sensitive, m1 still honestly red. Real boot unchanged at 555 chars with zero markers. * chore(gates): re-sweep and re-freeze after the unreadable-is-not-empty fix green:24 red:2 wip:2. All six frozen gates re-earned their mutation check against the corrected emitter and runner. Drain list unchanged: the two declared reds. * test: freeze the status-verb gate red, before bin/fm-status.sh exists Red-before evidence, by the command the ledger harness runs: bash tests/fm-status-verb.test.sh -> EXIT=1, 'bin/fm-status.sh must exist' The defect is observed, not hypothetical: a crewmate told by its brief to report with 'echo "done: ..." >> ~/firstmate/state/<id>.status' was refused five times in one task, because the profile denies Edit(~/firstmate/**) and a shell redirect into that tree is classified as an edit. The status file was never created, so the crewmate/firstmate channel was silently dead for the whole task. A test cannot invoke the real permission layer, so the policy is modelled as it behaves - a predicate over the command string - and the gate proves three things together: the model REFUSES the redirect form (so it has teeth), PERMITS the verb form, and running the verb actually appends. Any one alone would be vacuous. * feat(status): reporting is a verb, not a redirect; and fix two CI reds QUARTERDECK REJECT (attempt 2). Three findings, all real, all fixed. 1. bin/fm-status.sh - the new scope, and the cause of the silence. A crewmate told by its brief to report with echo "done: ..." >> ~/firstmate/state/<id>.status was refused five times in one task: the profile denies Edit(~/firstmate/**) and the harness classifies a shell redirect into that tree as an edit. The status file was never created; no report ever reached firstmate. The channel was silently dead while the work got done. The append now happens INSIDE the script, which the profile permits - verified empirically before building on it. Resolves the home through FM_HOME/FM_STATE_OVERRIDE like the other 23 bin/ scripts, appends rather than truncates, writes exactly one line per call so an embedded newline cannot forge extra wakes, and refuses an empty message or a traversing id. fm-brief.sh now teaches the verb in all three brief kinds and teaches the redirect form in none. gate-status-verb freezes it: the modelled policy must REFUSE the redirect and PERMIT the verb, and the verb must actually append - any one alone would be vacuous. tests/fm-secondmate-safety.test.sh asserted on the '>>' form; its intent was that briefs carry the shell-quoted FM_HOME-scoped path, which still holds, so its three assertions now check that path plus the verb naming the right id. 2. m4 was ~40% flaky, so the time budget was unproven. An independent verifier measured 1.512s and 1.587s against the 1.5s ceiling where I had measured 1.378s. Nudging a constant would not have been a fix, so the bound is now stated as an inequality and RENDER_RESERVE sized to satisfy it: total <= TOTAL_BUDGET - RENDER_RESERVE + post_deadline_cost, so the ceiling holds iff the reserve exceeds everything after the last deadline check - killing a wedged helper, rendering, teardown. Startup cancels out. It is post-deadline cost that stretches under load, so the reserve must cover its LOADED value: 0.3 -> 0.45, per-helper 0.6 -> 0.45. Re-measured over 20 wedged runs under 8-way CPU load: worst 1.280s. m4 passes 5/5 where it previously failed ~2/5. 3. CI was red twice over. - fm-boot-m0 exits 2 without the ledger CLI, which ci.yml never installs. Exit 2 is already this repo's convention for a missing prerequisite (fm-loop-l2 uses it), and tests/run-all.sh simply did not honour it. It now reports exit 2 as an announced prerequisite SKIP carrying the test's own explanation, while exit 1 still fails the suite. gate-ci-declared-red gains case G covering both halves. - fm-status-verb.test.sh was committed red-first against a script that was not yet committed, and was absent from the ledger so could never be skipped. The script is committed here and the gate registered. Simulated CI with neither ledger nor loop-audit on PATH: 46 ran, 3 skipped, 0 failed, exit 0. * chore(gates): sweep and freeze with gate-status-verb green:25 red:2 wip:2. Seven frozen gates, all mutation-verified, none vacuous. Drain list unchanged: the two declared reds. * fix(boot): a wedged helper is killed as a group, and cleanup stops draining QUARTERDECK REJECT (attempt 3). m4 was not fixed, and the previous 'clean run' was luck. Diagnosed rather than re-tuned: two implementation defects, both real, and the 1.5s ceiling did NOT need to move. 1. THE LEAK, and why the gate could not be trusted at all. subprocess killed the helper it started but not what that helper spawned, so every wedged boot orphaned two processes to init. Five boots per gate run, accumulating across runs: 194 live at once, load average 186 - and the boots this gate was timing then took ~2s against its own 1.5s ceiling. The gate was measuring its own side effects. That is why one run passed and six in a row did not, and why my earlier single-run measurement was meaningless: I had timed the emitter directly on a machine my own test had already degraded. Helpers now run in their own process group (start_new_session) and are killed as a group. Measured: 2 orphans per hostile boot before, 0 across 10 gate runs after. This was a PRODUCTION defect too, not only a test artifact - a boot hook that leaks a process whenever a helper wedges degrades the machine it runs on. 2. EXPENSIVE CLEANUP, which my own group-kill fix introduced. Draining a killed helper's pipes with communicate() costs up to its timeout per helper - 0.4s at two helpers, spent after the deadline was already gone. That put cold runs at 1.512-1.620s and is why 8/10 still was not 10/10. The output is unwanted (the call returns UNAVAILABLE either way), so the group is killed, reaped, and its fds closed WITHOUT draining. m4 gains the assertion that made this visible, checked BEFORE the timing verdict because a leak invalidates the timing: a hostile run must leak zero processes. A gate that poisons its own experiment is worse than no gate. PROOF, on this machine at load ~134 (its ambient state - I verified the load is not mine: interpreter start is 0.05-0.13s and bash 0.012-0.026s throughout): 10 consecutive m4 runs: 10 passed, 0 failed, 0 processes leaked worst single hostile boot across 8 timed runs: 1.104s against the 1.5s ceiling Also caught in passing: an earlier edit of mine left prose after a docstring's closing quotes, so the emitter was a SyntaxError and produced no output at all. It is exactly the silent-blind-boot failure this file exists to prevent, and it would have shipped had the leak measurement not printed stderr. * chore(gates): sweep and freeze after the process-group fix * fix(boot): helpers run concurrently, and m4 stops timing its own instrumentation Your reproduction was right and my 10/10 was not reproducible. Diagnosed with measurement rather than re-tuned, and the 1.5s ceiling still did not move. TWO SEPARATE CAUSES, one in the gate and one in the code. 1. THE GATE WAS TIMING ITS OWN INSTRUMENTATION - your hypothesis, confirmed. It stamped the start of each timed boot with one `python3 -c` and the end with another, so the second interpreter's startup was charged to the emitter. Measured over 60 samples at load ~180, bracketing each boot with bash's free EPOCHREALTIME clock as a control: TRUE boot p50 1.035s p90 1.174s max 1.370s breaches of 1.5s: 0 AS MEASURED p50 1.097s p90 1.291s max 1.863s breaches of 1.5s: 3 instrumentation p50 0.059s p90 0.163s max 0.608s Every breach in that sample was instrumentation. One interpreter now runs the whole loop and times each boot around the subprocess itself, so nothing but the boot is inside the window. It also prints each boot time, so the evidence is in the gate output rather than in a summary. 2. BUT THE CODE WAS ALSO GENUINELY TOO SLOW, which the fixed measurement then exposed honestly: with the instrumentation removed, real boots still reached 1.523s and 1.612s. Two 0.45s helper timeouts were being paid SERIALLY - 0.9s of deliberate waiting - and on a machine at load ~150 the tail cleared 1.5s. The helpers are independent reads, so they now run CONCURRENTLY under one shared deadline computed before either starts. The phase costs the max of their timeouts instead of the sum, and a third helper would cost no extra elapsed time at all. Chosen over shrinking the per-helper allowance so a slow-but-alive helper keeps the same generosity; the design named this option explicitly. Boot times fell from 1.05-1.61s to 0.55-1.11s. Also fixed a race in the leak assertion I added last round: killing is asynchronous, so an instantaneous sample can catch a process mid-death and call it a leak. It now settles briefly, bounded, before judging - a genuine orphan never comes back down, so a real leak still fails. Verified: 30 consecutive hostile boots never elevated the count at all. 10 consecutive gate runs green at load ~150, worst single boot 1.114s, zero processes leaked. Full suite: 47 ran, 2 skipped, 0 failed. * chore(gates): sweep and freeze after the concurrent-helper fix * fix(boot): say 'steering unknown' instead of guessing, and isolate m5 Two defects surfaced by m5 freezing VACUOUS - a false green, caught because ledger freeze runs a test differently than I had been running it. 1. A CONFIDENT FALSEHOOD IN THE EMITTER. When the lock helper returned UNAVAILABLE, is_steering() answered False and the block stated 'Another session is steering this home; this one is observing.' It does not know that. It could not read the lock. This is exactly the class of defect m5 exists to prevent, and it is reachable in ordinary operation: a loaded machine can exhaust the helper budget, the lock read is skipped, and the block then asserts something false about who owns the fleet. It now says 'steering unknown' and carries an explicit UNAVAILABLE marker naming the reason. 2. m5 WAS NOT ISOLATED FROM THE BUDGET. Its subject is silence, not timing - m4 owns the budget. On the shipped budget a loaded machine exhausted it before the lock helper ran, the emitter correctly degraded to Tier 1, and every m5 assertion about the digest failed for a reason unrelated to silence. Under that made the normal arm fail, so the mutation check could not bite and the gate froze VACUOUS. m5 now pins a generous budget: a gate must control everything except the one thing it tests. Verified under the harness's OWN invocation (/bin/bash -c, as execSync uses it), not just my shell - that difference is what hid this: m0 n=0 m=1 | m1 n=1 m=1 (declared red) | m2 n=0 m=1 m4 n=0 m=1 | m5 n=0 m=1 | ci n=0 m=1 | verb n=0 m=1 * fix(gates): m4 asserts what the code controls, with the measurements to justify it The single 'max < 1.5s' threshold was asserting a property of the MACHINE, not of the emitter. Measured, 40 boots under deliberate 8-way CPU saturation on top of an already-loaded host: p50 1.178s p90 2.385s p95 2.628s max 2.761s over 1.5s: 13 over 3.0s: 0 The same boots at ambient load ran 0.55-1.11s. The emitter's deliberate spend is bounded and small; total wall clock is dominated by CPU availability, which no budget logic can bound. That is why the gate kept flipping. Re-set with justification, and deliberately NOT a relaxation - it is three assertions where there was one: MEDIAN < 1.5s the design's number, kept as what a wall-clock figure can honestly be on a shared machine. Margin is real (0.65s ambient, 1.178s saturated) and it moves sharply if the deliberate spend regresses. MAX < 5.0s guards the actual hazard: the DECLARED hook timeout is 10s and overrunning it injects nothing. 5s is 2x headroom under that and 1.8x above the worst saturated boot ever measured. The pre-split design lands at ~10.3s and is still caught. CONCURRENCY new, and deterministic rather than a timing margin. Both helpers wedged with a 5s allowance: concurrent costs ~5.4s, serial ~10.4s, and the 7.5s bound sits between them with over 2s of clearance either way. THIS is what catches a regression to serial helpers - the old max threshold never tested it. A 3s allowance was tried first and rejected on measurement: parallel measured 3.72s typical but spiked to 4.62s under load, too close to a midpoint bound. Verified under the harness's own invocation: normal 0/0/0, mutated 1 (median 5.292s vs the 1.5s bound). * chore(gates): sweep and freeze - seven frozen, none vacuous * fix(gates): m4 asserts load-immune properties, not the machine's scheduler Your 3/10 reproduction was right, and after fixing the leak, the instrumentation and the serial helper phase it STILL flaked - 1/10 on a run where load spiked to 205 and boots reached 4.5s. I stopped tuning the threshold at that point, because tuning it until green is the false green this ledger exists to prevent. The measurements say the observable was wrong, not the number: ambient load ~150 boots 0.55-1.11s load ~205 boots 1.41-4.49s 8-way saturation, 40 samples: p50 1.178s p90 2.385s max 2.761s over 1.5s: 13 over 3.0s: 0 Same code throughout. Elapsed time is dominated by CPU availability, which no budget logic can bound, so a second-scale wall-clock threshold asserts a property of the MACHINE. Two observables the code genuinely controls replace it, both with order-of-magnitude margins that load cannot flip: 1a THE DEADLINE IS ENFORCED. Helpers wedged for 999s must not be waited for. Enforced the boot costs ~1s; unenforced, 999s. The 15s bound has 3x clearance below and 66x above, and still catches the pre-split design (two serial 5s timeouts, ~10.3s). 1b THE HELPERS ARE CONCURRENT. Each stub now records the instant it starts. Concurrent, the two starts are milliseconds apart; serial, the second starts a whole allowance after the first. Measured spread 0.386s against a 5s serial signature - 13x, and a structural observation rather than a timing threshold. The old max-threshold never tested concurrency at all. Every boot's elapsed time is still printed in the gate output, so the real numbers stay visible to a human even though pass/fail no longer rests on them. Mutation now hands the emitter a budget with no teeth (20s per helper under a 60s ceiling); it waits the wedged helpers out and breaches the 15s bound. Verified: normal ok, mutated 'worst of 5 boots took 25.493s'. * perf(gates): m4 stops early once a boot blows the bound The verdict is decided by the first breach; running the remaining boots only made a failing gate slow. Mutation arm drops from ~100s to ~20s, with no change to a passing run. * chore(gates): sweep and freeze - seven frozen, none vacuous, WIP is the two declared reds * docs(gates): stop the ledger claiming what m4 no longer proves; record condition 2 honestly Two fixes from the pre-check, both the same defect class this branch has been fixing all along: stated behaviour drifting from checked behaviour. 1. THE LEDGER WAS LYING. m4's observable still read 'exits 0 in under 1.5s wall clock' after the test stopped asserting it, and the title still said 'stays inside its wall-clock budget'. The ledger is what this repo points at when it says 'proven', so stale prose there is worse than stale prose anywhere else. Both rewritten to state exactly what the test now checks - enforced deadline with its clearances, concurrency by start-time spread, zero leaks, the cap, peer completeness, zero peer execs - and to say plainly that elapsed wall clock is measured and printed but NOT asserted, with the reason. The stale header comment in the test is corrected to match. 2. END-GOAL CONDITION 2 HAD NO HOME. Its 'under a second' half rode entirely on the wall-clock threshold that was removed, and m4 is the hostile path anyway, so it was never the right place for a normal-path latency claim. m4 gains a normal-path section that machine-checks the two clauses that CAN be: no network sweep (curl/wget/nc/ssh shimmed to log and fail, log must stay empty) and the structural properties that determine latency (at most two helper execs for the whole boot however many peers, none attributable to a peer). Latency is measured and printed - 0.23s against the live primary home, 0.39-1.06s on the fixture - but not asserted. The spec now records, in a table with the measurements, that the sub-second figure is NOT machine-checked and why: the same unchanged emitter measures 0.55s and 4.5s depending only on host load, so gating it asserts a property of the machine. END-GOAL demands machine-checked rather than asserted, so this is a real gap, stated as one. Recorded alongside it, because it surfaced while proving this: on a host at load ~130 the NORMAL boot can spend its whole 1.5s budget on interpreter startup and skip both helpers, rendering watcher and steering UNAVAILABLE. The degradation works and is never silent, but the budget is tight enough that ordinary boots degrade on a busy machine. Raising it trades injected detail against the declared 10s hook timeout - the captain's call, not mine. * chore(gates): sweep and freeze - seven frozen, none vacuous * fix(boot): close a killpg race that orphaned ~1 process in 50 boots The leak assertion tripped once in a 10-run sweep with an orphan that survived the settle, so it was a real residual leak and not only asynchronous reaping. Cause: os.getpgid(proc.pid) races the child's setsid(). Between fork and setsid the child still sits in OUR process group, so getpgid can return that, and the kill would land on this process rather than on the helper - or, when it failed, fall back to killing only the direct child and leave the grandchild orphaned. start_new_session makes the child its own group leader, so its pgid IS its pid. Targeting os.killpg(proc.pid) directly removes the race entirely and can never address our own group: if the group does not exist yet the call simply fails harmlessly. The group is then swept a second time after the child is reaped, which catches a group that formed just after the first attempt. Verified: 72 consecutive hostile boots across both budget configurations, zero processes leaked. * chore(gates): sweep and freeze after the killpg race fix * chore(gates): re-freeze g-boot-digest (spurious vacuous under load) * test: reap stray stub processes on every exit path, via a committed helper Cleanup only ever ran on the happy path. A run that failed an earlier assertion left through `fail` and abandoned whatever it had spawned - and failing runs are exactly when strays accumulate, being the runs least likely to tidy up. That is how 194 orphaned 'sleep 999' processes once piled up and carried the machine's load average to 186, which then slowed the very boots the budget gate was timing. I cleaned that up at the time with an ad-hoc 'ps | awk | kill' pipeline typed at a prompt. It worked, but nobody could review it, and a process-killing command nobody reviews is not the thing to leave as the answer. tests/fm-reap-strays.sh is that command, committed and readable. It is deliberately narrow, because it kills processes. A stray must have a FULL argv of exactly 'sleep 999' (not a substring match that could catch an unrelated sleep), be owned by the invoking user, be orphaned to init, AND be absent from a snapshot taken before the run - so it is provably this run's own debris and anything already alive is protected by construction. That last condition is what makes it safe to wire into a trap. m4 now snapshots at start and reaps from an EXIT trap, following the pattern tests/lib.sh documents: define your own trap, call fm_test_cleanup from inside it so the temp dirs are still removed. Verified: the reaper detects an orphan, PROTECTS it when snapshotted, and reaps it only when it is not; and after running m4's mutation arm - which fails, so it exits through the sad path - zero strays remain. * test: fix the temp-dir leak at its root, and clear the backlog reviewably tests/lib.sh registered temp dirs in an array and installed its EXIT trap INSIDE fm_test_tmproot - which is called as $(fm_test_tmproot x), so both happened in a command-substitution subshell. Two consequences, neither theoretical: - the subshell's trap fired the instant the function returned and deleted the directory it had just created, so every caller received a path that did not exist. Tests survived only because each mkdir -p's its own subpaths, and it is what made an earlier m1 edit fail until I added a redundant mkdir. - the parent's array stayed empty, so fm_test_cleanup never removed anything. 2,308 temp directories had accumulated across all 47 suites before this was found - most of them not from this branch's tests. Registration now crosses the subshell boundary through a FILE, and the trap is installed at source time, which is the parent's shell. Measured: a full suite run used to leave ~+60 directories behind and now leaves +1. tests/fm-reap-strays.sh gains tmpdirs / reap-tmpdirs to clear the backlog, with the same narrowness as its process modes: only directories directly under TMPDIR, only names matching the suites' mktemp template, only the invoking user's, and only older than an age guard (default 60m) so a suite running right now - here or in another worktree - is never touched. 2,210 stale dirs removed; the 112 recent ones were correctly left alone. Full suite after the change: 47 ran, 2 skipped, 0 failed. * chore(gates): sweep and freeze after the cleanup fixes * fix: a broken test is never a skip, and an unreadable fleet dir is never empty QUARTERDECK REJECT (attempt 3). Both findings reproduced before fixing. 1. A SYNTAX ERROR WAS LAUNDERED INTO A GREEN BUILD. Reproduced: a test whose entire body was `fi` exits 2, and tests/run-all.sh reported 'SKIP zz-broken.test.sh - prerequisite missing' with the suite exiting 0. bash exits 2 on a syntax error as well as on this repo's missing-prerequisite convention, and I had inferred the skip from the exit code alone. That is a fail-open in the one place this runner promises to fail closed, and worse than the problem it solved: a test that cannot run at all is precisely what CI exists to catch. A skip must now be CLAIMED, not inferred. Two independent defences: - every test is syntax-checked with `bash -n` BEFORE it runs; one that does not parse is a hard failure named as a broken test - exit 2 is a skip only when the output carries 'PREREQUISITE MISSING:'; any other exit 2 fails, and says why it was not treated as a skip fm-boot-m0 and fm-loop-l2 now declare the marker. Verified all three outcomes: syntax error FAILS, undeclared exit 2 FAILS, declared exit 2 SKIPS. gate-ci-declared-red freezes it as case H. 2. AN UNREADABLE FLEET DIR REPORTED AS AN EMPTY FLEET. Reproduced: chmod 000 on the fleet dir rendered '(no shared fleet view yet - only this home is reported)' with no marker. Peers may exist and be invisible, so that line claims the fleet is just this home when nothing is known - in the one section whose whole job is answering what the fleet is doing. Same defect as counting an unreadable state/ to zero, one level out. Three states now separated: ABSENT stays ordinary absence and says so plainly (there is no writer for it yet); UNREADABLE gets an explicit UNAVAILABLE marker stating that peers may exist and this is NOT a report that the fleet is only this home; readable-but-empty is genuinely no peers. Also fixed alongside it, same class one level down: a peer record that parses but lacks in_flight/needs_decision was rendered as '0 in flight, 0 need a decision', reporting live work as none. It is now marked UNAVAILABLE naming the missing fields. m5 freezes both, and still asserts that ordinary absence raises no marker - or the marker would mean nothing. * chore(gates): sweep and freeze after the reject fixes * fix(boot): a blocking file cannot hang the hook, and truncation is never silent QUARTERDECK REJECT. Both reproduced before fixing, both mine, and both were raised by the foreign lens in an earlier round while I fixed only its other items - that is the miss, not just the bug. 1. A FIFO HUNG THE ENTIRE BOOT. read() called bare open() with no regular-file check, and open() on a FIFO waits for a writer that never comes. Measured against a 0.05s normal boot: state/.wake-queue = FIFO killed at 10s, NO OUTPUT fleet/peer-x.json = FIFO killed at 10s, NO OUTPUT Zero injected context and no marker - the exact failure this file spends thirty lines swearing it prevents - and it defeated the budget completely, because the deadline only ever bounded subprocess helpers, never a read. So gate m4 was green while asserting it holds its budget by enforced deadline. fleet/*.json is the case that matters most: that directory is written by OTHER firstmate instances, so one bad entry would blind every other captain's boot, permanently and silently. safe_read() opens with O_NONBLOCK - which returns immediately even on a FIFO - then fstats the descriptor it actually holds and rejects anything that is not a regular file, naming what it was. Checking with stat() BEFORE opening would leave a window for the path to be swapped; inspecting the open descriptor does not. Every read now routes through it: statuses, metas, the wake queue, peer records, projects, secondmates, backlog. Measured after: every FIFO case returns in 0.06-0.07s with a marker naming the file. 2. TRUNCATION WAS SILENT, AND DROPPED THE LINE THAT MATTERED. read() took exactly READ_LIMIT bytes and could not tell a file ending there from one that does not - while its own comment promised that truncation is reported and never silent. Measured: a 400KB status file rendered with no marker and dropped its real last line, a done: line, so the boot reported a FINISHED task as still working. Confidently incomplete is the same lie as unreadable-reads-as-empty. One byte past the limit is now requested and its presence proves the truncation, which is reported and replaces the stale prefix rather than being presented as the task's live state. Also surfaced while wiring it: an unreadable .meta went through read_or(), which discards the reason, so the task rendered window=? kind=? mode=? with no explanation. meta_of() now returns its problem too. Gated both. m4 gains a blocking-file arm covering .wake-queue, .meta, .status and a fleet peer, bounded at 15s - a hang is 30s+ or infinite and a healthy boot is under a second, so the bound discriminates by orders of magnitude and cannot be flipped by load. m5 gains a truncation arm asserting the marker appears, the file is named, the stale prefix is NOT presented as live state, and a normal status raises no marker so the marker keeps its meaning. * chore(gates): sweep and freeze after the blocking-read fixes * no-mistakes(review): fix: a status report never lands in the wrong home * no-mistakes(review): fix: degraded boot paths never guess or vanish * no-mistakes(review): test: prove all three steering verdicts, never guess one * no-mistakes(review): test: prove the status verb in all three brief kinds * no-mistakes(review): test: model only the permission rule that really refuses * no-mistakes(document): docs: sync boot context, status verb, and test runner * no-mistakes(lint): silence deliberate SC2016 and pin lib.sh source in tests * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…headless fallback (#6) * test(mux): freeze gates h1-h7 red before wiring the seam Seven gates, both directions, frozen red against the pre-seam bin/ before a line of it changes. They are frozen in one commit rather than seven because gates/ledger.json is a shared file and fm/fmx-boot-e3 lands ahead of this branch; one mechanical ledger edit keeps that rebase trivial. Each gate's red-before evidence is recorded below, captured by running these tests against the pre-seam tree. Herdr direction (the default the captain asked for): h1 driver selection - herdr whenever a server is REACHABLE, loud tmux fallback, FM_MUX honoured in both directions h2 workspace scoping - resolved by project label, FM_HERDR_WORKSPACE overrides, absent is created, an unscoped tab create is unreachable h3 fm-spawn lands a named tab in the project's workspace and records the target and its driver in the meta h4 fm-send delivers with acknowledgment and fm-peek reads back, both routed by the driver recorded in the meta h5 LIVE herdr round-trip, end to end against a real server Tmux direction (the explicit fallback, which must stay a true rollback): h6 FM_MUX=tmux emits the pre-seam tmux call sequence byte for byte h7 LIVE tmux round-trip, end to end against a real server RED-BEFORE EVIDENCE, run against the pre-seam bin/: h1 not ok - default: no FM_MUX and a reachable server must select herdr (expected 'herdr', got 'tmux') h2 not ok - a project must resolve to ITS workspace, not the focused one (expected 'wM', got '') h3 not ok - the spawn did not succeed (missing: 'spawned fleet-view-q4') h4 not ok - the steer was not delivered through the acknowledged prompt path h5 SKIP - the pre-seam lib has no reachability predicate to call, so the gate cannot even determine whether a server is up; the skip is itself the absence being gated h6 ok - parity: FM_MUX=tmux emits the pre-seam call sequence (expected: this half pins OLD behaviour and must pass before and after) not ok - the tmux path did not record its driver h7 not ok - window_exists did not see a window that tmux really has The fakes are not yes-machines. tests/mux-helpers.sh reproduces the constraints verified against herdr 0.8.2 - agent names rejected unless ^[a-z][a-z0-9_-]{0,31}$, agent_blocked before any input is sent, agent_prompt_stalled after the text has already gone in, and a refusal of any tab create with no --workspace - so a gate cannot pass here against behaviour the real binary would reject. * feat(mux): herdr by default, resolved workspaces, and the shell-pane verbs The seam shipped complete and nothing called it. Before it can be called, three things in it were wrong or missing. DRIVER SELECTION. It chose herdr only when HERDR_ENV=1. Live evidence kills that predicate: HERDR_ENV is unset in firstmate's own process while a herdr server is running and the captain is watching it, so every crewmate would have landed in an invisible tmux session under the rule meant to put them in herdr. The predicate is now REACHABILITY - the binary is present and `herdr session list` reports a running server. That is the same check bin/fm-herdr-workspaces.sh already encoded, so it moves here and that script defers to it rather than keeping a second copy. An unconditional herdr default would strand every context with no server - cron, CI, a plain SSH session - so the seam still degrades to tmux. It now says so on stderr, once per process, naming the driver and the reason. A silent degrade is the worse bug: a captain who believes he is watching the fleet in herdr while the crew land somewhere invisible is worse off than one who knows. WORKSPACE SCOPING. `herdr tab create` was called with no --workspace, so a tab landed in whatever workspace happened to be focused. That is luck, not targeting. The workspace is now resolved explicitly by project label, overridden by FM_HERDR_WORKSPACE (label or id, refused loudly if it names nothing live), and CREATED when absent - the same thing fm-herdr-workspaces.sh --apply does, so the spawn path and the reconcile path agree. Ids are never hardcoded: `wJ` is a runtime value that does not survive a restart. NEW VERBS. Wiring fm-spawn needed what the seam did not have: scope, window_exists, run, run_launch, cwd, launch_failed, label, send_key, and resolve. run and run_launch differ only under tmux, and only because the pre-seam code did - `treehouse get` went out as one `send-keys <cmd> Enter` call while the launch line went out as `send-keys -l`, a sleep, then `Enter`. Both shapes are preserved exactly so FM_MUX=tmux is a true rollback rather than a similar-looking code path. NAMING. herdr 0.8.2 rejects an agent name that is not ^[a-z][a-z0-9_-]{0,31}$, so the convention is <project-short>-<work>, one hyphen, under 28 chars. The sanitizer and the validity check live here; fm-herdr-workspaces.sh remains where the convention is documented. SEND RETURN CODES. herdr accepts a submission and only then waits for a state change, so agent_prompt_stalled arrives AFTER the text has gone in - observed live against a claude TUI herdr reported as idle while it was still booting. That is now code 4, delivered-but-unconfirmed, not a failure: calling it a failure would make firstmate re-send a steer the crewmate already has, which is the worse of the two available errors and the opposite of the tmux path's standing rule that only a positively confirmed swallow counts as not-sent. * fix(herdr): the pane-name separator is a hyphen, because herdr rejects a slash fm-herdr-workspaces.sh --name built `<project>/<work>` and handed it to `herdr agent rename`. Verified against herdr 0.8.2, that call can never succeed: $ herdr agent rename wN:p2 'probe/mux-recon' {"error":{"code":"invalid_agent_name","message":"agent name must start with a lowercase letter and contain only lowercase letters, digits, '-' or '_' (1-32 characters)"}} So every name the helper produced was rejected and every pane it "named" kept a name nobody chose - the opposite of the standing order it exists to enforce. The convention becomes <project-short>-<work>: afs-resource-registry, mac-config-cutover-guard. Same information, one separator herdr accepts. The helper also now proves the rename landed instead of assuming it, and names both slots - the tab label the captain reads and the agent address herdr steers by - with the same string, because they are the same name. Its test previously passed against a fake herdr that accepted anything, which is exactly how a gate goes green while the fleet ends up unaddressable. The fake now enforces herdr's real name rule, so restoring the slash turns it red. The reachability predicate moves to bin/fm-mux-lib.sh and this script defers to it: one fact, one owner. * feat(spawn|send|peek): drive panes through the multiplexer seam The three entry points firstmate uses to create, steer, and observe a direct report now address panes through bin/fm-mux-lib.sh instead of calling tmux directly. With no FM_MUX set and a herdr server reachable, a spawned crewmate is a tab in that project's workspace, named for the work it is doing. FM_MUX=tmux reproduces the previous behaviour byte for byte (gate h6). fm-spawn Window creation, the treehouse-get wait, the cwd poll, shell readiness, the launch, and launch verification all go through seam verbs. --name <work> supplies the work half of the pane name; absent, the task id's random suffix is dropped and the rest is used, so `boot-activation-k3` in archify becomes `archify-boot-activation`. Batch dispatch refuses --name, since one name cannot address several panes. Untouched, deliberately: the worktree-isolation assertion, the turn-end hooks, the intake gate, and every state-file semantic. The meta gains two lines and changes none - mux= (the driver that minted window=) and name=. fm-send / fm-peek Both resolve the target AND its driver from the meta, never from ambient resolution. That is the load-bearing property: window= is opaque, so a crewmate spawned as a herdr tab must stay steerable from a process where herdr no longer resolves, and a tmux crewmate must never be probed with herdr verbs because a server happens to be up. A meta with no mux= predates the seam and is a tmux target by definition. A raw `session:window` argument carries no meta and is documented as a tmux address, so it stays one unless FM_MUX says otherwise - ambient resolution there would answer a tmux address with herdr verbs and quietly return nothing, which is how the composer-ghost suite caught this during development. Under herdr, fm-send delivers with `agent prompt --wait`: one acknowledged call in place of the type-then-retry-Enter dance, a distinct refusal when the agent is blocked at an approval dialog, and a warning rather than a false failure when delivery is unconfirmed. The tmux path keeps its verified-submit logic unchanged. tests/lib.sh pins FM_MUX=tmux for every suite. Since the seam now defaults to herdr whenever a server is reachable, an unpinned run on a developer machine with herdr up leaves the fakes untouched and creates real tabs in real workspaces - which is exactly what happened once during this work, leaving a stray `design-home` workspace behind. Suites exercising the herdr driver set FM_MUX themselves. * test(mux): fake both multiplexers in the spawn gate, not just herdr h3's FM_MUX=tmux case only had a fake herdr on PATH, so it drove the REAL tmux server and left `firstmate:fm-fallback-t8` in the captain's live session. The next run then failed on the duplicate-window guard - correctly, which is how it surfaced. Same hazard the tests/lib.sh pin closed for herdr, in the other direction: a spawn test must be unable to reach ANY live multiplexer, so it fakes both. The fake tmux also now answers #S and #{pane_current_path} separately; answering both with one value would have resolved the treehouse-get wait instantly to a bogus worktree and made the isolation assertion meaningless. * docs(agents): record the standing herdr rule, and flip gates h1-h7 green AGENTS.md now carries the rule itself, so it survives without the captain having to restate it: agents are spawned in herdr, that is the default, FM_MUX=tmux is the explicit fallback for headless contexts, and an unreachable server degrades loudly rather than silently. It also records what the three entry points now do, how a crewmate is routed by the driver recorded in its meta, and the pane-naming convention with the reason the separator is a hyphen. It also records what is NOT migrated. fm-watch.sh, fm-teardown.sh and fm-ff-lib.sh still read window= and call tmux on it, so for a herdr-spawned crewmate their stale-pane detection and window kill are inert. Supervision itself is unaffected - status files, heartbeats and per-task checks are multiplexer-agnostic - but a closed herdr tab needs tidying by hand until those callers move onto the seam. Migrating them was out of scope for this slice and is stated rather than left to be discovered. Gates h1-h7 flip to green. Each was observed red against the pre-seam implementation before any of bin/ changed; that implementation is the mutation these gates were verified against, and the evidence is in the red commit. * test(mux): re-gate h1 red - herdr is the only default, headless is an escalation Captain's rule (data/captain.md, "Where agents run", 2026-08-28) supersedes the reachability predicate this branch was built on: "Headless is not an automatic fallback and must never be selected by a reachability rule. Any tier may *recommend* a headless run, but it asks me first... Silent degradation to headless is the specific failure this rule exists to prevent." The previous h1 pinned the wrong shape. It made the degradation LOUD, which is better than silent but is still not asking - the code chose headless on the fleet's behalf and then told the captain about it. Under the rule as written, reachability may not decide at all. h1 now pins the opposite: with no FM_MUX the driver is herdr under every combination of server state, binary presence and HERDR_ENV; selection consults nothing (a rule that asks is a rule that can answer 'headless'); and an unreachable herdr escalates through a precondition check that names the problem, whose decision it is, and what the captain would say to authorise a headless run. FM_MUX stays the explicit human-chosen override in both directions. h4 also gains the ordering case that a colon-bearing target is an explicit address checked before the fm-<id> name shape, as the pre-seam resolve() had it. RED-BEFORE EVIDENCE, against the reachability implementation on this branch: not ok - an unreachable server must NOT auto-select headless (expected 'herdr', got 'tmux') * fix(mux): herdr is the only automatic choice; an unreachable herdr escalates Applies the captain's ruling (data/captain.md, "Where agents run") and the correction that followed it. The reachability-with-loud-fallback design was wrong: the rule is not "degrade loudly", it is "do not degrade". Announcing a degradation on stderr still means the machine chose headless on the captain's behalf, and a warning in a log nobody is reading at 2am does not close the gap between believing he is watching the fleet and actually watching it. The selector, in full: fm_mux_driver() { if [ -n "${FM_MUX:-}" ]; then printf '%s' "$FM_MUX"; return 0; fi printf 'herdr' } Nothing is consulted. There is no branch that can produce a headless driver on its own, because a rule that asks is a rule that can answer "headless". fm_mux_announce_fallback is deleted rather than renamed - a helper whose name blesses the forbidden behaviour has no business surviving the ruling. Reachability is still checked, by fm_mux_require_available, which ESCALATES: non-zero, naming the reason (no server running, or herdr not on PATH), saying whose decision it is, and giving the exact thing the captain would say to authorise a headless run. fm-spawn calls it before creating anything and stops on failure, so an unreachable herdr yields no window, no tab and no meta. An explicit FM_MUX=herdr is still checked - choosing herdr does not make an unreachable server reachable - while an explicitly chosen non-herdr driver is a decision already made and is never second-guessed. AGENTS.md's degrade paragraph is replaced rather than amended, so the tree does not carry two conflicting descriptions. The stale prose in fm-spawn's header goes with it. Gates, re-cut in both directions and mutation-checked as the correction asked: h1 herdr under every combination of server state, binary presence and HERDR_ENV; selection probes nothing; the escalation is actionable; FM_MUX honoured silently in both directions h3 with herdr unreachable and no FM_MUX the spawn FAILS non-zero and creates no window, no tab and no meta; FM_MUX=tmux still makes a tmux window MUTATION CHECK - the reachability fallback restored into fm_mux_driver: h1 not ok - an unreachable server must NOT auto-select headless (expected 'herdr', got 'tmux') rc=1 h3 not ok - the spawn succeeded with no herdr server rc=1 Both green again on revert. The previous gates passed while the rule was being broken, which is what made them the wrong measurement; these catch it. * fix(mux): review findings - stale test, dead warning branch, mislabelled rc Three real defects from an independent review of this branch, plus three smaller ones. All six were found by reading the code, not by a failing gate, so each that could regress silently now has a gate. 1. tests/fm-mux-lib.test.sh still asserted the reachability fallback that the captain's rule removed, so the tree did not pass its own suite and CI would have gone red on `tests/*.test.sh`. Inverted to match: neither a stopped server nor HERDR_ENV may change the answer. 2. fm-spawn read `$?` inside `if ! fm_mux_label ...; then`, where it is the negation's status - always 0 - so the rc=2 branch was unreachable dead code and a pane herdr REFUSED to name was silently left unlabelled. That is the exact defect the naming work exists to remove, reported by a warning that could never print. The status is now captured directly, and gated: h3 drives a pane get with no tab_id and asserts the operator is told. 3. fm_mux_herdr_send's agent_not_found branch fell through to blind shell delivery and returned that call's status - 0 on success. Code 0 means ACKNOWLEDGED, so a blind steer sent in the beat before herdr classifies an agent was reported identically to a confirmed one, erasing the distinction the whole driver exists to provide. It returns 4 now, landing in fm-send's existing "took the steer but did not acknowledge it" path. Gated in h4. 4. The prompt-failure match was the bare substring "error", which a future success payload carrying a field named e.g. last_error would trip. Anchored to the top-level {"error" envelope. 5. fm_mux_herdr_field's split-on-{ extraction is exact for the shapes it reads, but tab create nests root_pane.tab_id beside tab.tab_id and only works because those are identical today. The assumption is now written down next to the code that rests on it. 6. fm-herdr-workspaces.sh --name treated fm_mux_herdr_label's rc=1 - "tab named, no agent detected yet", explicitly non-fatal - as a refusal, reporting a failure that had not happened. It now distinguishes the two. Also confirmed while triaging: tests/fm-afk-inject-e2e.test.sh is flaky independent of this branch (2 of 3 runs pass; it failed on different scenarios across runs and passes on an unmodified checkout of main). The away-mode daemon uses bin/fm-tmux-lib.sh directly and touches neither fm-send.sh nor the seam. * refactor(herdr): collapse the seam into one herdr-native bin/fm-herdr.sh Captain-directed. bin/fm-mux-lib.sh and bin/fm-herdr-workspaces.sh were split across a driver-selection seam that no longer has two drivers to choose between, so they become one file: bin/fm-herdr.sh. The name is deliberate. NOT a bare `herdr`, which would shadow the real binary at ~/.local/bin/herdr and make which one a call site got depend on PATH order. WHAT GOES AWAY * driver selection - fm_mux_driver, fm_mux_dispatch and the whole per-verb dispatch layer, rather than being reduced to a constant. A rule enforced by a branch is a rule that can be branched around; a rule enforced by there being no branch cannot. FM_MUX is read by nothing. * the opaque-target discipline. It existed so a tmux session:window and a herdr id could travel through one contract without either caller knowing which it held. With one surface there is nothing to hide behind: a target is a herdr PANE ID and is named as one. WHAT STAYS EXACTLY AS IT WAS herdr is the only surface; an unreachable herdr STOPS the spawn and escalates, never degrades. Pane names are project-first with a single hyphen under 28 chars, gated against the real binary. Workspace resolution is explicit and never focus-luck. THE DRAIN - the constraint this was not allowed to break Crewmates spawned before the cutover live in tmux windows. They must stay READABLE (or the watcher is blind), STEERABLE (or a supervisor can watch work go wrong and not correct it) and CLOSABLE (or teardown strands work carrying unlanded commits). So bin/fm-tmux-lib.sh is retired for NEW use - nothing spawns onto it, fm-spawn no longer even sources it - and KEPT while any pre-cutover meta exists. The drain is DERIVED FROM THE METAS, never a hardcoded inventory, and that is not a stylistic choice. The list this migration was handed named four windows; the tree actually holds seven crewmates with metas - it missed cellarsky-sm, csky-port-r7 and afs-sm in the main firstmate home, and named two whose metas live in a home the author had not enumerated. A list has to be right. A meta-derived rule is correct under any inventory, needs no census, and closes by itself: fm_herdr_drain_pending answers "is there anything left to drain", and no live meta anywhere carries a mux= line today, so all seven are covered. CALLER SWEEP - option (a), taken deliberately The dangerous case does not arise: fm-tmux-lib.sh is NOT deleted, so fm-supervise-daemon.sh and fm-context-watch.sh are untouched and keep working. What was deleted are the two libraries this branch itself introduced, and every reference to them is swept in this same commit - tests, AGENTS.md, docs/scripts.md, docs/architecture.md, the two skills, and a superseded banner on docs/plans/cmux-herdr-surface-split.md. Code and prose move together; this fleet has had three documentation-vs-code drifts today and this is not a fourth. GATES h6 AND h7 ARE RE-CUT, NOT RE-WORDED They asserted tmux parity, which silently changes meaning under a drain: parity with a retired path proves nothing. They are new gates with new ids, because a gate whose text stays put while its meaning drifts is worse than no gate. h6 nothing new is created on tmux; the drain keeps read, send and close reachable; fm-tmux-lib.sh is retired but kept; drain-pending is computed h7 LIVE: a real pre-cutover tmux pane is still readable through the real fm-peek.sh, steerable through the real fm-send.sh (text and control key), routed from its meta, and closable MUTATION CHECKS, all three caught: 1. drain branch removed from fm-peek -> h6 red ("no drain branch; a live pre-cutover crewmate is unreadable") AND h7-live red against a real pane 2. drain made list-derived instead of meta-derived -> h6 red 3. a tmux spawn path put back into fm-spawn -> h6 red and h1 red All green again on revert. FIXED WHILE COLLAPSING, caught by the gates: the JSON field extractor lost a quote in the move, so every workspace id, pane id and cwd came back empty; and `local FM_HOME` in the reconcile CLI blanked the caller's home before the default could read it. Both now have cases in tests/fm-herdr-lib.test.sh and tests/fm-herdr-cli.test.sh. * test: keep suites off the live herdr server, and fix the fixtures the cutover broke THE HAZARD, twice realised. herdr is now the only surface, so any suite driving fm-spawn/fm-send/fm-peek without faking `herdr` reaches the captain's REAL server. That is not hypothetical: this run left stray `design-home`, `spawn-proj` and `alpha` workspaces in his sidebar, and an earlier one left `firstmate:fm- fallback-t8` in his tmux session. Both were found only because a later run tripped a duplicate guard. THE NET. tests/lib.sh used to pin FM_MUX=tmux, which is meaningless now and should not come back - headless is not a flag anyone gets to flip. Instead it prepends a PATH shim: a `herdr` that refuses loudly with exit 97 and says how to fake it. A suite with its own fakebin never sees the shim; a suite that forgot gets an obvious, greppable failure instead of silently touching the fleet. The two live gates set FM_TEST_ALLOW_LIVE_HERDR=1 before sourcing, stated at the top of each file where it is visible. FIXTURES REPOINTED. fm-tangle-guard and the secondmate lifecycle e2e drove fm-spawn with a fake tmux only. They now fake herdr too, and their assertions follow the delivery rather than the transport: the launch line is recorded by the fake herdr instead of the fake tmux, and the send case asserts the pane recorded in the meta is the one addressed - still beating a foreign same-named window, which was always the property under test. h7's fixture needed FM_COMPOSER_IDLE_RE. Its pane holds a bare SHELL, not an agent TUI, so the composer detector read an idle `%` prompt as unsubmitted text and fm-send reported a swallowed Enter for a line that had actually landed. That variable is the documented per-harness knob for exactly this; production behaviour is unchanged. * test(herdr): freeze the two rejected defects red, plus the hazard beneath one Quarterdeck rejected attempt 1 on two counts. Both reproduced here before any fix, and both had a gate-shaped hole where a gate should have been. 1. A FAILED STEER REPORTED AS ACKNOWLEDGED DELIVERY. fm_herdr_prompt ran `herdr agent prompt ... || true`, discarding the exit status, then matched stdout against a short list of known error strings; anything unmatched fell through to `*) return 0` - the function's own contract for "delivered AND acknowledged". Reproduced with a stub exiting 7: $ stub prints 'socket closed: connection reset by peer', exit 7 fm_herdr_prompt pane-xyz 'git reset --hard origin/main' -> RC=0 fm-send.sh ... -> RC=0, silent This is the worst failure this seam can have. The acknowledgment is the whole reason herdr replaced blind keystrokes; reporting a dropped socket as a confirmed steer is worse than tmux ever was, because tmux never claimed to know. Gated in fm-herdr-lib and h4: a nonzero exit is a failure whatever the message says, and an error envelope fails even on a zero exit. 2. "TEARDOWN CLOSES IT" WAS NEVER IMPLEMENTED. fm-teardown.sh passed the meta's window= - a herdr pane id for every post-cutover crewmate - to `tmux kill-window`, swallowed the failure with `|| true`, deleted the meta and printed "teardown complete". The tab leaked, untracked and possibly still running an agent, while firstmate believed the work was cleaned up. fm_herdr_close existed and worked; teardown never called it. The only prior coverage was h6's grep asserting `tmux kill-window` is still PRESENT for the drain - which proves the drain survives and says nothing about whether a herdr pane is ever closed. A gate can be green while the thing it is named for has never been done. New gate h8 covers the routing, the failure report, that fm-teardown actually reaches the helper, and a LIVE case that a real tab is really gone. 3. AND THE HAZARD UNDERNEATH (1). On agent_not_found the steer was forwarded to `herdr pane run`, whose job is running SHELL COMMAND LINES, and reported as merely unacknowledged - exit 0 with a warning. A crewmate instruction like `git reset --hard origin/main` would EXECUTE in the worktree. Verified: agent prompt w9:p2 git reset --hard origin/main --wait --timeout 15000 pane run w9:p2 git reset --hard origin/main The h4 case that existed BLESSED this ("an undetected agent gets a blind delivery, reported as unacknowledged"). Blind delivery into a TUI composer and into a shell are not the same risk; that case is inverted. RED-BEFORE EVIDENCE, against the code as rejected: lib not ok - herdr failed with exit 7 and the steer was reported as delivered AND acknowledged h4 not ok - a steer into a pane with no agent was reported as delivered h8 not ok - a herdr pane was not closed with herdr * fix(herdr): the exit status is authoritative, and teardown actually closes the pane Both Quarterdeck reject reasons, fixed and mutation-checked. 1. A FAILED STEER IS NO LONGER REPORTED AS ACKNOWLEDGED DELIVERY. fm_herdr_prompt now captures herdr's exit status instead of erasing it with `|| true`. The specific outcomes (blocked, stalled, no agent) are matched first because each says something a bare status cannot; after that a nonzero exit is a FAILURE whatever the message says, and an error envelope fails even on a zero exit so the fix does not swap one blind spot for the other. before: stub exit 7, 'socket closed' -> RC=0, fm-send exits 0 silently after: RC=1, "prompt failed for pane-xyz (herdr exited 7): socket closed", fm-send exits 1 2. TEARDOWN CLOSES THE PANE, ON THE SURFACE THAT CREATED IT. New fm_herdr_close_pane <target> <mux> routes a mux=herdr pane to fm_herdr_close and a pre-cutover window to tmux kill-window, and REPORTS a close it could not do rather than swallowing it under "teardown complete". fm-teardown.sh reads mux= from the meta and routes both its own window and its --force child windows through it. Proven live end to end, not asserted: spawned a real crewmate (pane w1G:p2, tab muxdemo-teardown-proof), ran the real fm-teardown.sh, and the pane is gone with zero tabs of that name left. 3. AND THE HAZARD UNDERNEATH (1): a steer is never executed as a shell command. On agent_not_found the steer used to be forwarded to `herdr pane run`, which submits a shell command line - so `git reset --hard origin/main` would have RUN in the worktree, reported as merely unacknowledged with exit 0. There is now no fallback: code 5, "no agent detected; the steer was NOT delivered, nothing was executed, peek the pane". Verified the `pane run` call is gone. Also fixed, from the same reviews: * fm_herdr_label returned 1 for EVERY agent-rename failure, so a duplicate name, permission or server error was mislabelled as benign startup lag and left a pane unnamed while spawn reported success. Only agent_not_found is lag now; anything else returns 2 and is reported. * The live gates could skip and still be read as proof. FM_TEST_REQUIRE_LIVE=1 turns a skip into a failure, so a verifier can demand real proof; without it they still skip loudly so CI stays green. Proven both ways. * Recovery instructions contradicted the runtime - AGENTS.md section 5 said to check tmux panes and that "all truth lives in tmux", and fm-captain-bootstrap.sh injected the same into every new session. Both now say to route by the meta's mux=. That is the fourth documentation-vs-code drift flagged on this fleet; it is closed in the same commit as the code. Deliberately NOT changed: fm-watch.sh and the context watchdog still cannot see herdr panes. The brief scopes that out ("do not migrate the rest of the ~100 direct tmux calls") and AGENTS.md discloses it; the verifier agreed it is disclosed scope, not a defect of this work. It should land next, with the drain. MUTATION CHECKS, all caught, all green again on revert: A restore `|| true` + fallthrough-to-success -> lib red, h4 red B restore the shell-exec fallback -> lib red C teardown kills the herdr pane with tmux -> h8 red Suite 49/50 (fm-loop-l2 red on main before this branch); shellcheck 17 findings, same classes as main's 20, none new. * fix(herdr): startup fails loudly when the fleet cannot dispatch Three Quarterdeck findings, all reproduced first, all mutation-checked. 1. THE FLIP SIDE OF MAKING HERDR MANDATORY. fm_herdr_require stops a spawn and escalates when no server is reachable - correct - but bin/fm-bootstrap.sh never learned about herdr. TOOLS listed everything except it and install_cmd had no case, so a fresh or headless host printed nothing, exited 0, and then hard-failed on every dispatch. The tier whose whole job is saying "you cannot work yet" said nothing, and the failure arrived later, one escalation per task, for the captain to interpret. herdr is now in TOOLS with an install_cmd, so an absent binary gets the same `MISSING: <tool> (install: ...)` line as every other tool. An installed binary with NO SERVER is a different problem with a different fix, so it gets its own line rather than being folded in - telling someone to reinstall a tool they already have is not help: no binary: MISSING: herdr (install: brew install herdr # or the ...) no server: NEEDS_HERDR_SERVER: no herdr server is running - crewmates cannot be spawned until one is. Start or attach it with `herdr`. Both are documented in AGENTS.md section 3 with what firstmate does about them, and bootstrap uses fm_herdr_up rather than a second copy of the check. 2. THE CAPTAIN DIGEST SHOWED AN EMPTY FLEET. fm-captain-bootstrap.sh inventoried only `tmux list-windows`, so a restarted supervisor saw nothing while herdr crew were live, and was then taught a lifecycle that "opens tmux window fm-<id>" - which spawn no longer does. It now lists both, labelled ("herdr:" and "tmux (draining):"), and teaches the herdr lifecycle including that an unreachable server stops the spawn. Gated behaviourally, not by grep: the real hook runs against a fake server holding a named agent and the gate asserts that agent reaches the captain's context. 3. MY OWN STRICT-MODE GUARD WAS BACKWARDS. The FM_TEST_REQUIRE_LIVE guard I added last round checked the switch BEFORE reachability in h8, so strict mode failed even with the server up - refusing to run the very case it exists to force. Last round's report claimed it was "proven both ways"; for h8 it was proven only in the direction that happened to be right. Reachability is checked first now, and the guard is itself gated: h8 re-enters its own live case under strict mode and asserts it passes with a server up. Proven in all four combinations (live/strict pass, live/lenient pass, dead/strict fail, dead/lenient skip). Also closed: the rest of the doc drift the lens listed. AGENTS.md:96, :358, :512 still described the tmux lifecycle and said "tmux is the ground truth", and :606 claimed teardown's kill is inert for herdr panes - which this branch fixed last round. tests/fm-bootstrap.test.sh gained a herdr stub so its silence cases stay silent now that herdr is required. MUTATION CHECKS, all caught, all green on revert: D drop herdr from TOOLS -> h9 red E drop the NEEDS_HERDR_SERVER check -> h9 red F captain digest back to tmux-only -> h9 red G h8 guard checks the switch before reachability -> h8 red Suite 50/51 (fm-loop-l2 red on main before this branch); shellcheck 17 findings, same classes as main's 20, none new. * fix(herdr): the acknowledgment is real for a busy agent, or it is not claimed Quarterdeck reject 3, and the sharpest of the three. From the binary's own help: "It does not track turns: if the agent is already working, that active turn's completion may match." fm_herdr_prompt called `--wait` with no `--until` and no check of the agent's state before submitting. So steering a crewmate that was ALREADY WORKING - the ordinary case, because that is exactly when a steer is needed - could match the PREVIOUS turn's completion and return 0, "delivered AND acknowledged, the agent consumed it", before the new steer had started. AGENTS.md asserted that as fact while the binary documented the opposite, and no gate covered it: h4 froze only that the --wait FLAG is present, and h5's live case runs against an idle agent. THE FIX reads the state first, because which acknowledgment is even available depends on it: * NOT WORKING - --wait is trustworthy: herdr requires an observed state change before it matches, so a match means the agent moved because of this prompt. --until is now explicit so a change to herdr's default cannot silently alter what "acknowledged" means. * WORKING - --wait's verdict is about the wrong turn, so it is not used. The steer is submitted and firstmate watches for the current turn to END and a new one to BEGIN, which only the queued steer causes. Seeing that is acknowledgment (0); not seeing it inside the budget is reported as delivered-but-unconfirmed (4), never invented as success. PROVEN LIVE against a real claude agent, both directions, same scenario: old code, pre-steer status=working, turn ends inside the budget -> RC=0 (acknowledged off the FIRST turn's completion - the defect) fixed, same setup -> RC=4 "was mid-turn; the steer is queued but no new turn started within the budget. Delivery is UNCONFIRMED, not acknowledged" and the steer DID land afterwards, so unconfirmed-assume-sent was the truthful answer fixed, busy agent with a budget long enough to see the new turn start -> RC=0 in 2s, and the reply arrived - so 0 is not merely conservative TWO GATES I HAD WRITTEN WERE ASSERTING THE OPPOSITE OF THE TRUTH and are re-cut: the lib suite and h4 required --wait on EVERY prompt, which for a working agent is requiring precisely the signal that cannot be trusted. They now cover the idle path, and the fake herdr no longer invents a --wait rule the binary does not have. Busy delivery is gate h10. Supporting finding: tests/fm-herdr-h9-startup ran a REAL package install. A recording shim caught `FAKE-BREW-INVOKED: install herdr` - it reached `eval "brew install herdr"`, and CI runs every tests/*.test.sh on ubuntu-latest with linuxbrew present. It now resolves install_cmd through a lib-only source instead, and a new case shims brew/apt-get/npm and fails if detection touches any of them: a test must not change the host it runs on. Supporting finding: AGENTS.md's disclosure said "supervision still works... and nothing unsafe happens", which flattened two unequal gaps. Stale detection has a fallback (status files, heartbeat review). The CONTEXT WATCHDOG HAS NONE - no herdr crewmate is ever selected for a compaction checkpoint, so one that reaches its context ceiling dies with no handoff. The disclosure now says which is which and that the watchdog moves first. Also: the rule cited data/captain.md, which is gitignored and per-home, so this repo's own section is now the statement of record. Suite 51/52 (fm-loop-l2 red on main before this branch); shellcheck 17 findings, same classes as main's 20, none new. * no-mistakes(review): fix herdr digest names, busy state, peek source, key errors * no-mistakes(review): fix pane-id routing, silent peek, session probe, temp leak * no-mistakes(review): fix test shim leak, read stderr, session probe, docs * no-mistakes(review): fix herdr session probe, already-gone close, stale ledger names * no-mistakes(review): fix test session hermeticity, drain close, stale docs * no-mistakes(review): pin herdr session into panes, fix name suffix stripper * no-mistakes(review): document herdr cutover, pin digest session, gate name derivation * no-mistakes(review): guard tests off live tmux, fix stale gate refs * no-mistakes(review): fix live-tmux guard argv parsing, optional fake CALLS * no-mistakes(review): fix drain-close fail-open, gate test hermeticity and shell-readiness * no-mistakes(review): anchor herdr pane-id match, verify hermeticity chain, fix temp-root leak * no-mistakes(review): fix leading-digit name strip, record ctxwatch waiver * no-mistakes(document): sync docs with herdr crew-surface cutover * no-mistakes(lint): quiet shellcheck SC2016/SC2013 in new herdr tests * chore(gates): migrate the ledger to upstream's array shape in one step fm/fmx-boot-e3 landed while this branch was in review and converted gates/ledger.json from an object to an ARRAY, with a new m0-ledger-shape gate asserting it. That is also the fix for the tooling mismatch this branch has been disclosing all along: the installed `ledger` CLI wants an array, so gates/verify.sh could not run at all here - or on main - before now. WHY THE MIGRATION IS ONE COMMIT RATHER THAN SPREAD ACROSS THE REBASE. Resolving the shape change commit-by-commit corrupted the file on the first attempt, and not through a bad conflict resolution: this branch's LATER ledger commits apply cleanly as TEXT onto the converted array and silently insert object syntax into it, producing invalid JSON with no conflict to notice. So the rebase kept this branch's own object form at every step - internally valid at every commit - and the shape change happens once, here, where it can be checked. The merge is mechanical and lossless. Upstream's 27 gates are taken VERBATIM, including the seven it added (m0, m1, m2, m4, m5, gate-ci-declared-red, gate-status-verb) and its own re-verification timestamps; the only differences in the gates both sides carry were upstream's last_verified stamps and one observable it extended, none of them this branch's edits. This branch's ten h-gates are appended. 27 + 10 = 37, and gates/LEDGER.md is regenerated from the ledger so the table and the drain count cannot drift from it by hand - that count has already been wrong once here. NOT included, deliberately: `bash gates/verify.sh` now runs (green:35 red:2 wip:2, and both red are gate-l2-loop-audit-level and m1-hook-registered, red on main before this branch - all ten h-gates are green). But it REWRITES the ledger as it verifies, and doing so promoted upstream's frozen gates to green. That is a semantic change to another branch's records that this branch has no business making, so the verifier's rewrite is not committed here. * no-mistakes(review): keep task record when teardown cannot close pane * no-mistakes(review): clear orphan-pane record once teardown closes the pane * no-mistakes(review): key teardown completion line on close, not record write * no-mistakes(document): sync remaining docs with herdr cutover and teardown records * no-mistakes(document): document herdr pane run as the agent relaunch path * no-mistakes: apply CI fixes
…blocking notes (#7) * feat(wardroom): give the intake council a severity bar so it can proceed (gates t1) The council shipped able to block and unable to pass: 0 proceeds, 59 revises, 37 escalations across its whole history. Every brief that ever entered intake left by captain escalation or a bypass. Raising the revise cap to 3 and then 4 bought more rounds of findings and still no proceed - the cap was never the constraint, the missing exit condition was. Three causes compounded. The reviewers were told to "roast" the brief, and a reviewer asked to roast will always produce findings. The thinkers got proceed|revise|escalate with no definition of when a finding is severe enough to block, so any imperfection mapped to revise - while the lens prompt alone was offered "or say no blocking findings". And synthesis requires unanimity, so one tidiness clause from either lens vetoed the spawn. The corpus shows the mechanism: revise lines are compound, one genuine blocker concatenated with four to eight nitpicks. The two classes separate by grammar, not topic - a blocker states premise then consequence and cites the artifact that proves it; a note is a bare imperative with neither. The fix is a severity bar plus a verdict word to carry what does not meet it. Thinkers gain PANEL: proceed-with-notes. Blocking is now defined: the crewmate would FAIL, DO HARM, or BUILD THE WRONG THING. Everything else is a note that rides along with the proceed - onto the proceed: line, attributed to its lens, and into a non-blocking section of the review file. The findings are not lost just because they did not block. Unanimity is deliberately unchanged: one revise from either lens still stops the spawn. The defect was what counted as a blocker, not who had to agree. Fail-closed is unchanged and slightly tightened - verdict parsing moves into fm_intake_verdict, which recognises exactly four words and maps anything else to escalate, where the old prefix match would have taken "PANEL: proceeds" for a proceed. A missing PANEL line, a dead thinker and an unreadable review all still escalate. FM_INTAKE_OVERRIDE semantics and the cap values are untouched. Also adds the regression that was missing. Nothing watched the one number that would have said the council was broken, so nothing did, for 96 verdicts - a gate that can never pass is indistinguishable from a gate that is broken. fm_intake_health reads the whole corpus and reports a structurally zero proceed rate as a fault in the bar, not evidence that every brief was bad. fm-intake calls it on every decision it records, including the at-cap escalate that was 30 of the 37 escalations. Below a sample floor it stays quiet. Gates, red-first and mutation-verified, with the models stubbed out through FM_INTAKE_CMD so they test synthesis and not a model's mood: - gate-t1-severity-proceeds: two non-blocking findings proceed with the notes carried; one blocker from either lens still revises; any escalate still escalates; a missing or malformed PANEL line still escalates. - gate-t1-proceed-rate-nonzero: the detector, with its sample floor. Spec amended rather than duplicated: docs/specs/2026-07-03-wardroom-intake.md. Review follow-ups folded in: - fm_intake_verdict now requires the verdict word to be the WHOLE word - only an empty tail or the " - reason" tail may follow it - so "PANEL: proceed/revise", "PANEL: proceed?" and "PANEL: proceed_with_notes" escalate instead of being truncated into a clean proceed. A trailing space or CR is transport noise and is still read. - fm_intake_health judges a ROLLING WINDOW (FM_INTAKE_HEALTH_WINDOW, default 20) instead of the all-time corpus. Nothing prunes .intake files, so read cumulatively the first proceed would have disarmed the detector forever and the regression it exists to catch could only ever have been caught once. - AGENTS.md no longer claims the crewmate gets the non-blocking notes: nothing delivers them, so it says plainly that firstmate folds them into the brief before spawning. Wiring delivery is queued work. - Re-froze the seven unrelated gates the verify sweep demoted to green. * no-mistakes(review): correct notes-delivery claim in prompt; pick most blocking PANEL verdict * no-mistakes(review): filter template echoes from PANEL verdicts; bound health min * no-mistakes(document): sync intake docs with severity bar and health detector
* fix(quarterdeck): the verifier honours gates/accepted-red.md
bin/fm-verify.sh did not know gates/accepted-red.md existed. Its verifier
prompt demanded an all-green ledger, which is unsatisfiable while a declared
baseline exists - and this repo declares two. Acceptance therefore turned on
whether the LLM verifier happened to reason about the baseline on a given run,
so correct work was rejected non-deterministically, after the build, the
pipeline, and CI. CI honoured the baseline; the verifier contradicted it.
The rule now has one implementation: fm_gates_classify in bin/fm-gates-lib.sh,
lifted out of tests/run-all.sh rather than rewritten. It is pure - a read of
gates/ledger.json and gates/accepted-red.md, root taken as an argument - and
never invokes gates/verify.sh or the ledger CLI, which would re-run every gate
and rewrite the ledger it is classifying.
Classification is separated from policy. run-all.sh skips a test and is
otherwise unchanged; fm-verify.sh rejects or escalates, structurally, ahead of
the lens and the verifier, so an unacceptable ledger costs neither model. Reject
covers undeclared reds and a test_ref naming a file that is gone; escalate
covers a missing or unreadable ledger and an unrecognised status. A repo with no
gates/ dir proceeds and never escalates - most projects have no ledger at all.
The verifier prompt's gate clause is deleted and replaced by a fence telling the
model the decision is not its to make: two authorities over one decision is what
produced the contradiction. fm-brief.sh's GATE_CHECK and the council spec now
cite the classifier instead of restating it.
Deviation from the brief, stated openly: acceptable is green, frozen, or
red-and-declared. The brief withdrew frozen on the premise that it "is not a
status", but the committed ledger holds 7 frozen gates, and CONTRIBUTING.md
records that ledger verify DEMOTES frozen to green - frozen is strictly stronger
than green, and ledger verify's own empty-drain definition of done excludes it.
Excluding frozen would have rejected every ship task in this repo. The premise
looks like a snapshot taken right after a verify sweep, before the re-freeze.
Also fixes an apostrophe in the verifier prompt heredoc that killed fm-verify at
run time: it is an unquoted heredoc inside a command substitution, bash re-scans
the body, and bash -n does not catch it. Now guarded by gate-q9.
New gates gate-q8-gate-classifier and gate-q9-verify-honours-declared-red, each
observed red before green and then frozen; the 7 gates ledger verify demoted
along the way are re-frozen at their prior status.
* fix(quarterdeck): a non-array gates value is BADLEDGER, never coerced
The classifier inherited an `if isinstance(gates, dict): gates =
list(gates.values())` coercion when the parsing logic was lifted out of
tests/run-all.sh, and that coercion was a fail-open in the one place the file
promises to fail closed. {"gates": {}} became an empty list, produced zero
rows, and classified OK - so bin/fm-verify.sh printed "gates: acceptable" and
proceeded over a ledger it had read no gates from at all. A populated object is
the same defect in a more convincing disguise.
CONTRIBUTING.md states that `gates` must be a JSON array and that any other
shape makes every `ledger` subcommand abort before doing any work, and frozen
gate m0-ledger-shape freezes exactly that. A shape the harness itself calls
fatal is not one this classifier may quietly repair. The coercion is removed:
any non-array `gates` is now BADLEDGER, which fm-verify escalates and
run-all.sh treats as "skip nothing".
An EMPTY array stays valid and acceptable - a ledger with no gates is
well-formed, merely empty - and that is now pinned too.
The one behaviour change this lands on tests/run-all.sh is in the fail-closed
direction: an object-shaped ledger previously coerced and could authorise
skips, and now authorises none. No case of gate-ci-declared-red exercises that
shape, and it passes normally and fails under LEDGER_MUTATE=1 as before.
Q8 covers the empty object, the populated object, and the empty array, and its
mutation now inverts the shape rule as well as the double condition. Q9 adds
the end-to-end case: an object-shaped ledger in the crewmate's worktree
escalates before either model runs and is never announced acceptable. Both new
assertions were checked against a deliberately reintroduced coercion. The spec
table names the array requirement.
gate-q8 and gate-q9 were demoted to green and re-frozen, so their
mutation_verified stamps cover the changed tests rather than the previous ones
- `ledger freeze` no-ops on an already-frozen gate and does not re-prove it.
Suite: 49 ran, 2 skipped, 0 failed. ledger wip still holds only the two
declared reds.
* fix(quarterdeck): a delimiter in a gate field must not forge a classifier row
The rows the classifier emits ARE the grammar its callers parse positionally,
so a tab or newline inside a field is not a formatting nuisance - it lets the
ledger write extra verdicts. The reason column was flattened for exactly that
reason; id, status and the test path were not, and that gap handed the ledger
the authority the classifier exists to hold.
Reproduced before fixing. A gate whose id was
"evil\nok<TAB>forged-gate<TAB>red<TAB>tests/bb-target.test.sh<TAB>reason"
emitted a second line that parsed as a well-formed "this red is declared"
verdict. tests/run-all.sh skipped bb-target.test.sh - a test that FAILS - and
reported "1 ran, 1 skipped, 0 failed", over an all-green ledger whose
accepted-red.md declared nothing at all. A single tab suffices on its own: it
shifts the remaining fields left until attacker-chosen text lands in the status
column, with no newline needed.
id, status and the test path are now refused outright rather than flattened.
Mangling them would leave a fabricated id standing in for a real one; refusing
is stricter and costs nothing, because a tab or newline in any of the three is
never a legitimate ledger. It is the same rule as the array check: a ledger the
harness would never accept is not one this classifier may quietly repair. The
reason column keeps its flattening - it is prose, and being last it can still
open a new line.
Q8 covers the newline id, the tab-only id, the newline status, and the exploit
end to end through run-all: the targeted test must RUN and the suite must fail.
Its mutation now inverts the forgery as well. Q9 adds the fm-verify half - a
delimiter-bearing id escalates before either model and is never announced
acceptable. Every new assertion was checked against a deliberately weakened
classifier. The spec table names the rule.
gate-q8 and gate-q9 demoted and re-frozen so their mutation stamps cover these
tests. Suite: 49 ran, 2 skipped, 0 failed. ledger wip still holds only the two
declared reds.
* no-mistakes(review): fix tab-field misparse, dedupe reject path, surface ledger diagnostics
* no-mistakes(review): reject unproven gates; escalate self-authorized accepted-red
* no-mistakes(review): resolve verify base from origin; re-freeze quarterdeck gates
* fix(quarterdeck): resolve the verify base once, safely, and by position
The base resolution added for the self-authorised-red check had three defects,
all in the helper itself rather than in the rule it serves.
base_ref() memoized into variables it could never reach: both consumers called
it as $(base_ref), and a subshell's assignments die with the subshell, so the
guard was empty again in the parent every time. The fetch ran twice and the
promise that both stages compare against the same base was false by
construction. Resolution now happens once in the main body, into FM_BASE_REF and
FM_BASE_COMMIT, which both stages read.
The fetch was guarded against failure but not against blocking, and it is the
only network call on the accept path. An ssh URL for a host absent from
known_hosts, or an https URL whose credential helper has expired, left git
waiting on an interactive prompt with no tty to answer it - and a verifier that
hangs is worse than one that escalates. GIT_TERMINAL_PROMPT=0, ssh batch mode
with a connect timeout, and an http slow-transfer cap remove the hang; an
optional timeout(1) is defence in depth where one exists.
Preferring origin/<default> unconditionally mirrored the stale-base defect it
was meant to fix. A declaration committed to the local default and not yet
pushed sits ahead of origin, so a merge base taken there predates it and the
inherited line reads as forged. The base is now the furthest-forward candidate.
That is safe in one direction only, which is why it is the right rule and not
mere permissiveness: a merge base is always an ancestor of HEAD, and a
declaration written in this branch's own commit exists on no default-branch
candidate, so it can never appear at any candidate's merge base. Moving the base
forward removes false escalations; it cannot manufacture a pass.
Q9 gains case P (a local default ahead of origin proceeds; fails exit 3 under
unconditional origin preference) and case Q (an unreachable origin degrades and
the run completes; fails exit 128 if the fetch aborts). gate-q9 re-frozen, since
freeze is re-earned rather than sticky; gate-q8's scope is untouched, so its
freeze stands. Suite 62 ran, 2 skipped, 0 failed; shellcheck clean; census
unchanged at 41 gates.
* no-mistakes(review): scope ledger debt, harden verify base, refuse bare gates dir
* no-mistakes(review): scope ledger debt by field, split verify bases, refuse bad test_ref
* no-mistakes(review): read declared-red set as awk input; normalize ledger test paths
* no-mistakes(review): announce unknown gate scope; re-earn q9 freeze stamp
* no-mistakes(review): escalate on unrecognised gate classifier header
* no-mistakes(document): sync docs with Quarterdeck gate adjudication
…un (#10) Three briefs written on 2026-08-28/29 specified work no crewmate could do - moving a gitignored file, fixing briefs under gitignored data/, and "testing" a pool-leasing command under FM_HOME=$(mktemp -d), which leaks a durable lease into the captain's live pool. The intake council caught all three, but only after a full cycle of three model lenses. A fourth reached a live crewmate: a brief carrying the retired >> status redirect, whose reports the permission profile silently refused for ten minutes while the pane looked idle. All four are structural - decidable from the brief text, the project's own ignore rules, and a list of known-hazardous commands, with no model in the loop. bin/fm-preflight-lib.sh decides them, and bin/fm-spawn.sh consults it after the brief-exists check and ahead of the council: nothing is created before the refusal, and scouts are covered where the council exempts them. Secondmates are exempt - their home is a firstmate home, where data/ and state/ are theirs to operate. The refusal names the offending path or command and says why; one that only said "invalid brief" would cost another cycle to diagnose, which is half of what this saves. Fail closed, but not noisily wrong. A lookalike is not a match: firstmate-notes is not under firstmate, fm-spawn.sh.bak is not fm-spawn.sh, and a MENTION of a pool-leasing script is not an invocation of it - a brief that asks a crewmate to change one of those scripts is ordinary work. Every invocation form must be written as code, because English can put "bash" or a ./ path directly before a script name while saying the opposite of run it; the FM_HOME=$(mktemp -d) half needs no code context and is the truer signal for the third defect anyway. git is the authority for gitignored, so a tracked path a pattern happens to match still reads as visible. The known false-refusal class - a brief that must cite an invisible path - is documented, and FM_PREFLIGHT_OVERRIDE=1 is the loud, logged way through. Gates gate-t1-brief-preflight-rules and gate-t1-brief-preflight-spawn-gate, with a committed fixture per rule plus a clean brief and a lookalike brief that must pass - the two that stop this becoming a gate nothing gets past. Spec: docs/specs/2026-08-31-brief-preflight.md.
…ne crewmate stale (#11) * feat(supervision): follow crewmates onto herdr, and stop calling a done crewmate stale fm/muxwire-h2 moved the crew runtime onto herdr and left fm-watch.sh and the context watchdog sensing through tmux, so for every post-cutover crewmate the fleet lost both supervision guarantees: nothing noticed a wedged agent, and nothing auto-compacted a bloating one. Swap the sensor, keep the loop. The watcher's loop, singleton lock, coalescing, backoff and durable wake queue are untouched; the watchdog's thresholds, fresh-handoff guard, cooldown and per-secondmate scoping are untouched. Routing is by what state/<id>.meta records, so the mixed fleet keeps working and every tmux path is unchanged. - bin/fm-sense-lib.sh: the sensing, split out of the watcher so it is testable. One `herdr api snapshot` per cycle reads the whole fleet; `agent_status: unknown` is the only value that means "stopped without reporting". The snapshot is parsed, never grepped - terminal_title is crewmate-authored text that can otherwise forge a wake. - The context watchdog's three touchpoints route by surface, MEASURE stamps a herdr crewmate managed through FM_HERDR_PANE (pinned by fm-spawn, since herdr sets no env in a pane), and fire-time ownership is re-confirmed against this home's own metas rather than a session name. - fm-supervise-daemon.sh moves in the same change: a herdr pane id encodes no task id, so window_to_task/window_for_task map through the meta and away-mode escalates a wedged herdr crewmate instead of reading "no window" as "torn down". - Awaiting a verdict becomes a supervision state. A `done:` claim raises no stale wake while a verify cycle runs, after an approve, or at the attempt cap; a reject hands the ball back and still wakes. Four false "stalled" reports across 2026-08-28..31 came from not having it. Also fixes gate-h5-herdr-live-roundtrip, red on the merge base: for a pane with no agent, `pane read --source recent` succeeds with empty output, so a populated pane read as silent. Gates w1-w7, spec docs/specs/2026-08-31-supervision-on-herdr.md. Ledger green. * fix(supervision): a herdr wedge wakes every time, and the session pin reaches the snapshot Two defects the Quarterdeck reproduced in the new sensor, each one defeating the guarantee the branch exists to restore. 1. Stale detection fired once per PANE, not once per wedge. `.stale-<key>` means "this stalled state was already reported"; under tmux the value was a content hash, so a fresh wedge always carried a fresh value and the marker aged out by accident. A herdr observation is categorical - the literal `unknown` every time - so the second wedge on a pane matched its own first wedge's marker and was suppressed for the life of the pane. That lands squarely on the documented recovery path: stuck-crewmate-recovery relaunches the agent IN THE SAME PANE. An episode now ends, and clears the marker, on either of the two things that end one: the observation changing, and the pane ceasing to hold an agent. The second matters because a relaunch passes through no-agent, so a crewmate can go wedged -> relaunched -> wedged with no healthy sample in between. Raising nothing for an unlisted pane is unchanged; remembering nothing is the part that was wrong. 2. `fm_sense_herdr_statuses` called `herdr api snapshot` bare, so `fm_herdr_session` never exported FM_HERDR_SESSION and a pinned fleet was polled through `default`. That comes back EMPTY rather than failing, so every recorded pane hit the silent not-listed path: total, unannounced blindness for any non-default fleet. fm-sense-lib now sources fm-herdr.sh and re-resolves the pin before each read. Gate w1 grows three cases, each proven to catch its own defect independently: a recovered-then-rewedged pane wakes again; a relaunch through no-agent does not poison the marker; and the pin reaches the snapshot verb. The test fake's `api snapshot` is now session-scoped, as the real verb is, so a wrong-session probe returns no agents instead of answering every session alike. Also restores the AGENTS.md gap-section rewrite, which an aborted edit dropped from the previous commit. Suite 69 ran / 0 failed. Ledger green; wip is only the two declared accepted-red. * test(supervision): make the watcher gates load-proof, and their silences non-vacuous The w1/w2 defects reported in the reject were already fixed in 02e59377; both re-proved end to end through bin/fm-watch.sh against this HEAD, and against abbc4ce3 for contrast. What was NOT sound was the gates themselves. w1 and w2 flaked RED under a full `ledger verify` sweep while passing standalone 24 runs running. They drive the real watcher and were bounded at 6s wall-clock, so 47 suites back to back starved a wake that was on its way. A gate that goes red because the box was busy is a false alarm that teaches everyone to ignore it - the same "check nobody trusts" failure the reject cited. Raising the budget alone would only trade false reds for false greens, because every case asserting that NOTHING happens passes trivially if the watcher never reached its decision. So the budget goes to 30s AND every negative case now demands evidence the decision point was reached: fm_watch_assert_sensed the per-target counter advanced - the wake/no-wake branch actually ran fm_watch_assert_ran a full cycle completed; the available evidence for a kind=secondmate, exempted before any marker is touched fm_watch_assert_reset the pane's episode was ended because it stopped holding an agent Each assertion is itself proven to fire on an empty state dir, and each of w1's two original findings is re-proved to fail the gate when its fix is reverted. ledger verify green twice running; suite 69 ran / 0 failed; w1/w2/w3/w6 re-mutation-verified. * fix(gates): re-freeze the nine gates `ledger verify` demoted `ledger verify` is not read-only and freeze is not sticky: a passing frozen gate is rewritten as green on every sweep, so the many verify runs this branch made silently stripped the mutation-proof guard from nine gates that are frozen on origin/main and entirely outside this work - g-boot-digest, m0-ledger-shape, m2-boot-emitter-is-read-only, m4-boot-budget-hostile, m5-digest-never-silent, gate-ci-declared-red, gate-status-verb, gate-q8-gate-classifier and gate-q9-verify-honours-declared-red. Each is re-frozen through `ledger freeze <id>`, which re-earns the freeze by running its test normally and under LEDGER_MUTATE and demanding the first pass and the second fail. The frozen set now matches origin/main exactly - nine, none missing, none added - and no gate's status has regressed against it. CONTRIBUTING.md line 62 already states this rule. Not reading it is the root cause; `frozen:0` was in my very first verify output and I did not act on it. * fix(ctxwatch): read meta paths as lines, not words (CI shellcheck) CI runs bare `shellcheck bin/*.sh tests/*.sh`, which reports every severity. I had been checking with `-S warning`, which hides `info`, so SC2013 on the `for m in $(grep -l ...)` loop in _ctx_target_in_session passed locally and failed the Lint job. The fix is the one shellcheck points at rather than a suppression, and it is a real correctness gain: FM_HOME is a user-chosen directory, so a meta path can contain a space, and word-splitting a filename into fragments would answer "not ours" for a pane that is - a fire gate silently refusing to compact a crewmate it owns. `shellcheck bin/*.sh tests/*.sh` (CI's exact command, no severity filter) now exits 0. Suite 71 ran / 0 failed. w4 and w5 re-run and re-mutation-checked, and g2/g4/g6 - the tmux-side context-watchdog gates - re-run.
…chine tests/fm-wardroom-t1-preflight-rules.test.sh was green on one laptop and red in CI, and the ledger read green the whole time. The cause was a fixture, not the preflight. primary-checkout.md hardcoded /Users/<captain>/firstmate as one of its three offenders. That path is under a checkout root only where $HOME happens to be /Users/<captain>, so on every other machine the offender was under no root at all, was correctly not reported, and the assertion naming it failed. A gate that asserts an absolute path asserts a property of the machine it runs on, which our own standard forbids; this is what that costs. The refusal does NOT collapse distinct offenders into a glob. The glob it prints is the text the brief itself wrote - "fix the 68 stale briefs under data/*/brief.md" - quoted back with its line number, which is the offender. All three offenders were always emitted as three separate TSV findings; in CI only two of them existed. The gate now counts the records, so "three findings" can no longer be confused with "one finding containing three substrings", and a dropped offender cannot hide behind a substring assertion again. - Fixtures write __PRIMARY__ wherever an absolute checkout root is meant; the suite substitutes a root it chose itself. - $HOME is synthetic for every preflight run, so the ~/firstmate form the other fixtures use no longer reaches for the machine's real layout. - Both primary-checkout cases - the refusal and the lookalike - run under two roots differing in depth and name, asserting against the root the run used. - The lookalike's sharpest near-miss is now <root>-old rendered from the root under test. As a literal it was a real boundary test on one machine and vacuously true everywhere else; it now catches a boundary regression anywhere. Proved red first: with render() pinned back to one machine's root the suite fails naming the dropped offender, and with under()'s boundary check reduced to a bare prefix test the lookalike case fails under both roots. Green under three different $HOMEs. CONTRIBUTING.md records the rule alongside the other test-hygiene rules. bash tests/run-all.sh: 71 ran, 2 skipped, 0 failed. gates/verify.sh: green:48 frozen:0 red:2 unproven:0 wip:2 - the two declared reds, gate-l2-loop-audit-level and m1-hook-registered. The nine gates the sweep demoted are re-frozen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PzTiy6wHvj4nUJFnwXiuu8
|
Too many files changed for review (152 files, 100 file limit). Bypass the limit by tagging |
Author
|
Opened against the wrong repository by mistake — this change belongs on the fork it was branched from, not here. Closing; apologies for the noise. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
mainis red in CI ontests/fm-wardroom-t1-preflight-rules.test.sh. It came in with #10 and was merged without a CI check, so this is repairing a known-bad merge.What was actually wrong
A fixture, not the preflight.
tests/fixtures/preflight/primary-checkout.mdhardcoded/Users/<captain>/firstmate/data/old-task-q4/report.mdas one of its three offenders. That path is under a checkout root only where$HOMEis/Users/<captain>. Everywhere else it is under no root at all, so the preflight correctly did not report it, and the assertion naming it failed. The gate was green on exactly one laptop on earth.That is our own standard broken: a test must never assert a property of the machine it runs on. CI's
/home/runner/firstmatewas never the problem; the fixture pinning one machine was.On the reported "collapsed glob": there is no collapse. The glob in the refusal is the text the brief itself wrote — "fix the 68 stale briefs under
data/*/brief.md" — quoted back with its line number, which is exactly what naming the offender means. All three offenders were always emitted as three separate TSV findings; in CI only two of them existed. The two-line output in the CI log was a dropped offender, not a merged one.Verified directly:
The gate now counts the records, so "three findings" can no longer be confused with "one finding that happens to contain three substrings" — which is the assertion that would have caught the original defect, and the one a substring check structurally cannot make.
The fix
__PRIMARY__wherever an absolute checkout root is meant; the suite substitutes a root it chose itself.$HOMEis synthetic for every preflight run, so the~/firstmateform the other fixtures use no longer reaches for the machine's real layout either.<root>-oldrendered from the root under test. As a literal it was a real boundary test on one machine and vacuously true everywhere else; it now catches a boundary regression anywhere.bin/fm-preflight-lib.shis unchanged — the preflight was doing the right thing.Gated
Red first, twice. With
render()pinned back to one machine's root the suite fails naming the dropped offender. Withunder()'s boundary check reduced to a bare prefix test, the lookalike case fails under both roots. Green under three different$HOMEs.bash tests/run-all.sh— 71 ran, 2 skipped, 0 failedbash gates/verify.sh—green:48 frozen:0 red:2 unproven:0 wip:2; the two reds are the declared ones,gate-l2-loop-audit-levelandm1-hook-registered, and nothing else. The nine gates the sweep demoted are re-frozen, so the ledger diff againstorigin/mainis timestamps plus one observable edit.shellcheck bin/*.sh tests/*.sh— clean.origin/main(93162ce).Expected to stay red, and not from this branch
CI's
PR must be raised via no-mistakescheck fails on every branch now, because firstmate moved this project todirect-PRmode and raises PRs itself. That is a policy gap between the delivery mode and the workflow, not something this change touches or could fix.So: the suite and the gate ledger pass; the no-mistakes provenance check remains red for every PR until that workflow is reconciled with
direct-PRmode.Note
CONTRIBUTING.mdgains the durable rule next to the other test-hygiene rules. It went there rather thanAGENTS.mdbecause that is where gate- and test-authoring rules already live.🤖 Generated with Claude Code
https://claude.ai/code/session_01PzTiy6wHvj4nUJFnwXiuu8