Skip to content

fix(bin): treat Claude Code's default external-imports flags as never asked, not declined - #4387

Merged
kunchenguid merged 2 commits into
kunchenguid:mainfrom
Rangezi:fix/claude-trust-default-import-flags
Sep 14, 2026
Merged

kunchenguid merged 2 commits into
kunchenguid:mainfrom
Rangezi:fix/claude-trust-default-import-flags

Conversation

@Rangezi

@Rangezi Rangezi commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Intent

Upstream fix for #4378, which the maintainer's triage bot confirmed and labeled ready-for-pr (patch in the issue is evidence, not a merge). bin/fm-claude-trust.sh refuses the whole Claude trust registration when the project-root entry in ~/.claude.json has hasClaudeMdExternalIncludesApproved === false, assuming Claude Code writes false only on an explicit 'No, disable'. In Claude Code 2.1.270 the default project entry carries Approved and WarningShown both false before the dialog is ever shown, and the dialog writes WarningShown: true for either answer; on the reporter's machine this refused spawns in 7 of 10 registered projects. Deliberate decision: only Approved === false together with WarningShown === true is a decline; false/false (never asked) must behave exactly like an absent flag — trust registered, no import consent manufactured, so a CLAUDE.md that really imports external files still stops on the dialog. Scope is intentionally narrow: correct the decline predicate and its comment block only, no new trust surface, no change to worktree-mode import-flag handling. The new test test_project_root_entry_default_import_flags_are_not_a_decline seeds Claude Code's default entry shape and expects exit 0 with trust only; it fails on b182d0f with the refusal and passes with the fix, and the existing decline fixture (false/true shape) must stay green. tests/fm-claude-trust.test.sh passes 31/31 locally and bin/fm-lint.sh is clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. The PR should close #4378.

What Changed

  • bin/fm-claude-trust.sh: declinedExternalImports now counts a decline only when hasClaudeMdExternalIncludesApproved === false and hasClaudeMdExternalIncludesWarningShown === true. Claude Code's default project entry has both flags set to false, and the script used to read that as a decline and refuse registration. That false/false shape now behaves like an absent flag: trust gets registered, no import consent is written, and the import dialog still shows. The header comment and the inline comments now describe the new rule.
  • tests/fm-claude-trust.test.sh: new test test_project_root_entry_default_import_flags_are_not_a_decline. It seeds Claude Code's default project entry shape and expects exit 0, with trust but no import consent on both the worktree entry and the project-root entry. The comment on the existing decline test now names the false/true pair.
  • .agents/skills/harness-adapters/references/harness/claude.md: the harness doc now uses the same two-flag decline rule and says that both flags false means the project was never asked.

Closes #4378

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The change only narrows a single refusal predicate, from Approved===false to Approved===false && WarningShown===true, and it can never manufacture import consent. I checked the new predicate against the installed Claude Code 2.1.270 binary. There the default project entry DX sets both flags to !1 (false), and the dialog handler zmt writes {Approved:f, WarningShown:!0} for either answer. The show-dialog check edr skips the dialog only when Approved or WarningShown is truthy, so false/false and false/absent both mean never asked. The new test seeds the exact DX shape and would fail against the old predicate. The existing false/true decline fixture is unchanged, and the pipeline-authored doc fix at claude.md:28-29 states the corrected predicate accurately without going beyond the requested sentences.

Testing

First ran the targeted trust test file: 31/31 at HEAD. The same file run against base b182d0f fails on the new test with the reported refusal, while the existing decline fixture stays green. Then I ran the trust script directly on real git worktrees with isolated Claude config stores, at HEAD and at base, across seven entry shapes. HEAD refuses only false/true. Base also refuses false/false and false/absent. Trust-only and consent-carry-forward results match base. Real Claude Code 2.1.270 in tmux showed three things: its default project entry is Approved=false, WarningShown=false even though the import dialog was never shown; "No, disable" writes false/true and "Yes, allow" writes true/true. After HEAD registers a false/false project, claude launched into the worktree skips the trust dialog and still stops on the external-imports dialog, so no consent is created. A real false/true decline written by Claude is still refused and the store is left unchanged. Finally, fm-spawn itself (with stubbed tmux and claude) launches the worker for false/false at HEAD and refuses at base, and refuses the decline at both. The unit-test comparison and the stubbed fm-spawn run were not live product runs, so those two scenarios are recorded as untested. All text transcripts are in the evidence directory; there is no UI surface, so no screenshots. The scratch directories, the tmux socket and the base archive were removed, and the worktree is clean.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Claude Code 2.1.270, opened once in a repo with no CLAUDE.md, writes a project entry with both external-imports flags false (the premise of the fix) ✅ pass live claude-2.1.270-default-project-entry.txt: the entry has hasTrustDialogAccepted=true, Approved=false, WarningShown=false; the import dialog was never shown
Spawn registration on a false/false entry that Claude itself wrote: HEAD exits 0 and records trust on the worktree and project root without adding import consent; base b182d0f refuses ✅ pass live live-claude-default-entry-then-decline.txt §1; cli-shape-matrix-head-vs-base.txt (default-false-false: HEAD exit=0, worktree import flags unset, project flags stay false; BASE exit=1)
After HEAD registers a false/false project whose CLAUDE.md imports an outside file, real claude launched in the worktree skips the trust dialog and still stops on the 'Allow external CLAUDE.md file im… ✅ pass live live-claude-fixture-shape-trust-skip-import-dialog.txt (project root trust was false before registration; first screen is the import dialog; choosing Yes goes straight to the REPL); live-claude-defaul…
Adversarial: a genuine 'No, disable' answered in real claude writes Approved=false, WarningShown=true, and HEAD still refuses the next worktree spawn without touching the store ✅ pass live live-claude-default-entry-then-decline.txt §3-4 (entry false/true; HEAD exit=1 'already declined external CLAUDE.md imports'; store sha256 unchanged)
Adversarial shapes that are not a decline (Approved=false with WarningShown absent, WarningShown=true with Approved absent, a non-boolean WarningShown) register trust and never set Approved=true ✅ pass live cli-shape-matrix-head-vs-base.txt (approved-false-warning-absent, warningshown-true-approved-absent, decline-string-true-not-boolean: HEAD exit=0, no import flag set to true on the worktree, project A…
Regression: standing 'Yes, allow' consent (true/true) is carried to a new worktree, and real claude launched there reaches the REPL with no dialog; an entry with no flags stays trust-only, same as bas… ✅ pass live live-claude-fixture-shape-trust-skip-import-dialog.txt (wt2 trust/Approved/WarningShown all true, REPL first screen); cli-shape-matrix-head-vs-base.txt (consent-true-true, flags-absent identical at HE…
The new test fails on base b182d0f with the refusal and passes at HEAD, and the existing decline fixture stays green on both ⏸️ untested no The earlier payload only ran the unit test file (tests/fm-claude-trust.test.sh) at HEAD and against a base archive. A unit test run is not a live product run, so it did not show a live result. The liv…
fm-spawn of a claude worker into a project with the default false/false entry launches the worker at HEAD (base refuses), and a false/true decline is refused at both ⏸️ untested no The earlier payload ran bin/fm-spawn.sh only with the test-fixture fake tmux backend and fake claude binary, so it did not show a live result. A live check needs fm-spawn run on a real tmux session wi…
Evidence: Targeted trust test file at HEAD (31 ok)

Source: Targeted trust test file at HEAD (31 ok)

ok - fm-claude-trust.sh: a fresh task worktree is trusted
ok - fm-claude-trust.sh: a fresh registration trusts the project root without manufacturing import consent
ok - fm-claude-trust.sh: carries forward a project's already-granted import consent to the worktree entry
ok - fm-claude-trust.sh: preserves unrelated keys on the project-root entry
ok - fm-claude-trust.sh: refuses to override a project's declined external-imports consent
ok - fm-claude-trust.sh: a never-asked default external-imports pair is not treated as a decline
ok - fm-claude-trust.sh: repeat registration is idempotent
ok - fm-claude-trust.sh: refuses the primary checkout
ok - fm-claude-trust.sh: an exported CDPATH cannot defeat the scope refusal
ok - fm-claude-trust.sh: inherited git environment overrides cannot defeat the scope refusal
ok - fm-claude-trust.sh: refuses a home directory the git checks would accept
ok - fm-claude-trust.sh: refuses the Claude config directory
ok - fm-claude-trust.sh: refuses a relative CLAUDE_CONFIG_DIR
ok - fm-claude-trust.sh: refuses a directory that is not a git worktree
ok - fm-claude-trust.sh: refuses a path that does not exist
ok - fm-claude-trust.sh: refuses a worktree belonging to another project
ok - fm-claude-trust.sh: refuses a subdirectory of the worktree
ok - fm-claude-trust.sh: a project argument that is itself a linked worktree resolves to the primary checkout
ok - fm-claude-trust.sh: preserves unrelated store content
ok - fm-claude-trust.sh: refuses a store symlinked to another user's file
ok - fm-claude-trust.sh: follows a store symlink to this user's own file and leaves the link intact
ok - fm-claude-trust.sh: refuses an unparseable store and leaves it untouched
ok - fm-claude-trust.sh: a missing node is refused rather than degraded
ok - fm-claude-trust.sh: a scope refusal stays fail-closed without node
ok - fm-spawn.sh: a claude spawn pre-trusts its worktree and launches with the brief
ok - fm-spawn.sh: a trust-refused claude spawn leaves no task state behind
ok - fm-spawn.sh: a claude secondmate spawn pre-trusts a standalone-clone home
ok - fm-spawn.sh: a claude secondmate spawn pre-trusts a leased worktree home
ok - fm-claude-trust.sh: home-level trust is refused for everything but a home seeded for this secondmate
ok - fm-claude-trust.sh: worktree mode still refuses a secondmate home
ok - fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded
Evidence: HEAD's test file run against base b182d0f: new test fails with the refusal, decline fixture passes

Source: HEAD's test file run against base b182d0f: new test fails with the refusal, decline fixture passes

# HEAD 6ae7979 tests/fm-claude-trust.test.sh run inside a git-archive of base b182d0f (base bin/fm-claude-trust.sh)
# Expectation: existing decline fixture passes; new default-flags test fails with the refusal (suite aborts on first failure)
$ (cd <base-tree> && bash tests/fm-claude-trust.test.sh)
ok - fm-claude-trust.sh: a fresh task worktree is trusted
ok - fm-claude-trust.sh: a fresh registration trusts the project root without manufacturing import consent
ok - fm-claude-trust.sh: carries forward a project's already-granted import consent to the worktree entry
ok - fm-claude-trust.sh: preserves unrelated keys on the project-root entry
ok - fm-claude-trust.sh: refuses to override a project's declined external-imports consent
not ok - a never-asked default entry must not be refused as a decline: error: project entry for <tmp>/project-default-flags/project in <tmp>/project-default-flags/claude-config/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/project-default-flags/wt' and project '<tmp>/project-default-flags/project' in '<tmp>/project-default-flags/claude-config/.claude.json': expected exit 0, got 1
exit=1

# Same file at HEAD: see fm-claude-trust.test.log (31 ok, exit 0)
Evidence: Trust script CLI across 7 project-entry shapes, HEAD vs base

Source: Trust script CLI across 7 project-entry shapes, HEAD vs base

=================== HEAD (6ae7979) ===================
## head-default-false-false  script=fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":false}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-default-false-false/wt
trusted (project root): <tmp>/case-head-default-false-false/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=false WarningShown=false

=================== BASE (b182d0f) ===================
## base-default-false-false  script=base-fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":false}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
error: project entry for <tmp>/case-base-default-false-false/project in <tmp>/case-base-default-false-false/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/case-base-default-false-false/wt' and project '<tmp>/case-base-default-false-false/project' in '<tmp>/case-base-default-false-false/cfg/.claude.json'
exit=1  store-changed=no
  worktree: <no entry>
  project: trust=false Approved=false WarningShown=false

=================== HEAD (6ae7979) ===================
## head-decline-false-true  script=fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
error: project entry for <tmp>/case-head-decline-false-true/project in <tmp>/case-head-decline-false-true/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/case-head-decline-false-true/wt' and project '<tmp>/case-head-decline-false-true/project' in '<tmp>/case-head-decline-false-true/cfg/.claude.json'
exit=1  store-changed=no
  worktree: <no entry>
  project: trust=undefined Approved=false WarningShown=true

=================== BASE (b182d0f) ===================
## base-decline-false-true  script=base-fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
error: project entry for <tmp>/case-base-decline-false-true/project in <tmp>/case-base-decline-false-true/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/case-base-decline-false-true/wt' and project '<tmp>/case-base-decline-false-true/project' in '<tmp>/case-base-decline-false-true/cfg/.claude.json'
exit=1  store-changed=no
  worktree: <no entry>
  project: trust=undefined Approved=false WarningShown=true

=================== HEAD (6ae7979) ===================
## head-approved-false-warning-absent  script=fm-claude-trust.sh
seeded project entry flags: {"Approved":false}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-approved-false-warning-absent/wt
trusted (project root): <tmp>/case-head-approved-false-warning-absent/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=false WarningShown=undefined

=================== BASE (b182d0f) ===================
## base-approved-false-warning-absent  script=base-fm-claude-trust.sh
seeded project entry flags: {"Approved":false}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
error: project entry for <tmp>/case-base-approved-false-warning-absent/project in <tmp>/case-base-approved-false-warning-absent/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/case-base-approved-false-warning-absent/wt' and project '<tmp>/case-base-approved-false-warning-absent/project' in '<tmp>/case-base-approved-false-warning-absent/cfg/.claude.json'
exit=1  store-changed=no
  worktree: <no entry>
  project: trust=undefined Approved=false WarningShown=undefined

=================== HEAD (6ae7979) ===================
## head-flags-absent  script=fm-claude-trust.sh
seeded project entry flags: {}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-flags-absent/wt
trusted (project root): <tmp>/case-head-flags-absent/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=undefined WarningShown=undefined

=================== BASE (b182d0f) ===================
## base-flags-absent  script=base-fm-claude-trust.sh
seeded project entry flags: {}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-base-flags-absent/wt
trusted (project root): <tmp>/case-base-flags-absent/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=undefined WarningShown=undefined

=================== HEAD (6ae7979) ===================
## head-consent-true-true  script=fm-claude-trust.sh
seeded project entry flags: {"Approved":true,"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-consent-true-true/wt
trusted (project root): <tmp>/case-head-consent-true-true/project
exit=0  store-changed=yes
  worktree: trust=true Approved=true WarningShown=true
  project: trust=true Approved=true WarningShown=true

=================== BASE (b182d0f) ===================
## base-consent-true-true  script=base-fm-claude-trust.sh
seeded project entry flags: {"Approved":true,"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-base-consent-true-true/wt
trusted (project root): <tmp>/case-base-consent-true-true/project
exit=0  store-changed=yes
  worktree: trust=true Approved=true WarningShown=true
  project: trust=true Approved=true WarningShown=true

=================== HEAD (6ae7979) ===================
## head-warningshown-true-approved-absent  script=fm-claude-trust.sh
seeded project entry flags: {"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-warningshown-true-approved-absent/wt
trusted (project root): <tmp>/case-head-warningshown-true-approved-absent/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=undefined WarningShown=true

=================== BASE (b182d0f) ===================
## base-warningshown-true-approved-absent  script=base-fm-claude-trust.sh
seeded project entry flags: {"WarningShown":true}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-base-warningshown-true-approved-absent/wt
trusted (project root): <tmp>/case-base-warningshown-true-approved-absent/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=undefined WarningShown=true

=================== HEAD (6ae7979) ===================
## head-decline-string-true-not-boolean  script=fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":"true"}
$ CLAUDE_CONFIG_DIR=<isolated> fm-claude-trust.sh <worktree> <project>
trusted: <tmp>/case-head-decline-string-true-not-boolean/wt
trusted (project root): <tmp>/case-head-decline-string-true-not-boolean/project
exit=0  store-changed=yes
  worktree: trust=true Approved=undefined WarningShown=undefined
  project: trust=true Approved=false WarningShown=true

=================== BASE (b182d0f) ===================
## base-decline-string-true-not-boolean  script=base-fm-claude-trust.sh
seeded project entry flags: {"Approved":false,"WarningShown":"true"}
$ CLAUDE_CONFIG_DIR=<isolated> base-fm-claude-trust.sh <worktree> <project>
error: project entry for <tmp>/case-base-decline-string-true-not-boolean/project in <tmp>/case-base-decline-string-true-not-boolean/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/case-base-decline-string-true-not-boolean/wt' and project '<tmp>/case-base-decline-string-true-not-boolean/project' in '<tmp>/case-base-decline-string-true-not-boolean/cfg/.claude.json'
exit=1  store-changed=no
  worktree: <no entry>
  project: trust=undefined Approved=false WarningShown=true
Evidence: Real Claude Code 2.1.270 writes a false/false default project entry

Source: Real Claude Code 2.1.270 writes a false/false default project entry

# Scenario: Claude Code 2.1.270's own default project entry (isolated CLAUDE_CONFIG_DIR, real binary in tmux)
$ claude --version
2.1.270 (Claude Code)

## Store 'projects' while the trust dialog is on screen: {} (nothing written yet)

## Screen after choosing 'Yes, I trust this folder' in a repo with NO CLAUDE.md (import dialog never shown):
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏸ manual mode on · ? for shortcuts · ← for agents

## Project entry Claude Code wrote after that session exited (last* session-metric keys omitted):
{
  "<tmp>/premise/repo": {
    "allowedTools": [],
    "mcpContextUris": [],
    "mcpServers": {},
    "enabledMcpjsonServers": [],
    "disabledMcpjsonServers": [],
    "hasTrustDialogAccepted": true,
    "hasClaudeMdExternalIncludesApproved": false,
    "hasClaudeMdExternalIncludesWarningShown": false
  }
}
Evidence: Live: base refuses and HEAD registers on Claude's own entry; real claude shows the import dialog; a real 'No, disable' is still refused

Source: Live: base refuses and HEAD registers on Claude's own entry; real claude shows the import dialog; a real 'No, disable' is still refused

# Scenario: reporter's shape on a real Claude-written entry -> HEAD registers, real claude skips trust, import dialog still renders; a real 'No, disable' is still honored

## 1. Project already opened once in Claude Code 2.1.270 (entry written by Claude itself: trust=true, Approved=false, WarningShown=false).
##    Project then gains a committed CLAUDE.md importing <tmp>/premise/external/shared.md; firstmate spawns worktree task-1.

$ base b182d0f bin/fm-claude-trust.sh <tmp>/premise/wt <tmp>/premise/repo   (on a copy of the store)
error: project entry for <tmp>/premise/repo in <tmp>/premise/cfg-basecopy/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/premise/wt' and project '<tmp>/premise/repo' in '<tmp>/premise/cfg-basecopy/.claude.json'
exit=1

$ HEAD 6ae7979 bin/fm-claude-trust.sh <tmp>/premise/wt <tmp>/premise/repo
trusted: <tmp>/premise/wt
trusted (project root): <tmp>/premise/repo
exit=0   (recorded output of the run made before launching claude)

## 2. Real claude 2.1.270 launched in tmux into <tmp>/premise/wt with that store - first screen (no trust dialog; import dialog renders):
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  Allow external CLAUDE.md file imports?
  This project's CLAUDE.md imports files outside the current working directory. Never allow this for third-party repositories.
  External imports:
    <tmp>/premise/external/shared.md
  Important: Only use Claude Code with files you trust. Accessing untrusted files may pose security risks https://code.claude.com/docs/en/security
  ❯ No, disable external imports
    Yes, allow external imports
  Enter to confirm · Esc to cancel

## 3. Pressed Enter on the default 'No, disable external imports'. Project entry Claude wrote:
{"hasTrustDialogAccepted":true,"hasClaudeMdExternalIncludesApproved":false,"hasClaudeMdExternalIncludesWarningShown":true}

## 4. Firstmate spawns worktree task-2 against that genuine decline:
$ HEAD 6ae7979 bin/fm-claude-trust.sh <tmp>/premise/wt2 <tmp>/premise/repo
error: project entry for <tmp>/premise/repo in <tmp>/premise/cfg/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/premise/wt2' and project '<tmp>/premise/repo' in '<tmp>/premise/cfg/.claude.json'
exit=1  store sha256 before=700756e114c88fa5 after=700756e114c88fa5 (unchanged)
Evidence: Live: fixture shape (trust=false, false/false): trust dialog skipped, import dialog renders, 'Yes, allow' consent carries forward to the next worktree

Source: Live: fixture shape (trust=false, false/false): trust dialog skipped, import dialog renders, 'Yes, allow' consent carries forward to the next worktree

# Scenario: never-asked default pair with trust=false (the new test's fixture shape) -> real claude 2.1.270 in the spawned worktree

## Seeded project-root entry: hasTrustDialogAccepted=false, hasClaudeMdExternalIncludesApproved=false, hasClaudeMdExternalIncludesWarningShown=false
## Project has committed CLAUDE.md importing <tmp>/fixture/external/shared.md (outside the tree)

$ HEAD bin/fm-claude-trust.sh <tmp>/fixture/wt <tmp>/fixture/repo
trusted: <tmp>/fixture/wt
trusted (project root): <tmp>/fixture/repo
exit=0

## Real claude launched into <tmp>/fixture/wt - first screen (trust dialog NOT shown, import dialog renders because no consent was manufactured):
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  Allow external CLAUDE.md file imports?
  This project's CLAUDE.md imports files outside the current working directory. Never allow this for third-party repositories.
  External imports:
    <tmp>/fixture/external/shared.md
  Important: Only use Claude Code with files you trust. Accessing untrusted files may pose security risks https://code.claude.com/docs/en/security
  ❯ No, disable external imports
    Yes, allow external imports
  Enter to confirm · Esc to cancel

## Chose 'Yes, allow external imports' -> REPL reached with no further dialog:
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏸ manual mode on · ? for shortcuts · ← for agents

## Project-root entry Claude wrote: trust=true Approved=true WarningShown=true (either answer sets WarningShown=true)

$ HEAD bin/fm-claude-trust.sh <tmp>/fixture/wt2 <tmp>/fixture/repo   (second spawn, standing consent now exists)
trusted: <tmp>/fixture/wt2
trusted (project root): <tmp>/fixture/repo
exit=0
## wt2 entry after registration: trust=true Approved=true WarningShown=true (consent carried forward)

## Real claude launched into <tmp>/fixture/wt2 - first screen (no trust dialog, no import dialog, straight to REPL):
  ▝▝ ▝▝    <tmp>/fixture/wt2
⚠ Remote managed settings failed to load (authentication rejected (401)) · no remote policy applied · /status for details
                                                                                                                                              ● high · /effort
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏸ manual mode on · ? for shortcuts · ← for agents
Evidence: fm-spawn claude with default vs decline stores, HEAD vs base (fake tmux and fake claude)

Source: fm-spawn claude with default vs decline stores, HEAD vs base (fake tmux and fake claude)

=== HEAD 6ae7979: fm-spawn.sh (real) with fake tmux backend + fake claude binary, isolated CLAUDE_CONFIG_DIR ===
## head-default (repo root: 01M2E24HBM1J4AD0J95AX3KWH2) seeded project entry Approved=false WarningShown=false
$ fm-spawn.sh nmspawnhead-default2428083 <project> claude --mode no-mistakes --yolo off
warning: <tmp>/spawn-head-default/home/data/nmspawnhead-default2428083/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
spawned nmspawnhead-default2428083 harness=claude kind=ship mode=no-mistakes yolo=off window=firstmate:fm-nmspawnhead-default2428083 worktree=<tmp>/spawn-head-default/wt
exit=0
launch command sent: claude --dangerously-skip-permissions
  <tmp>/spawn-head-default/project trust=true Approved=false WarningShown=false
  <tmp>/spawn-head-default/wt trust=true Approved=undefined WarningShown=undefined

## head-decline (repo root: 01M2E24HBM1J4AD0J95AX3KWH2) seeded project entry Approved=false WarningShown=true
$ fm-spawn.sh nmspawnhead-decline2428911 <project> claude --mode no-mistakes --yolo off
warning: <tmp>/spawn-head-decline/home/data/nmspawnhead-decline2428911/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: project entry for <tmp>/spawn-head-decline/project in <tmp>/spawn-head-decline/claude-config/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/spawn-head-decline/wt' and project '<tmp>/spawn-head-decline/project' in '<tmp>/spawn-head-decline/claude-config/.claude.json'
error: could not pre-register Claude workspace trust for <tmp>/spawn-head-decline/wt; refusing to launch a claude worker that would wedge on the trust dialog; inspect window firstmate:fm-nmspawnhead-decline2428911
exit=1
launch command sent: <none>
  <tmp>/spawn-head-decline/project trust=false Approved=false WarningShown=true

=== BASE b182d0f (git archive): same driver ===
## base-default (repo root: base-tree) seeded project entry Approved=false WarningShown=false
$ fm-spawn.sh nmspawnbase-default2429394 <project> claude --mode no-mistakes --yolo off
warning: <tmp>/spawn-base-default/home/data/nmspawnbase-default2429394/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: project entry for <tmp>/spawn-base-default/project in <tmp>/spawn-base-default/claude-config/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/spawn-base-default/wt' and project '<tmp>/spawn-base-default/project' in '<tmp>/spawn-base-default/claude-config/.claude.json'
error: could not pre-register Claude workspace trust for <tmp>/spawn-base-default/wt; refusing to launch a claude worker that would wedge on the trust dialog; inspect window firstmate:fm-nmspawnbase-default2429394
exit=1
launch command sent: <none>
  <tmp>/spawn-base-default/project trust=false Approved=false WarningShown=false

## base-decline (repo root: base-tree) seeded project entry Approved=false WarningShown=true
$ fm-spawn.sh nmspawnbase-decline2429850 <project> claude --mode no-mistakes --yolo off
warning: <tmp>/spawn-base-decline/home/data/nmspawnbase-decline2429850/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: project entry for <tmp>/spawn-base-decline/project in <tmp>/spawn-base-decline/claude-config/.claude.json already declined external CLAUDE.md imports; refusing to override that consent
error: refusing to pre-register Claude trust: could not record trust for '<tmp>/spawn-base-decline/wt' and project '<tmp>/spawn-base-decline/project' in '<tmp>/spawn-base-decline/claude-config/.claude.json'
error: could not pre-register Claude workspace trust for <tmp>/spawn-base-decline/wt; refusing to launch a claude worker that would wedge on the trust dialog; inspect window firstmate:fm-nmspawnbase-decline2429850
exit=1
launch command sent: <none>
  <tmp>/spawn-base-decline/project trust=false Approved=false WarningShown=true

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ .agents/skills/harness-adapters/references/harness/claude.md:28 - This doc still says an explicit decline is the flag at ===false and that such a decline makes the whole registration refuse. That was true before this change and is wrong after it. Commit 94d90a3 makes a false/false entry (Claude Code's default, never asked) register trust normally. Now only Approved===false together with WarningShown===true refuses. Someone who reads this harness reference for a spawn into a project with the default entry will expect a refusal and a wedge on the trust dialog. What actually happens is that trust is recorded and the worker stops on the import dialog. Fix: change the parenthetical to hasClaudeMdExternalIncludesApproved===false with hasClaudeMdExternalIncludesWarningShown===true, and optionally add that a false/false default is treated like an absent flag. This is a one-line doc correction and changes no behavior.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Claude Code 2.1.270, opened once in a repo with no CLAUDE.md, writes a project entry with both external-imports flags false (the premise of the fix) ✅ pass live claude-2.1.270-default-project-entry.txt: the entry has hasTrustDialogAccepted=true, Approved=false, WarningShown=false; the import dialog was never shown
Spawn registration on a false/false entry that Claude itself wrote: HEAD exits 0 and records trust on the worktree and project root without adding import consent; base b182d0f refuses ✅ pass live live-claude-default-entry-then-decline.txt §1; cli-shape-matrix-head-vs-base.txt (default-false-false: HEAD exit=0, worktree import flags unset, project flags stay false; BASE exit=1)
After HEAD registers a false/false project whose CLAUDE.md imports an outside file, real claude launched in the worktree skips the trust dialog and still stops on the 'Allow external CLAUDE.md file im… ✅ pass live live-claude-fixture-shape-trust-skip-import-dialog.txt (project root trust was false before registration; first screen is the import dialog; choosing Yes goes straight to the REPL); live-claude-defaul…
Adversarial: a genuine 'No, disable' answered in real claude writes Approved=false, WarningShown=true, and HEAD still refuses the next worktree spawn without touching the store ✅ pass live live-claude-default-entry-then-decline.txt §3-4 (entry false/true; HEAD exit=1 'already declined external CLAUDE.md imports'; store sha256 unchanged)
Adversarial shapes that are not a decline (Approved=false with WarningShown absent, WarningShown=true with Approved absent, a non-boolean WarningShown) register trust and never set Approved=true ✅ pass live cli-shape-matrix-head-vs-base.txt (approved-false-warning-absent, warningshown-true-approved-absent, decline-string-true-not-boolean: HEAD exit=0, no import flag set to true on the worktree, project A…
Regression: standing 'Yes, allow' consent (true/true) is carried to a new worktree, and real claude launched there reaches the REPL with no dialog; an entry with no flags stays trust-only, same as bas… ✅ pass live live-claude-fixture-shape-trust-skip-import-dialog.txt (wt2 trust/Approved/WarningShown all true, REPL first screen); cli-shape-matrix-head-vs-base.txt (consent-true-true, flags-absent identical at HE…
The new test fails on base b182d0f with the refusal and passes at HEAD, and the existing decline fixture stays green on both ⏸️ untested no The earlier payload only ran the unit test file (tests/fm-claude-trust.test.sh) at HEAD and against a base archive. A unit test run is not a live product run, so it did not show a live result. The liv…
fm-spawn of a claude worker into a project with the default false/false entry launches the worker at HEAD (base refuses), and a false/true decline is refused at both ⏸️ untested no The earlier payload ran bin/fm-spawn.sh only with the test-fixture fake tmux backend and fake claude binary, so it did not show a live result. A live check needs fm-spawn run on a real tmux session wi…
  • bash tests/fm-claude-trust.test.sh at HEAD (31 ok, exit 0)
  • HEAD's tests/fm-claude-trust.test.sh run inside git archive b182d0f: the decline fixture passes, and test_project_root_entry_default_import_flags_are_not_a_decline fails with the refusal
  • bin/fm-claude-trust.sh &lt;worktree&gt; &lt;project&gt; at HEAD and base on real git worktrees with an isolated CLAUDE_CONFIG_DIR, across 7 project-entry shapes: false/false, false/true, false/absent, absent/absent, true/true, absent/true, false/"true"
  • Real claude 2.1.270 in tmux with an isolated CLAUDE_CONFIG_DIR/HOME and a fake API key: accepted trust in a repo with no CLAUDE.md, then read the project entry Claude wrote
  • Committed a CLAUDE.md with an import from outside the tree, ran base (refused) and HEAD (registered) on Claude's own entry, launched real claude into the worktree and captured the import dialog, answered 'No, disable', read the entry, then re-ran HEAD for a second worktree (refused, store unchanged)
  • Test-fixture shape (trust=false, false/false): HEAD registration, real claude launched into the worktree (import dialog, no trust dialog), answered 'Yes, allow', second-worktree registration carried consent forward, real claude reached the REPL with no dialog
  • Real bin/fm-spawn.sh ... claude --mode no-mistakes --yolo off using the test fixtures (fake tmux backend and fake claude binary), with false/false and false/true stores, at HEAD and at base from git archive
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Rangezi and others added 2 commits September 13, 2026 18:57
…asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

First look on main b182d0f908b78d08c7ccb8dce3775bdca8c5d657. Tip 6ae7979b5d2cf6b35bf70c6da01e246459551469. Fork PR (Rangezi).

What I inspected (diff vs main + issue #4378)

contract-class: restore — spawn/trust already promised not to treat never-asked defaults as decline; restores correct consent detection.

Attestation: MATCH (body head_sha == tip).

Checks: first-time fork workflows approved this pass — CI 34777072795, Require no-mistakes 34777072752. Waiting-ci for green tip CI+NM.

VISION.md (each rule)

  • One captain, one interface: aligns — trust/spawn gate stays on fm-claude-trust; no alternate UI.
  • Authority is explicit and never inferred: aligns — decline only when WarningShown=true with false; defaults are not inferred as consent/refusal.
  • Scripts own the mechanics agents own the judgment: aligns — script owns flag decode; human owns real decline.
  • A restart is a non-event: aligns — fresh defaults must not permanently refuse spawn.
  • Delegation with a spine: aligns — existing spawn spine; predicate only.
  • The fleet outlives any vendor: aligns — vendor default flags decoded carefully, not treated as human intent.
  • Scope: aligns — narrow decline-predicate fix; no new trust product.

Security FYI (tip only): none material — narrows refusal; never manufactures import consent.

Merge-eligible N. Waiting-ci. Do not merge. No Firstmate flag.

@dbordenxcures dbordenxcures left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. I verified the premise this rests on against real data rather than reasoning about it, and the measurement is worth recording here because it also sizes the bug.

On my machine, ~/.claude.json holds 169 project entries:

hasClaudeMdExternalIncludesApproved hasClaudeMdExternalIncludesWarningShown entries
false false 90
absent absent 77
true true 2
false true 0

So the false/false pair this PR reclassifies as never-asked is the single most common shape in the store - 90 of 169 entries, over half the machine, including directories I have never been asked about at all. The pair that the old predicate was actually written for, false with WarningShown===true, appears zero times. That is the whole defect: declinedExternalImports was matching the default entry, so the registration refused on the majority of projects and the spawn wedged on the trust dialog. It matches the reported symptom of workers being unable to launch on almost any clone.

The narrowed predicate is also safe in the direction that matters, which is why I am comfortable approving it on a premise I could not fully sample. I have no false/true entries, so I cannot confirm from my own store that an explicit "No, disable" sets WarningShown===true. If that assumption were wrong, the consequence is bounded: the consent write still requires an explicit hasClaudeMdExternalIncludesApproved===true, so a misclassified decline would land the trust flag only and never manufacture import consent. The new test asserts exactly that on both entries. Trust and consent stay separate, which is the property I care about - nothing here can spend a consent answer on my behalf.

Verified locally at the PR head with this change applied to my working tree: tests/fm-claude-trust.test.sh passes 31 of 31, exit 0, including the new never-asked case. CI is green at 14 of 14.

One documentation point, non-blocking: the reference file now states the false/false rule in prose, and the script header states it too. The script header is the right owner for the contract; if that prose ever needs editing again, the reference should shrink to a pointer rather than keep a second copy in step.

This is currently blocking dispatch across my whole fleet, so I would like it landed.

@Rangezi

Rangezi commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

The premise you could not sample is covered by live evidence in this PR's no-mistakes Test section: in real Claude Code 2.1.270, answering "No, disable" wrote Approved=false / WarningShown=true, and HEAD then refused the next worktree spawn with the store's sha256 unchanged. On the reference file: the sentence restating the decline rule predates this PR and became wrong once the predicate changed, so the PR corrects it in place. Shrinking the reference to a pointer would rewrite maintainer text beyond this fix; happy to leave that to a follow-up if the maintainer wants it.

@kunchenguid
kunchenguid merged commit 267441d into kunchenguid:main Sep 14, 2026
14 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @Rangezi — really appreciate you taking the time on this.

Tip was 6ae7979b5d2cf6b35bf70c6da01e246459551469 (attestation MATCH). Squash landed as 267441d72d2add4551dfa8a9ba6657f33665db22. Closes #4378. CI 34777072795 SUCCESS / NM 34777072752 SUCCESS. contract-class restore. Overlap leftover #4363 closed in favor of this PR.

@Rangezi
Rangezi deleted the fix/claude-trust-default-import-flags branch September 14, 2026 18:41
jjtylr added a commit to jjtylr/firstmate that referenced this pull request Sep 15, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: nateliuroberts <nate@cipherlab.ai>
Co-authored-by: Jon Roosevelt <jon@arcs.health>
Co-authored-by: AnPod <drejc83@gmail.com>
Co-authored-by: NewAiCoder-bot <iamacodernow-bot@theinbtw.com>
Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Co-authored-by: NewAiCoder <iamacodernow@theinbtw.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Rangezi <46404232+Rangezi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
jjtylr added a commit to jjtylr/firstmate that referenced this pull request Sep 15, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: nateliuroberts <nate@cipherlab.ai>
Co-authored-by: Jon Roosevelt <jon@arcs.health>
Co-authored-by: AnPod <drejc83@gmail.com>
Co-authored-by: NewAiCoder-bot <iamacodernow-bot@theinbtw.com>
Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Co-authored-by: NewAiCoder <iamacodernow@theinbtw.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Rangezi <46404232+Rangezi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
sanis added a commit to sanis/firstmate that referenced this pull request Sep 19, 2026
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that wa…
Ye1806431561 pushed a commit to Ye1806431561/firstmate that referenced this pull request Sep 20, 2026
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
sanis added a commit to sanis/firstmate that referenced this pull request Sep 21, 2026
…dpoint reclaim (#28)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In t…
doitdigital0495 added a commit to doitdigital0495/firstmate that referenced this pull request Sep 21, 2026
…ent checks (#20)

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
th…
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.


* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
kaku-san added a commit to kaku-san/firstmate that referenced this pull request Sep 24, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both direction…
rub-a-dub-dub added a commit to rub-a-dub-dub/firstmate that referenced this pull request Sep 27, 2026
…gences (#29)

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the pa…
adibirzu added a commit to adibirzu/firstmate that referenced this pull request Sep 30, 2026
* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visi…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fm-claude-trust.sh reads Claude Code's default external-imports flags (false/false) as a human decline and refuses the spawn

3 participants