Skip to content

fix(bin): read the latest status event so buried declarations and open decisions aren't lost - #3753

Merged
kunchenguid merged 26 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/firstmate-secondmate-buried-block-latest-line-read
Sep 17, 2026
Merged

kunchenguid merged 26 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/firstmate-secondmate-buried-block-latest-line-read

Conversation

@tiago-peixoto

@tiago-peixoto tiago-peixoto commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Remove the incidental .github/workflows/ci.yml change from #3753.
The branch added one top-level fm-fleet-snapshot-view test and changed the expected suite count from 18 to 19, but that workflow change is not required product behavior.
Preserve all regression coverage by integrating the new assertions into an existing semantically related top-level test, or by another minimal arrangement that keeps all behavior exercised while the suite still reports 18 top-level tests.
Update the existing pull request only through the no-mistakes pipeline, verify the workflow file has zero diff, checks pass, and the pull request is mergeable, then report the new exact head.
Do not merge the pull request.

What Changed

  • last_status_line in bin/fm-classify-lib.sh now returns the latest recognized status event instead of the last non-blank line, so trailing notes can no longer hide a blocked/needs-decision/done line. It reads a 200-line tail and falls back to the whole file. A new status_current_line returns the newest still-open decision, or the latest event when nothing is open. fm-crew-state.sh, fm-watch.sh and fm-inactive-reconcile.sh now use these shared readers instead of their own last-line reads.
  • The decision fold now looks at the task kind. A done:/failed: line from a ship or scout closes its stale open decisions, but a secondmate's does not. Only lines with a colon or a full [key=...] tag count as decisions. status_open_decisions takes an optional kind (the fleet snapshot passes it), fm-captain-hold.sh drops its own origin_open_decisions wrapper, and the open-decisions cursor moves to version 8 with the kind included. Old cursors are rebuilt.
  • The daemon-socket-down override in fm-crew-state.sh now applies only while that blocker is still the log's latest event. A later event hands the reading back to the run record. Regression tests are added or updated across the classify, crew-state, fleet-snapshot view, inactive-reconcile, watch-triage, captain-hold and cursor suites. The new fleet-snapshot assertions go into an existing test, so the CI workflow is unchanged. Related docs (docs/architecture.md, the fmx-respond skill) are updated.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: .github/workflows/ci.yml now has zero diff against the base, and the terminal-cleanup assertions were merged into the related completed-scout test with no assertion dropped, so the suite still lists 18 top-level tests; the rest of the branch's status-reader changes were already fixed in earlier review rounds and look consistent.

Testing

I confirmed ci.yml is identical to the base commit and still expects 18 tests. I ran the fleet-snapshot suite under the stock /bin/bash 3.2, the same way CI does: it passed with exactly 18 top-level tests. I also broke the product on purpose in scratch copies to confirm the merged checks still work: turning off only failed:ship clearing fails the moved terminal-cleanup assertion. Pushing to the PR, CI status, whether the PR can merge, and the new head commit are handled by later pipeline phases, not this one.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
The branch leaves .github/workflows/ci.yml identical to base, so CI still expects 18 snapshot/fleet-view tests ✅ pass live ci-workflow-zero-diff-and-count.txt (git diff --stat 795e5e4 64e0762 -- .github/workflows/ci.yml is empty; line 429 still checks -eq 18)
Running the fleet-snapshot suite under stock macOS /bin/bash 3.2, as CI does, reports exactly 18 passing top-level tests and no failures ✅ pass live /bin/bash tests/fm-fleet-snapshot-view.test.sh → exit 0, 18 ok -, 0 not ok (fleet-snapshot-view-bash32.txt)
Break-it check: the terminal-cleanup checks merged into the completed-scout test still fail when ship failed stops clearing decisions ✅ pass live Scratch copy with failed:ship removed from the clearing rule → not ok - terminal snapshot revived a completed decision or lost a current one (mutation-disable-failed-ship-supersession.txt)
Existing PR updated through the pipeline, checks pass, PR mergeable, new head reported ⏸️ untested no This test phase is not allowed to push, run CI, or open or update PRs. The outer no-mistakes pipeline's push/CI phases handle that and report the final head and whether the PR can merge.
Evidence: CI workflow zero diff + suite count under /bin/bash 3.2

Source: CI workflow zero diff + suite count under /bin/bash 3.2

git diff --stat 795e5e4 64e0762 -- .github/workflows/ci.yml:
(empty = zero diff)

          snapshot_output=$(/bin/bash tests/fm-fleet-snapshot-view.test.sh)
          printf '%s\n' "$snapshot_output"
          snapshot_count=$(printf '%s\n' "$snapshot_output" | grep -c '^ok - ')
          [ "$snapshot_count" -eq 18 ] || {
            echo "::error::expected 18 snapshot/fleet-view tests, got $snapshot_count"
            exit 1
          }

/bin/bash tests/fm-fleet-snapshot-view.test.sh ok count:
18
0
Evidence: fm-fleet-snapshot-view suite transcript under stock /bin/bash 3.2

Source: fm-fleet-snapshot-view suite transcript under stock /bin/bash 3.2

ok - empty fleet snapshot and view use explicit absence markers
ok - fixture snapshot covers task rows, backlog rows, pointers, and stable ordering
ok - home-summary excludes kind=secondmate from unowned_current and terminal_in_flight
ok - undated captain holds age after a configurable threshold, decided only from structured fields
ok - captain-hold buckets are total, mutually exclusive, and never decided by prose
ok - main_inventory discloses orphan/unstructured and clears when inventory is consistent
ok - backlog normalization preserves strict roles and resolves every blocker compatibly
ok - snapshot event hints follow reconciled current state
ok - durable fold keeps an open decision past a later unrelated event
ok - a live secondmate endpoint preserves unrelated open decisions
ok - durable captain-held transfer closes the duplicate live status decision
ok - durable fold clears a decision only on a keyed resolution
ok - a completed scout's stale decision surfaces as a report pointer, not pending
ok - a scout still parked at a decision stays pending (terminal clear does not over-fire)
ok - snapshot includes durable scout reports after teardown
ok - snapshot parses tasks-axi rows and respects operational overrides
ok - fleet view renders the snapshot without secondmate peek guidance
ok - fleet view renders secondmate agent liveness
Evidence: Deliberate bug: failed:ship clearing turned off, caught by merged checks

Source: Deliberate bug: failed:ship clearing turned off, caught by merged checks

ok - empty fleet snapshot and view use explicit absence markers
ok - fixture snapshot covers task rows, backlog rows, pointers, and stable ordering
ok - home-summary excludes kind=secondmate from unowned_current and terminal_in_flight
ok - undated captain holds age after a configurable threshold, decided only from structured fields
ok - captain-hold buckets are total, mutually exclusive, and never decided by prose
ok - main_inventory discloses orphan/unstructured and clears when inventory is consistent
ok - backlog normalization preserves strict roles and resolves every blocker compatibly
ok - snapshot event hints follow reconciled current state
ok - durable fold keeps an open decision past a later unrelated event
ok - a live secondmate endpoint preserves unrelated open decisions
ok - durable captain-held transfer closes the duplicate live status decision
ok - durable fold clears a decision only on a keyed resolution
not ok - terminal snapshot revived a completed decision or lost a current one: {
  "schema": "fm-fleet-snapshot.v1",
  "generated": "2026-09-17T03:34:19Z",
  "fm_home": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup",
  "roots": {
    "fm_root": "/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/tmp.qiThK7ZQig",
    "state": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state",
    "data": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data",
    "config": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/config",
    "projects": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/projects"
  },
  "backlog": {
    "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/backlog.md",
    "present": false,
    "records": []
  },
  "tasks": [
    {
      "id": "scout-done",
      "kind": "scout",
      "harness": "claude",
      "mode": "",
      "yolo": "",
      "project": "",
      "spawn_gen": null,
      "backend": "tmux",
      "remote": null,
      "paths": {
        "meta": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/scout-done.meta",
          "present": true
        },
        "status_log": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/scout-done.status",
          "present": true,
          "kind": "event_history",
          "last_event": {
            "state": "note",
            "note": "cleanup complete",
            "raw": "note: cleanup complete"
          }
        },
        "worktree": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/projects/task",
          "present": true
        },
        "home": {
          "path": null,
          "present": false
        },
        "report": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/scout-done/report.md",
          "present": false
        }
      },
      "secondmate_projects": [],
      "current_state": {
        "state": "unknown",
        "source": "none",
        "detail": "no current-state source available",
        "raw": "state: unknown · source: none · no current-state source available",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "endpoint": {
        "target": "firstmate:fm-scout-done",
        "exists": true,
        "agent_alive": "not_checked",
        "status": "unknown",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "pr": {
        "url": null,
        "source": "absent",
        "head": null
      },
      "hints": {
        "pending_decision": false,
        "blocked_event": false,
        "open_decisions": [],
        "scout_report_present": false,
        "last_event_text": "note: cleanup complete"
      },
      "actions": {
        "watch": "bin/fm-peek.sh fm-scout-done",
        "steer": "bin/fm-send.sh fm-scout-done '<instruction>'",
        "return_channel_note": null
      },
      "backlog": null
    },
    {
      "id": "scout-failed",
      "kind": "scout",
      "harness": "claude",
      "mode": "",
      "yolo": "",
      "project": "",
      "spawn_gen": null,
      "backend": "tmux",
      "remote": null,
      "paths": {
        "meta": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/scout-failed.meta",
          "present": true
        },
        "status_log": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/scout-failed.status",
          "present": true,
          "kind": "event_history",
          "last_event": {
            "state": "note",
            "note": "cleanup complete",
            "raw": "note: cleanup complete"
          }
        },
        "worktree": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/projects/task",
          "present": true
        },
        "home": {
          "path": null,
          "present": false
        },
        "report": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/scout-failed/report.md",
          "present": false
        }
      },
      "secondmate_projects": [],
      "current_state": {
        "state": "unknown",
        "source": "none",
        "detail": "no current-state source available",
        "raw": "state: unknown · source: none · no current-state source available",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "endpoint": {
        "target": "firstmate:fm-scout-failed",
        "exists": true,
        "agent_alive": "not_checked",
        "status": "unknown",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "pr": {
        "url": null,
        "source": "absent",
        "head": null
      },
      "hints": {
        "pending_decision": false,
        "blocked_event": false,
        "open_decisions": [],
        "scout_report_present": false,
        "last_event_text": "note: cleanup complete"
      },
      "actions": {
        "watch": "bin/fm-peek.sh fm-scout-failed",
        "steer": "bin/fm-send.sh fm-scout-failed '<instruction>'",
        "return_channel_note": null
      },
      "backlog": null
    },
    {
      "id": "secondmate-done",
      "kind": "secondmate",
      "harness": "claude",
      "mode": "",
      "yolo": "",
      "project": "",
      "spawn_gen": null,
      "backend": "tmux",
      "remote": null,
      "paths": {
        "meta": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/secondmate-done.meta",
          "present": true
        },
        "status_log": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/state/secondmate-done.status",
          "present": true,
          "kind": "event_history",
          "last_event": {
            "state": "note",
            "note": "cleanup complete",
            "raw": "note: cleanup complete"
          }
        },
        "worktree": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/projects/task",
          "present": true
        },
        "home": {
          "path": null,
          "present": false
        },
        "report": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/secondmate-done/report.md",
          "present": false
        }
      },
      "secondmate_projects": [],
      "current_state": {
        "

... [7157 bytes truncated] ...

us",
          "present": true,
          "kind": "event_history",
          "last_event": {
            "state": "note",
            "note": "cleanup complete",
            "raw": "note: cleanup complete"
          }
        },
        "worktree": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/projects/task",
          "present": true
        },
        "home": {
          "path": null,
          "present": false
        },
        "report": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/ship-failed/report.md",
          "present": false
        }
      },
      "secondmate_projects": [],
      "current_state": {
        "state": "parked",
        "source": "status-log",
        "detail": "choose a route",
        "raw": "state: parked · source: status-log · choose a route",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "endpoint": {
        "target": "firstmate:fm-ship-failed",
        "exists": true,
        "agent_alive": "not_checked",
        "status": "unknown",
        "observed_at": "2026-09-17T03:34:19Z",
        "freshness": "fresh"
      },
      "pr": {
        "url": null,
        "source": "absent",
        "head": null
      },
      "hints": {
        "pending_decision": true,
        "blocked_event": true,
        "open_decisions": [
          {
            "key": "access",
            "verb": "blocked",
            "summary": "waiting"
          },
          {
            "key": "choice",
            "verb": "needs-decision",
            "summary": "choose a route"
          }
        ],
        "scout_report_present": false,
        "last_event_text": "note: cleanup complete"
      },
      "actions": {
        "watch": "bin/fm-peek.sh fm-ship-failed",
        "steer": "bin/fm-send.sh fm-ship-failed '<instruction>'",
        "return_channel_note": null
      },
      "backlog": null
    }
  ],
  "main_inventory": {
    "valid": true,
    "reason": null,
    "orphan_in_flight": [],
    "unstructured_current_count": 0
  },
  "contributions": {
    "known": 0,
    "checked": 0,
    "counts": {
      "captain": 0,
      "fleet": 0,
      "maintainer": 0,
      "nobody": 0
    },
    "unmeasured": 0,
    "complete": false,
    "proven_clear": false,
    "stale_verdicts": 0,
    "missing_verdicts": 0,
    "unreadable_records": 1,
    "valid_until": 900,
    "captain": [],
    "captain_omitted": 0
  },
  "scout_reports": [],
  "secondmate_current": {
    "registry": {
      "present": false,
      "available": true,
      "complete": true,
      "reason": null,
      "provenance": "registered-table",
      "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.myFacY/terminal-cleanup/data/secondmates.md",
      "freshness": {
        "status": "fresh",
        "observed_at": "2026-09-17T03:34:19Z"
      },
      "records": [],
      "input_truncated": false,
      "records_truncated": false,
      "reasons": [],
      "lines_in_window": 0,
      "records_in_window": 0
    },
    "records": [
      {
        "id": "secondmate-done",
        "home": null,
        "host": null,
        "remote": false,
        "registered": false,
        "spawn_gen": null,
        "current": {
          "state": "unknown",
          "reason": "secondmate metadata is not registered"
        },
        "invalidity": null,
        "reconcile_inventory": null,
        "provenance": {
          "selected": "parent-event-fallback",
          "structured_home": null,
          "parent_event_role": "fallback-only-not-current"
        },
        "freshness": {
          "status": "historical-event",
          "observed_at": "2026-09-17T03:34:19Z",
          "age_seconds": 0
        },
        "active_children": [],
        "decisions_open": [],
        "holds": [],
        "queued": [],
        "landed": [],
        "endpoints": [],
        "counts": {
          "active_children": 0,
          "decisions_open": 0,
          "holds": 0,
          "queued": 0,
          "landed": 0,
          "endpoints": 0
        },
        "omitted": [],
        "parent_event": {
          "raw": "note: cleanup complete",
          "note": "cleanup complete",
          "age_seconds": 0,
          "open_activities": [],
          "open_decisions": [
            {
              "key": "access",
              "verb": "blocked",
              "summary": "waiting"
            },
            {
              "key": "choice",
              "verb": "needs-decision",
              "summary": "choose a route"
            }
          ],
          "activity_scan": {
            "records": [],
            "available": true,
            "input_truncated": false,
            "retained_truncated": false,
            "reasons": [],
            "lines_in_window": 4,
            "records_in_window": 0
          }
        },
        "terminal_evidence": {
          "provenance": "parent-direct-report-terminal",
          "trust": "untrusted-supplement",
          "captured": true,
          "observed_at": "2026-09-17T03:34:19Z",
          "freshness": "fresh",
          "reason": null,
          "lines": 2,
          "bytes": 12,
          "event_note_seen": false,
          "contradiction": false
        },
        "contradiction": false
      },
      {
        "id": "secondmate-failed",
        "home": null,
        "host": null,
        "remote": false,
        "registered": false,
        "spawn_gen": null,
        "current": {
          "state": "unknown",
          "reason": "secondmate metadata is not registered"
        },
        "invalidity": null,
        "reconcile_inventory": null,
        "provenance": {
          "selected": "parent-event-fallback",
          "structured_home": null,
          "parent_event_role": "fallback-only-not-current"
        },
        "freshness": {
          "status": "historical-event",
          "observed_at": "2026-09-17T03:34:19Z",
          "age_seconds": 0
        },
        "active_children": [],
        "decisions_open": [],
        "holds": [],
        "queued": [],
        "landed": [],
        "endpoints": [],
        "counts": {
          "active_children": 0,
          "decisions_open": 0,
          "holds": 0,
          "queued": 0,
          "landed": 0,
          "endpoints": 0
        },
        "omitted": [],
        "parent_event": {
          "raw": "note: cleanup complete",
          "note": "cleanup complete",
          "age_seconds": 0,
          "open_activities": [],
          "open_decisions": [
            {
              "key": "access",
              "verb": "blocked",
              "summary": "waiting"
            },
            {
              "key": "choice",
              "verb": "needs-decision",
              "summary": "choose a route"
            }
          ],
          "activity_scan": {
            "records": [],
            "available": true,
            "input_truncated": false,
            "retained_truncated": false,
            "reasons": [],
            "lines_in_window": 4,
            "records_in_window": 0
          }
        },
        "terminal_evidence": {
          "provenance": "parent-direct-report-terminal",
          "trust": "untrusted-supplement",
          "captured": true,
          "observed_at": "2026-09-17T03:34:19Z",
          "freshness": "fresh",
          "reason": null,
          "lines": 2,
          "bytes": 12,
          "event_note_seen": false,
          "contradiction": false
        },
        "contradiction": false
      }
    ],
    "total_registered": 0,
    "total": 2,
    "shown": 2,
    "truncated": 0
  },
  "secondmate_landed": {
    "records": [],
    "truncated": [],
    "unreadable": [
      "<secondmate-done: unavailable>",
      "<secondmate-failed: unavailable>"
    ],
    "partial": []
  },
  "secondmate_guidance": {
    "note": "For kind=secondmate, bearings selects validated structured state from that registered home; parent events and bounded terminal evidence are fallback-only supplements and never current-state authority."
  }
}
Evidence: Deliberate bug: all ship/scout clearing turned off, caught by original assertion

Source: Deliberate bug: all ship/scout clearing turned off, caught by original assertion

ok - empty fleet snapshot and view use explicit absence markers
ok - fixture snapshot covers task rows, backlog rows, pointers, and stable ordering
ok - home-summary excludes kind=secondmate from unowned_current and terminal_in_flight
ok - undated captain holds age after a configurable threshold, decided only from structured fields
ok - captain-hold buckets are total, mutually exclusive, and never decided by prose
ok - main_inventory discloses orphan/unstructured and clears when inventory is consistent
ok - backlog normalization preserves strict roles and resolves every blocker compatibly
ok - snapshot event hints follow reconciled current state
ok - durable fold keeps an open decision past a later unrelated event
ok - a live secondmate endpoint preserves unrelated open decisions
ok - durable captain-held transfer closes the duplicate live status decision
ok - durable fold clears a decision only on a keyed resolution
not ok - a completed scout report must be a pointer, not a pending decision: {
  "schema": "fm-fleet-snapshot.v1",
  "generated": "2026-09-17T03:34:03Z",
  "fm_home": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout",
  "roots": {
    "fm_root": "/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/tmp.COBQWxbOEx",
    "state": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/state",
    "data": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/data",
    "config": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/config",
    "projects": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/projects"
  },
  "backlog": {
    "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/data/backlog.md",
    "present": false,
    "records": []
  },
  "tasks": [
    {
      "id": "lavish-103",
      "kind": "scout",
      "harness": "claude",
      "mode": "scout",
      "yolo": "",
      "project": "firstmate",
      "spawn_gen": null,
      "backend": "tmux",
      "remote": null,
      "paths": {
        "meta": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/state/lavish-103.meta",
          "present": true
        },
        "status_log": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/state/lavish-103.status",
          "present": true,
          "kind": "event_history",
          "last_event": {
            "state": "done",
            "note": "report ready at data/lavish-103/report.md",
            "raw": "done: report ready at data/lavish-103/report.md"
          }
        },
        "worktree": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/projects/scout-wt",
          "present": true
        },
        "home": {
          "path": null,
          "present": false
        },
        "report": {
          "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/data/lavish-103/report.md",
          "present": true
        }
      },
      "secondmate_projects": [],
      "current_state": {
        "state": "parked",
        "source": "status-log",
        "detail": "adopt approach A or B for Lavish issue 103",
        "raw": "state: parked · source: status-log · adopt approach A or B for Lavish issue 103",
        "observed_at": "2026-09-17T03:34:03Z",
        "freshness": "fresh"
      },
      "endpoint": {
        "target": "firstmate:fm-lavish-103",
        "exists": true,
        "agent_alive": "not_checked",
        "status": "unknown",
        "observed_at": "2026-09-17T03:34:03Z",
        "freshness": "fresh"
      },
      "pr": {
        "url": null,
        "source": "absent",
        "head": null
      },
      "hints": {
        "pending_decision": true,
        "blocked_event": false,
        "open_decisions": [
          {
            "key": "default",
            "verb": "needs-decision",
            "summary": "adopt approach A or B for Lavish issue 103"
          }
        ],
        "scout_report_present": true,
        "last_event_text": "done: report ready at data/lavish-103/report.md"
      },
      "actions": {
        "watch": "bin/fm-peek.sh fm-lavish-103",
        "steer": "bin/fm-send.sh fm-lavish-103 '<instruction>'",
        "return_channel_note": null
      },
      "backlog": null
    }
  ],
  "main_inventory": {
    "valid": true,
    "reason": null,
    "orphan_in_flight": [],
    "unstructured_current_count": 0
  },
  "contributions": {
    "known": 0,
    "checked": 0,
    "counts": {
      "captain": 0,
      "fleet": 0,
      "maintainer": 0,
      "nobody": 0
    },
    "unmeasured": 0,
    "complete": false,
    "proven_clear": false,
    "stale_verdicts": 0,
    "missing_verdicts": 0,
    "unreadable_records": 1,
    "valid_until": 900,
    "captain": [],
    "captain_omitted": 0
  },
  "scout_reports": [
    {
      "id": "lavish-103",
      "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/data/lavish-103/report.md",
      "kind": "scout"
    }
  ],
  "secondmate_current": {
    "registry": {
      "present": false,
      "available": true,
      "complete": true,
      "reason": null,
      "provenance": "registered-table",
      "path": "/private/var/folders/r8/cylyt7xd7t50y9x5wc05my380000gn/T/fm-fleet-snapshot.GknGdF/completed-scout/data/secondmates.md",
      "freshness": {
        "status": "fresh",
        "observed_at": "2026-09-17T03:34:03Z"
      },
      "records": [],
      "input_truncated": false,
      "records_truncated": false,
      "reasons": [],
      "lines_in_window": 0,
      "records_in_window": 0
    },
    "records": [],
    "total_registered": 0,
    "total": 0,
    "shown": 0,
    "truncated": 0
  },
  "secondmate_landed": {
    "records": [],
    "truncated": [],
    "unreadable": [],
    "partial": []
  },
  "secondmate_guidance": {
    "note": "For kind=secondmate, bearings selects validated structured state from that registered home; parent events and bounded terminal evidence are fallback-only supplements and never current-state authority."
  }
}
Evidence: Deliberate change: kind argument removed, no effect because kind comes from .meta

Source: Deliberate change: kind argument removed, no effect because kind comes from .meta

ok - empty fleet snapshot and view use explicit absence markers
ok - fixture snapshot covers task rows, backlog rows, pointers, and stable ordering
ok - home-summary excludes kind=secondmate from unowned_current and terminal_in_flight
ok - undated captain holds age after a configurable threshold, decided only from structured fields
ok - captain-hold buckets are total, mutually exclusive, and never decided by prose
ok - main_inventory discloses orphan/unstructured and clears when inventory is consistent
ok - backlog normalization preserves strict roles and resolves every blocker compatibly
ok - snapshot event hints follow reconciled current state
ok - durable fold keeps an open decision past a later unrelated event
ok - a live secondmate endpoint preserves unrelated open decisions
ok - durable captain-held transfer closes the duplicate live status decision
ok - durable fold clears a decision only on a keyed resolution
ok - a completed scout's stale decision surfaces as a report pointer, not pending
ok - a scout still parked at a decision stays pending (terminal clear does not over-fire)
ok - snapshot includes durable scout reports after teardown
ok - snapshot parses tasks-axi rows and respects operational overrides
ok - fleet view renders the snapshot without secondmate peek guidance
ok - fleet view renders secondmate agent liveness

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
The branch leaves .github/workflows/ci.yml identical to base, so CI still expects 18 snapshot/fleet-view tests ✅ pass live ci-workflow-zero-diff-and-count.txt (git diff --stat 795e5e4 64e0762 -- .github/workflows/ci.yml is empty; line 429 still checks -eq 18)
Running the fleet-snapshot suite under stock macOS /bin/bash 3.2, as CI does, reports exactly 18 passing top-level tests and no failures ✅ pass live /bin/bash tests/fm-fleet-snapshot-view.test.sh → exit 0, 18 ok -, 0 not ok (fleet-snapshot-view-bash32.txt)
Break-it check: the terminal-cleanup checks merged into the completed-scout test still fail when ship failed stops clearing decisions ✅ pass live Scratch copy with failed:ship removed from the clearing rule → not ok - terminal snapshot revived a completed decision or lost a current one (mutation-disable-failed-ship-supersession.txt)
Existing PR updated through the pipeline, checks pass, PR mergeable, new head reported ⏸️ untested no This test phase is not allowed to push, run CI, or open or update PRs. The outer no-mistakes pipeline's push/CI phases handle that and report the final head and whether the PR can merge.
  • git diff --stat 795e5e4 64e0762 -- .github/workflows/ci.yml (empty) and git show 39c74c5, which confirms the checks were moved into another test, not deleted
  • /bin/bash tests/fm-fleet-snapshot-view.test.sh under GNU bash 3.2.57 with tasks-axi and jq, the same way CI runs it: exit 0, 18 ok - lines, no not ok
  • Break-it check in a scratch copy: removed the &#34;$kind&#34; argument in bin/fm-fleet-snapshot.sh. The suite still passes, because the function looks up the kind from the task's .meta file when no argument is given, so this change alters nothing
  • Break-it check in a scratch copy: turned off all ship/scout done/failed clearing in bin/fm-classify-lib.sh. The suite fails at the original completed-scout assertion
  • Break-it check in a scratch copy: turned off only the failed:ship clearing. The suite fails with not ok - terminal snapshot revived a completed decision or lost a current one, which comes from the merged checks
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@tiago-peixoto
tiago-peixoto marked this pull request as ready for review September 5, 2026 00:04
@greptile-apps

greptile-apps Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge because the previously reported socket-blocker conflation is fixed and no blocking failure remains.

No blocking failure remains.

Reviews (6): Last reviewed commit: "no-mistakes(ci): Greptile P1 (bin/fm-cre..." | Re-trigger Greptile

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

First look on HEAD 71b3606b05e516c3d44bc2f1138c46902d858517 (Fixes #3186). Diff reviewed: shared last_status_line scans recognized verbs (incl. resolve/captain-held), status_current_line keeps open blockers ahead of unrelated appends, crew-state drops its private last-line helper. Tests cover multiline pause, trailing blanks, buried blocker. Greptile SUCCESS. Workflow-zero vs main. No security tip.

Attestation: MISSING (fork cannot supply). CI/NM: first-time fork workflows approved this pass (33931235722 CI, 33931235592/33931743475/33932040384 Require no-mistakes); queued after approve.

VISION.md (each rule)

  • One captain, one interface: aligns — stops false undeclared-idle / quiet-worker noise.
  • Authority is explicit: aligns — tightens the reader; no new grant.
  • Scripts own the mechanics: aligns — status-log parse is exact script work.
  • A restart is a non-event: aligns — durable declared verb survives multiline append.
  • Delegation with a spine: aligns — restores existing current-state contract.
  • The fleet outlives any vendor: aligns — harness-agnostic log.
  • Scope: aligns — watcher/crew-state; field incident → regression.

Contract-class: restore. Merge-eligible N until attestation MATCH and CI/NM green. Waiting on the author for attestation rebind (not a captain hold). Do not auto-merge yet.

@tiago-peixoto
tiago-peixoto force-pushed the fm/firstmate-secondmate-buried-block-latest-line-read branch from 53658f8 to 5269b73 Compare September 5, 2026 11:43
@tiago-peixoto tiago-peixoto changed the title fix: preserve declared status through trailing log lines fix(bin): preserve declared status through trailing log lines Sep 5, 2026
@tiago-peixoto
tiago-peixoto force-pushed the fm/firstmate-secondmate-buried-block-latest-line-read branch from 5269b73 to 85b1b3f Compare September 6, 2026 08:45
@tiago-peixoto tiago-peixoto changed the title fix(bin): preserve declared status through trailing log lines fix: resolve crew status from the latest recognized event and keyed decision fold Sep 6, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is waiting on the captain (not the author).

Tip 2026-09-06 review on HEAD 85b1b3f3cebaaeedc2364977bc3d318e6e2a14e3 (Fixes #3186): attestation MATCH; CI 34022741596 + tip NM 34023636572 SUCCESS; Greptile SUCCESS. Contract-class restore (status_current_line / last_status_line keep declared blockers and recognized verbs across multiline status records). Workflow delta is fleet-view expected-count 18→19 only — not a gate weaken. Security clean.

Merge blocker: squash via the triage gh OAuth app fails with refusing to allow an OAuth App to create or update workflow .github/workflows/ci.yml without workflow scope. Branch ruleset also marks unattributed fork changes BLOCKED (admin can bypass). Otherwise ready for auto-merge as restore.

Captain: please land with a token that has workflow scope (or apply the expected-count bump another way), then this can close #3186.

@tiago-peixoto
tiago-peixoto force-pushed the fm/firstmate-secondmate-buried-block-latest-line-read branch from 85b1b3f to 444a3b9 Compare September 6, 2026 15:52
@tiago-peixoto tiago-peixoto changed the title fix: resolve crew status from the latest recognized event and keyed decision fold fix(bin): resolve crew status from the latest recognized event, not the last line Sep 6, 2026
Comment thread bin/fm-crew-state.sh Outdated
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is waiting on the captain (not the author).

Newer activity since stamp 2026-09-06T10:17:40Z (head 85b1b3f3): tip advanced to 925147c3366493e7570c4fb77d9b626da87b7849 — docs/header contract corrections + Greptile P1 fix (stale socket-blocker must not latch after a later event). Attestation MATCH on new tip. Behavior CI + Greptile SUCCESS; tip Require no-mistakes SUCCESS (101523308417; one earlier synchronize FAIL is superseded). MERGEABLE but mergeStateStatus BLOCKED.

Contract-class: restore — status/crew readers already promised current declared state; multiline continuation prose was breaking that contract (Fixes #3186). Workflow delta remains fleet-view expected-count 18→19 only — not a gate weaken.

VISION.md (each rule): One captain/interface — aligns (stops false quiet/undeclared-idle). Authority — aligns. Scripts/agents — aligns (exact status parse). Restart non-event — aligns (durable verb survives). Delegation spine — aligns. Fleet outlives vendor — aligns. Scope — aligns (watcher/crew-state; field incident → regression).

Merge blocker unchanged in kind: gh pr merge --squash refused — base branch policy prohibits the merge (workflow file touch + ruleset). Prior OAuth workflow scope issue still applies for this app. Otherwise ready for auto-merge as restore.

Captain: please land with a token that has workflow scope (or --admin bypass), then this can close #3186. Do not rebase from firstmate while the captain hold stands.

@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Checking in, five days on. As far as I can verify there is nothing outstanding from my side, but I would rather ask than assume.

Re-verified against the live PR just now:

  • Head is still 925147c3366493e7570c4fb77d9b626da87b7849. Nothing has been pushed since your 2026-09-06 triage, and nothing will be while your hold stands.
  • The attestation in the body still binds that exact head, with review, test and document recorded complete.
  • Of the 17 check runs on that head, 16 pass. The single red entry is the earlier synchronize run of "PR must be raised via no-mistakes" (run 34045513277), which evaluated body compliance about two seconds before the attestation was written into the body. The two later runs of that same gate on the same commit pass (34045515620, and check run 101523308417 as cited in your triage) - consistent with your note that "one earlier synchronize FAIL is superseded".

Your triage put the remaining blocker on branch policy plus workflow scope for the squash, since the diff touches .github/workflows/ci.yml.

Is there anything needed from me to unblock it? On the workflow file specifically: the only change there is the snapshot/fleet-view expected count, 18 to 19, which is the counterpart to the fleet-view test this PR adds, so dropping it on its own would make that guard fail. If your policy would rather this PR not touch a workflow file at all, I can land the count bump as its own small PR first, or drop the added test, whichever you prefer. Happy to follow your lead.

@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

Another brief check-in.

Head is still 925147c3366493e7570c4fb77d9b626da87b7849 - unchanged since your 2026-09-06 waiting-captain triage and our 2026-09-11 note. Nothing pushed from our side while your hold stands.

As far as we can verify: attestation still binds that head; the only red body-compliance entry is the superseded pre-attestation synchronize run. Mergeable stays true; mergeStateStatus remains BLOCKED on the workflow-scope / branch-policy class you already named.

Is there anything needed from us, or does it still only need the scoped merge on your side?

andrewesweet added a commit to andrewesweet/firstmate that referenced this pull request Sep 17, 2026
…fixes (#19)

* fix(bin): clear parent pending-replies on local secondmate retirement (kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (kunchenguid#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

---------

Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
rega10 added a commit to rega10/firstmate that referenced this pull request Sep 17, 2026
* fix(bin): report verified PR state for passed runs (kunchenguid#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes kunchenguid#4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes kunchenguid#4469) (kunchenguid#4627)

* fix: restore published contribution follow-up (Fixes kunchenguid#4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (kunchenguid#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (kunchenguid#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (kunchenguid#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (kunchenguid#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (kunchenguid#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* no-mistakes(document): Clarify typed dispatch gate ownership

---------

Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
Co-authored-by: Rene Garza Jr. <rega1011@RG-Mac-mini.local>
vifar added a commit to vifar/firstmate that referenced this pull request Sep 17, 2026
… commits, adopt 31 upstream commits) (#13)

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

---------

Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
vifar added a commit to vifar/firstmate that referenced this pull request Sep 18, 2026
* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* no-mistakes(ci): Fixed Lint SC2034 in tests/fm-inactive-reconcile.test.sh (unused attempt → _) and tests/fm-spawn-pool-base-freshen.test.sh (shared-record LOCAL_ONLY_TIP disable). Fixed Behavior portable serial 4 by aligning tests/fm-wake-daemon-lifecycle-e2e.test.sh with classify_stale's post-c6e3cc1 contract: recorded meta + FM_FAKE_CREW_STATE working evidence so transient stale self-handles. Verified locally: lifecycle e2e passes; shellcheck -x on the three files exits 0. PR must be raised via no-mistakes is an attestation/pipeline check, not a code defect—no code change for it

* no-mistakes(document): Document Claude foreign-owner session-lock consumers

---------

Co-authored-by: Pablo Ontiveros <pablo.ontiveros@gmail.com>
Co-authored-by: Umer <umeranjum17@gmail.com>
Co-authored-by: Yasuhito Takamiya <yasuhito@hey.com>
Co-authored-by: Marsjohn-11 <74795701+Marsjohn-11@users.noreply.github.com>
Co-authored-by: tbillings28 <todd@toddbillings.com>
Co-authored-by: Todd Billings <todd@usdvcapital.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Amin Roudaki <roudaky@gmail.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Sebastian <80847374+thelad-dev@users.noreply.github.com>
Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
Co-authored-by: Juan José González Giraldo <juanjose.eng@gmail.com>
Co-authored-by: Cody <72239807+codyjohnsontx@users.noreply.github.com>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
sanis added a commit to sanis/firstmate that referenced this pull request Sep 19, 2026
* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that wa…
Ye1806431561 pushed a commit to Ye1806431561/firstmate that referenced this pull request Sep 20, 2026
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
sanis added a commit to sanis/firstmate that referenced this pull request Sep 21, 2026
…dpoint reclaim (#28)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* fix(bin): select suites that read a changed top-level test fixture (#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs https://github.com/kunchenguid/firstmate/issues/4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In t…
doitdigital0495 added a commit to doitdigital0495/firstmate that referenced this pull request Sep 21, 2026
…ent checks (#20)

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
th…
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
ionnich added a commit to ionnich/firstmate that referenced this pull request Sep 22, 2026
* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path

* fix(bin): create captain-hold rows when Beads requires due (#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: disable compact adviser for spawned agents (#4877)

* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee

* fix(bin): preserve Claude lock ownership after helper recycling (#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names t…
kunchenguid added a commit that referenced this pull request Sep 23, 2026
* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
bcriswell added a commit to bcriswell/firstmate that referenced this pull request Sep 23, 2026
* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path

* fix(bin): create captain-hold rows when Beads requires due (#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: disable compact adviser for spawned agents (#4877)

* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee

* fix(bin): preserve Claude lock ownership after helper recycling (#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation

* feat: park main under the away posture on Pi (#4889)

* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation

* ci: standardize workflow timeouts into three tiers (#4910)

* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy

* fix(bin): keep supervisor status closes from waking the same home (#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs

* fix(bin): stop labeling Herdr as experimental (#4972)

* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording

* fix(bin): treat a live no-mistakes run as current after rebase (#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict

* fix(bin): prevent long worker launch command truncation (#4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs #4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs #4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>

* test: authorize isolated Herdr lab validation (#4998)

* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.

* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): unify Lavish host and disconnect handling (#5060)

* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees

* feat: act on captain's away words during AFK supervision (#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation

* fix(bin): render the remote charter's steering-inbox path host-local (#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes #5012

* no-mistakes(document): document remote charter's host-local steering inbox

* feat: route Lavish feedback directly to owning workers (#5099)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle

* fix(bin): fit pull observation within the contribution poll budget (#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget

* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>

* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>

* feat(bin): append optional home-local include to briefs (#5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>

* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout

* fix(bin): require a non-draft pull request before a PR-based done report (#5141)

* fix(bin): require a no…
dscott98 added a commit to dscott98/firstmate that referenced this pull request Sep 23, 2026
* feat: cut validation token cost without lowering review rigor

Classify documentation-only changes as direct-PR surfaces on
no-mistakes-prod-only projects, fold small pre-merge follow-ups into
the open task's branch for one validation of the final head, document
the machine-global per-purpose model routing no-mistakes supports
(economical agent_config default with strong review_agents pins and its
unsupported per-step boundary), and pin the repo-side boundary with a
deterministic contract test.

* no-mistakes(review): Prioritize full validation over every lighter delivery category

* no-mistakes(document): Remove machine-specific claims from validation routing documentation

* fix: submit stuck inbox doorbells instead of skipping them (kunchenguid#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation

* docs: state the captain-approved model-routing boundary precisely

Strong models serve review and review fixes; economical models serve all
other validation work and fixes, which is the approved policy rather than
an unsupported claim that difficult non-review fixes get stronger models.

* feat: add opt-in fleet activity ledger (kunchenguid#5375)

* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed

* fix: validate public follow-up deliverables and wake on rejection (kunchenguid#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation

* feat: add Devin CLI crewmate and scout adapter (kunchenguid#5380)

* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>

* fix(bin): recognize passed-with-skips as a passing outcome (kunchenguid#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.

* fix(bin): refuse unavailable backend adapters before sourcing (kunchenguid#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning

* test: repair base-red liveness, export-DOM, and wake-queue self-tests (kunchenguid#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>

* fix: keep watcher status classification bounded to new log spans (kunchenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage

* test: close pr-check watcher test gaps (original flake already fixed by kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859.

* feat: record fleet status immediately and emit PR-ready events (kunchenguid#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing

* test: synchronize foreign queue stall checks with watcher progress (kunchenguid#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg

* test: isolate the bearings render fixture from the shared Lavish store (kunchenguid#5391)

The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.

* fix(bin): prune a torn-down task's wake rows at teardown (kunchenguid#5390)

* fix(bin): prune a torn-down task's wake rows at teardown

Prune pending durable wake rows (.wake-queue) for a task when it is torn
down, clearing stale wakes for its target window, signal wakes for its status
or turn-ended files, and task-specific check wakes.

Fixes kunchenguid#3419.
Adjacent to kunchenguid#5252.

- bin/fm-wake-lib.sh: add fm_wake_queue_prune_task
- bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown
- tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task

* no-mistakes(document): docs: note teardown prunes a task's wake rows

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>

* test: align portable test expectations with resolved host paths and fixture readiness (kunchenguid#5392)

* Make portable tests match resolved host paths

Summary:
- Match macOS full Node command paths by basename in the Gemini behavior test.
- Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom.

Testing:
- bin/fm-lint.sh
- bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh

Related:
- None

* no-mistakes(review): Mirror production PATH helper rules per directory group in tests

* no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep

* no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm

---------

Co-authored-by: Sandeep Salwan <salwansa@amazon.com>

* no-mistakes(document): Correct capture inbox ownership reference

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Denis Beliaev <battler73@yandex.ru>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: NATHAN Menkin <nate@atxlakescapes.com>
Co-authored-by: blackxwhite88 <shakir.shahruddin@gmail.com>
Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
Co-authored-by: Sandeep Salwan <salwansandeep5@gmail.com>
Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
KaranSantra added a commit to KaranSantra/firstmate that referenced this pull request Sep 23, 2026
)

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path

* fix(bin): create captain-hold rows when Beads requires due (#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: disable compact adviser for spawned agents (#4877)

* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee

* fix(bin): preserve Claude lock ownership after helper recycling (#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation

* feat: park main under the away posture on Pi (#4889)

* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation

* ci: standardize workflow timeouts into three tiers (#4910)

* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy

* fix(bin): keep supervisor status closes from waking the same home (#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs

* fix(bin): stop labeling Herdr as experimental (#4972)

* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording

* fix(bin): treat a live no-mistakes run as current after rebase (#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict

* fix(bin): prevent long worker launch command truncation (#4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs #4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs #4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>

* test: authorize isolated Herdr lab validation (#4998)

* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.

* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead o…
dsantosg1103 added a commit to dsantosg1103/firstmate that referenced this pull request Sep 24, 2026
* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path

* fix(bin): create captain-hold rows when Beads requires due (#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: disable compact adviser for spawned agents (#4877)

* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee

* fix(bin): preserve Claude lock ownership after helper recycling (#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation

* feat: park main under the away posture on Pi (#4889)

* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation

* ci: standardize workflow timeouts into three tiers (#4910)

* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy

* fix(bin): keep supervisor status closes from waking the same home (#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs

* fix(bin): stop labeling Herdr as experimental (#4972)

* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording

* fix(bin): treat a live no-mistakes run as current after rebase (#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict

* fix(bin): prevent long worker launch command truncation (#4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs #4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs #4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>

* test: authorize isolated Herdr lab validation (#4998)

* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.

* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a …
mituso89 pushed a commit to mituso89/firstmate that referenced this pull request Sep 26, 2026
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
rub-a-dub-dub added a commit to rub-a-dub-dub/firstmate that referenced this pull request Sep 27, 2026
…gences (#29)

* fix(bin): treat Claude Code's default external-imports flags as never asked, not declined (#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): keep operator-address labels out of no-mistakes intent (#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes https://github.com/kunchenguid/firstmate/issues/3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments

* fix: classify OpenCode ellipsis hint as idle (#4451)

* fix(composer): recognize Grok 1.0.5's oversized titled bottom border as a proven empty composer (#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage

* fix(bin): translate Stop hook timeout signals into durable auto-arm failure (#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list

* fix(spawn): establish Claude task channel authority (#4464)

* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference

* fix(bin): refuse fm-control.sh exit when the composer holds unproven or pending text (#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now

* fix(spawn): establish crewmate identity first (#4481)

* fix(bin): reconcile redundant secondmate divergence during updates (#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md

* feat: enable gpt-5.6-luna max reasoning for crew dispatch (#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference

* feat(calm): render smooth Unicode swell with asymmetric two-color sail (#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。

* fix(bin): supersede stale scout delivery text in brief.md on promotion (#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch

* fix(bin): make captain holds work on hosts with an older JSON::PP, and stop cleanup dropping accents from a held body (#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc

* fix(bin): read codex 0.154's idle braille starfield rows as composer furniture (#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>

* fix(bin): refuse empty text steers in fm-send (#4259)

* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.

* fix(calm): paint the working ship one yellow over all-blue water (#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.

* fix(bin): stop aging a second mate's active turn from its launch (#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch

* feat(bin): add read-only PR blocker and reviewer discovery commands (#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes #3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests

* fix(bin): teach validation-round pauses in generated briefs (#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples

* docs(readme): add star history chart (#4558)

* fix(bin): refuse teardown when a task's endpoint close fails (#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree

* feat(calm): add flag-gated Claude Code Calm mode (#4565)

* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence

* fix(bin): honour a declared wait before wedge-escalating a quiet pane (#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs

* fix(bin): report verified PR state for passed runs (#4624)

* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes #4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library

* fix: restore published contribution follow-up (Fixes #4469) (#4627)

* fix: restore published contribution follow-up (Fixes #4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower

* fix(bin): make remote report transfers explicit and fail-open (#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks

* fix(calm): preserve substantive mid-turn responses (#4655)

* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`

* fix(bin): preserve PR merge polls across volume remounts (#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes #4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from #556, reused by the #932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope

* fix(bin): keep contribution records when the poll budget runs out (follow-up to #4627) (#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.

* fix(bin): clear parent pending-replies on local secondmate retirement (#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>

* fix(bin): accept Orca's composite worktree id when tearing down a task (#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the pa…
RooseveltAdvisors pushed a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 29, 2026
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
keenvc added a commit to keenvc/firstmate that referenced this pull request Sep 30, 2026
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): unify Lavish host and disconnect handling (#5060)

* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees

* feat: act on captain's away words during AFK supervision (#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation

* fix(bin): render the remote charter's steering-inbox path host-local (#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes #5012

* no-mistakes(document): document remote charter's host-local steering inbox

* feat: route Lavish feedback directly to owning workers (#5099)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle

* fix(bin): fit pull observation within the contribution poll budget (#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget

* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>

* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>

* feat(bin): append optional home-local include to briefs (#5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>

* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout

* fix(bin): require a non-draft pull request before a PR-based done report (#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes #4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata

* fix: support quota-axi schema 6 snapshots (#4904)

* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility

* test: fix Claude session-start drain live E2E (#5165)

* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references

* ci: pin the no-mistakes required check to v1.80.1 (#5195)

Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.

* fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified

* fix(bin): allow cleanup of windowless legacy task records (#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* ci: exempt kunchenguid from the no-mistakes required check (#5256)

* fix(bin): surface launches parked on an interactive prompt as not-started (#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop

* fix: record away posture immediately on /afk (#5260)

* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence

* fix(bin): recognize passed-with-override as a passing outcome (#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor

* fix: clean up workers after their pull requests land (#5317)

* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership

* fix: surface green no-mistakes PRs awaiting merge (#5327)

* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback

* fix: derive Lavish polling route from board session (#5334)

* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims

* fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes #4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root

* fix(bin): stop each keyed answer from re-waking this home (#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment

* fix: deliver failed public follow-ups with updated AXI floors (#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed

* fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers

* fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier

* test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue #3793).
The original 0.25s window after confirmation was widened to 80 polls in
#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.

* fix: stop watchers reliably during blocked polls (#5362)

* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation

* fix: submit stuck inbox doorbells instead of skipping them (#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation

* feat: add opt-in fleet activity ledger (#5375)

* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed

* fix: validate public follow-up deliverables and wake on rejection (#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation

* feat: add Devin CLI crewmate and scout adapter (#5380)

* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>

* fix(bin): recognize passed-with-skips as a passing outcome (#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.

* fix(bin): refuse unavailable backend adapters before sourcing (#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning

* test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>

* fix: keep watcher status classification bounded to new log spans (#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage

* test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6.

* feat: record fleet status immediately and emit PR-ready events (#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing

* test: synchronize foreign queue stall checks with watcher progress (#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg

* test: isolate the bearings render fixture from the shared Lavish store (#5391)

The listener resolves its server from that store before it polls. Without a session for this bo…
gk-io-dev added a commit to gk-io-dev/firstmate that referenced this pull request Sep 30, 2026
…bility fixes (#3)

* feat(bin): add opt-in typed dispatch resolution (#4692)

* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior

* fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check

* fix(bin): launch codex crewmates with codex's hook layer disabled (#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid/firstmate#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record

* fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by #4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records

* fix: select authoritative no-mistakes runs (#4476)

* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: https://github.com/kunchenguid/firstmate/issues/3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation

* fix: distinguish captain outcomes from no-op updates (#4738)

* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since #2339 (2026-08-13), #4655 changed only the Claude Code mod, and
#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.

* fix(bin): let non-owner Claude Stops exit safely (#4777)

* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit

* fix(bin): survive bash 3.2 empty-array expansion in watcher churn absorb (#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.

* Make the foreign-owner turn-end repro create a Linux-readable session lock. (#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: require complete captain-facing final responses (#4779)

* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks

* fix: preserve substantive mid-turn text in Pi Calm (#4788)

* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation

* fix: harden mail checks and rebalance full-coverage CI (#4800)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation

* fix(bin): answer Kimi 2.0.0 folder-trust dialog during spawn (#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca

* fix(bin): report a dead-agent record once instead of escalating forever (#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: #4412, #4482, #4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path

* fix(bin): create captain-hold rows when Beads requires due (#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: disable compact adviser for spawned agents (#4877)

* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee

* fix(bin): preserve Claude lock ownership after helper recycling (#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in #3902, #2314, #3398, and #4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation

* feat: park main under the away posture on Pi (#4889)

* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation

* ci: standardize workflow timeouts into three tiers (#4910)

* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy

* fix(bin): keep supervisor status closes from waking the same home (#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs

* fix(bin): stop labeling Herdr as experimental (#4972)

* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording

* fix(bin): treat a live no-mistakes run as current after rebase (#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict

* fix(bin): prevent long worker launch command truncation (#4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs #4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs #4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>

* test: authorize isolated Herdr lab validation (#4998)

* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.

* docs(vision): accept vendor-semantics and 9k AGENTS ceiling (#4873) (#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): unify Lavish host and disconnect handling (#5060)

* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees

* feat: act on captain's away words during AFK supervision (#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation

* fix(bin): render the remote charter's steering-inbox path host-local (#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes #5012

* no-mistakes(document): document remote charter's host-local steering inbox

* feat: route Lavish feedback directly to owning workers (#5099)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle

* fix(bin): fit pull observation within the contribution poll budget (#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget

* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>

* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>

* feat(bin): append optional home-local include to briefs (#5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>

* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.c…
keenvc added a commit to keenvc/firstmate that referenced this pull request Sep 30, 2026
…ession start (#12)

* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): unify Lavish host and disconnect handling (#5060)

* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees

* feat: act on captain's away words during AFK supervision (#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation

* fix(bin): render the remote charter's steering-inbox path host-local (#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes #5012

* no-mistakes(document): document remote charter's host-local steering inbox

* feat: route Lavish feedback directly to owning workers (#5099)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle

* fix(bin): fit pull observation within the contribution poll budget (#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget

* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>

* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>

* feat(bin): append optional home-local include to briefs (#5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>

* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout

* fix(bin): require a non-draft pull request before a PR-based done report (#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes #4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata

* fix: support quota-axi schema 6 snapshots (#4904)

* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility

* test: fix Claude session-start drain live E2E (#5165)

* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references

* ci: pin the no-mistakes required check to v1.80.1 (#5195)

Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.

* fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified

* fix(bin): allow cleanup of windowless legacy task records (#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* ci: exempt kunchenguid from the no-mistakes required check (#5256)

* fix(bin): surface launches parked on an interactive prompt as not-started (#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop

* fix: record away posture immediately on /afk (#5260)

* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence

* fix(bin): recognize passed-with-override as a passing outcome (#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor

* fix: clean up workers after their pull requests land (#5317)

* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership

* fix: surface green no-mistakes PRs awaiting merge (#5327)

* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback

* fix: derive Lavish polling route from board session (#5334)

* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims

* fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes #4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root

* fix(bin): stop each keyed answer from re-waking this home (#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment

* fix: deliver failed public follow-ups with updated AXI floors (#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed

* fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers

* fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier

* test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue #3793).
The original 0.25s window after confirmation was widened to 80 polls in
#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.

* fix: stop watchers reliably during blocked polls (#5362)

* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation

* fix: submit stuck inbox doorbells instead of skipping them (#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation

* feat: add opt-in fleet activity ledger (#5375)

* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed

* fix: validate public follow-up deliverables and wake on rejection (#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation

* feat: add Devin CLI crewmate and scout adapter (#5380)

* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>

* fix(bin): recognize passed-with-skips as a passing outcome (#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.

* fix(bin): refuse unavailable backend adapters before sourcing (#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning

* test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>

* fix: keep watcher status classification bounded to new log spans (#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage

* test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6.

* feat: record fleet status immediately and emit PR-ready events (#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing

* test: synchronize foreign queue stall checks with watcher progress (#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg

* test: isolate the bearings render fixture from the shared Lavish store (#5391)

The listener resolves its server from that store before it polls…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants