Skip to content

feat(bin): merge 47 upstream firstmate commits into the fork - #12

Merged
jbalke merged 55 commits into
mainfrom
fm/fm-reconcile-upstream-2026-09-18
Sep 20, 2026
Merged

jbalke merged 55 commits into
mainfrom
fm/fm-reconcile-upstream-2026-09-18

Conversation

@jbalke

@jbalke jbalke commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain asked what was new on firstmate upstream. He was told plainly
that there are 47 new commits on the upstream remote kunchenguid/firstmate
that this fork does not have, that the fork is also 60 commits ahead, and that
reaching upstream is therefore a merge and not a fast-forward the ordinary
updater can do. He chose to merge upstream now. He still merges the PR himself.

Measured at intake, 2026-09-18:

  • 47 upstream commits this fork lacks, dated 2026-09-14 to 2026-09-17.
  • 60 local commits upstream lacks.
  • Three-dot diff HEAD...upstream/main: 164 files changed, 17,429
    insertions, 2,701 deletions.
  • Last merge base: b182d0f9 fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361), 2026-09-13.

The ask is to bring those 47 commits in while preserving this fork's
deliberate local behaviour. The fork's 60 commits are divergence on purpose,
not drift.

Several of the upstream commits fix faults that matter to this home, which is
part of why the merge was chosen:

  • 9bc051ff reports a dead-agent record once instead of escalating forever.
  • 334fa122 reads the latest status event so buried declarations and open
    decisions are not lost.
  • a2216406 tells real captain outcomes apart from no-op updates.
  • e213343c survives bash 3.2 empty-array expansion in the watcher churn
    absorb. This home runs macOS, where bash 3.2 is the system shell.
  • baede47d preserves PR merge polls across volume remounts, and 7111081c
    reports verified PR state for passed runs.

What Changed

  • Merges 47 upstream kunchenguid/firstmate commits (merge base b182d0f9) into the fork as a conflict-resolved merge that keeps the fork's 60 local commits of deliberate divergence. This adds new bin/fm-contributions.sh + bin/fm-contributions.jq, bin/fm-dispatch-resolve.sh, bin/fm-pr-state.sh, bin/fm-pr-reviewers.sh, bin/fm-env-lib.sh and the flag-gated Claude Code Calm mod under .claude/mods/firstmate-calm/, plus upstream fixes to the watcher (bash 3.2 empty-array expansion in churn absorb), dead-agent reporting, latest-status-event reads for buried declarations and open decisions, captain outcome vs no-op discrimination, and PR merge poll durability across volume remounts.
  • Post-merge corrections to the contributions path: fm-contributions.jq gains one shared nameable_task rule so unnameable owners are dropped by known() before any forge read instead of being rejected per-caller mid-poll, fm-contributions.sh strips trailing slashes from DATA before its literal prefix comparison, and an unusable backlog repo: hint now falls back to no-project rather than failing the poll for every other owner.
  • Corrects the Codex idle-starfield classification rule in docs/verification/runtime-backends.md to cover only its eight single-dot Braille particles (other Braille, including progress spinners, stays typed content), and points the live e2e guard at _fm_composer_row_is_codex_particles on a trimmed row.

Risk Assessment

�� Medium: A 47-commit, 164-file upstream merge that activates substantial new behaviour in the watcher, spawn, composer, and contribution paths of a deliberately divergent fork is inherently more than cosmetic, but every conflict resolution is coherent, all six intent-named upstream fixes are present and intact, fork-local divergences survive, no orphaned symbols or markers remain, and the only residual source concern is a comment-level cross-reference.

Testing

I stood up isolated firstmate homes and drove the real bin/fm-*.sh CLIs rather than only running assertions. The fork's post-merge contribution work is shown working in a hand-driven transcript: a captain backlog carrying one good row plus rows whose owner id the durable layer cannot name spends zero forge reads on the URL only the unnameable owner holds, writes no durable directory for it, arms no check that could never be cleared, and leaves Bearings reporting complete/proven_clear true - the permanent-incompleteness failure the earlier review rounds were about. An adversarial differential over fifteen ids confirms the jq rule the coverage projection now shares with poll accepts an id exactly when both shell predicates do, including the tasks and -dash cases that previously disagreed. The bash 3.2 watcher fix was driven under this host's real /bin/bash 3.2.57 with the CI lane's PATH, with a negative control that first reproduces the missing_keys[@]: unbound variable crash, and the stock-bash snapshot consumers pass there too. The other upstream fixes the intent names - dead-agent reported once, latest status event, verified PR state for passed runs, merge polls across remounts - were each driven through the suites that own them, and every file the merge had to reconcile from both sides was exercised through its own suite, all green. Two things I could not drive: the captain-outcome change is AGENTS.md and pi.md prose (its author states no executable contract evaluates MAIN's reply, so there is nothing to execute - I confirmed the text landed intact), and the live codex idle-composer guard cannot reach an idle screen because codex asks for folder trust on the ephemeral gate worktree path; I did verify the d82afe3 part of it live, with a negative control showing the pre-fix line calling a function upstream's rename had removed. No screenshots apply here: firstmate is a terminal CLI with no rendered UI surface, so the reviewer-visible evidence is CLI transcripts and forge-call logs.

  • Live validation: â�� go - 14 of 16 scenarios driven live against the product
Scenario Result Live Evidence
A backlog row whose owner id the durable layer cannot name spends no forge read, writes no record, and never reaches the wake queue � pass live bash tests/fm-contributions.test.sh (test_unusable_task_id_keeps_poll_alive, test_unusable_task_id_raises_no_captain_wake) plus the hand-driven transcript in contributions-cli-transcript.log showing�
Bearings still reports contribution coverage complete and proven_clear while an unnameable owner sits in the backlog, in the home and through a child home summary � pass live fm-bearings-snapshot.sh --json | jq .contributions in contributions-cli-transcript.log returning complete=true, proven_clear=true; test_home_summary_coverage covers the child-home demotion path
Adversarial: poll and the coverage projection never disagree about which owner is nameable, across ids each shell predicate treats differently � pass live nameable-owner-rule-differential.log - 15 ids including tasks, -dash, .hidden, a b, a/b, ../escape, non-ASCII and a shell-injection string, all agreeing with the intersection of fm_pr_task�
fm-contributions.sh arm --if-owned arms no check when the backlog's only contribution row has an unnameable owner � pass live contributions-cli-transcript.log section 5: exit 0, state/*.check.sh absent; same assertion in test_unusable_task_id_keeps_poll_alive
Adversarial: a backlog row whose repo: hint is prose rather than a registry name groups under no-project instead of ending the poll for every other owner � pass live contributions-cli-transcript.log shows repo: kunchenguid/firstmate writing to data/tasks/_none/upstream/contributions.json while the good owner's record is still refreshed; `bash tests/fm-contributi�
Adversarial: a data root spelled with a trailing slash still records observations, and a symlinked data root is still refused through the same spelling � pass live contributions-cli-transcript.log section 6 - record written with the trailing slash, then fm-contributions: data directory unavailable once the root is a symlink; test_data_root_trailing_slash
On this macOS home's stock bash 3.2, a second churning turn-end inside an open deferral window is absorbed instead of killing the watcher cycle � pass live bash32-churn-regression.log - negative control with the e213343 guard reverted dies with missing_keys[@]: unbound variable under real /bin/bash 3.2.57, merged code passes
The merged snapshot consumers still run under this home's stock /bin/bash 3.2.57 � pass live macos-stock-bash32-snapshot-lane.log - fm-fleet-snapshot-view and fm-bearings-snapshot, 80 cases, both exit 0 with bash resolving to /bin/bash 3.2.57
A record whose endpoint is dead reports itself once, re-arms if the endpoint returns, and live or unproven endpoints still wedge-escalate � pass live upstream-dead-agent-report-once.log - 5/5 targeted fm-watch-triage cases from 9bc051f
A buried status declaration or open decision is read from the latest status event rather than lost behind newer history � pass live upstream-latest-status-event.log - 123 cases across fm-classify-decision-key, fm-wake-drain-open-decisions-cursor, fm-inactive-reconcile, fm-send-resolve-key, fm-captain-hold-lifecycle
A run that passed reports verified PR/MR state (open, merged, GitLab, unreadable identity) instead of claiming merged � pass live upstream-crew-state-pr-and-status.log - 141 cases including the terminal passed run ... group from 7111081
A PR merge poll survives a volume remount and a validated merged poll notifies once then retires � pass live upstream-pr-merge-poll-remount.log - 37 cases in fm-pr-check-security from baede47
The fork's deliberate local behaviour still holds in every file the merge had to reconcile from both sides � pass live fork-local-behaviour-merge-conflicts.log / -2.log / -3.log - 605 cases across composer, brief, control relaunch, fleet snapshot, gotmp, task delivery, spawn dispatch profile, zellij backend, teardown,�
The live codex starfield guard calls a classifier that still exists after upstream's rename � pass live live-codex-idle-composer-guard.log - the pre-d82afe3 line produces _fm_composer_row_is_braille_furniture: command not found against the installed codex 0.155.0; the merged line runs clean and report�
An idle codex composer classifies as empty through both the tmux and cursorless styled read profiles �� untested no codex records folder trust only in the user-level ~/.codex/config.toml, which trusts ~/dev/john/firstmate but not this ephemeral gate worktree path; `codex -c 'projects."<path>".trust_l�
MAIN answers a finished captain-requested deliverable with an outcome naming what finished, not Captain, shipshape. �� untested no No runtime surface. The upstream commit message states it outright: "No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner�
Evidence: Evidence index

Source: Evidence index

# Live validation evidence - fm/fm-reconcile-upstream-2026-09-18

Change under test: `d5967e12..1e2e5e5` - merge of 47 upstream `kunchenguid/firstmate`
commits into this fork (`8fa641e`), plus five post-merge review-round fixes to the
contribution path and the live codex composer guard.

Everything below was driven against the real firstmate CLIs (`bin/fm-*.sh`) on this
macOS host, in isolated temporary homes.

| File | What it shows |
|---|---|
| `contributions-cli-transcript.log` | Hand-driven CLI transcript: a captain backlog with one good row and three unnameable/prose-hinted rows, through `fm-contributions.sh poll`, `arm --if-owned` and `fm-bearings-snapshot.sh --json`. Shows **0** forge reads spent on the URL owned only by an unnameable id, no durable directory created, coverage `complete/proven_clear = true`, the prose repo hint falling back to the no-project container, and a trailing-slash data root writing records while a symlinked root is still refused. |
| `drive-contributions.sh`, `gh-fixture.sh` | The driver and offline forge stub used to produce the transcript above. |
| `bash32-churn-regression.log` | The `e213343c` bash 3.2 fix, driven under this host's real `/bin/bash 3.2.57` with the CI lane's PATH. Negative control first (guard reverted -> `missing_keys[@]: unbound variable`, watcher cycle dies), then the merged code passing. |
| `macos-stock-bash32-snapshot-lane.log` | The macOS stock-bash CI lane's snapshot consumers (`fm-fleet-snapshot-view`, `fm-bearings-snapshot`) run under real `/bin/bash 3.2.57` with `tasks-axi 0.2.5`. |
| `upstream-dead-agent-report-once.log` | `9bc051ff` - a dead/missing endpoint reports once, re-arms when it comes back, and live/unproven endpoints still escalate. |
| `upstream-latest-status-event.log` | `334fa122` - latest-status-event folding: decision keys, wake-drain cursor, inactive reconcile, send key resolution, captain-hold lifecycle. |
| `upstream-crew-state-pr-and-status.log` | `7111081c` - `terminal passed` runs report verified PR/MR state (open, merged, GitLab, unreadable identity) instead of claiming merged. |
| `upstream-pr-merge-poll-remount.log` | `baede47d` - merge polls survive volume remounts; validated merged polls notify once and retire. |
| `upstream-watch-triage-partial.log` | Watcher triage suite, stopped at 85/244 cases with 0 failures (each case runs a real watcher at live poll intervals). The cases owned by `9bc051ff` and `e213343c` were run individually to completion - see the two rows above/below. |
| `fork-local-behaviour-merge-conflicts.log`, `-2.log`, `-3.log` | The suites owning every file the merge had to reconcile from both sides (composer, brief, control relaunch, fleet snapshot, gotmp, task delivery, spawn dispatch profile, zellij backend, teardown, teardown endpoint safety, backlog handoff, backlog atomicity, public followup/promote, herdr backend smoke). |
| `nameable-owner-rule-differential.log` | Adversarial differential proving the jq `nameable_task` rule the coverage projection now shares with poll accepts an id exactly when both shell predicates (`fm_pr_task_id_valid`, `fm_task_data_valid_id`) do - across `tasks`, `-dash`, `.hidden`, `a b`, `a/b`, `../escape`, non-ASCII and a shell-injection string. |
| `live-codex-idle-composer-guard.log` | `d82afe3` starfield guard. The negative control shows the pre-fix line calling `_fm_composer_row_is_braille_furniture`, removed by upstream's rename -> `command not found`. The guard as a whole cannot reach its verdict on this host: codex parks on a folder-trust prompt for the ephemeral gate worktree path (see below). |

## Not drivable here

`tests/fm-composer-codex-idle-live-e2e.test.sh` launches the installed `codex`
in an isolated tmux server and requires an idle composer. On this host codex
0.155.0 shows an update-available modal (dismissed by the guard's Escape) and
then a **folder-trust prompt** for
`~/.no-mistakes/worktrees/6b5c7015552c/01M2ZFHFZER4D9XE67D807JABW`,
which `~/.codex/config.toml` has never trusted (it trusts
`~/dev/john/firstmate`). Trust is recorded only in that
user-level file, and `codex -c 'projects."<path>".trust_level="trusted"'` does
not override it. Writing it is a global user-config change this run's workspace
boundary forbids.

To enable: run `codex` once in this worktree (or in the checkout the guard will
run from) and answer *"1. Yes, continue"*, then re-run the guard. In CI the
guard's `fm_live_gate default-on FM_COMPOSER_CODEX_IDLE_LIVE codex tmux` skips
because codex is not installed, so this is a local-host condition only.
Evidence: CLI transcript: unnameable owners cost nothing

Source: CLI transcript: unnameable owners cost nothing

$ cat forge/calls.log api repos/o/r/pulls/8 api repos/o/r/issues/8/comments?per_page=100 --paginate --slurp api repos/o/r/pulls/8/reviews?per_page=100 --paginate --slurp ... reads against issues/9 (owned only by the unnameable "-dash"): 0 $ fm-bearings-snapshot.sh --json | jq .contributions { "known": 1, "checked": 1, "complete": true, "proven_clear": true, "counts": { "captain": 0, "fleet": 0, "maintainer": 1, "nobody": 0 } } $ fm-contributions.sh arm --if-owned # backlog has ONLY the unnameable row no check armed - nothing can ever clear it, so nothing was armed

================ backlog the captain left ================
# Backlog

## Queued
- [ ] delivery - Contribution delivery https://github.com/o/r/pull/8 (repo: sample) (kind: ship)
- [ ] tasks - Filed https://github.com/o/r/pull/8 (repo: sample) (kind: ship)
- [ ] -dash - Filed https://github.com/o/r/issues/9 (repo: sample) (kind: ship)
- [ ] upstream - Filed https://github.com/o/r/pull/8 (repo: kunchenguid/firstmate) (kind: ship)

================ 1. poll the forge ================

$ fm-contributions.sh poll
(exit 0)

================ 2. which forge reads were actually spent ================

$ cat forge/calls.log
api repos/o/r/pulls/8
api repos/o/r/issues/8/comments?per_page=100 --paginate --slurp
api repos/o/r/pulls/8/reviews?per_page=100 --paginate --slurp
api repos/o/r/pulls/8/comments?per_page=100 --paginate --slurp
api repos/o/r/commits/aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa/check-runs?filter=all&per_page=100 --paginate --slurp
api repos/o/r/commits/aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa/statuses?per_page=100 --paginate --slurp
api repos/o/r
pr view https://github.com/o/r/pull/8 --json headRefOid,reviewDecision

reads against issues/9 (owned only by the unnameable "-dash"): 0

================ 3. durable records written ================

$ find data -name contributions.json
data/delivery/contributions.json
data/tasks/_none/upstream/contributions.json

$ jq .records[0].checked_at data/delivery/contributions.json
2026-09-16T08:00:00Z

$ jq .task data/tasks/_no-project/upstream/contributions.json  # prose repo hint fell back to no-project
_none

/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T//fm-drive.Z1pEQm/data/tasks/_none:
upstream

/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T//fm-drive.Z1pEQm/data/tasks/_none/upstream:
contributions.json

================ 4. what Bearings reports to the captain ================

$ fm-bearings-snapshot.sh --json | jq .contributions
{
  "known": 1,
  "checked": 1,
  "complete": true,
  "proven_clear": true,
  "counts": {
    "captain": 0,
    "fleet": 0,
    "maintainer": 1,
    "nobody": 0
  }
}

================ 5. arm --if-owned on an unnameable-only home ================

$ fm-contributions.sh arm --if-owned   # backlog has ONLY the unnameable row
(exit 0)

$ ls state/*.check.sh
ls: /var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T//fm-drive-unowned.uew6J6/state/*.check.sh: No such file or directory
no check armed - nothing can ever clear it, so nothing was armed

================ 6. data root spelled with a trailing slash ================

$ FM_DATA_OVERRIDE="$HOME/data/" fm-contributions.sh poll
(exit 0)

$ jq .records[0].checked_at data/delivery/contributions.json
2026-09-16T08:00:00Z

$ same trailing slash, but the data root is now a symlink
fm-contributions: data directory unavailable
refused as expected: YES
Evidence: bash 3.2 churn regression with negative control

Source: bash 3.2 churn regression with negative control

host stock shell: GNU bash, version 3.2.57(1)-release (arm64-apple-darwin25) === NEGATIVE CONTROL: e213343c guard reverted to bare ${missing_keys[@]} === bin/fm-watch.sh: line 560: missing_keys[@]: unbound variable not ok - a churning turn-end inside an open deferral window was not absorbed exit=1 === WITH e213343c AS MERGED (HEAD 1e2e5e5) === ok - a churning turn-end inside an already-open deferral window is absorbed without re-marking exit=0

host stock shell: GNU bash, version 3.2.57(1)-release (arm64-apple-darwin25)
PATH pinned as the macos-stock-bash CI lane does, so fm-watch.sh's '#!/usr/bin/env bash' resolves to /bin/bash 3.2:
  PATH=/bin:/usr/bin:/usr/sbin:/sbin:/opt/homebrew/bin
/bin/bash
GNU bash, version 3.2.57(1)-release (arm64-apple-darwin25)

=== NEGATIVE CONTROL: e213343c guard reverted to bare ${missing_keys[@]} ===
~/.no-mistakes/worktrees/6b5c7015552c/01M2ZFHFZER4D9XE67D807JABW/bin/fm-watch.sh: line 560: missing_keys[@]: unbound variable
not ok - a churning turn-end inside an open deferral window was not absorbed: 
exit=1

=== WITH e213343c AS MERGED (HEAD 1e2e5e5) ===
ok - a churning turn-end inside an already-open deferral window is absorbed without re-marking
exit=0
Evidence: Shared nameable-owner rule differential

Source: Shared nameable-owner rule differential

ID fm_pr_task_id_valid fm_task_data_valid_id jq:nameable VERDICT 'delivery' yes yes yes agree 'tasks' yes no no agree '-dash' yes no no agree '.hidden' no yes no agree 'a b' no yes no agree '../escape' no no no agree '$(touch /tmp/pwned)' no no no agree ok - poll and the coverage projection share one nameable-owner rule across the whole probe set

Adversarial differential of the shared nameable-owner rule (1e2e5e5, bin/fm-contributions.jq):

ID               fm_pr_task_id_valid fm_task_data_valid_id jq:nameable VERDICT
'delivery'       yes            yes                yes        agree
'ok_id'          yes            yes                yes        agree
'ok.id'          yes            yes                yes        agree
'ok-id'          yes            yes                yes        agree
'_under'         yes            yes                yes        agree
'UPPER9'         yes            yes                yes        agree
'tasks'          yes            no                 no         agree
'-dash'          yes            no                 no         agree
'.hidden'        no             yes                no         agree
'a b'            no             yes                no         agree
'a/b'            no             no                 no         agree
'../escape'      no             no                 no         agree
'tasks2'         yes            yes                yes        agree
'üñî'         no             yes                no         agree
'$(touch /tmp/pwned)' no             no                 no         agree

ok - poll and the coverage projection share one nameable-owner rule across the whole probe set
exit=0
Evidence: Driver script used for the CLI transcript

Source: Driver script used for the CLI transcript

#!/usr/bin/env bash
# Hand-driven transcript of the fork's post-merge contribution changes.
# Stands up a real firstmate home and runs the real CLIs against it.
set -u
ROOT=$1
HOME_DIR=$(mktemp -d "${TMPDIR:-/tmp}/fm-drive.XXXXXX")
NOW=2026-09-16T08:00:00Z
HEAD_A=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa

say() { printf '\n\033[1m$ %s\033[0m\n' "$*"; }
run() { say "$*"; "$@"; printf '(exit %s)\n' "$?"; }

mkdir -p "$HOME_DIR"/{data,state,config,projects,fakebin,forge,root/bin,wt}
printf '#!/bin/sh\nexit 1\n' > "$HOME_DIR/fakebin/tmux"
printf '#!/bin/sh\nexit 0\n' > "$HOME_DIR/fakebin/no-mistakes"
printf '#!/bin/sh\nexit 0\n' > "$HOME_DIR/root/bin/fm-guard.sh"
chmod +x "$HOME_DIR/fakebin/"* "$HOME_DIR/root/bin/fm-guard.sh"
printf 'worktree=%s/wt\nkind=ship\n' "$HOME_DIR" > "$HOME_DIR/state/delivery.meta"
chmod 600 "$HOME_DIR/state/delivery.meta"
printf '%s\n' "$HEAD_A" > "$HOME_DIR/forge/head"
for f in comments reviews inline labels events; do printf '[]\n' > "$HOME_DIR/forge/$f.json"; done
cp "$(dirname "$0")/gh-fixture.sh" "$HOME_DIR/fakebin/gh"; chmod +x "$HOME_DIR/fakebin/gh"

# The captain's backlog, exactly as an operator would leave it: one good row,
# and three rows whose owner id the durable layer cannot name.
cat > "$HOME_DIR/data/backlog.md" <<'MD'
# Backlog

## Queued
- [ ] delivery - Contribution delivery https://github.com/o/r/pull/8 (repo: sample) (kind: ship)
- [ ] tasks - Filed https://github.com/o/r/pull/8 (repo: sample) (kind: ship)
- [ ] -dash - Filed https://github.com/o/r/issues/9 (repo: sample) (kind: ship)
- [ ] upstream - Filed https://github.com/o/r/pull/8 (repo: kunchenguid/firstmate) (kind: ship)
MD

mkdir -p "$HOME_DIR/data/delivery"
jq -n --arg head "$HEAD_A" --arg at 2026-09-15T08:00:00Z '
  {schema:"fm-contributions.v1",task:"delivery",records:[{
    url:"https://github.com/o/r/pull/8",kind:"pr",checked_at:$at,error:null,pending:[],seen:[],verdict:null,
    observation:{head:$head,state:"open",draft:false,mergeable:"mergeable",review_decision:"APPROVED",
      can_merge:false,checks:[],reviews:[],events:[]}}]}' > "$HOME_DIR/data/delivery/contributions.json"

with_home() {
  PATH="$HOME_DIR/fakebin:$PATH" FORGE="$HOME_DIR/forge" \
    FM_HOME="$HOME_DIR" FM_ROOT_OVERRIDE="$HOME_DIR/root" FM_STATE_OVERRIDE="$HOME_DIR/state" \
    FM_DATA_OVERRIDE="${FM_DATA:-$HOME_DIR/data}" FM_CONFIG_OVERRIDE="$HOME_DIR/config" \
    FM_CONTRIBUTIONS_NOW="$NOW" "$@"
}

echo "================ backlog the captain left ================"
cat "$HOME_DIR/data/backlog.md"

echo
echo "================ 1. poll the forge ================"
say "fm-contributions.sh poll"
with_home "$ROOT/bin/fm-contributions.sh" poll; printf '(exit %s)\n' "$?"

echo
echo "================ 2. which forge reads were actually spent ================"
say "cat forge/calls.log"
cat "$HOME_DIR/forge/calls.log"
printf '\nreads against issues/9 (owned only by the unnameable "-dash"): %s\n' \
  "$(grep -c 'issues/9' "$HOME_DIR/forge/calls.log")"

echo
echo "================ 3. durable records written ================"
say "find data -name contributions.json"
( cd "$HOME_DIR" && find data -name contributions.json | sort )
say "jq .records[0].checked_at data/delivery/contributions.json"
jq -r '.records[0].checked_at' "$HOME_DIR/data/delivery/contributions.json"
say "jq .task data/tasks/_no-project/upstream/contributions.json  # prose repo hint fell back to no-project"
jq -r '.task' "$HOME_DIR/data/tasks/_no-project/upstream/contributions.json" 2>/dev/null \
  || ls -R "$HOME_DIR/data/tasks" 2>/dev/null

echo
echo "================ 4. what Bearings reports to the captain ================"
say "fm-bearings-snapshot.sh --json | jq .contributions"
PATH="$HOME_DIR/fakebin:$PATH" FM_HOME="$HOME_DIR" FM_ROOT_OVERRIDE="$ROOT" \
  FM_STATE_OVERRIDE="$HOME_DIR/state" FM_DATA_OVERRIDE="$HOME_DIR/data" \
  FM_CONFIG_OVERRIDE="$HOME_DIR/config" FM_BEARINGS_NOW="$NOW" \
  "$ROOT/bin/fm-bearings-snapshot.sh" --json \
  | jq '.contributions | {known,checked,complete,proven_clear,counts}'

echo
echo "================ 5. arm --if-owned on an unnameable-only home ================"
UNOWNED=$(mktemp -d "${TMPDIR:-/tmp}/fm-drive-unowned.XXXXXX")
mkdir -p "$UNOWNED"/{data,state,config,fakebin,root/bin}
printf '#!/bin/sh\nexit 0\n' > "$UNOWNED/root/bin/fm-guard.sh"; chmod +x "$UNOWNED/root/bin/fm-guard.sh"
printf '# Backlog\n\n## Queued\n- [ ] tasks - Filed https://github.com/o/r/pull/8 (repo: sample) (kind: ship)\n' \
  > "$UNOWNED/data/backlog.md"
say "fm-contributions.sh arm --if-owned   # backlog has ONLY the unnameable row"
FM_HOME="$UNOWNED" FM_ROOT_OVERRIDE="$UNOWNED/root" FM_STATE_OVERRIDE="$UNOWNED/state" \
  FM_DATA_OVERRIDE="$UNOWNED/data" FM_CONFIG_OVERRIDE="$UNOWNED/config" \
  FM_CONTRIBUTIONS_NOW="$NOW" "$ROOT/bin/fm-contributions.sh" arm --if-owned
printf '(exit %s)\n' "$?"
say "ls state/*.check.sh"
ls "$UNOWNED/state/"*.check.sh 2>&1 || echo "no check armed - nothing can ever clear it, so nothing was armed"

echo
echo "================ 6. data root spelled with a trailing slash ================"
say 'FM_DATA_OVERRIDE="$HOME/data/" fm-contributions.sh poll'
jq '.records[0].checked_at="2026-09-15T08:00:00Z"' "$HOME_DIR/data/delivery/contributions.json" > "$HOME_DIR/u.json"
mv "$HOME_DIR/u.json" "$HOME_DIR/data/delivery/contributions.json"
FM_DATA="$HOME_DIR/data/" with_home "$ROOT/bin/fm-contributions.sh" poll >/dev/null; printf '(exit %s)\n' "$?"
say "jq .records[0].checked_at data/delivery/contributions.json"
jq -r '.records[0].checked_at' "$HOME_DIR/data/delivery/contributions.json"

say 'same trailing slash, but the data root is now a symlink'
mv "$HOME_DIR/data" "$HOME_DIR/data-target"; ln -s "$HOME_DIR/data-target" "$HOME_DIR/data"
FM_DATA="$HOME_DIR/data/" with_home "$ROOT/bin/fm-contributions.sh" poll 2>&1 | tail -3
printf 'refused as expected: %s\n' "$( FM_DATA="$HOME_DIR/data/" with_home "$ROOT/bin/fm-contributions.sh" poll >/dev/null 2>&1 && echo NO || echo YES )"

rm -rf "$HOME_DIR" "$UNOWNED"
Evidence: macOS stock bash 3.2 snapshot lane

Source: macOS stock bash 3.2 snapshot lane

shell resolved by '#!/usr/bin/env bash': /bin/bash -> GNU bash, version 3.2.57(1)-release (arm64-apple-darwin25)
tasks-axi: 0.2.5
--- /bin/bash tests/fm-fleet-snapshot-view.test.sh ---
ok - empty fleet snapshot and view use explicit absence markers
ok - fixture snapshot covers task rows, backlog rows, pointers, and stable ordering
ok - home-summary excludes kind=secondmate from unowned_current and terminal_in_flight
ok - undated captain holds age after a configurable threshold, decided only from structured fields
ok - captain-hold buckets are total, mutually exclusive, and never decided by prose
ok - main_inventory discloses orphan/unstructured and clears when inventory is consistent
ok - backlog normalization preserves strict roles and resolves every blocker compatibly
ok - snapshot event hints follow reconciled current state
ok - durable fold keeps an open decision past a later unrelated event
ok - a live secondmate endpoint preserves unrelated open decisions
ok - durable captain-held transfer closes the duplicate live status decision
ok - durable fold clears a decision only on a keyed resolution
ok - a completed scout's stale decision surfaces as a report pointer, not pending
ok - a scout still parked at a decision stays pending (terminal clear does not over-fire)
ok - snapshot includes durable scout reports after teardown
ok - snapshot parses tasks-axi rows and respects operational overrides
ok - fleet view renders the snapshot without secondmate peek guidance
ok - fleet view renders secondmate agent liveness
ok - a malformed home summary degrades only that home, not the whole snapshot
ok - a malformed remote home summary degrades only that home, not the whole snapshot
ok - a summary read that fails after validation degrades only that home
exit=0
--- /bin/bash tests/fm-bearings-snapshot.test.sh ---
ok - task teardown during metadata capture is omitted without aborting the snapshot
ok - current state and decision hints share one captured status observation
ok - reused live state is discarded when task generation changes
ok - large local snapshot overlaps local reads with byte-identical serial and concurrent projections
ok - remote ledgers collect concurrently under one budget, reuse aged cache, and cancel wedged collectors
ok - a missing remote ledger stays explicitly unreadable without remote summary computation
ok - Domain Alpha structured state overrides a stale parent Phase 7 event
ok - GNU stat file reads select -c without BSD filesystem-report pollution
ok - parent activity evidence is bounded and disclosed
ok - Bearings excludes a status-only child decision
ok - a structured child captain hold reaches Captain's Call
ok - missing, invalid, unreadable, malformed, and unavailable-child homes stay explicit unknowns
ok - oversized ledgers stay strict unknown
ok - secondmate and per-home child counts are bounded, disclosed, and explicitly expandable
ok - parent decisions remain untrusted contradiction evidence
ok - parent evidence reconciliation distinguishes matching holds, blocks, and decisions
ok - nonprogressing child states are explicit and inconsistent terminal rows invalidate
ok - registry unavailability and bounded truncation remain explicit
ok - repeated snapshots keep the same current landed baseline and ignore prior reports
ok - default output is bounded, local-only, and marks omitted surfaces
ok - TOON and JSON are parity representations of the same model
ok - landed includes secondmate-managed merges alongside main-home merges
ok - landed accepts only kind-owned delivery artifacts while answered questions stay out
ok - kind fallback matches tasks-axi word boundaries and preserves scout deliveries
ok - kindless v1 summaries retain report artifacts from fresh and cached ledgers
ok - default landed selection balances one dominant home with sparse homes
ok - landed selection refills capacity after sparse homes exhaust
ok - landed selection uses deterministic home order when homes exceed the cap
ok - landed selection preserves deterministic home and internal tie ordering
ok - landed selection handles no landed items
ok - --all-landed keeps the complete global landed output
ok - landed stays bounded with per-home + overall caps and omitted[] disclosure
ok - Bearings keeps a live blocker in structured live state and never converts it to Charted Next queue work
ok - action-free items (working/done/queued/landed) do not leak into Captain's Call
ok - main orphan in-flight stays out of Underway and is disclosed in omitted/gates
ok - main unstructured current is disclosed while structured siblings still project
ok - counterfactual meta clears main inventory warning and projects the live task
ok - working captain holds retain main and secondmate bucket surfaces
ok - active children reach Underway independently of a home captain hold
ok - blank legacy summary names use their durable identifier
ok - newest filed gates are selected before snapshot bounds
ok - Underway rows carry the durable task name and gates carry their filed date
ok - mixed secondmate roles, partial state, and captain readiness project independently
ok - main and secondmate captain actionability use the same blocker readiness
ok - a completed scout with decision-like report prose is a pointer, not pending
ok - an authoritative captain hold surfaces end-to-end
ok - current report pointers surface
ok - queued body prose never hides an item from the board
ok - --include-prs is the only path that fetches, and it enriches correctly
ok - a partial GitHub failure degrades gracefully
ok - Perl fallback bounds stalled GitHub calls without coreutils timeout
ok - all fleet-sized sections are capped with counted opt-in expansion
ok - captain-held tasks of any kind reach Captain's Call, deferral is honored, and landed excludes answered calls
ok - structured hold buckets decide Captain's Call, Charted Next, and the reveal
ok - a blocked deferred hold stays gated when its date arrives and is revealed by --all-decisions
ok - revealed deferred holds display their deferral reason while live calls stay unannotated
ok - live PR enrichment caps repositories with counted expansion
ok - per-repository open-PR caps are disclosed with an expansion knob
ok - projection and TOON rendering failures exit nonzero with diagnostics
exit=0
Evidence: 9bc051f dead-agent reported once

Source: 9bc051ff dead-agent reported once

--- FM_TEST_ONLY=test_gone_endpoint_reports_once_instead_of_escalating_forever tests/fm-watch-triage.test.sh ---
ok - a record whose endpoint is dead or missing reports itself once and is never re-escalated
exit=0
--- FM_TEST_ONLY=test_gone_report_rearms_when_the_endpoint_comes_back tests/fm-watch-triage.test.sh ---
ok - the once-only gone report re-arms when the endpoint comes back, and reports a later death again
exit=0
--- FM_TEST_ONLY=test_identical_dead_display_of_a_successor_still_reports tests/fm-watch-triage.test.sh ---
ok - a successor's byte-identical dead display reports in full, and the same incarnation still absorbs
exit=0
--- FM_TEST_ONLY=test_second_death_after_a_same_window_relaunch_reports_in_full tests/fm-watch-triage.test.sh ---
ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent
exit=0
--- FM_TEST_ONLY=test_live_and_unproven_endpoints_still_wedge_escalate tests/fm-watch-triage.test.sh ---
ok - a live wedged agent, an unattributable one, and an unreadable endpoint escalate unchanged
exit=0
Evidence: 7111081 verified PR state for passed runs

Source: 7111081c verified PR state for passed runs

ok - captured AXI replacement status replays through crew-state
ok - captured AXI parked status replays through crew-state
ok - captured AXI failed status replays through crew-state
ok - captured capped inventory replays selection, ambiguity, and unavailable lookup
ok - captured status formats reject a synthetic authority transition
ok - captured completed status yields to synthetic subsequent development
ok - active run-step is authoritative
ok - stale needs-decision over active run is superseded
ok - stale blocked over active run is superseded
ok - daemon/timeout blocked claim over a live fixing run reads as run alive
ok - socket refusal or missing socket over a stale fixing run reports blocked
ok - socket refusal over a terminal attributed run reports blocked
ok - socket-down evidence outranks a live run only while it is the log's latest event
ok - broken-pipe blocker over a live run keeps the plain superseded reading
ok - genuine daemon-down blocked line still reports blocked
ok - a busy secondmate keeps its open blocker until that exact key closes
ok - the most recently opened decision supplies the reported state and detail
ok - ship and scout terminal declarations supersede stale decisions
ok - latest status retains legacy completion events and shared captain matching
ok - latest status subprocess work stays bounded and still reads past a long prose tail
ok - genuine parked run is not flagged superseded
ok - scalar gate parked run is not flagged superseded
ok - gate block parked run is not flagged superseded
ok - ci-ready status log beats monitoring run
ok - ci-monitoring run with checks already green surfaces done
ok - top-level ci status uses ci log green marker
ok - terminal no-checks ci-monitor marker surfaces done
ok - base-advance rearm after green stays working
ok - pending no-checks ci-monitor marker stays working
ok - ci-monitoring run with checks not yet green stays working
ok - a fresh issue after an earlier green reading is not masked
ok - stale checks-green status log does not mask CI relapse
ok - ci fixing is not overridden by an earlier green marker
ok - top-level fixing is not overridden by a stale ci running row
ok - top-level fixing is not overridden by a stale done log
ok - terminal passed run is authoritative
ok - terminal passed run uses matching retirement receipt without forge
ok - terminal passed no-forge mode preserves local receipt evidence
ok - terminal passed run with open PR does not claim merged
ok - terminal passed run PR overrides stale task metadata
ok - terminal passed run without readable PR identity reports unknown
ok - terminal passed run reads open GitLab MR state
ok - terminal passed run reads merged GitLab MR state
ok - terminal passed run handles failed GitLab read
ok - terminal failed run is authoritative
ok - orphaned ci monitor after green reads as held-for-merge done
ok - status-only failed orphaned ci monitor after green reads done
ok - genuinely failing CI keeps the failed verdict
ok - a second failed step disqualifies the orphaned-monitor reclassification
ok - cross-branch run is attributed via the real runs list
ok - socket refusal over a coarse active run reports blocked
ok - failed ledger record reads unknown only when the daemon is provably down
ok - cross-branch attribution picks the branch's most recent row
ok - a newer failure is not hidden by a live sibling
ok - runs-list selection keeps the newer failure over an older live row
ok - an older unfetched live sibling does not hide a newer failure
ok - two terminal rows keep the existing newest-first precedence
ok - an unclassifiable status row keeps the ledger's newest-first precedence
ok - a terminal run with no live sibling is unchanged
ok - coarse run does not probe another branch's ci log
ok - another branch's run is ignored, falls back
ok - no run + a busy semantic record reads working, attributed to its source
ok - a converted adapter never reads working from rendered footer text
ok - grok still reads working through its isolated rendered-tail fallback
ok - herdr's native busy verdict reads working with no record present
ok - a herdr CLI that fails to answer reads unknown/unreachable, never gone
ok - an alive endpoint whose scrollback read failed stays working
ok - a husk pane (agent gone) still reads gone for reclaim
ok - a mid-tool-call crew stays working because its record outranks herdr's generation state
ok - an idle record with idle agent_status stays not-busy (no regression for a human-blocked agent)
ok - no run + idle pane uses the status-log verb
ok - no run + idle pane parses keyed status syntax
ok - no run + idle pane on a paused: status reports state: paused with its reason
ok - no run + idle pane honors the configured paused verb
ok - a trailing resolved: event does not corrupt state render (idle stays idle)
ok - dead window ignores stale status log
ok - a tmux that fails to answer reads unknown/unreachable, never gone
ok - closed pane still reports a terminal run-step
ok - closed pane still reports an active run-step
ok - no timeout command uses perl bound
ok - scout skips the run lookup
ok - torn-down worktree is handled gracefully
ok - fm-crew-state remote: alive endpoint falls through to the routed status log
ok - fm-crew-state remote: an idle alive endpoint reads alive, never gone or dead
ok - fm-crew-state remote: an unreachable host reads unknown-remote, never gone or dead
ok - fm-crew-state remote: the remote host's own dead verdict is reported truthfully
ok - missing meta is handled gracefully
ok - crew_is_provably_working absorbs a validating crew found only via the runs-list fallback
ok - crew_is_provably_working still surfaces a genuinely stopped crew (safety property preserved)
ok - usage error exits 2
ok - historical same-branch rewritten head is not attributed as current
ok - active run with valid descendant fix head remains current
ok - local work advanced past run head invalidates attribution
ok - pipeline-owned active run binds without head equality and beats the failed row
ok - a genuinely failed run with no later run is not hidden
ok - coarse scan anchors the unresolvable active row instead of falling to an older one
ok - coarse scan with a mismatched anchor stays unknown and lets the pane answer
ok - the exemption requires branch_sync.state=pipeline_owned
ok - the exemption never applies to a terminal run
ok - missing run head falls back instead of matching by branch
ok - active fix round with an unfetched pipeline head reads working
ok - unanchored unverifiable active row is never attributed
ok - unresolvable terminal row never reads as current
ok - runs-list continuation attribution works when axi answers another branch
ok - herdr stale registration over a shell-only pane reads agent gone, not alive
ok - herdr stale working record never reports a shell-only pane busy
ok - capped overview retains both competing same-branch run ids
ok - same-branch identity survives both runs falling outside the overview
ok - complete inventory preserves the replacement gate without writes
ok - missing complete-inventory failure reports unknown
ok - corrupt complete-inventory failure reports unknown
ok - schema complete-inventory failure reports unknown
ok - repo complete-inventory failure reports unknown
ok - count complete-inventory failure reports unknown
ok - R6 complete selection ignores unrelated branch semantics
ok - R6 requested branches use exact identity without a whitelist
ok - R6 capped inventory ignores unrelated semantics and names both ids
ok - R6 complete identity lookup precedes partial-row semantic rejection
ok - R6 capped inventory preserves quoted requested-branch identity
ok - R6 structural completeness and requested-run validation remain enforced
ok - R5 complete inventory without Python keeps the replacement gate
ok - R5 complete ambiguity without Python names both ids
ok - R5 capped lookup without Python preserves available ids
ok - R5 capped lookup without SQLite support preserves available ids
ok - R1 both directions of inventory liveness disagreement read unknown
ok - R2 uninitialized busy workers retain pane reporting
ok - R2 uninitialized idle workers retain status reporting
ok - R3 historical inventory yields to the current busy pane
ok - R3 historical inventory yields to current worker status
ok - superseded cancelled run preserves the replacement review gate
ok - competing live runs report unknown with both run ids
ok - newer failed run remains failed beside an older live run
ok - missing run selection reports unknown with candidate ids
ok - wrong-id run selection reports unknown with candidate ids
ok - wrong-branch run selection reports unknown with candidate ids
ok - wrong-head run selection reports unknown with candidate ids
ok - missing-status run selection reports unknown with candidate ids
ok - malformed-table run selection reports unknown with candidate ids
ok - inventory-error run selection reports unknown with candidate ids
ok - selected-error run selection reports unknown with candidate ids
ok - legacy conflicting run records report unknown
all fm-crew-state tests passed
exit=0
Evidence: Live codex idle composer guard (blocked, with d82afe3 negative control)

Source: Live codex idle composer guard (blocked, with d82afe3 negative control)

=== NEGATIVE CONTROL: starfield report loop reverted to pre-d82afe3 ===
tests/fm-composer-codex-idle-live-e2e.test.sh: line 127: _fm_composer_row_is_braille_furniture: command not found

(with d82afe3 as merged, no such error; the guard's own verdict is blocked by codex's folder-trust prompt for the gate worktree path)

Pipeline

Updates from git push no-mistakes

� **intent** - passed

� No issues found.

� **Rebase** - passed

� No issues found.

�� **Review** - 1 info
  • â�¹ï¸� bin/fm-contributions.jq:18 - The shared nameable-owner rule is spelled once in jq (nameable_task) and once, implicitly, as the intersection of two general-purpose shell predicates. The jq comment names both shell predicates, but neither fm_task_id_path_safe (bin/fm-pr-lib.sh:97) nor fm_task_data_valid_id (bin/fm-task-data-lib.sh:40) points back at the jq copy, and write_record (bin/fm-contributions.sh:179) still calls fm_pr_task_id_valid ... || fail, which aborts the entire poll rather than skipping one owner. Since 1e2e5e5 removed the per-use guards, the jq filter in a different file is the ONLY thing keeping that fail unreachable. Concrete divergence risk in each direction: (a) tighten fm_task_id_path_safe (it has 15+ call sites in fm-pr-lib.sh and is edited for reasons unrelated to contributions) so it rejects an id jq still accepts, and the next poll that reaches that owner ends with fail &#39;invalid contribution task&#39;, leaving every later URL in known.tsv unobserved after the forge read was already spent - the exact escalating shape rounds 1-3 addressed; (b) loosen either shell predicate so it accepts an id jq rejects, and that contribution silently disappears from both poll and coverage with no ERRORS count and no disclosure. The round-3 decision explicitly required the two spellings be "adjacent and cross-referenced so a later edit cannot change one without seeing the other"; that is satisfied in the jq->shell direction only. Remedy is comment-only and matches the decision's "do not add guards" boundary: add a back-reference at fm_task_id_path_safe and fm_task_data_valid_id naming nameable_task in bin/fm-contributions.jq, and one line at bin/fm-contributions.sh:179 recording that this fail is unreachable only because known() already filtered the owner.
�� **Test** - 1 info
  • â�¹ï¸� tests/fm-composer-codex-idle-live-e2e.test.sh:33 - The live codex composer guard (tests/fm-composer-codex-idle-live-e2e.test.sh) hard-fails rather than gate-skips when run from a no-mistakes gate worktree on this host. codex 0.155.0 first shows an update-available modal (the guard's single Escape dismisses it) and then a folder-trust prompt for the ephemeral worktree path, which ~/.codex/config.toml has never trusted (it trusts ~/dev/john/firstmate). The guard documents a trust dialog as a legitimate hard failure, so this is an environment condition, not a defect in this change, and CI is unaffected because fm_live_gate skips when codex is absent. To enable it locally: run codex once in the gate worktree and answer "1. Yes, continue", then re-run the guard. Writing that trust entry requires modifying ~/.codex/config.toml, a global user-level tool config this run's workspace boundary forbids; codex -c &#39;projects.&#34;&lt;path&gt;&#34;.trust_level=&#34;trusted&#34;&#39; was tried and does not override it.
  • Live validation: â�� go - 14 of 16 scenarios driven live against the product
Scenario Result Live Evidence
A backlog row whose owner id the durable layer cannot name spends no forge read, writes no record, and never reaches the wake queue � pass live bash tests/fm-contributions.test.sh (test_unusable_task_id_keeps_poll_alive, test_unusable_task_id_raises_no_captain_wake) plus the hand-driven transcript in contributions-cli-transcript.log showing�
Bearings still reports contribution coverage complete and proven_clear while an unnameable owner sits in the backlog, in the home and through a child home summary � pass live fm-bearings-snapshot.sh --json | jq .contributions in contributions-cli-transcript.log returning complete=true, proven_clear=true; test_home_summary_coverage covers the child-home demotion path
Adversarial: poll and the coverage projection never disagree about which owner is nameable, across ids each shell predicate treats differently � pass live nameable-owner-rule-differential.log - 15 ids including tasks, -dash, .hidden, a b, a/b, ../escape, non-ASCII and a shell-injection string, all agreeing with the intersection of fm_pr_task�
fm-contributions.sh arm --if-owned arms no check when the backlog's only contribution row has an unnameable owner � pass live contributions-cli-transcript.log section 5: exit 0, state/*.check.sh absent; same assertion in test_unusable_task_id_keeps_poll_alive
Adversarial: a backlog row whose repo: hint is prose rather than a registry name groups under no-project instead of ending the poll for every other owner � pass live contributions-cli-transcript.log shows repo: kunchenguid/firstmate writing to data/tasks/_none/upstream/contributions.json while the good owner's record is still refreshed; `bash tests/fm-contributi�
Adversarial: a data root spelled with a trailing slash still records observations, and a symlinked data root is still refused through the same spelling � pass live contributions-cli-transcript.log section 6 - record written with the trailing slash, then fm-contributions: data directory unavailable once the root is a symlink; test_data_root_trailing_slash
On this macOS home's stock bash 3.2, a second churning turn-end inside an open deferral window is absorbed instead of killing the watcher cycle � pass live bash32-churn-regression.log - negative control with the e213343 guard reverted dies with missing_keys[@]: unbound variable under real /bin/bash 3.2.57, merged code passes
The merged snapshot consumers still run under this home's stock /bin/bash 3.2.57 � pass live macos-stock-bash32-snapshot-lane.log - fm-fleet-snapshot-view and fm-bearings-snapshot, 80 cases, both exit 0 with bash resolving to /bin/bash 3.2.57
A record whose endpoint is dead reports itself once, re-arms if the endpoint returns, and live or unproven endpoints still wedge-escalate � pass live upstream-dead-agent-report-once.log - 5/5 targeted fm-watch-triage cases from 9bc051f
A buried status declaration or open decision is read from the latest status event rather than lost behind newer history � pass live upstream-latest-status-event.log - 123 cases across fm-classify-decision-key, fm-wake-drain-open-decisions-cursor, fm-inactive-reconcile, fm-send-resolve-key, fm-captain-hold-lifecycle
A run that passed reports verified PR/MR state (open, merged, GitLab, unreadable identity) instead of claiming merged � pass live upstream-crew-state-pr-and-status.log - 141 cases including the terminal passed run ... group from 7111081
A PR merge poll survives a volume remount and a validated merged poll notifies once then retires � pass live upstream-pr-merge-poll-remount.log - 37 cases in fm-pr-check-security from baede47
The fork's deliberate local behaviour still holds in every file the merge had to reconcile from both sides � pass live fork-local-behaviour-merge-conflicts.log / -2.log / -3.log - 605 cases across composer, brief, control relaunch, fleet snapshot, gotmp, task delivery, spawn dispatch profile, zellij backend, teardown,�
The live codex starfield guard calls a classifier that still exists after upstream's rename � pass live live-codex-idle-composer-guard.log - the pre-d82afe3 line produces _fm_composer_row_is_braille_furniture: command not found against the installed codex 0.155.0; the merged line runs clean and report�
An idle codex composer classifies as empty through both the tmux and cursorless styled read profiles �� untested no codex records folder trust only in the user-level ~/.codex/config.toml, which trusts ~/dev/john/firstmate but not this ephemeral gate worktree path; `codex -c 'projects."<path>".trust_l�
MAIN answers a finished captain-requested deliverable with an outcome naming what finished, not Captain, shipshape. �� untested no No runtime surface. The upstream commit message states it outright: "No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner�
  • bash tests/fm-contributions.test.sh (43 cases, includes test_unusable_task_id_keeps_poll_alive, test_unusable_task_id_raises_no_captain_wake, test_unusable_project_hint_keeps_poll_alive, test_data_root_trailing_slash, test_home_summary_coverage)
  • Hand-driven CLI transcript: fm-contributions.sh poll, fm-contributions.sh arm --if-owned, fm-bearings-snapshot.sh --json against an isolated home with a malformed backlog (evidence/drive-contributions.sh)
  • Adversarial differential of the shared nameable-owner rule: bin/fm-contributions.jq's nameable_task vs fm_pr_task_id_valid and fm_task_data_valid_id over 15 ids including tasks, -dash, .hidden, a b, a/b, ../escape, non-ASCII and a shell-injection string
  • PATH=/bin:/usr/bin:/usr/sbin:/sbin:/opt/homebrew/bin FM_TEST_ONLY=test_turn_ended_churn_existing_marker_absorbed /bin/bash tests/fm-watch-triage.test.sh under real /bin/bash 3.2.57, with a negative control that reverts the e213343c guard
  • /bin/bash tests/fm-fleet-snapshot-view.test.sh and /bin/bash tests/fm-bearings-snapshot.test.sh under real /bin/bash 3.2.57 with tasks-axi 0.2.5 (the macos-stock-bash CI lane)
  • FM_TEST_ONLY=test_gone_endpoint_reports_once_instead_of_escalating_forever|test_gone_report_rearms_when_the_endpoint_comes_back|test_identical_dead_display_of_a_successor_still_reports|test_second_death_after_a_same_window_relaunch_reports_in_full|test_live_and_unproven_endpoints_still_wedge_escalate bash tests/fm-watch-triage.test.sh
  • bash tests/fm-crew-state.test.sh (141 cases, includes the terminal passed verified-PR-state group)
  • bash tests/fm-pr-check-security.test.sh (37 cases, merge polls across volume remounts)
  • bash tests/fm-classify-decision-key.test.sh, fm-wake-drain-open-decisions-cursor, fm-inactive-reconcile, fm-send-resolve-key, fm-captain-hold-lifecycle
  • bash tests/fm-composer-lib.test.sh, fm-brief, fm-control-relaunch, fm-fleet-snapshot-view, fm-gotmp, fm-task-delivery, fm-spawn-dispatch-profile, fm-backend-zellij, fm-teardown, fm-teardown-endpoint-safety, fm-backlog-handoff, fm-backlog-atomicity, fm-public-followup, fm-backend-herdr-smoke (every file the merge reconciled from both sides)
  • bash tests/fm-composer-codex-idle-live-e2e.test.sh against the installed codex 0.155.0 in an isolated tmux server, plus a negative control reverting the d82afe3 starfield line
  • bash tests/fm-watch-triage.test.sh (stopped at 85/244 cases, 0 failures; each case runs a real watcher at live poll intervals)
� **Document** - passed

� No issues found.

� **Lint** - 1 issue found � auto-fixed �
  • â� ï¸� linter found issues (exit code 1)

� Fix applied.
� Re-checked - no issues remain.

� **Push** - passed

� No issues found.


Broad suite run and failure attribution

The permitted broad run completed once: 168 scripts, 13 failed, 5 gate-skipped, 14,615,824 ms.
The real-Herdr and live-harness-optin families were excluded because live agent lifecycle operations were outside this task's authorization.
Selected gate skips covered unavailable cmux, zellij, TypeScript compiler and native Windows Node, plus a stock-renderer case requiring a newer Pi version.
Those surfaces remain for their authorized CI lanes.

Focused composer, brief, task delivery, grouped contributions, relaunch and gotmp checks passed.
The corrected Zellij fixture passed its complete suite.
Executable smoke passed for session-start, spawn --help, wake-drain and an actual update against isolated local-origin fixtures.
Final ShellCheck 0.11.0, actionlint 1.7.12, documentation audience checks (102 surfaces, 418 links), whitespace checks and conflict-index checks passed.

Failure attribution

Controls used clean tracked pre-merge main at d5967e12e416dcf0f7df949cd83d7d1ca6db0cf1.
Where home identity mattered, controls added synthetic untracked identity, local-parent and dispatch fixtures without copying live operational state.

Broad failure Result
fm-backend-zellij Fixed fixture home/root isolation; complete rerun passed without weakening teardown.
fm-calm-pi-extension Identical focused readiness failure on pre-merge main with Pi 0.84.1; left unchanged.
fm-extension-binding Same primary missing-Node remote-binding failure on pre-merge main; later recursive lock-path diagnostics remain unexplained.
fm-gate-refuse Same teardown assertion fails on pre-merge main with the synthetic durable parent record; ordinary clean control passes.
fm-omp-harness Same Codex-versus-Claude ancestry assertion fails on pre-merge main.
fm-pr-check-security Unexplained: focused pre-merge replacement-authority control passes; this does not establish an environmental cause.
fm-remote-secondmate-lifecycle-e2e Unresolved: pre-merge control fails earlier at blocked inheritance and does not reach the original retirement assertion.
fm-remote-transport-lanes Unexplained: complete pre-merge control passes; no quiet-host proof of environmental causation.
fm-secondmate-harness Same focused config-push pointer-delivery assertion fails on pre-merge main.
fm-spawn-batch Same failure on pre-merge main with active synthetic dispatch configuration; ordinary clean control passes.
fm-teardown Same first-case refusal on pre-merge main with the synthetic durable parent record; ordinary clean full suite passes.
fm-wake-queue Same foreign-queue stall assertion fails on pre-merge main, with different checkpoint output; common root cause remains unproven.
fm-watch-triage Pre-merge control also exceeds 900 seconds after the same last reported checks (exit 124, 900,241 ms); cause of slowness remains unproven.

The host remained heavily loaded during controls.
No failure is dismissed as timing or flakiness, and no quiet-machine environmental verdict is claimed.
Pre-existing failures were recorded without unrelated fixes.
The unresolved failures above remain validation limitations; this merge does not claim a green broad suite.

NewAiCoder and others added 30 commits September 14, 2026 03:18
…kunchenguid#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (kunchenguid#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs kunchenguid#4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes kunchenguid#3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send

A marked secondmate request sent with an empty message delivered only
marker and correlation bytes and minted a pending-reply expectation the
parent could never see resolved, stalling the fleet with no loud error
(kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text
path, mirroring the existing --resolve-key refusal.

* chore: retain ambient Pi-lens autoformat as its own commit

Formatting-only edits produced by ambient Pi-lens autoformat during the
msg-loss investigation, kept separate from the behavioural change in
c23acba6 so the fix stays reviewable on its own.

AGENTS.md is deliberately excluded: its only autoformat edit stripped the
trailing space from the documented FM_OPERATIONAL_PREFIX value, which
bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11
records as permanent compatibility. Documenting that constant without its
trailing space makes the doc wrong about the contract, so that one line was
restored rather than retained.
…chenguid#4554)

On rose-pine-moon the two-color water (cyan crests over blue troughs) read as
a pink stripe over aqua, the yellow left sail and mast clashed with the red
right sail, and the hull carried a blue interior run. Every water cell is now
blue so the swell reads through glyph height alone, and both sail halves, the
mast, and the whole hull are one yellow run. Geometry, cadence, animation,
direction flip, resize clamping, and the narrow fallback are unchanged.

Update the unit and real-TUI color assertions to the new palette and the Calm
docs that described the old one.
…chenguid#4270)

* fix(watch): stop aging a second mate's active turn from its launch

The parent watcher's second-mate wake-loop stall check exempts a mate that
is demonstrably inside an active turn, but secondmate_in_active_turn asked
busy_turn_over_age first and returned "not in a turn" whenever that said
the bound was crossed.

busy_turn_over_age ages from state/<task>.turn-ended, falling back to
state/<task>.meta. A second mate's turns end in its own home, so the
parent never gets a turn-ended mark for it and the fallback ages the
mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago
was therefore permanently "over age", the busy pane was never consulted,
and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false
wake-loop stall.

The gate now bounds the busy exemption by <idle> - how long the queue's
drain position has not moved - which is evidence this home actually
holds. A busy mate stays exempt while the queue has been frozen for less
than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so
the bound that stops a busy pane from proving liveness forever is kept
rather than removed. busy_turn_over_age is untouched; its remaining
callers are the ordinary crew busy-pane bound.

The regression pins the case that actually broke: a mate whose launch
record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn
must not escalate, while the same mate with its queue frozen past the
bound still publishes exactly one notification. The existing coverage
only exercised a freshly launched mate, which passes either way.

Reaching that alert now costs a pane capture inside the gate, so the
three checkpoints in this suite that assert an alert move from a 1s to a
4s bound - the value the neighbouring active-turn cases already use. The
bound is a ceiling, not a wait: the checkpoint returns on the first
actionable wake. On a loaded machine a 1s bound missed the alert
repeatedly; at 4s it did not miss in 20 runs under the same load.

* no-mistakes(review): scope the second-mate active-turn regression test's coverage claim

* no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278)

* feat(bin): add read-only PR blocker and reviewer-discovery commands

Two focused, opt-in commands that read GitHub and never write to it.

fm-pr-state.sh reports what still blocks one pull request from the
author's side: a closed or merged state, draft state, unknown or
conflicting mergeability, absent or failing required checks, and a
blocking CHANGES_REQUESTED decision explained by each reviewer's latest
verdict, marked STALE when it was left at a superseded head. A pull
request that only awaits an approval is not reported as blocked, and
advisory checks are omitted. Every reading is taken against one exact
head; a push that lands mid-read invalidates the whole result rather
than mixing two snapshots.

fm-pr-reviewers.sh suggests reviewers from the most recent commits to
the pull request's exact changed paths, counting each commit once,
resolving handles through GitHub's own commit author.login mapping, and
excluding the author and Bot accounts.

Both stay read-only: no review request, no approval, no merge.
Unresolved review-thread state is left unreported because the REST API
does not expose it and unattended commands may not use GraphQL.

Closes kunchenguid#3731

* no-mistakes(review): accept only PR URLs and stop at terminal state

* no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate

* no-mistakes(review): stop attributing readings to unverified heads

* no-mistakes(review): narrow readiness contract to checks that have reported

* no-mistakes(review): read the pull request once, drop the head guard

* no-mistakes(document): scope pr-forge isolation proof to its measured members

* no-mistakes(document): record uncovered pr-forge members and their pending proof

* docs(isolation-proof): re-prove pr-forge at its full membership

tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the
pr-forge family in this branch, and script_allows_concurrency grants
four workers by family membership alone, so both ran concurrently on a
proof measured before they existed.

Re-proved the family at all eight members: two consecutive runs, 0
failures, each begun with the one-minute load average below 6.0 so the
result measures isolation rather than contention. A third run taken
between them is disclosed rather than recorded, because it started
while the previous run's workers were still decaying.

The new durations are not comparable with the six-member measurement
above them, so they are not presented as evidence about the two new
members, and that record's 1.72x four-worker figure is left as a
statement about its own run rather than restated as current.

* no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752)

* fix(bin): teach validation-round pauses in briefs

* no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510)

* fix(teardown): refuse a cleanup whose endpoint close failed

bin/fm-teardown.sh discarded both the exit status and the stderr of every
fm_backend_kill call, so a close that genuinely failed was indistinguishable
from one that succeeded. Teardown continued past it, deleted the task's durable
records, returned its worktree, and reported the cleanup as completed. The
deleted metadata is the only record of which endpoint belongs to the task, so
such a close did not merely leave a stray session behind, it stranded one:
nothing was left on disk naming it.

The adapters could not carry that signal either. Driven against the real code,
every backend arm returned 0 for a genuine failure exactly as it did for an
already-exited endpoint, so there was nothing for the four call sites to
propagate even once they stopped swallowing it.

The tmux arm now resolves a close that did not succeed against the window's
exact recorded identity, since kill-window fails the same way for a window that
is gone and one that is still there. The Orca arm reports a close its missing
CLI never attempted. Both stay silent for an endpoint that is already
legitimately gone, and the remaining arms are unchanged: their close-command
timing cannot be established without the real Zellij, Orca, and cmux binaries,
and a gate that refused ordinary cleanup of an already-exited session would be
worse than the defect. docs/verification/runtime-backends.md records what each
backend can prove.

A reported close failure now reaches teardown's existing retain-and-stop
refusal before the records naming the endpoint are removed, matching where the
Herdr confirmed-gone gates already sit for the same hazard, and the retained
records let a rerun finish once the close works.

* no-mistakes(review): refuse unreadable tmux close re-read; honor --force override

* no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close

* no-mistakes(document): document endpoint-close refusal in its backend and retirement owners

* no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag

Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that
brings Calm to Claude Code: the sailboat replaces the stock working row through a
Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration,
and canonically classified operational user rows draw at zero height. /calm is
registered by the hooks module itself and toggles the same per-home config/calm
preference the Pi extension uses, so one choice applies on either harness; rows
redraw retroactively on toggle and stay hidden across claude --continue.

The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS
flag is on. Nothing sets that flag in any settings file, and the plugin carries no
command file, skill, agent, or classic hook, so it is a complete no-op while the
flag is off. The trusted project auto-loads it through an .agents/skills symlink,
the only path Claude Code scans for project plugins.

Extract the working-ship geometry, bounce track, cadences, and freeze/resume state
into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module
imports from outside the plugin folder) and have the Pi widget paint that core's
frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify
operational rows through a port of bin/fm-operational-input.sh's classify command
guarded by a corpus parity test against the shell owner.

Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering,
Raster packing, policy, classifier parity), the mod's own claude plugin test suites
behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op,
the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272.

Docs: record the version-scoped Claude Code evidence and the three bounded gaps in
docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md,
and make the shared preference, layout, and contributor notes harness-neutral.

* no-mistakes(review): Preserve colliding final replies and strengthen parser parity

* no-mistakes(review): Preserve final replies and strengthen canonical parity checks

* no-mistakes(review): Require exact function-hooks opt-in before Calm activation

* no-mistakes(review): Clarify Calm module loading and activation boundaries

* no-mistakes(review): Reset Calm presentation state across session starts

* no-mistakes(document): Refresh Calm session lifecycle documentation

* feat(calm): paint the Claude Code working ship in Claude's own theme colors

The captain picked the "Claude native" palette for the Claude Code mod's Raster:
every water cell takes the spinner blue of the active theme family (#93a5ff dark,
#5769f7 light) and the whole boat takes the Claude orange of the stock spinner
(#d77757), one water color and one boat color. The family follows the `theme`
setting's prefix, read at load through $.config.list and re-read on a
config.set of that row, with `auto` and custom themes falling back to the dark
set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte.

Rename the shared sprite's color classes from hue names to `water` and `boat`,
since each harness now maps them to its own colors; geometry, motion, cadence,
and the activation gate are untouched.

Tests cover both palettes' packing and the family rule under Node, and the
plugin kit drives every theme value, a theme change mid-session, the Calm-off
pass-through, and inertness of the menu read while the flag is off. The docs
describe the Claude Code colors and record the guard passing on 2.1.273.

* no-mistakes(review): Use light palette for unresolved Claude themes

* no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes kunchenguid#4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627)

* fix: restore published contribution follow-up (Fixes kunchenguid#4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope
mremond and others added 25 commits September 16, 2026 12:19
…llow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
…kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
…nchenguid#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by kunchenguid#4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: kunchenguid#3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and
kunchenguid#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
kunchenguid#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
@jbalke
jbalke merged commit a9be36c into main Sep 20, 2026
20 checks passed
@jbalke
jbalke deleted the fm/fm-reconcile-upstream-2026-09-18 branch September 20, 2026 15:31
@jbalke jbalke mentioned this pull request Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.