Skip to content

fix(bin): follow the checkout's configured upstream instead of a hardcoded origin - #2726

Closed
joliverMI wants to merge 19 commits into
kunchenguid:mainfrom
joliverMI:fm/fm-fleet-follows-fork
Closed

joliverMI wants to merge 19 commits into
kunchenguid:mainfrom
joliverMI:fm/fm-fleet-follows-fork

Conversation

@joliverMI

Copy link
Copy Markdown

Intent

Make the fleet's self-update follow the repository we actually develop in, not the upstream it was forked from. A firstmate checkout can have "origin" pointing at the public template it was forked from while its own branch tracks a separate fork it actually develops on (e.g. this checkout: origin=joliverMI/firstmate is the fork we develop on, upstream=kunchenguid/firstmate is the public template), and the two can genuinely diverge. bin/fm-update.sh compared against origin/main unconditionally and never advanced a home tracking a different upstream past the divergence point. Scope:

  1. bin/fm-update.sh must resolve the update source from the branch's configured upstream (git tracking branch) rather than hardcoding origin/main; where no upstream is configured, keep today's origin/main behavior and say so explicitly rather than guessing.
  2. Audit every other place in bin/ that assumes origin is the development remote (fleet sync and anything else that fetches, compares, or reports divergence); fix them together under the same resolution, or state explicitly why a given script is already correct (e.g. project-clone scripts that intentionally always use origin because project clones have no fork/upstream split).
  3. Do not change push or PR targeting behavior in this task; that remains owned by standing repo convention and is currently working correctly.
  4. Determine whether the right fix for the remote code root (a clone whose own origin points at upstream rather than the fork) is repointing that clone's origin at the fork, or teaching the remote update path which remote to use, and recommend one with reasoning; do not apply a repoint as part of this change.
  5. Prove the outcome end to end, not just via code inspection: after the fix, a fleet update must leave fork-only files (bin/fm-dashboard.sh and fm-spawn.sh's --card flag) present in a remote secondmate home that updates through the fixed path.
    Constraints: fast-forward-only safety is absolute and unchanged - never force, never stash, never discard unlanded work, never merge; a dirty or genuinely diverged target is still skipped and reported, not forced. Do not touch anything under projects/. The PR targets joliverMI/firstmate, never kunchenguid/firstmate.

What Changed

  • Added bin/fm-dev-remote-lib.sh as the single owner of development-remote resolution: resolve_update_base reads the default branch's configured upstream (%(upstream:remotename) / %(upstream:remoteref)) and falls back to origin/<default> with an explicit note when no remote upstream is configured. bin/fm-ff-lib.sh now fetches and fast-forwards through it (per-remote fetch_once keying, default_branch takes a remote argument), so /updatefirstmate advances a checkout that tracks a fork instead of comparing against a diverged origin/main; the fast-forward-only guards (no force, stash, merge, or dirty/diverged advance) are untouched.
  • Applied the same resolution to the other paths that fetch, diff, or report divergence: fm-review-diff.sh (base ref and refs/pull/<n>/head fetch), fm-teardown.sh (cached resolve_dev_remote for PR-object and landed-content checks), fm-spawn.sh (pooled-worktree base freshening), and the changed-file base in fm-lint.sh and fm-test-run.sh. fm-fleet-sync.sh is documented as deliberately origin-only because it walks project clones under projects/, which have no fork/upstream split; push and PR targeting are unchanged.
  • Added tests/fm-remote-update-follows-fork.test.sh, which drives fm-remote-secondmate-control.sh update against a fork/upstream fixture and asserts fork-only files land in the secondmate home, plus new cases in the lint, review-diff, test-run, and spawn-pool-base-freshen suites and fm-test-run.sh selection entries for the new library and test. Docs and skills were reworded from "origin" to "configured upstream", and docs/remote-secondmates.md records the recommendation to set the remote code root's branch upstream rather than repoint its origin.

Risk Assessment

✅ Low: The change is broad (one new shared resolver plus six consumers touching self-update, pooled-worktree provisioning, teardown's landed-work gate, and diff-base selection), but every resolution failure mode falls back to the prior origin behavior or refuses rather than forcing, fast-forward-only safety is untouched, each fixed path gained a real behavioral regression test, and only two narrow fail-safe issues survived verification.

Testing

Ran the change's targeted suites (fork-follow e2e, fm-update, fm-lint, fm-review-diff, fm-test-run, fm-spawn-pool-base-freshen, fm-gotmp, fm-teardown, fm-secondmate-sync, fm-bootstrap, fm-fleet-sync) — all pass — and, because passing unit tests alone do not show the fleet actually following the fork, drove the real remote-update entrypoint end to end against a fixture built from this repository's own history, A/B-ing the pre-fix and post-fix bin/ over identical fixtures: the old path reported "already current" and left the secondmate home without bin/fm-dashboard.sh or --card, while the fixed path reported "updated 4f0a4f5..a28663a"/"synced:" and delivered both, with --card proven by its real argument-validation error rather than help text. I also captured the explicit no-upstream origin/main fallback note and confirmed fast-forward-only safety (diverged and dirty roots skipped, nothing forced, stashed, or discarded). No visual artifact applies: this change is git plumbing behind a CLI update path with no rendered surface, so the reviewer-visible evidence is the CLI transcript. Two failures I hit are environment-only and pre-existing at the base commit (locale-sensitive coverage guard, missing ruby); the working tree is clean and all evidence lives in the evidence directory.

Evidence: Fleet update follows the fork: before/after CLI transcript (fork-only dashboard + --card reach the remote secondmate home)

Source: Fleet update follows the fork: before/after CLI transcript (fork-only dashboard + --card reach the remote secondmate home)

### Fixture (real firstmate history) public template kunchenguid/firstmate main = 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5) development fork joliverMI/firstmate main = a28663a no-mistakes(review): repair gotmp fixture and teardown test selection fork-only commits ahead of the template: 12 ### BEFORE the fix (bin/ at b98e098) fm-remote-secondmate-control.sh update ios (as /updatefirstmate dispatches it to the host) current: 4f0a4f544dc0051dfaa06436a7d5e8e6ea0b4689 home HEAD : 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5) bin/fm-dashboard.sh : ABSENT fm-spawn.sh --card : ABSENT -> error: unknown harness '--card='; pass a raw launch command to use an unverified adapter ### AFTER the fix (bin/ at a28663a) fm-remote-secondmate-control.sh update ios (as /updatefirstmate dispatches it to the host) synced: a28663aa13d756fb7433f4cb130dc5c4b321ef04 home HEAD : a28663a no-mistakes(review): repair gotmp fixture and teardown test selection bin/fm-dashboard.sh : PRESENT -> fm-dashboard.sh - the ONLY way an agent touches the Admiral's Fleet Dashboard. fm-spawn.sh --card : PRESENT -> error: --card requires a non-empty value ### The update line the captain actually sees (bin/fm-update.sh against the same code root) before fix: firstmate: already current after fix: firstmate: updated 4f0a4f5..a28663a (instructions changed: AGENTS.md, bin, .agents/skills) ### No upstream configured: keeps origin/main and says so firstmate: no upstream configured for main; using origin/main firstmate: already current ### Fast-forward-only safety is unchanged diverged code root -> firstmate: skipped: diverged from fork/main diverged code root -> unlanded commit still at HEAD, nothing forced or discarded dirty code root -> firstmate: skipped: dirty working tree dirty code root -> uncommitted edit still in the worktree, nothing stashed

### Fixture (real firstmate history)
public template  kunchenguid/firstmate  main = 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5)
development fork joliverMI/firstmate    main = a28663a no-mistakes(review): repair gotmp fixture and teardown test selection
fork-only commits ahead of the template: 12

### BEFORE the fix (bin/ at b98e098)
  fm-remote-secondmate-control.sh update ios   (as /updatefirstmate dispatches it to the host)
  current: 4f0a4f544dc0051dfaa06436a7d5e8e6ea0b4689
  home HEAD                 : 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5)
  bin/fm-dashboard.sh       : ABSENT
  fm-spawn.sh --card        : ABSENT -> error: unknown harness '--card='; pass a raw launch command to use an unverified adapter

### AFTER the fix  (bin/ at a28663a)
  fm-remote-secondmate-control.sh update ios   (as /updatefirstmate dispatches it to the host)
  synced: a28663aa13d756fb7433f4cb130dc5c4b321ef04
  home HEAD                 : a28663a no-mistakes(review): repair gotmp fixture and teardown test selection
  bin/fm-dashboard.sh       : PRESENT -> fm-dashboard.sh - the ONLY way an agent touches the Admiral's Fleet Dashboard.
  fm-spawn.sh --card        : PRESENT -> error: --card requires a non-empty value

### The update line the captain actually sees (bin/fm-update.sh against the same code root)
  before fix: firstmate: already current
  after  fix: firstmate: updated 4f0a4f5..a28663a (instructions changed: AGENTS.md, bin, .agents/skills)

### No upstream configured: keeps origin/main and says so
  firstmate: no upstream configured for main; using origin/main
  firstmate: already current

### Fast-forward-only safety is unchanged
  diverged code root -> firstmate: skipped: diverged from fork/main
  diverged code root -> unlanded commit still at HEAD, nothing forced or discarded
  dirty code root    -> firstmate: skipped: dirty working tree
  dirty code root    -> uncommitted edit still in the worktree, nothing stashed
Evidence: Reproduction script for the end-to-end fork-follow demo (builds the fixture from real repo history, A/Bs pre-fix vs fixed bin/)

Source: Reproduction script for the end-to-end fork-follow demo (builds the fixture from real repo history, A/Bs pre-fix vs fixed bin/)

#!/usr/bin/env bash
# End-to-end demo: does a fleet self-update follow the fork the checkout tracks?
set -u
WT=/home/joliv/.no-mistakes/worktrees/4cc5c0885385/01M0GZ5GAYE1MMH0ZHM3GTDHTH
LAB=/tmp/fm-fork-lab
rm -rf "$LAB"; mkdir -p "$LAB"
export GIT_AUTHOR_NAME=fmtest GIT_AUTHOR_EMAIL=fmtest@example.invalid
export GIT_COMMITTER_NAME=fmtest GIT_COMMITTER_EMAIL=fmtest@example.invalid
export FM_GATE_REFUSE_BYPASS=1

B=$(git -C "$WT" rev-parse 3a972e3^)   # real ancestor: before bin/fm-dashboard.sh and --card existed
F=$(git -C "$WT" rev-parse HEAD)       # fork tip under test

echo "### Fixture (real firstmate history)"
echo "public template  kunchenguid/firstmate  main = $(git -C "$WT" log -1 --format='%h %s' "$B")"
echo "development fork joliverMI/firstmate    main = $(git -C "$WT" log -1 --format='%h %s' "$F")"
echo "fork-only commits ahead of the template: $(git -C "$WT" rev-list --count "$B".."$F")"
echo

git init -q --bare "$LAB/template.git"; git -C "$LAB/template.git" symbolic-ref HEAD refs/heads/main
git init -q --bare "$LAB/fork.git";     git -C "$LAB/fork.git" symbolic-ref HEAD refs/heads/main
git -C "$WT" push -q "$LAB/template.git" "$B:refs/heads/main"
git -C "$WT" push -q "$LAB/fork.git" "$F:refs/heads/main"

# pre-fix bin/ (base commit b98e098), run from its own directory
mkdir -p "$LAB/prefix"
git -C "$WT" archive b98e098 bin | tar -x -C "$LAB/prefix"

build_fixture() {   # remote host: code root whose origin is the template but whose main tracks the fork
  rm -rf "$LAB/remote-root" "$LAB/remote-home"
  git clone -q --origin origin "$LAB/template.git" "$LAB/remote-root"
  git -C "$LAB/remote-root" remote add fork "$LAB/fork.git"
  git -C "$LAB/remote-root" fetch -q fork
  git -C "$LAB/remote-root" branch -q --set-upstream-to=fork/main main
  git clone -q "$LAB/template.git" "$LAB/remote-home"
  printf 'ios\n' > "$LAB/remote-home/.fm-secondmate-home"
}

report_home() {
  local h="$LAB/remote-home"
  echo "  home HEAD                 : $(git -C "$h" log -1 --format='%h %s' HEAD)"
  if [ -f "$h/bin/fm-dashboard.sh" ]; then
    echo "  bin/fm-dashboard.sh       : PRESENT -> $(bash "$h/bin/fm-dashboard.sh" --help 2>&1 | sed -n '1p')"
  else
    echo "  bin/fm-dashboard.sh       : ABSENT"
  fi
  if bash "$h/bin/fm-spawn.sh" --help 2>&1 | grep -q -- '--card <card-id>'; then
    echo "  fm-spawn.sh --card        : PRESENT -> $(bash "$h/bin/fm-spawn.sh" x /tmp --mode direct-PR --yolo off --card= 2>&1 | sed -n '1p')"
  else
    echo "  fm-spawn.sh --card        : ABSENT -> $(bash "$h/bin/fm-spawn.sh" x /tmp --mode direct-PR --yolo off --card= 2>&1 | sed -n '1p')"
  fi
}

run_update() {   # $1 = bin dir to run the fleet update from
  FM_ROOT_OVERRIDE="$LAB/remote-root" FM_HOME="$LAB/remote-home" \
    bash "$1/fm-remote-secondmate-control.sh" update ios 2>&1 | sed 's/^/  /'
}

for variant in prefix fixed; do
  case $variant in
    prefix) BIN="$LAB/prefix/bin"; label="BEFORE the fix (bin/ at b98e098)" ;;
    fixed)  BIN="$WT/bin";         label="AFTER the fix  (bin/ at $(git -C "$WT" rev-parse --short HEAD))" ;;
  esac
  build_fixture
  echo "### $label"
  echo "  fm-remote-secondmate-control.sh update ios   (as /updatefirstmate dispatches it to the host)"
  run_update "$BIN"
  report_home
  echo
done

echo "### The update line the captain actually sees (bin/fm-update.sh against the same code root)"
for variant in prefix fixed; do
  case $variant in
    prefix) BIN="$LAB/prefix/bin"; label="before" ;;
    fixed)  BIN="$WT/bin";         label="after " ;;
  esac
  build_fixture
  out=$(FM_ROOT_OVERRIDE="$LAB/remote-root" FM_HOME="$LAB/remote-root" bash "$BIN/fm-update.sh" 2>&1 | grep '^firstmate:')
  echo "  $label fix: $out"
done
echo

echo "### No upstream configured: keeps origin/main and says so"
build_fixture
git -C "$LAB/remote-root" branch -q --unset-upstream main
FM_ROOT_OVERRIDE="$LAB/remote-root" FM_HOME="$LAB/remote-root" bash "$WT/bin/fm-update.sh" 2>&1 | grep '^firstmate:' | sed 's/^/  /'
echo

echo "### Fast-forward-only safety is unchanged"
build_fixture
git -C "$LAB/remote-root" commit -q --allow-empty -m 'unlanded local work'
DIV=$(git -C "$LAB/remote-root" rev-parse HEAD)
FM_ROOT_OVERRIDE="$LAB/remote-root" FM_HOME="$LAB/remote-root" bash "$WT/bin/fm-update.sh" 2>&1 | grep '^firstmate:' | sed 's/^/  diverged code root -> /'
[ "$(git -C "$LAB/remote-root" rev-parse HEAD)" = "$DIV" ] \
  && echo "  diverged code root -> unlanded commit still at HEAD, nothing forced or discarded"

build_fixture
printf 'uncommitted\n' > "$LAB/remote-root/AGENTS.md"
FM_ROOT_OVERRIDE="$LAB/remote-root" FM_HOME="$LAB/remote-root" bash "$WT/bin/fm-update.sh" 2>&1 | grep '^firstmate:' | sed 's/^/  dirty code root    -> /'
grep -q uncommitted "$LAB/remote-root/AGENTS.md" \
  && echo "  dirty code root    -> uncommitted edit still in the worktree, nothing stashed"
- Outcome: ⚠️ 2 warnings across 1 run (8m7s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Push main to origin, or rebase your branch onto origin/main, before gating.

⚠️ **Review** - 2 infos
  • ⚠️ bin/fm-review-diff.sh:110 - The PR-head fetch was repointed from origin to the resolved development remote, but refs/pull/&lt;n&gt;/head only exists in the repository the PR was opened against — which by standing repo convention is origin (CONTRIBUTING.md:18-30: "clone the parent repo or set your local origin back to the parent" ... "it pushes the branch to your fork and opens the PR against the parent repo"). The intent says "Do not change push or PR targeting behavior in this task; that remains owned by standing repo convention and is currently working correctly." Concrete failure: in exactly the code-root shape this commit recommends and its own fixture builds (origin = upstream template, main tracking fork/main), a firstmate self-dev task with pr=.../pull/9 makes fetch_pull_head fetch refs/pull/9/head from the fork. resolve_pr_head (bin/fm-review-diff.sh:120-136) prefers any successful fetch unconditionally and never cross-checks the fetched SHA against the recorded pr_head=, so an unrelated PR docs: clarify live supervision constraints #9 on the fork silently becomes the compare tip and the review shows the wrong diff with no warning. bin/fm-teardown.sh:826 has the same repoint, though it fails safe there (the recorded head SHA won't be present, so ensure_commit_object returns 1 and teardown falls back to the content check). Update lineage and PR base repo are independent concepts that only coincide when origin==fork; recommend leaving the PR-ref fetch on origin (or resolving it from the no-mistakes push/PR config) and keeping resolve_update_base for base/diff refs only.
  • ⚠️ bin/fm-dev-remote-lib.sh:31 - resolve_update_base derives the remote by splitting %(upstream:short) on the first /, which is not a remote name in every case. When &lt;default&gt; tracks a local branch (branch.main.remote = ., e.g. after git branch --set-upstream-to=dev main), %(upstream:short) prints a bare branch name with no slash, so ${upstream%%/*} and ${upstream#*/} both yield that branch name and RESOLVE_BASE_REF becomes dev/dev. A non-default fetch refspec that maps refs/heads/main on remote fork to refs/remotes/foo/bar breaks it the same way. Every caller then treats a non-existent remote as authoritative: ff_target (bin/fm-ff-lib.sh:314) reports skipped: no dev remote and the firstmate self-update stops advancing at all, and freshen_spawn_worktree_base (bin/fm-spawn.sh:1775) hard-refuses to launch any pooled task worktree — both cases previously worked off origin/&lt;default&gt;. Read the remote and branch directly instead: %(upstream:remotename) and %(upstream:remoteref) (strip refs/heads/), and fall back to the origin path when remotename is empty or ..
  • ⚠️ bin/fm-test-run.sh:996 - bin/fm-dev-remote-lib.sh is not enumerated in families_for_changed_path, so it falls through to the bin/*) branch and is mapped by families_for_test_reference — a literal grep of every test for the basename. Only tests/fm-test-run.test.sh mentions the file (the three cp lines added in this commit), so editing the shared resolver selects pure-contract-unit and nothing else. That silently skips every other suite that now depends on it: tests/fm-review-diff.test.sh and tests/fm-teardown.test.sh (pr-forge), tests/fm-update.test.sh (session-bootstrap), tests/fm-secondmate-sync.test.sh (secondmate), the spawn suites (backend-dispatch), and the new tests/fm-remote-update-follows-fork.test.sh (unclassified). A regression in resolve_update_base would pass bin/fm-test-run.sh --changed, which contradicts that map's own contract ("Conservative path → family map. Over-selects rather than under-selects"). Add bin/fm-dev-remote-lib.sh to the explicit case list next to bin/fm-ff-lib.sh with the families of its real consumers.
  • ⚠️ bin/fm-spawn.sh:1754 - freshen_spawn_worktree_base still hard-requires origin before it resolves the development remote: an unconditional full git fetch origin (line 1754) and git remote set-head origin --auto (line 1758), both fatal on failure. The intent's scope item 2 asks to fix "every other place in bin/ that assumes origin is the development remote" under the same resolution, and this is the function the commit message names as the one that stranded the task worktree. Concretely: a checkout whose development remote is fork and whose origin (upstream template) is unreachable, renamed, or absent now refuses to launch any pooled task worktree even though fork/main is fetchable — and in the healthy case it pays a redundant full fetch of a remote the base no longer comes from. Resolve the remote first, fetch/set-head that remote, and only consult origin for the default-branch name when the resolved remote can't supply one.
  • ℹ️ bin/fm-teardown.sh:823 - ensure_commit_object only needs a remote, but now calls default_branch purely to feed resolve_update_base, and returns 1 when it fails. In a project clone with no refs/remotes/origin/HEAD and no local main/master (bin/fm-teardown.sh:626-640), a merged PR that previously confirmed via fetch origin refs/pull/&lt;n&gt;/head can no longer be confirmed; content_in_default then fails on the same missing default, so work_is_landed returns false and teardown refuses work that has in fact landed. The direction is safe but the reachable set is newly narrowed. Resolve DEV_REMOTE once at top level after DEFAULT is already computed (the pattern bin/fm-review-diff.sh:78-80 uses) rather than re-deriving it per call inside both ensure_commit_object and content_in_default.
  • ℹ️ docs/scripts.md:90 - bin/fm-dev-remote-lib.sh is a new shared *-lib.sh but is not listed in docs/scripts.md, which README.md:218 calls "the bin/ toolbelt reference" and which enumerates every other sourced helper including fm-ff-lib.sh (line 90) and fm-tangle-lib.sh (line 86). Add a row so the new single owner of development-remote resolution is discoverable next to the fast-forward helper that sources it.
  • ℹ️ docs/remote-secondmates.md:214 - Intent item 5 requires proving "end to end, not just via code inspection" that "a fleet update must leave fork-only files (bin/fm-dashboard.sh and fm-spawn.sh's --card flag) present in a remote secondmate home that updates through the fixed path." The delivered proof is a synthetic fixture (tests/fm-remote-update-follows-fork.test.sh) that constructs a code root with fork added and main set to track fork/main, using a stand-in file named fm-dashboard.sh. The change itself deliberately does not repoint the real remote code root (intent item 4 forbids applying it), and the new doc line states the residual gap plainly: "a code root whose upstream resolves to a different repository than the primary checkout's own fork updates from the wrong lineage with no error." So the real remote home only receives fork-only files once its origin is repointed or its main upstream is configured out of band. Worth confirming with the captain whether the fixture proof satisfies item 5 or whether a live run against the real remote secondmate is expected before this lands.

🔧 Fix: harden dev-remote resolution and its consumers
5 issues (1 warning, 4 infos) still open:

  • ⚠️ bin/fm-test-run.sh:990 - The new bin/fm-dev-remote-lib.sh changed-path case claims to select "spawn's pooled worktree base" coverage via the backend-dispatch family, but tests/fm-spawn-pool-base-freshen.test.sh has no entry in family_for_basename (lines 137-235) and therefore resolves to unclassified, which no case emits. Concretely: edit resolve_update_base so freshen_spawn_worktree_base picks the wrong remote, run bin/fm-test-run.sh --changed, and backend-dispatch selects fm-spawn-batch/-dispatch-profile/-worktree-settle but never fm-spawn-pool-base-freshen.test.sh - the suite docs/architecture.md:162 names as the owner of the pooled-worktree base-freshness regression coverage, and the very suite this fix round extended with test_fork_tracking_pool_refreshes_from_fork_not_origin. The regression passes --changed silently, contradicting that map's own "over-selects rather than under-selects" contract. Emit printf &#39;%s\n&#39; &#34;__script__:fm-spawn-pool-base-freshen.test.sh&#34; from this case (the __script__: marker is already handled at line 862 and consumed at line 1107), or classify the suite into backend-dispatch in family_for_basename.
  • ℹ️ bin/fm-dev-remote-lib.sh:29 - The header comment states that with a non-default fetch refspec %(upstream:...) "prints a tracking path whose first component is not a remote name at all" and that such a checkout therefore "take[s] the announced origin fallback". That is true of the old %(upstream:short) split but false of the code it now documents: %(upstream:remotename) returns the real remote name and %(upstream:remoteref) returns refs/heads/&lt;branch&gt; regardless of the fetch refspec (verified on git 2.43), so resolve_update_base returns &lt;remote&gt;/&lt;branch&gt; with an empty NOTE. With e.g. remote.fork.fetch = +refs/heads/*:refs/remotes/mirror/*, ff_target's origin mode runs the plain fetch_once (populating refs/remotes/mirror/*), then fails the existence check at bin/fm-ff-lib.sh:331 and reports skipped: fork/main does not exist instead of the fallback the comment promises. It fails safe, but the comment misdescribes the behavior a maintainer would rely on. Either correct the comment to say only the local-branch (.) and no-upstream cases fall back, or additionally guard the returned ref.
  • ℹ️ docs/architecture.md:161 - "no worker starts until its clean task worktree matches the fetched tip of origin's resolved default branch" is now stale: freshen_spawn_worktree_base (bin/fm-spawn.sh:1752-1799) resolves the remote from the default branch's configured upstream and can provision from fork/main with origin never contacted. The change updated the "Self-updates stay safe" section of this same file (lines 327-328) and docs/remote-secondmates.md, so this line is an omission rather than a deliberate carve-out. Reword to "the development remote's resolved default branch (the configured upstream, falling back to origin)".
  • ℹ️ bin/fm-test-run.sh:1011 - bin/fm-ff-lib.sh maps to only pure-contract-unit, and because that is an explicit case it shadows the bin/* reference scan. This change materially altered ff_target's base resolution, fetch_once's dedup key, and default_branch's signature - whose actual owning suites are tests/fm-secondmate-sync.test.sh (family secondmate, sources fm-ff-lib.sh directly and asserts FF_STATUS/FF_INSTR) and tests/fm-update.test.sh (family session-bootstrap). Neither is selected by --changed when only fm-ff-lib.sh is edited. The mapping is pre-existing, but the fix round corrected exactly this class for the new sibling lib while leaving the now-broader-blast-radius one wrong; add secondmate and session-bootstrap to this case.
  • ℹ️ bin/fm-teardown.sh:656 - resolve_dev_remote, and the rerouting of ensure_commit_object (line 852) and content_in_default (line 950) onto it, changed which remote decides whether a worktree's work has landed - the check that gates discarding unpushed commits. Every other consumer touched by this change got a fork-tracking regression case in the same commit (fm-review-diff, fm-spawn-pool-base-freshen, fm-lint, fm-test-run, fm-remote-update-follows-fork), but tests/fm-teardown.test.sh was not extended, so the diverged-fork content_in_default comparison and the resolve-once caching (including the DEV_BRANCH empty fallback at line 951) have no coverage. The failure direction is safe (an unresolvable remote makes content_in_default return non-zero and teardown refuse), which is why this is informational rather than blocking.

🔧 Fix: correct dev-remote docs and test-selection coverage
3 issues (1 error, 2 infos) still open:

  • 🚨 tests/fm-gotmp.test.sh:79 - bin/fm-teardown.sh:195-196 now sources $SCRIPT_DIR/fm-dev-remote-lib.sh unconditionally under set -eu (line 160), but tests/fm-gotmp.test.sh builds its fake root by symlinking each sourced sibling one by one and was not updated. Concrete failure: make_fake_root (line 47) does ln -s &#34;$TEARDOWN&#34; &#34;$fake/bin/fm-teardown.sh&#34; and symlinks fm-backend/fm-tmux-lib/fm-nm-run-lib/fm-lock-lib/fm-control-lib/fm-classify-lib/fm-wake-lib/fm-gate-refuse-lib/fm-pr-lib/fm-public-followup-lib/fm-x-lib/fm-secondmate-registry-lib/fm-secondmate-parent-lib but not fm-dev-remote-lib.sh; SCRIPT_DIR resolves to $fake/bin (dirname of BASH_SOURCE[0], the symlink path), so the . of the missing file returns 1 and set -e aborts. test_teardown_removes_tasktmp_dir (line 124) runs bash &#34;$fake/bin/fm-teardown.sh&#34; &#34;$id&#34; || fail &#34;teardown exited non-zero with a valid tasktmp&#34; and fails, as do the sibling cases. The second inline fixture builder (around line 157, used by test_teardown_skips_gracefully_without_tasktmp) has the same omission. Fix: add ln -s &#34;$ROOT/bin/fm-dev-remote-lib.sh&#34; &#34;$fake/bin/fm-dev-remote-lib.sh&#34; to both builders. This is the only fixture affected - every other test either uses $ROOT/bin directly, symlinks the whole bin dir (tests/fm-session-start.test.sh:602,649), cp -Rs it (tests/fm-teardown.test.sh:1667), or was already updated (the three cp &#34;$RUNNER&#34; sites in tests/fm-test-run.test.sh).
  • ℹ️ bin/fm-test-run.sh:950 - The bin/fm-teardown.sh changed-path case emits only pr-forge and dashboard, but tests/fm-gotmp.test.sh - which runs the real teardown binary end to end against a hand-built fake bin/ and is therefore the suite most sensitive to teardown's sourced-sibling set - is classified session-bootstrap (line 183). So a teardown-only edit that adds or removes a sourced lib never selects the suite it breaks; that is exactly why the fixture drift above is invisible to bin/fm-test-run.sh --changed for teardown edits. (This change happens to be caught only because the new bin/fm-dev-remote-lib.sh case at line 990 does emit session-bootstrap.) The round-2 fix corrected this same class for bin/fm-ff-lib.sh (line 1011); add session-bootstrap to the teardown case for the same reason.
  • ℹ️ docs/scripts.md:90 - The fm-ff-lib.sh row still reads "Shared guarded fast-forward helper for origin pulls and local secondmate syncs", which this change makes inaccurate: ff_target's origin base mode now resolves the base through resolve_update_base and pulls from the checkout's configured upstream (e.g. fork/main), with origin/&lt;default&gt; only as the announced fallback. The change updated the neighbouring prose in docs/architecture.md (lines 161, 327-328), docs/remote-secondmates.md, and .agents/skills/updatefirstmate/SKILL.md, and added the fm-dev-remote-lib.sh row directly beneath this one, so this is an omission rather than a deliberate carve-out. Reword to "for configured-upstream pulls (falling back to origin) and local secondmate syncs".

🔧 Fix: repair gotmp fixture and teardown test selection
2 infos still open:

  • ℹ️ bin/fm-teardown.sh:950 - resolve_dev_remote resolves the remote from $PROJ (bin/fm-teardown.sh:660 -> resolve_update_base &#34;$PROJ&#34; &#34;$name&#34;), but both consumers then use that name as a remote in $WT: ensure_commit_object (line 852) runs git -C &#34;$WT&#34; remote get-url &#34;$DEV_REMOTE&#34; and content_in_default (line 952) runs git -C &#34;$WT&#34; fetch &#34;$DEV_REMOTE&#34; .... For task teardown those are the same repository, but for a secondmate they are not: fm-spawn.sh:634-636 records worktree=$home and project=$root, i.e. the persistent secondmate home and the code root are separate clones. Concrete failure: a code root whose main tracks fork/main while the secondmate home clone has only origin. A secondmate home holding a local commit not on any remote reaches work_is_landed -> content_in_default; git -C &#34;$WT&#34; remote get-url fork fails, so it drops to the elif and compares against the home's possibly stale local refs/heads/main instead of a freshly fetched default, and teardown REFUSES work that has in fact landed on the fork. Previously the hardcoded origin existed in both repositories, so this path always fetched. The fail direction is safe (refuse, never discard), which is why this is informational. Resolve the remote from $WT - the repository whose objects are actually being fetched and compared - rather than from $PROJ, or fall back to origin when $DEV_REMOTE is not a remote of $WT.
  • ℹ️ bin/fm-spawn.sh:1788 - The second resolve_update_base &#34;$worktree&#34; &#34;$default&#34; is keyed on the REMOTE's default-branch name (line 1781, read from refs/remotes/$remote/HEAD after set-head --auto), not on the local branch whose upstream produced $remote in the first place. When those names differ, refs/heads/$default has no upstream locally and the resolver falls back to origin, silently undoing the fork resolution the first pass made. Concrete path: a checkout whose local main tracks fork/main while fork's own HEAD is develop - pass 1 resolves fork, set-head fork --auto makes $default=develop, refs/heads/develop does not exist, so remote flips to origin and the pooled worktree is reset --hard to origin/develop (the diverged template lineage this change exists to stop using), or refuses if origin is unreachable. RESOLVE_BASE_NOTE is discarded here, so nothing in the error or the success path explains why origin was used. Suggest keeping pass 1's remote when pass 2 falls back to origin while pass 1 resolved something else, and echoing RESOLVE_BASE_NOTE when it is non-empty.
⚠️ **Test** - 2 warnings
  • ⚠️ bin/fm-test-run.sh:606 - bin/fm-test-run.sh --check-coverage exits 1 under this machine's LANG=en_US.UTF-8: its comm invocations run in the ambient locale while their inputs are sorted with LC_ALL=C, so comm reports "file 2 is not in sorted order". This aborts tests/fm-test-run.test.sh mid-suite. It is NOT a regression — it reproduces identically from a clean checkout of base commit b98e098, and both base and target pass with LC_ALL=C bin/fm-test-run.sh --check-coverage. I re-ran the suite under LC_ALL=C to validate this change rather than editing the shared runner, since the bug is pre-existing and unrelated to the fork-follow intent; the author may want to pin LC_ALL=C on those comm calls in a separate change.
  • ⚠️ tests/fm-test-run.test.sh:674 - tests/fm-test-run.test.sh's CI-workflow YAML check hard-fails when ruby is absent ("not ok - ruby is required to parse .github/workflows/ci.yml as YAML"). ruby is not installed in this sandbox and installing system packages is out of bounds here, so this one case cannot be exercised locally. It is environment-only, independent of the change under test, and every other case in that file passes.
  • bash tests/fm-remote-update-follows-fork.test.sh — new e2e: remote code root tracking a fork updates from it and the fork-only file reaches the secondmate home
  • bash tests/fm-update.test.sh — /updatefirstmate mechanics incl. off-default/detached/unsafe-home skips
  • bash tests/fm-lint.test.sh, bash tests/fm-review-diff.test.sh, bash tests/fm-spawn-pool-base-freshen.test.sh, bash tests/fm-gotmp.test.sh — the other resolve_update_base consumers
  • LC_ALL=C bash tests/fm-test-run.test.sh — includes "--changed defaults to main's configured upstream, not a hardcoded origin/main" (17 ok; only the ruby-dependent CI-YAML case fails)
  • bash tests/fm-fleet-sync.test.sh — confirms project-clone sync (deliberately origin-only) is unregressed
  • LC_ALL=C bash tests/fm-teardown.test.sh, tests/fm-secondmate-sync.test.sh, tests/fm-bootstrap.test.sh — remaining fm-ff-lib consumers
  • Manual A/B e2e: fixture from real history (template=4f0a4f5 pre-dashboard, fork=a28663a), ran FM_ROOT_OVERRIDE=&lt;code-root&gt; FM_HOME=&lt;home&gt; bin/fm-remote-secondmate-control.sh update ios with pre-fix bin/ (extracted from b98e098) vs fixed bin/, then checked the home for bin/fm-dashboard.sh --help and fm-spawn.sh ... --card= flag behavior
  • Manual: bin/fm-update.sh against a code root with git branch --unset-upstream main to confirm the announced origin/main fallback
  • Manual safety: bin/fm-update.sh against a diverged code root (extra local commit) and a dirty code root, verifying skip messages and that neither the unlanded commit nor the uncommitted edit was touched
  • bash /tmp/fm-base-checkout/bin/fm-test-run.sh --check-coverage — reproduced the locale failure at base commit b98e098 to confirm it is pre-existing
⚠️ **Document** - 1 info
  • ℹ️ bin/fm-remote-secondmate-control.sh:261 - bin/fm-remote-secondmate-control.sh:261 still dies with "remote code root did not complete a safe origin update" after the code root moved to configured-upstream resolution; the message names the wrong remote when a code root tracks a fork. I left it because it is a user-facing runtime string rather than documentation, and this phase may not change executable behavior. Suggested wording: "remote code root did not complete a safe update".

🔧 Fix: reword remaining origin-as-dev-remote strings and comments
1 info still open:

  • ℹ️ tests/fm-secondmate-sync.test.sh:4 - tests/fm-secondmate-sync.test.sh header/case comments (lines 4, 14, 236, 809) still describe the local-HEAD sync as "no origin fetch" and say standalone-clone homes converge through "/updatefirstmate's origin fetch". The assertions themselves are correct (they prove no fetch happens at all), so this is comment prose only. I left it because this phase is explicitly forbidden from editing tests; the equivalent wording in bin/ and the skill was updated to "remote fetch" / "fetch-based". Suggested follow-up in a phase allowed to touch tests: same one-line reword.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

joliverMI and others added 19 commits August 16, 2026 18:28
…ng state (#2)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(bin): render the captain's four-section status board from existing state

fm-status-board.sh prints one HTML page (Needs you to continue, In
progress, Waiting, Recently completed) rendered entirely from
fm-bearings-snapshot.sh's already-correct cross-home classification and
fm-fleet-snapshot.sh's untruncated per-item detail, with no second store
for agents to keep in sync.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
Tearing down a task deletes state/<id>.meta, the record bin/fm-pr-merge.sh
resolves the PR through. A branch can be fully pushed and reachable from a
remote (so the existing dirty/unpushed/landed checks find nothing to
refuse) while its own PR still sits open, stranding a mergeable PR with no
guarded path to land it - this happened twice in one night.

The new check runs after the existing dirty/unpushed/landed checks inside
validate_worktree_teardown_safety, so their refusal messages take
precedence and never contradict this one. It fires only when a pr= is
actually recorded (scout, local-only, and secondmate teardowns never
record one). A gh lookup error refuses loudly too, naming the PR firstmate
could not confirm, rather than tearing down blind. --force already skips
the dirty/landed checks and now explicitly skips this one too, since it is
the same "captain says discard" escape hatch - the usage comment spells
out what that means here (the PR needs merging or closing by hand,
outside fm-pr-merge.sh's guarded path).

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* fix(procevent): deliver captured results at most once

A Lavish-captured answer reached a secondmate five times because
forwarding it ran fm-send.sh and only afterward, manually, remembered to
call `handled` - so every re-announcement of the still-unacknowledged
capture (a restart, a compaction, a re-read wake) repeated the forward.
The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves
capture is exactly-once, but nothing paired a downstream effect with its
acknowledgement atomically.

Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`,
which checks handled status, runs the command, and marks the generation
handled as one call under the source lock, so a repeated delivery attempt
against an already-delivered generation runs the command zero times and a
failed command stays eligible for retry instead of being dropped. Point
the process-event-sources skill and the runner's operating contract at it
for exactly this class of forward.

Five delivery attempts against one captured generation now run the
downstream command once instead of five times - the read-and-discard
cost of the other four is eliminated structurally. The historical
incident's own token cost was not measured at the time and is not
reconstructable after the fact.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
…#4)

Adds a fifth board section between "Needs you to continue" and "In
progress" for finished work that is deployed and usable, not merely
merged. It renders only when firstmate has written a "ready: <where
to see it>" note on a landed backlog row, confirming deployment and
naming the surface - never inferred from merge/landed state alone, so
a merged-but-not-yet-deployed change never appears here. The item
shows its pointer text, never a PR link or repo path.

Approval buttons (part 2) are not implemented in this PR: investigation
found Lavish's artifact bridge (window.lavish.queuePrompt) can queue an
exact, single-item approval string, but transmission still needs the
existing manual "Send to Agent" tap by Lavish's own design, and the
delivery path could not be confirmed end-to-end from this environment.
Ship the verified section now and report the button investigation as a
plan.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
#5)

The board named locations the captain could not reach from his phone: a
pull request he does not review, a raw data/<id>/report.md path, a "fuller
detail lives in the X home record" cross-home line with nothing to tap, and
a bearings-truncated in-progress summary with no full text behind it.

Pull requests and GitHub links are now never rendered anywhere on the page.
A report links only when its backlog note carries an explicit report-url:
marker firstmate wrote after actually publishing it (the same convention as
the existing ready: marker) - investigated whether Lavish's existing route
could serve reports directly and confirmed it cannot: Lavish is a
publish-one-artifact-at-a-time surface, not a path-mapped static file
server, so there is no passive route to piggyback on. A Ready item's
"where to see it" text is flagged honestly instead of shown verbatim when
it looks like a filesystem path or a pull-request URL. A cross-home item's
bounded detail says plainly that the rest is not reachable from the page,
instead of naming an unreachable home record. An in-progress row now
prefers its fuller current-state text over the shortened one-line summary.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515)

* feat(dashboard): add the Admiral's Fleet Dashboard task board

Replaces the generated Lavish status page with a purpose-built board:
a zero-dependency Python backend (stdlib http.server + sqlite3), a
vendored React frontend styled to match Spectra, an agent-facing CLI
(bin/fm-dashboard.sh) that is the only way agents touch it, and the
fleet-dashboard skill teaching the fleet auditor and other agents how
to use it.

The board owns its own persistent records rather than mirroring any
backlog - its six-status vocabulary and four review tabs don't map
onto tasks-axi state, and the fleet auditor needs a separate claim to
check against live reality. See docs/dashboard.md for the full
architecture, the deploy story, and the drift-risk tradeoff that
choice implies.

Link posting rejects GitHub/PR links and local-only hosts structurally
(standing order 17), not just by agent memory. The page shows a loud,
persistent banner when it cannot reach its own API, and the auditor's
"never run" state is rendered distinctly from a clean sweep, so an
absence of confirmation is never shown as a confirmed-good state.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
Testing and needs-attention were conflated: both put a finished-enough
card in front of the Admiral, but only needs-attention actually blocks
on him. A card can now be set to needs-attention with a --reason that
renders directly on the card, sorts above every other status, and gets
its own loud always-visible section on the page. The fleet auditor's
sweep now treats a needs-attention card's age as a finding, since that
status - unlike testing - is not supposed to sit unanswered.

Existing databases migrate in place; the new column is nullable and
existing cards, history, and statuses are untouched.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…orce Audit button (#8)

The auditor's 15-minute cadence never actually ran: it registered a durable
check in its own home, but nothing polled it because that home had no live
supervision session. Host cron (already running independent of any chat
session) now drives bin/fm-fleet-audit-tick.sh, which heartbeats every tick
and sweeps via bin/fm-fleet-audit-sweep.sh once the dashboard's own interval
setting says it's due, self-healing a lock abandoned by a crashed sweep. The
Force Audit button uses the identical sweep script through a detached
subprocess so a scheduled and a forced run can never stack. Also fixes the
"unassigned" agent noise on every card (shows the captain instead) and adds
a Go to card jump from each discrepancy log entry.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…ard (#9)

Every dashboard card's agent and backlog_ref sat empty, so status
transitions only happened when someone remembered to make them by
hand - worst of all at teardown, the exact moment the worker that
would have remembered is being removed.

fm-spawn.sh --card <card-id> now records dashboard_card=<id> in the
task's own state/<id>.meta and best-effort links the card's ref/agent
to the task identity, advancing a not_started card to working.
fm-teardown.sh reads that identity back and advances the card to
testing once a real (non-forced) landing succeeds, leaving scout,
secondmate, and already-complete cards untouched.

No card is the normal case and stays a complete no-op. A bad card id
or an unreachable dashboard never fails the spawn or the teardown; it
warns loudly and best-effort records fm-dashboard.sh audit-log
--fleet so a failed link is visible to the fleet auditor rather than
silently dropped.

All 21 live cards still carry no agent/ref - 20 belong to a project
this task has no access to verify task identities for, and the one
firstmate-owned card is already complete with no live task to link -
so none were backfilled; see the PR description for the per-card list.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…nd sends (#12)

* fix(bin): tmux endpoint liveness reports dead targets as alive

fm_backend_target_exists's tmux branch trusted display-message's exit
status alone. tmux silently falls back to the caller's own active pane
(or an unrelated live session/window whose name is an unambiguous
PREFIX of the missing target) while still exiting 0, so a dead tmux
endpoint reads alive whenever the caller itself runs inside tmux -
which firstmate always does - or shares a socket with any other live
session/window. Pin the tmux probe to an exact target: session:window
pairs use list-panes -t "=$session:=$window" (the same exact-match
form fm_backend_tmux_kill already uses), bare %N pane ids / @n window
ids / $N session ids pass straight through since list-panes already
resolves them exactly, a bare (colon-less) target is checked against
exact session and window-name inventory since tmux's `=` pin does not
make a bare target exact the way it does the session component of a
session:window target, an all-digit window component resolves exactly
against the session's own window-index inventory (FM_SUPERVISOR_TARGET
_DEFAULT is "firstmate:0", an index not a name), and a dotted window
component is tried first as an exact window name and, if that fails,
falls back to pane-qualified addressing (window.pane or
window-index.pane) - both are valid, competing readings of the same
string in real tmux.

fm-crew-state.sh's pane_readable duplicated the same probe and now
defers to the fixed primitive. A task record naming a remote host is
no longer judged by the local tmux server at all in either caller:
fm-session-start.sh's digest reports those as an explicit "not checked
from here" verdict instead of probing a session that can never resolve
locally, and fm-crew-state.sh skips the local probe entirely for a
remote_host record so its status-log fallback - previously reached
only by accident, through the old probe's own bug - still runs.
fm-fleet-snapshot.sh, fm-control.sh, and fm-send.sh were checked and
already guard remote_host before reaching this probe.

herdr, zellij, orca, and cmux already address their target by id
through a structured inventory query, so they do not share this
defect shape.

Adds a real-tmux regression suite that runs the probe from inside a
live tmux pane (the one condition that reproduces the false positive,
since the fallback needs a live session to fall back to) against
nonexistent, prefix-colliding, dotted, pane-qualified, bare-name, and
window-index targets, and a crew-state case proving a remote
secondmate's status is still read when the local probe always fails.
All fail against the pre-fix code and pass after. Updates fake-tmux
test fixtures that modeled only display-message's semantics to model
the pinned probe instead.

Known gaps, deliberately not fixed here and tracked in one follow-up,
fm-tmux-agent-state-session-prefix-match: the recovery-grade sibling
fm_backend_tmux_agent_state (bin/backends/tmux.sh) still resolves its
session component by unpinned prefix match, same defect class,
unfixed; and fm_backend_tmux_kill can silently destroy a DIFFERENT
live window (not merely fail to remove the target one) when given a
dotted window name - verified live, also unfixed. No task id in this
fleet currently contains a dot, so nothing live is exposed by either
gap today. The follow-up will converge every tmux target-resolution
helper in this repo onto one shared resolver rather than each parsing
target strings independently.

Delivery note: this branch was built by hand directly on this fork's
own current main tip (git apply --3way of the reviewed diff) rather
than by rebasing, because this fork's history has diverged from the
upstream template repository it was created from and a rebase would
pull in commits that exist only on that upstream - this repo's
standing rule is that nothing is ever proposed to it.

* no-mistakes(review): refuse tmux key sends to targets that resolve inexactly

* no-mistakes(review): guard and exact-pin every tmux input-delivering send

* no-mistakes(review): resolve tmux probe and send targets through one resolver

* no-mistakes(review): refuse ambiguous bare tmux names, correct stale disclosures

* no-mistakes(test): add list-panes arm to shared secondmate fake tmux

* no-mistakes(document): align endpoint-verdict and supervisor-target docs with exact tmux resolution

* no-mistakes: apply CI fixes

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(dashboard): mechanically link handed-off backlog items to their card

Cards for work routed to a secondmate had no mechanism to advance: status
advances only through bin/fm-spawn.sh --card and bin/fm-teardown.sh, both
keyed off local task metadata that bin/fm-backlog-handoff.sh never creates,
so a card handed to a secondmate froze at not-started forever.

bin/fm-backlog-handoff.sh --card <card-id> gives a single handed-off item
the same best-effort link: once it lands in the secondmate's backlog, it
sets the card's ref/agent and advances a not_started card to working. The
durable identity lives in this home's own state/handoff-cards/<id> record
(never in the backlog item, which the handoff does not own) and is retired
per pair only once the board genuinely confirms that pair's link - never on
a merely-attempted one, so a failed link is retried by the next arrival
instead of silently orphaning the card. bin/fm-dashboard.sh's own curl calls
are now bounded so an unreachable board cannot stall a held handoff lock.

Includes a fix to the shared dashboard-card-link test's own server-start
wait: it previously trusted bin/fm-dashboard.sh's 1-second process-liveness
sleep as proof the server was ready, which flakes on a busier CI runner
(confirmed pre-existing on main, not introduced here). It now polls the
real /api/health reachability check before running any test.

* no-mistakes(review): bound dashboard curl calls, fix card-link identity, 404, and record races

* no-mistakes(review): confirm card links only on read state, keep unlinkable pairs

* no-mistakes(review): fail card-link ownership check closed, never link superseded cards

* no-mistakes(document): fix stale usage line for multi-item handoff

* fix(tests): show captured output on a failed exit-code assertion

expect_code only ever reported "expected exit 0, got 1", with no way to
see why - the same missing-evidence shape as the CI failure this is
meant to help diagnose. Accept the command's own captured output as an
optional argument and print it on failure, matching how
assert_contains/assert_not_contains already show theirs. Wired into
every expect_code call in tests/fm-dashboard-card-link.test.sh.

Also fold in the changed-file test-family mapping this branch's rebuild
had dropped: bin/fm-backlog-handoff.sh, bin/fm-spawn.sh, and
bin/fm-teardown.sh (the three entry points to the mechanical card link)
plus bin/fleet-dashboard/* now select the dashboard family, so the
suites protecting this change actually run when the files they protect
change.

* no-mistakes(review): sweep local card records on --resume-pending, reject zero timeout

* no-mistakes(review): hold back only still-staged pairs on --resume-pending

* no-mistakes(document): document handoff card record write and resume sweep

* no-mistakes(review): fail loudly on card record rewrite errors, clear stale warnings ledger

* no-mistakes(review): commit superseded marks after record write, gate ledger reset

* no-mistakes(review): reject separator-bearing --card ids before they reach the record

* no-mistakes(review): drop durable warnings ledger for in-process report suppression

* no-mistakes(review): gate unreadable-card warning through reason-keyed report mark

* no-mistakes(review): replace bash-4 assoc array with 3.2-portable report set

* no-mistakes(document): scope outbox recovery claims, fix card-link report bound

* fix(dashboard): stop flooding the fleet audit log for unretirable pairs

An unretirable pending pair (the board says a card id does not exist)
was writing a durable audit-log --fleet entry every time it was swept.
bin/fm-bootstrap.sh's --resume-pending sweep now runs unattended on
every session start, so a single permanently-unlinkable pair would
grow that entry on a cadence no operator controls - the exact
always-red-marker failure the fleet auditor's own design already
guards against.

The stderr warning stays exactly as built, still bounded in-memory per
process; only the fleet-log write is removed. Surfacing a permanently
unlinkable pair to the Admiral belongs to the fleet auditor's own
sweep, raised once as a finding, not to this handoff path on every boot.

* no-mistakes(review): handle unterminated final lines in handoff card state

* no-mistakes(document): explain unterminated-line handling in handoff card state

* fix(dashboard): never write the fleet audit log from --resume-pending

The no-such-card case already stopped writing to the fleet audit log
from bin/fm-bootstrap.sh's unattended --resume-pending sweep. This
applies the same rule to the general transport-failure case in
dashboard_link_card: a ref/agent/status write that fails still warns
on stderr everywhere, but only an operator-initiated invocation (no
--resume-pending) may write it to audit-log --fleet, and at most once
per invocation.

* no-mistakes(review): surface pending handoff card links at session start

* no-mistakes(document): document spawn/teardown fleet-log writes and card-link resume gating

* no-mistakes(document): note superseded pairs as a persisting card-link cause

* no-mistakes(ci): drain the denied arm attempt before flipping the lock

tests/fm-pi-watch-extension.test.sh's OpenCode session-lock case wrote a
foreign pid, fired session.idle, slept a fixed 120ms, then took the lock and
fired again. The event hook does not await its arm attempt, and that attempt
walks the process tree with one `ps` per level, so on a loaded runner the
denied attempt was still in flight at 120ms. ensureArm coalesces onto
launchInFlight, so the second event joined the stale read-only answer and
never armed - the shard's "expected exit 0, got 1".

Wait on the plugin's own coordinator handle instead of the clock: ensureArmed
joins an in-flight launch, so once it resolves no attempt is outstanding, and
its status is the denial under test. Also pass the node output to expect_code
so a future failure here shows why rather than an exit code alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#13)

* fix(bin): count a trust-registered custom check as supervision-needed

fm_supervision_status only looked for state/*.meta, state/x-watch.check.sh,
and state/procevent/*.source. A home whose only duty is a check registered
through bin/fm-check-register.sh (a state/<id>.check-trust binding paired
with its state/<id>.check.sh) matched none of those, so FM_SUP_NEEDED stayed
false and the Claude Stop auto-arm never armed a watcher for it.

Count a state/<id>.check-trust file paired with its state/<id>.check.sh as
supervision-needed, mirroring the existing presence-only style used for
process-event sources and the X-mode relay poll. Update the turn-end guard
and pull-warning banners to name a registered check specifically instead of
falling through to the X-mode wording, and update docs/turnend-guard.md and
its cross-references (docs/architecture.md, AGENTS.md, docs/scripts.md,
docs/subagent-guard.md, docs/supervision-protocols/grok.md, and
bin/fm-claude-stop-autoarm.sh's own header) to point at docs/turnend-guard.md
as this predicate's single owner rather than restating its input list.

* no-mistakes(document): note check-predicate rationale, generalize stale guard comments

* no-mistakes(document): record check-predicate verification evidence entry

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…14)

* fix(watcher): re-evaluate arm coalescing when the fleet lock changes

Two independent causes made the arm-readiness suite fail a different
assertion almost every run:

- ensureArm() in the OpenCode watch plugin reused a still-resolving
  earlier caller's beginArm() result unconditionally. Every ordinary
  session.idle produces two callers, so when the fleet lock was
  reacquired while an earlier attempt was mid-flight, the later
  caller inherited that attempt's stale read-only verdict and never
  armed. Fixed with premise-validated coalescing: a caller shares an
  in-flight attempt only while the lock file content it captured is
  still current; otherwise it evaluates fresh. Two callers on an
  unchanged lock still coalesce into one subprocess walk, so this
  costs nothing on the ordinary turn a serialized-everything fix
  would have doubled.

- Both adapters spawn their arm child through a login shell, which
  sources /etc/profile in addition to the account's own profile
  files; the system-wide half is not relocatable via HOME. main
  already raised the readiness timeout to absorb that cost (250ms to
  2000ms) and its own measurement shows a worst case of ~1740ms
  under contention against that budget - narrower headroom, not a
  removed confound. This adds FM_WATCH_ARM_NO_LOGIN_SHELL, a
  test-only opt-out that spawns the arm child under plain bash -c,
  removing the confound instead of padding around it. Production
  keeps the login shell as the unconditional default.

Also fixes three unhandled-EPIPE crash sites found while proving this
change under load: child.stdin.end() in fm-primary-turnend-guard.js,
fm-primary-turnend-guard.ts, and fm-operational-input.js raises an
unhandled 'error' event on the stdin stream (not the ChildProcess)
when the child exits before the write lands, which was crashing the
whole session process. Each site now no-ops that stream error since
the child's own close/error handlers already drive resolution.

Adds two regression tests (opencode coalescing on an unchanged lock,
login-shell default vs. opt-out for both adapters) and one EPIPE
regression test, and rewrites the existing OpenCode lock test to
force the stale-verdict race deterministically via a gated ps shim
instead of waiting for it. See docs/arm-readiness-determinism-proof.md
for the repeated-run proof and docs/configuration.md for the new
FM_WATCH_ARM_NO_LOGIN_SHELL entry.

Built directly on current main (45bd292); does not touch main's own
prior timeout raise or its test_pi_session_transition_generation_owner
fixture-ordering fix, both kept as-is.

* no-mistakes(review): correct stale arm-readiness proof claims and restore expect_code diagnostics

* no-mistakes(review): wait on observable arm rows and share gate shims

* no-mistakes(review): restore fallback diagnostics and bound unretired arm counts

* no-mistakes(review): correct lock-assertion cause attribution in proof doc

* no-mistakes(review): restore two-cause attribution and note login-shell residual bound

* no-mistakes(review): qualify stale readiness-window claim on login-shell test

* no-mistakes(review): correct stale readiness-budget claims in suite header

* no-mistakes(test): correct stale login-shell contention figures in determinism proof

* no-mistakes(document): scope proof-doc timeout raise and idle figure to sources

* no-mistakes(document): stamp determinism proof verification at a15d993 with provenance

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…not a hardcoded origin

A firstmate checkout can have "origin" pointing at the public template it was
forked from while its own branch tracks a separate fork it actually develops
on, and the two can genuinely diverge. bin/fm-update.sh (via bin/fm-ff-lib.sh's
ff_target) hardcoded origin/<default> as the fast-forward base, so a home
tracking a fork never advanced past the point the two remotes diverged.

Extracts the "which remote does this checkout actually develop on" resolution
into bin/fm-dev-remote-lib.sh (resolve_update_base: the branch's configured
upstream when set, falling back to origin/<default> and saying so otherwise)
and applies it everywhere a hardcoded origin assumption had the same failure
mode: bin/fm-spawn.sh's pooled-worktree base refresh (this is what stranded
this task's own worktree 28 commits behind on the wrong lineage, missing
bin/fm-dashboard.sh and fm-spawn.sh --card entirely - direct evidence caught
live), bin/fm-review-diff.sh's review base and PR-head fetch, bin/fm-lint.sh
and bin/fm-test-run.sh's --changed diff base, and bin/fm-teardown.sh's
landed-work fallback checks (ensure_commit_object, content_in_default).
bin/fm-fleet-sync.sh, bin/fm-merge-local.sh, bin/fm-home-seed.sh, and the
remote-home provisioning scripts were audited and left unchanged: they
correctly always use "origin" for project clones, which have no fork/upstream
split by design.

Recommends (does not apply - push/PR targeting is the captain's call):
repointing the remote code root's own "origin" at the fork it actually
develops on, matching the primary checkout; and, for no-mistakes PR
mergeability, keeping "origin" pointed at the fork for ordinary fleet work
and reserving the CONTRIBUTING.md --fork-url external-contributor recipe
(push fork, PR against origin) for a separate checkout used only for
genuine upstream contributions.

tests/fm-remote-update-follows-fork.test.sh proves the fix end to end: a
remote code root tracking a fork updates from it and a fork-only file
reaches the secondmate home, confirmed to fail against the unfixed code.
@joliverMI

Copy link
Copy Markdown
Author

Opened in error against upstream - this work belongs on our own fork and is being redirected there. Apologies for the noise.

@joliverMI joliverMI closed this Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants