Skip to content

feat(bin): restore fork fixes (Bitbucket Cloud PRs, Herdr containers, zsh) onto upstream main - #25

Merged
cloud-practitioner merged 21 commits into
mainfrom
fm/fm-fork-fixes-restore
Oct 1, 2026
Merged

cloud-practitioner merged 21 commits into
mainfrom
fm/fm-fork-fixes-restore

Conversation

@cloud-practitioner

@cloud-practitioner cloud-practitioner commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

Create a PR to merge all fixes from the closed (merged) PRs of the firstmate fork back into its main branch. Context: at 2026-10-01T03:52Z main of https://github.com/cloud-practitioner/firstmate was force-reset to upstream kunchenguid/firstmate main (aedb7bb), which removed every fork-only commit from main. The fork's old main tip 2c375b2 contains the merged fork PRs #1, #2, #7, #8, #9, #10, #11, #13, #14, #15, #16, #17, #20, #21, #22 and the upstream-merge PRs #18 and #19. Bring every one of those fixes back onto main while keeping upstream's new commits, as a "merge upstream" style PR like #18 and #19 so main moves forward with no force push. Where upstream already fixed the same thing, keep upstream's version and say so in the PR body. Closed PRs that were never merged (#3, #4, #5, #6, #12, #23) stay out.

The referenced fork PRs: #1 feat(bin): add Bitbucket Cloud as a first-class PR provider; #2 fix(backend): enable Herdr detection in devcontainers via HERDR_SOCKET_PATH fallback; #7, #8, #9, #10 Fix/herdr container detection; #11 feat(bin): support Bitbucket Cloud pull requests in the PR scripts; #13 chore: ignore .claude/settings.local.json; #14 fix(bin): let concurrent Herdr presentation recoveries wait for the session lock; #15 test: pin a neutral shell and nested agent shell in the Herdr control smoke test; #16 test: remove read-only strip hooks and staged launch dirs in fixture teardown; #17 fix(bin): load backend sibling libraries when fm-backend.sh is sourced from zsh; #18 and #19 merge upstream; #20 fix(bin): scope each task's temp root to its Firstmate home; #21 test: back Herdr live-duplicate agent registrations with a running process; #22 test: correct fixture helper comment on home-scoped task temp root. PR #12 is not included because its upstream version (kunchenguid#6103) is already on main as 46d58d6.

What Changed

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The fix-round commit 6201ec1 sends apply_pending_retained_artifact through the same boundary fm_backlog_done and fm_backlog_retain already use. That is fm_backlog_row_artifact_supported plus fm_backlog_pr_note_label. A retained Bitbucket, GitLab, or Gerrit URL that tasks-axi cannot link is now written as a labelled note on the captain-answer close, and supported GitHub/Forgejo URLs and reports still take the update path. No sibling site still special-cases only Gerrit.

Testing

This round focused on the scenario that failed in round 1. In a real throwaway Herdr lab, the zsh container case (HERDR_SOCKET_PATH only) now picks the lab session and a real workspace call reaches it. Running the same check against the pre-fix code still shows the old 'default' result. The guards also hold: a path that is not a socket is not detected as Herdr, an explicit HERDR_SESSION still takes priority, and PATH stays intact. At 4b2c1da I also re-ran the real-herdr container e2e test, the backend test file with its new zsh assertion, zsh backend sourcing, a live two-home Herdr spawn and teardown (each home keeps its own 0700 temp root, and tearing down one does not touch the other), live Bitbucket PR state reads and refusals, the full captain-hold lifecycle file, and the git ignore and ancestry checks. Everything passed. The lifecycle file takes about 15 minutes, so a first 600s timeout was length, not a hang. These are CLI changes with no UI, so the evidence is CLI transcripts. Lab sessions, homes and the treehouse pool were torn down, and the worktree is clean.

  • Live validation: ✅ go - 8 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A zsh shell that sources fm-backend.sh with only HERDR_SOCKET_PATH set targets the socket's session ✅ pass live herdr-socket-path-detection-4b2c1da.txt section 2 and section 6 (real workspace list on the lab session); the pre-fix contrast shows 'default'
A bash container caller with only HERDR_SOCKET_PATH detects herdr and binds to that socket's session ✅ pass live herdr-socket-path-detection-4b2c1da.txt section 1; herdr-container-session-e2e-4b2c1da.txt
Adversarial: a HERDR_SOCKET_PATH that is not a socket is not detected; an explicit HERDR_SESSION takes priority; a default socket resolves to default ✅ pass live herdr-socket-path-detection-4b2c1da.txt sections 3-5
Sourcing fm-backend.sh from zsh loads backend sibling libraries for every backend ✅ pass live zsh-backend-source-4b2c1da.txt
Two Firstmate homes spawning the same task id get separate home-scoped temp roots, and tearing down one leaves the other alone ✅ pass live task-temp-root-home-scoped-4b2c1da.txt
Bitbucket Cloud PR URLs are read through the PR scripts, with refusals for a missing token or a non-canonical URL ✅ pass live bitbucket-pr-state-4b2c1da.txt
A captain answer before cleanup replay records an unlinkable retained Bitbucket PR URL as a note (Gerrit keeps its own wording) ✅ pass live captain-hold-lifecycle-4b2c1da.txt (full test file drives the real fm-captain-hold.sh/fm-teardown.sh); manual round-1 transcript captain-hold-retained-bitbucket-note.txt
The merge moves main forward with no force push, keeps upstream commits, ignores .claude/settings.local.json, and leaves out never-merged PRs ✅ pass live gitignore-and-ancestry-4b2c1da.txt
Evidence: zsh/bash Herdr socket-path detection in a live lab at 4b2c1da

Source: zsh/bash Herdr socket-path detection in a live lab at 4b2c1da

# target commit: 4b2c1da
# lab session: fm-lab-detect2-2193055-32614
# lab socket: ~/.config/herdr/sessions/fm-lab-detect2-2193055-32614/herdr.sock  (is socket: yes)

## 1. container shape (bash): only HERDR_SOCKET_PATH forwarded, HERDR_ENV/HERDR_SESSION absent
detected=herdr signal=HERDR_SOCKET_PATH
herdr adapter loaded; session=fm-lab-detect2-2193055-32614

## 2. container shape (zsh sources fm-backend.sh)  [round-1 failure]
detected=herdr signal=HERDR_SOCKET_PATH
herdr adapter loaded; session=fm-lab-detect2-2193055-32614

## 3. adversarial: HERDR_SOCKET_PATH names a path that is not a socket
detected=none signal=none

## 4. adversarial (zsh): explicit HERDR_SESSION wins over the socket-derived name
session=fm-lab-explicit

## 5. zsh: default-session socket shape and PATH still intact after lookup
session=default
/usr/bin/dirname
/usr/bin/basename

## 6. real operational call from zsh container-shaped env reaches the lab server (workspace list)
resolved=fm-lab-detect2-2193055-32614
{"id":"cli:workspace:list","result":{"type":"workspace_list","workspaces":[]}}
## lab sessions now:
[{"name":"default","running":true},{"name":"fm-lab-detect2-2193055-32614","running":true}]
# teardown rc=0
Evidence: Real-herdr container session e2e at 4b2c1da

Source: Real-herdr container session e2e at 4b2c1da

# target 4b2c1da
ok - real herdr: resolved the isolated lab session's real control-socket path
ok - real herdr: fm_backend_herdr_session derives the exact lab session name from a real HERDR_SOCKET_PATH alone
ok - real herdr: container_ensure resolved purely from HERDR_SOCKET_PATH targets the exact real lab session
ok - real herdr: the container-shaped call reused the already-running server's exact socket, never starting a fresh disconnected one
exit=0
Evidence: zsh backend sibling loading at 4b2c1da

Source: zsh backend sibling loading at 4b2c1da

## target 4b2c1da (zsh)
fm_backend_source tmux rc=0; sibling fm_composer_strip_ansi: function; adapter fm_backend_tmux_capture: function
fm_backend_source herdr rc=0; sibling fm_composer_strip_ansi: function; adapter fm_backend_herdr_capture: function
fm_backend_source zellij rc=0; sibling fm_composer_strip_ansi: function; adapter fm_backend_zellij_capture: function
fm_backend_source orca rc=0; sibling fm_composer_strip_ansi: function; adapter fm_backend_orca_capture: function
fm_backend_source cmux rc=0; sibling fm_composer_strip_ansi: function; adapter fm_backend_cmux_capture: function
Evidence: Home-scoped task temp roots, live two-home Herdr spawn and teardown

Source: Home-scoped task temp roots, live two-home Herdr spawn and teardown

# target 4b2c1da; herdr lab session: fm-lab-tasktmp2-2505937-6124
# before: legacy shared root /tmp/fm-tt9 exists? no

## homeA: $ FM_HOME=<W>/homeA bin/fm-spawn.sh tt9 <project> "sh -c 'echo GOTMPDIR=$GOTMPDIR; sleep 60'" --mode no-mistakes --yolo off --backend herdr
warning: <W>/homeA/data/tt9/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
warning: herdr presentation parent is absent or ambiguous; using the ordinary flat layout without projection
spawned tt9 harness=sh kind=ship mode=no-mistakes yolo=off window=fm-lab-tasktmp2-2505937-6124:w1:p2 worktree=~/.treehouse/proj-84a894/1/proj
spawn rc=0
meta tasktmp=<W>/homeA/state/tt9.tasktmp
drwx------ node <W>/homeA/state/tt9.tasktmp
drwxr-xr-x node <W>/homeA/state/tt9.tasktmp/gotmp

## homeB: $ FM_HOME=<W>/homeB bin/fm-spawn.sh tt9 <project> "sh -c 'echo GOTMPDIR=$GOTMPDIR; sleep 60'" --mode no-mistakes --yolo off --backend herdr
warning: <W>/homeB/data/tt9/launch-brief.md records no ship branch; defaulting to legacy branch fm/tt9
warning: <W>/homeB/data/tt9/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
spawned tt9 harness=sh kind=ship mode=no-mistakes yolo=off window=fm-lab-tasktmp2-2505937-6124:w2:p2 worktree=~/.treehouse/proj-84a894/2/proj
spawn rc=0
meta tasktmp=<W>/homeB/state/tt9.tasktmp
drwx------ node <W>/homeB/state/tt9.tasktmp
drwxr-xr-x node <W>/homeB/state/tt9.tasktmp/gotmp
homeA pane output:
homeB pane output:

# after both spawns: shared /tmp/fm-tt9 exists? no

## homeA teardown
<W>/proj: skipped: no origin remote
teardown tt9 complete (window fm-lab-tasktmp2-2505937-6124:w1:p2, worktree ~/.treehouse/proj-84a894/1/proj)
Backlog: tt9 just finished (this home keeps no markdown backlog at <W>/homeA/data/backlog.md). Update <W>/homeA/data/backlog.md - move tt9 to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
teardown rc=0
homeA root exists? no
homeB root exists (must be yes)? yes

## homeB teardown
<W>/proj: skipped: no origin remote
teardown tt9 complete (window fm-lab-tasktmp2-2505937-6124:w2:p2, worktree ~/.treehouse/proj-84a894/2/proj)
Backlog: tt9 just finished (this home keeps no markdown backlog at <W>/homeB/data/backlog.md). Update <W>/homeB/data/backlog.md - move tt9 to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
teardown rc=0
homeB root exists? no
# herdr lab teardown rc=0
Evidence: Live Bitbucket PR state reads and refusals

Source: Live Bitbucket PR state reads and refusals

# target 4b2c1da
$ bin/fm-pr-state.sh https://bitbucket.org/iqxbusiness/supplier_online_orchestration_api/pull-requests/1
STATE: merged
exit=0
$ bin/fm-pr-state.sh https://bitbucket.org/iqxbusiness/ixia-marketplace/pull-requests/7
CHECKS: none reported yet
exit=0
$ bin/fm-pr-state.sh https://bitbucket.org/iqxbusiness/ixia-marketplace/pull-requests/4
STATE: declined
exit=0
# adversarial: missing credential
fm-pr-state: reading a Bitbucket pull request requires the NO_MISTAKES_BITBUCKET_API_TOKEN environment variable
exit=2
# adversarial: non-canonical URL
fm-pr-state: expected a GitHub or Bitbucket pull-request URL
exit=2
Evidence: Captain-hold lifecycle incl. retained Bitbucket PR note

Source: Captain-hold lifecycle incl. retained Bitbucket PR note

# target 4b2c1da: bash tests/fm-captain-hold-lifecycle.test.sh (full file, 15m03s, exit 0, 56 ok, 0 not ok)
ok - an answer before cleanup replay notes the retained Gerrit pull request
ok - an answer before cleanup replay notes the retained Bitbucket pull request
ok - cleanup keeps a captain-held Gerrit task open and records its change URL
Evidence: gitignore rule and merge ancestry at 4b2c1da

Source: gitignore rule and merge ancestry at 4b2c1da

$ git check-ignore -v .claude/settings.local.json
.gitignore:16:.claude/settings.local.json	.claude/settings.local.json
aedb7bb is an ancestor of HEAD 4b2c1da
2c375b2 is an ancestor of HEAD 4b2c1da
merge cad187e parents: 2c375b2c1444e3f85487abe12fde276b8e97422c aedb7bbf3b038672cef8b700df2b65ff4491dbfd 
# excluded unmerged PRs (#3 #4 #5 #6 #23) present on HEAD? (subjects with (#N) from fork range)
none
no-mistakes(test): Fix zsh Herdr socket session lookup shadowing PATH
no-mistakes(review): Note unlinkable retained PR URLs on captain answer
merge upstream
fix(bin): close Gerrit-landed backlog items with the change URL as a note (#6140)
feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (#4859)
fix(bin): keep the steering doorbell short under deep homes (#6240)
fix(bin): let a stale record on a reassigned slot retire records-only (#6213)
fix(bin): run no repository hook when core.hooksPath is empty (#6216)
Evidence: Pre-fix vs post-fix zsh lookup
pre-fix (HEAD~1 bin/): fm_backend_herdr_socket_session_name:4: command not found: dirname / basename -> default
post-fix (4b2c1da): session=fm-lab-detect2-2193055-32614
- Outcome: 🔧 2 issues found → auto-fixed ✅ across 2 runs (52m16s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ bin/fm-backlog-transition-lib.sh:541 - The merge resolution broadened the 'record an unlinkable PR as a note' rule. fm_backlog_done (bin/fm-backlog-transition-lib.sh:541) and fm_backlog_retain (:587-591) now handle any URL that fails fm_backlog_row_artifact_supported, labelled through fm_backlog_pr_note_label. The upstream sibling path in bin/fm-captain-hold.sh:960-964 (apply_pending_retained_artifact, which this branch carries unchanged from aedb7bb) still special-cases only Gerrit. Failing sequence: a captain-held Bitbucket task is torn down; cleanup is interrupted after it writes the mode=retain pending-close record (arg=--pr, arg=https://bitbucket.org/ws/repo/pull-requests/7) but before the deliverable line is written; the captain then answers before replay. fm_backlog_pr_is_gerrit_change returns false, then fm_backlog_row_artifact_supported fails the /pulls?/ regex, so the function returns 0 with RETAINED_CLOSE_ARGS empty. close_answered then runs a plain tasks-axi done, and replay retires the record. The Bitbucket (or GitLab) PR URL ends up nowhere on the closed item, while a Gerrit URL in the same spot gets a note. This contradicts fork PR feat(bin): add Bitbucket Cloud as a first-class PR provider #1/feat(bin): support Bitbucket Cloud pull requests in the PR scripts #11's rule that an unlinkable PR 'is still recorded on the closed item, as a note'. Fix at the same boundary the merge already uses: in fm-captain-hold.sh:960, replace the Gerrit-only test with [ &#34;${args[0]}&#34; = --pr ] &amp;&amp; ! fm_backlog_row_artifact_supported &#34;$id&#34; --pr &#34;${args[1]-}&#34; and set RETAINED_CLOSE_ARGS=(--note &#34;$(fm_backlog_pr_note_label &#34;${args[1]}&#34;) ${args[1]}&#34;). Same invariant, still worded Gerrit-only: docs/captain-hold-lifecycle.md:142 (only Gerrit 'appears only in the deliverable line') and :156 (only a 'retained Gerrit change URL' becomes a note on answer). Also add a Bitbucket variant next to the new upstream Gerrit case in tests/fm-captain-hold-lifecycle.test.sh.

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Test** - 2 issues found → auto-fixed ✅
  • ⚠️ bin/backends/herdr.sh:551 - fm_backend_herdr_socket_session_name declares local path=$1. In zsh, path is tied to PATH, so dirname and basename cannot be found and fm_backend_herdr_session silently returns 'default' for a zsh-sourced container caller. This bug was already on fork main 2c375b2, so the merge did not introduce it. Fix: rename the local, for example to socket_path, then extend the zsh case in tests/fm-backend.test.sh to assert that a named-session socket path resolves to that session's name.
  • 🚨 live validation verdict: no-go (11 of 11 scenarios were driven live against the product); failed: A zsh shell that sources fm-backend.sh with only HERDR_SOCKET_PATH set targets the socket's session
  • Live validation: ❌ no-go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Operator reads real Bitbucket PRs with fm-pr-state.sh; a missing token or non-canonical URL is refused ✅ pass live bitbucket-pr-state.txt
fm-pr-check.sh records a Bitbucket PR and arms a merge poll that stays silent while open and prints 'merged' once merged ✅ pass live bitbucket-pr-check-record.txt, bitbucket-pr-poll.txt
fm-pr-merge.sh refuses a declined Bitbucket PR before sending any merge request ✅ pass live bitbucket-pr-merge-refusal.txt
A container-shaped process with only HERDR_SOCKET_PATH detects Herdr and targets the socket's session (bash) ✅ pass live herdr-container-session-e2e.txt, herdr-socket-path-detection.txt
A zsh shell that sources fm-backend.sh with only HERDR_SOCKET_PATH set targets the socket's session ❌ fail live herdr-socket-path-detection.txt section 2
Sourcing fm-backend.sh from zsh loads each backend adapter (#17) ✅ pass live zsh-backend-source.txt
Equal task ids in two homes get private home-scoped temp roots, and teardown removes only its own (#20) ✅ pass live task-temp-root-home-scoped.txt
Concurrent Herdr recoveries serialize on the session lock, a stuck holder is refused (#14), and a live duplicate agent refuses launch (#21) ✅ pass live herdr-presentation-e2e.txt
A captain answer before an interrupted cleanup replays keeps 'PR <bitbucket url>' on the closed item ✅ pass live captain-hold-retained-bitbucket-note.txt
.claude/settings.local.json is ignored (#13) ✅ pass live gitignore-and-ancestry.txt
Main moves forward without a force push and the excluded PRs stay out ✅ pass live gitignore-and-ancestry.txt
  • bin/fm-pr-state.sh on real Bitbucket PRs (merged, open, declined) plus a missing-token refusal and a non-canonical URL refusal
  • FM_HOME=&lt;lab&gt; bin/fm-pr-check.sh &lt;id&gt; &lt;bitbucket url&gt; for an open PR and a merged PR, then ran each armed state/&lt;id&gt;.check.sh
  • FM_HOME=&lt;lab&gt; bin/fm-pr-merge.sh bbdeclined https://bitbucket.org/iqxbusiness/ixia-marketplace/pull-requests/4 (refusal only)
  • bash tests/fm-backend-herdr-container-session-e2e.test.sh
  • Own fm-lab-* Herdr session via bin/fm-herdr-lab.sh: detection and session resolution under bash and zsh
  • Real bin/fm-spawn.sh tt1 ... --backend herdr in two lab homes with the same task id, then bin/fm-teardown.sh tt1 in each
  • Real fm-captain-hold.sh hold -> interrupted fm-teardown.sh --force -> fm-captain-hold.sh answer with real tasks-axi, for Bitbucket at 6201ec1 and pre-fix cad187e, plus Gerrit and GitHub
  • zsh -c &#39;source bin/fm-backend.sh &amp;&amp; fm_backend_source &lt;backend&gt;&#39; at the target commit and at base aedb7bb
  • bash tests/fm-backend-herdr-presentation-e2e.test.sh (24 ok, exit 0)
  • bash tests/fm-pr-bitbucket.test.sh (19 ok)
  • git check-ignore -v .claude/settings.local.json; git merge-base --is-ancestor for aedb7bb and 2c375b2

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 8 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A zsh shell that sources fm-backend.sh with only HERDR_SOCKET_PATH set targets the socket's session ✅ pass live herdr-socket-path-detection-4b2c1da.txt section 2 and section 6 (real workspace list on the lab session); the pre-fix contrast shows 'default'
A bash container caller with only HERDR_SOCKET_PATH detects herdr and binds to that socket's session ✅ pass live herdr-socket-path-detection-4b2c1da.txt section 1; herdr-container-session-e2e-4b2c1da.txt
Adversarial: a HERDR_SOCKET_PATH that is not a socket is not detected; an explicit HERDR_SESSION takes priority; a default socket resolves to default ✅ pass live herdr-socket-path-detection-4b2c1da.txt sections 3-5
Sourcing fm-backend.sh from zsh loads backend sibling libraries for every backend ✅ pass live zsh-backend-source-4b2c1da.txt
Two Firstmate homes spawning the same task id get separate home-scoped temp roots, and tearing down one leaves the other alone ✅ pass live task-temp-root-home-scoped-4b2c1da.txt
Bitbucket Cloud PR URLs are read through the PR scripts, with refusals for a missing token or a non-canonical URL ✅ pass live bitbucket-pr-state-4b2c1da.txt
A captain answer before cleanup replay records an unlinkable retained Bitbucket PR URL as a note (Gerrit keeps its own wording) ✅ pass live captain-hold-lifecycle-4b2c1da.txt (full test file drives the real fm-captain-hold.sh/fm-teardown.sh); manual round-1 transcript captain-hold-retained-bitbucket-note.txt
The merge moves main forward with no force push, keeps upstream commits, ignores .claude/settings.local.json, and leaves out never-merged PRs ✅ pass live gitignore-and-ancestry-4b2c1da.txt
  • Live Herdr lab (bin/fm-herdr-lab.sh name/provision/teardown, session fm-lab-detect2-*): in bash and in zsh, env -i ... HERDR_SOCKET_PATH=&lt;lab sock&gt; then source bin/fm-backend.sh, run fm_backend_detect and fm_backend_herdr_session, and run a real herdr workspace list against the resolved session
  • Adversarial cases in the same lab: HERDR_SOCKET_PATH pointing at something that is not a socket, an explicit HERDR_SESSION taking priority, and a default-session socket shape with dirname/basename still on PATH
  • Ran the new regression command against the pre-fix bin/ from HEAD~1 (fails: dirname/basename not found, resolves 'default'), then at HEAD (passes)
  • timeout 300 bash tests/fm-backend-herdr-container-session-e2e.test.sh (real herdr)
  • bash tests/fm-backend.test.sh (includes the new zsh named-session assertion)
  • zsh fm_backend_source for tmux/herdr/zellij/orca/cmux, checking that the sibling and adapter functions load
  • Live two-home spawn in a Herdr lab: FM_HOME=&lt;labA|labB&gt; bin/fm-spawn.sh tt9 ... --backend herdr, then bin/fm-teardown.sh tt9 --force for each home
  • Live Bitbucket API: bin/fm-pr-state.sh on merged, open and declined PRs, plus the missing-token and non-canonical-URL refusals
  • bash tests/fm-captain-hold-lifecycle.test.sh (whole file, includes the Bitbucket retained-PR answer case)
  • git check-ignore -v .claude/settings.local.json, git merge-base --is-ancestor for aedb7bb and 2c375b2, and a check that the never-merged PRs #3/#4/#5/#6/#23 are absent
⚠️ **Document** - 1 info
  • ℹ️ docs/configuration.md:534 - Fork PR Fix/herdr container detection #10 restored a sentence saying the /afk supervisor's backend detection falls back to HERDR_SOCKET_PATH in containers. The sentence also had a stray ', ,'. But discover_supervisor_backend in bin/fm-supervisor-target-lib.sh has never checked HERDR_SOCKET_PATH, on the fork or upstream. I changed the doc to match the code, using upstream's wording. As a result, an /afk supervisor in a container with only HERDR_SOCKET_PATH still falls back to tmux unless FM_SUPERVISOR_BACKEND/FM_SUPERVISOR_TARGET are set. If supervisor container detection was meant to ship, that is a separate code change.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

cloud-practitioner and others added 21 commits September 29, 2026 17:07
* Fix Herdr backend detection in devcontainers

Add fallback detection via HERDR_SOCKET_PATH when HERDR_ENV=1 is not
injected into container environments. This handles the case where Herdr
launches crewmates or Pi processes inside devcontainers with an accessible
socket file.

The fix preserves all existing behavior:
- Normal Herdr shells: detects via HERDR_ENV=1 (unchanged)
- Containers with socket: detects via HERDR_SOCKET_PATH (new)
- No Herdr markers: falls back to default tmux (unchanged)

* no-mistakes(review): Document HERDR_SOCKET_PATH fallback in backend detection comments

* no-mistakes(document): Document HERDR_SOCKET_PATH container fallback in user-facing docs

* Fix herdr container session binding gap and add cmux precedence test

fm_backend_detect()'s HERDR_SOCKET_PATH fallback correctly reports the
herdr backend for a container that has only a forwarded/mounted socket,
but every downstream operational call (fm_backend_herdr_session and
everything built on it) resolved its target purely by HERDR_SESSION,
defaulting to "default" independent of that socket. Without an
equally-forwarded HERDR_SESSION naming the same session, a container's
herdr binary could silently start a brand-new, disconnected server
instead of reaching the real one - a worse failure mode than the
pre-fix broken-tmux fallback because it fails silently.

fm_backend_herdr_session() now derives the session name directly from
HERDR_SOCKET_PATH's verified sessions/<name>/herdr.sock shape when
HERDR_SESSION is unset, so detection and every operational call
provably agree on the same session and socket. An explicit
HERDR_SESSION still wins outright, and an unrecognized socket shape
(or the default session's own no-"sessions"-component socket) still
falls back to "default" exactly as before.

Also add the missing precedence test noted in review: the
HERDR_SOCKET_PATH fallback still wins over CMUX_WORKSPACE_ID.

Test coverage: pure-function unit tests for the new derivation logic
in tests/fm-backend-herdr.test.sh, the cmux precedence case in
tests/fm-backend.test.sh, and a new real-herdr end-to-end suite
(tests/fm-backend-herdr-container-session-e2e.test.sh) proving against
a live herdr binary that a devcontainer-shaped env with only
HERDR_SOCKET_PATH set resolves to and reaches the exact already-running
session, never a fresh disconnected one.

* Update architecture.md

---------

Co-authored-by: fm-bb-pr-support <fm-bb-pr-support@firstmate.local>
* fix: add .claude/settings.local.json to .gitignore

* feat(bin): support Bitbucket Cloud pull requests in the PR scripts

Parse https://bitbucket.org/<workspace>/<repo>/pull-requests/<n> URLs and
read, watch, record, merge, and prove landed Bitbucket Cloud pull requests
through the REST API with curl and jq, under the no-mistakes pipeline's own
NO_MISTAKES_BITBUCKET_EMAIL and NO_MISTAKES_BITBUCKET_API_TOKEN credential,
which reaches curl only on stdin.

The merge keeps every guard: open and not a draft, every unwaived build
green at the live head, the destination branch's required build count met,
an unreadable restriction set refused, attended-only waivers, captain-hold
and lock handling, a final head re-read inside the away-record lock because
the API cannot bind a head, and a read-back proof of the merged outcome.

* no-mistakes(review): Enforce Bitbucket merge checks and resolve review heads live

* no-mistakes(review): Map Bitbucket --rebase and fix review-diff fallback warning

* no-mistakes(test): Record non-GitHub/Forgejo PR links as backlog close notes

* no-mistakes(document): Fix stale forge wording and note unproved pr-forge member

* no-mistakes(document): Record nine-member pr-forge isolation proof

* no-mistakes(review): Narrow Bitbucket live-head warning assertion; drop .gitignore line
…ession lock (#14)

* fix(herdr): wait for a live session lock holder instead of refusing after 5s

Concurrent recoveries of one Herdr session from different homes serialize
on the shared presentation lock, and the holder keeps it across its whole
placement, including worktree allocation and launch. The fixed 50 x 0.1s
retry therefore refused the second recovery on slow machines.

fm_backend_herdr_presentation_session_lock_acquire is now the one wait
policy for spawn, teardown, and the serialized pane close: it waits for a
live holder and restarts its bounded per-holder wait
(FM_HERDR_PRESENTATION_LOCK_HOLDER_WAIT, default 120s) whenever the lock
changes hands, and refuses, naming the holder, only when one holder
outlasts it.

* no-mistakes(review): Limit long Herdr lock wait to concurrent recovery

* no-mistakes(document): Clarify Herdr recovery refusal follows bounded lock wait

* no-mistakes(ci): Root cause: the failing script in shard 7 was tests/fm-spawn-compact-adviser-disable-remote.test.sh. It failed with `fatal: failed to copy file to .../.git/objects/4b/...: No such file or directory`, followed by 'could not clone the remote Firstmate home'. This PR's Herdr lock change did not cause it; it is a flake in the test fixtures. The fixture tars the repo tree (about 730 objects), runs `git init`, commits, and then immediately does a local `git clone` of that repo (inside fm-remote-home-provision.sh). With Git 2.55 (the version on the runner), that commit starts a detached `git maintenance run --auto`, which packs the loose objects and deletes them while the clone is still copying them file by file. I reproduced this locally with GIT_TRACE: the loose objects were packed within about 2 seconds. Invariant: git must never rewrite a fixture repository in the background while a test is still reading it. Several suites share this pattern (whole-tree commit followed by a clone or seed): the remote-secondmate suites, bootstrap-network-parallel, cursor-primary, teardown, and others. So I fixed it at the one shared boundary every fixture git process already passes through, tests/git-config-helpers.sh. Its global layer now points at a new tracked file, tests/git-fixture.gitconfig (`[maintenance] auto = false`), instead of /dev/null. The host's global config is still ignored and the system layer stays off. Explicit `-c`, GIT_CONFIG_COUNT and a GIT_CONFIG_GLOBAL set by the caller all still take precedence. I added the new file to the changed-file map in bin/fm-test-run.sh, next to git-config-helpers.sh. Regression test: tests/fm-test-fixtures.test.sh gains `test_fixture_commit_starts_no_background_maintenance`. It makes a real 300-object fixture commit under GIT_TRACE, asserts that git started no `maintenance run --auto` or `gc --auto`, and then clones the fixture locally. It failed with the helper change stashed and passes with it. Checks run: fm-test-fixtures, fm-spawn-compact-adviser-disable-remote, fm-git-strip-ai-trailers, fm-test-isolation-proof and fm-remote-secondmate-parent-binding all pass. `fm-test-run.sh --check-coverage` is OK, `--changed` maps the new file, and bin/fm-lint.sh is clean
* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d from zsh (#17)

* fix(bin): make fm_backend_source sibling checks work under zsh

* no-mistakes(review): Rename zsh-special local path in fm_backend_source

* merge upstream (#19)

* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* no-mistakes(review): Load backend sibling libraries under zsh and tighten zsh test

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
… smoke test (#15)

* test(herdr): pin a neutral shell prompt in the control smoke pane

The real-Herdr control smoke test runs its task pane in the developer's
login shell. A prompt drawn with claude's composer glyph, such as a
starship-style `❯` zsh prompt, sits in the viewport as a bare agent
composer row holding text, so the final exit case read the composer as
pending and refused with the pending-text message instead of the
not-proven-empty one the test asserts. Both refusals are safe, but the
assertion depended on the developer's prompt rather than on the product.

Replace the pane shell with an rc-free bash using a shell-glyph prompt
and clear the screen before any case runs, so the composer verdict is
set by the test alone.

* no-mistakes(test): Stop the smoke test's shell from writing to bash history

* merge upstream (#19)

* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* no-mistakes(test): Run Herdr smoke agent under nested shell to keep registration

* no-mistakes(document): Note nested-shell requirement for Herdr stale-registration guard

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…teardown (#16)

* fix(tests): remove read-only strip hooks and staged launch dirs in test cleanup

The Herdr presentation e2e suite removed its temp root with a plain rm -rf,
which fails on the read-only state/<id>.git-hooks directories that the real
spawn installs, so every run leaked its /tmp tree. Spawns the suite never
tears down also left their /tmp/fm-<id>+<home-hash> launch directories behind.

Move fm_test_remove_tree into tests/fixture-tree-helpers.sh, sourced by both
tests/lib.sh and tests/herdr-test-safety.sh, and add
fm_test_remove_spawn_launch_dirs, which removes only the launch directories
scoped to the fixture's own homes. Use them in the Herdr suites that spawn
for real, assert in the presentation suite that nothing is left behind, and
cover both helpers in tests/fm-test-fixtures.test.sh.

* no-mistakes(review): Fix macOS /tmp find, concurrent flake, root guards

* no-mistakes(review): Reword launch-dir helper comment about per-task tmp root

* no-mistakes(review): Remove staged launch dirs in shared lib.sh test cleanup

* no-mistakes(review): Keep launch-dir cleanup helper safe under set -e

* no-mistakes(document): Document fixture-tree-helpers.sh in CONTRIBUTING shared helpers

* merge upstream (#19)

* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* no-mistakes(document): Fix shared test helper list wording in CONTRIBUTING

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
* fix(bin): scope each task's temp root to its Firstmate home

fm-spawn put every task's temp root at /tmp/fm-<id>, a path not tied to
the Firstmate home. Test spawns that were never torn down stranded it in
/tmp, and two homes spawning one task id (a main home and a secondmate,
or a test lab and the fleet) shared one root, so one home's teardown
deleted the other's live root.

The root now lives at state/<id>.tasktmp/ in the spawning home, next to
the other per-task state directories, so equal ids in different homes
never meet and a disposable home takes its roots with it. The pane's
GOTMPDIR export is shell-quoted now that the path follows the home.

A live task that recorded the legacy /tmp/fm-<id> keeps that root on
relaunch, so its record stays valid and teardown removes the root it
used. The live lab no longer removes /tmp/fm-<id> for its own ids.
bin/fm-spawn.sh's header owns the path contract.

* no-mistakes(review): Drop legacy child temp-root removal from forced teardown

* no-mistakes(review): Scope relaunch test fixtures' tasktmp roots to their homes
…ocess (#21)

* test(herdr): back the smoke live-duplicate agent with a running process

The duplicate-live-label case in tests/fm-backend-herdr-smoke.test.sh
registered an agent with `herdr pane report-agent` on a pane idling at
its top shell, then expected create_task to refuse the same label.
Herdr 0.9.3 releases such a registration within about a second once the
pane's foreground is its own top shell (a nested shell or any other
foreground process keeps it), so the case raced that release and failed
whenever create_task read the pane after it: agent_not_found classifies
the pane as an agent-free husk and create_task correctly replaces it.

The stale side is the test, not the backend. A pane whose only
registration sits over its bare top shell is not a live agent, and the
backend's husk classification is right to close and replace it; a real
agent holds the pane's foreground, so Herdr keeps its registration and
the backend still classifies it live and refuses. Older Herdr kept the
report, so the case only passed by reading it as a stale registration,
which the husk check also refuses.

The case now starts an agent-named foreground process (a `claude`
symlink to `sleep`, the fixture tests/fm-control-herdr-smoke.test.sh
already uses) before reporting, and asserts the pane classifies live
before the refusal check, so it proves the live-agent refusal rather
than passing on a stale or unreadable registration. The 0.9.3
measurement is recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Align smoke test comment with measured Herdr 0.9.3 behavior

* test(herdr): back the respawn-idem live-duplicate agent with a running process

tests/fm-backend-herdr-respawn-idem-e2e.test.sh made the same stale
assumption as the smoke suite's live-duplicate case: it registered an
agent with `herdr pane report-agent` on a pane idling at its top shell
and expected create_task to refuse the duplicate label. Herdr 0.9.3
releases that registration within about a second, so the case raced the
release and failed 2 of 6 runs against the installed Herdr 0.9.3.

Apply the same test-side fix: start an agent-named foreground process
(a `claude` symlink to `sleep`) before reporting, and assert the pane
classifies live before the refusal check. The backend's classification
is unchanged and correct.
Restore the fork's merged fixes onto upstream main (aedb7bb).

Conflict in bin/fm-backlog-transition-lib.sh: keep the fork's rule that any
pull request tasks-axi cannot link on a row (Bitbucket, GitLab, Gerrit) is
recorded on the close as a note, and keep upstream's 'Gerrit change <url>'
note wording for a Gerrit change.
…/fm-teardown.sh in Lint 2 (reason=memory rc=251, rss_kib=8390420). It hit the roughly 8 GiB usable heap left under the default per-root cap FM_LINT_ROOT_MEMORY_KIB=12582912 (12 GiB of address space). This PR did not introduce the failure: upstream aedb7bb on its own fails identically when pushed to fork main (run 36812483689, Lint 1, fm-teardown.sh reason=memory rss_kib=8390500). Upstream's own recent passing CI puts that root at 8275532 KiB, just under the limit. Locally, uncapped peak RSS was 8.39 GB at base and 8.54 GB at HEAD, so this PR's Bitbucket additions to the teardown source graph push it over consistently. Invariant: each canonical lint root must fit within the per-root address-space cap. That cap is set in one place only: the ROOT_MEMORY_KIB default in bin/fm-lint.sh, which CI and the pre-push gate both use. The lint comment rules out narrowing source-following, and the code already raised the cap once from 8 to 12 GiB when roots hit it, so I followed that precedent. Change: - bin/fm-lint.sh: default cap raised to 16777216 (16 GiB, about 10.7 GiB usable heap). I updated the header sentence and the sizing rationale with the new measurement and kept the earlier history. CI still lints one root per job, so worst-case resident memory is about 10.7 GiB plus overhead on the 16 GiB public runner. - tests/fm-lint.test.sh: the pinned-memory test's expected sidecar value is now 16777216. Verification: - Before the fix, `FM_LINT_REQUIRE_BOUNDS=1 FM_LINT_JOBS=1 bin/fm-lint.sh bin/fm-teardown.sh` reproduced the CI failure locally: rc=251, `shellcheck: out of memory`. - After the fix, the same command over bin/fm-teardown.sh, bin/fm-lint.sh and tests/fm-lint.test.sh exits rc=0 with no findings. - The full tests/fm-lint.test.sh suite passes with rc=0, including the memory-limit refusal and the pinned-memory default cases
@cloud-practitioner
cloud-practitioner merged commit 76a9f28 into main Oct 1, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant