Skip to content

feat(bin): sync upstream Firstmate into the fork and extend CI supersession tests - #27

Merged
BohnBawerick merged 202 commits into
mainfrom
fm/fm-upstream-fork-ci-sync
Sep 18, 2026
Merged

BohnBawerick merged 202 commits into
mainfrom
fm/fm-upstream-fork-ci-sync

Conversation

@BohnBawerick

@BohnBawerick BohnBawerick commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Upstream commits incorporated

Merges upstream main at 888871de5cdf875ba4f4c0d231da6efdf7bad9a8 into the fork, retaining both histories.
That adds 44 upstream-only commits relative to local main 425591d85757d6efdcf72a05002fae45c02baf66.
The reconciliation merge is 0f30207e94e09769af5b14ffb430edf75557c244; both parents are preserved and no history was rewritten.

Upstream has since moved on to 9bc051ff (three further commits: kunchenguid#4775, kunchenguid#4799, kunchenguid#4800). Those are deliberately not in this PR and remain follow-up work.

Fork-only commits preserved

Local main carries 153 commits absent from upstream and 146 absent from the fork base 308bdfab525e28d11edc16953d872cb8e5d3bf4c.
Head 7845743954b4c6a2e454aa1158532391c6c2b564 is 202 commits ahead of that fork base across 438 changed paths, with 630 tracked paths in the resulting tree.
308bdfab, 425591d8, 888871de, and 0f30207e are all confirmed ancestors of the head.

Preserved fork behavior includes the strict idle-shell ownership proof, complete intent with separate provenance labels, the quality controls, agy integration, session ownership, away supervision, and bounded test execution.
Run selection adopts upstream's creation-ordered, ID-addressed reads under the accepted reconciliation decision.

CI-savings coverage added

Carries the missing regression from 1c03c6ea576af387389df566c0e3ce92e90df239. The cancellation behavior itself already existed and is unchanged; only the executable proof was missing.
The workflow semantic tests now cover that ordinary pull-request CI cancels only superseded runs, while distinct PRs, non-PR events, no-mistakes body-event groups, release work, and every required result name survive.
Six deliberate workflow mutations are rejected by the regression.
Firstmate CI stays parallel; nothing is serialized behind lint.

Private-path audit

All 190 incoming commit trees were inspected: 631 unique paths, 2327 unique blobs. The 11 later pipeline commits were re-audited at each new head.
No .env, data/, state/, config/, projects/, .no-mistakes/, or other ignored material appears anywhere in the publication. No fleet-private records are present.
Credential-pattern scanning returns only an explicit synthetic AWS fixture already present in the test suite.
The final audit at head 78457439 reports private_path_hits: [], ignored_path_hits: [], and credential_pattern_candidates: [], with the upstream merge still an ancestor.

Validation

Documentation audience check, bin/fm-lint.sh, the focused workflow test, coverage partition guard, run-selection and ownership regressions, and the full watcher, turn-end, and workflow batch all pass locally.
The no-mistakes pipeline ran Intent, Rebase, Review, Test, Document, Lint, Push, PR, and CI; its attestation and per-step evidence are preserved in full below.
All 14 GitHub checks pass.

Four CI regressions surfaced by the merge were fixed under explicit instruction not to weaken, skip, or delete any test, and none was removed:

  • tests/fm-agy-harness.test.sh — the shell-only verdict now satisfies the fork's strict idle-shell proof, matching the fixture shape tests/fm-backend-herdr.test.sh already uses.
  • tests/fm-session-lock-ownership.test.sh — upstream fix(bin): let non-owner Claude Stops exit safely kunchenguid/firstmate#4777 changed --claude mode, so the fork's report-once decline now covers only non-Claude harnesses. The block-budget assertion an earlier timed-out agent had deleted was restored.
  • tests/fm-afk-inject-self-deadlock-e2e.test.sh — fixed a shutdown-flush race against the shared pidfile.
  • .pi/extensions/lib/fm-branch-dispatch.ts — replaced ES2023 Array.findLast with filter(...).at(-1) so the typecheck passes against its ES2022 target.

Two workflow values changed, both evidence-backed: the snapshot count moved 18 → 19 because the fork-only test_large_backlog_does_not_hit_jq_argument_limit is a real snapshot consumer that passes under bash 3.2, and the two portable parallel caps moved 10 → 15 minutes because the shards grew after the merge (lane 1 measured 586 s of script time and was cancelled 0.5 s after its last script passed). No test hangs; the 1.5x margin matches the serial shards.

Post-merge steps owned by firstmate

The worker does not perform any of these. It has not updated the primary copy, pushed to fork main, merged this request, restarted either mate, or run any fleet update script.

  1. Merge with FM_HOME="$PWD" bin/fm-pr-merge.sh fm-upstream-fork-ci-sync https://github.com/BohnBawerick/firstmate/pull/27 -- --merge from the primary Firstmate home, once the required checks and merge authority pass. Use a merge commit so both histories stay reachable.
  2. Fetch origin in the primary checkout and read the merged origin/main commit. On the primary, the work-PC code root, and its persistent secondmate home, verify a clean checkout and git merge-base --is-ancestor HEAD <merged-commit> before advancing anything. Resolve any divergence without dropping commits; the upstream updater's reset --keep path for redundant secondmate divergence is not authorized by this task.
  3. Run FM_HOME="$PWD" bin/fm-update.sh from the primary home, following .agents/skills/updatefirstmate/SKILL.md. Confirm the primary and work-PC update lines report the same target commit.
  4. Re-read AGENTS.md when the updater prints reread-firstmate: yes. Pass every ID from restart-secondmates: to FM_HOME="$PWD" bin/fm-secondmate-restart.sh <ids>; that persist-gated command must confirm a restart, since an unreachable or merely nudged mate is not a completed reload. Apply the documented fm-send.sh re-read message only to the separate nudge-secondmates: list.
  5. Verify the primary, remote code root, and work-PC persistent home report the same full commit ID, and that the replacement secondmate is running. Keep skipped, dirty, divergent, or unreachable targets open for reconciliation.

Intent

Captain intent authorized for --intent

Run a complete Firstmate update. First check the main upstream repository and bring in the updates that belong in our copy. Then upload our complete updated copy to the BohnBawerick fork. Make sure the primary first mate and the work-PC second mate end on the same level. Finally, apply the approved GitHub CI savings. Do not force, stash, discard, or lose any local or fork work.

Firstmate specification authorized for --intent

Before editing, read and follow ~/labs/axi-sandbox/firstmate/.agents/skills/firstmate-coding-guidelines/SKILL.md, CONTRIBUTING.md, and the repository's upstream-contribution and no-mistakes instructions. Work only in this isolated copy. The current evidence at intake is: local main is clean at 425591d85757d6efdcf72a05002fae45c02baf66; it contains the previous upstream merge and is 146 commits ahead of origin/main; a fresh fetch advanced upstream/main to 888871de5cdf875ba4f4c0d231da6efdf7bad9a8, leaving 44 upstream-only and 153 local-only commits; fm/fm-actions-savings contains the focused CI regression commit 1c03c6ea.

Fetch both remotes again and rebuild that graph from authoritative refs before acting. Inventory every origin-only, upstream-only, and local-only path. Verify that the intended fork publication contains no .env, private data/, state/, config/, projects/, .no-mistakes/, credentials, fleet-private records, or other ignored material. Stop if any private or uncommitted content would be lost or published.

Merge current upstream/main into the task branch using the repository's established upstream reconciliation pattern. Preserve all intentional fork-specific commits and both histories. Never force-push, rewrite shared history, rebase away local commits, stash, reset, or discard. Resolve conflicts from the authoritative owner and test evidence rather than choosing a side wholesale. Review the complete resulting diff and commit range against origin/main.

Carry the pending CI-savings regression from 1c03c6ea only if current upstream and the merged tree still lack equivalent coverage. The cancellation behavior itself already exists, so do not duplicate it. The intended remaining result is the smallest executable regression that proves ordinary pull-request CI cancels only superseded runs while preserving distinct PRs, non-PR events, no-mistakes body-event groups, release work, and every required result. Keep Firstmate CI parallel; do not serialize behind lint.

Run the repository documentation audience check, bin/fm-lint.sh, the focused workflow test, and the complete required local checks. Then run the no-mistakes pipeline. Its broad-commit rebase finding may be expected because the captain explicitly requested publishing the complete intended local copy to the fork, but do not answer your own finding; route it to firstmate with the exact commit counts and path audit evidence.

Open one PR against the BohnBawerick fork's main. The PR body must separate: upstream commits incorporated, fork-only commits preserved, CI-savings coverage added, private-path audit, validation, and the exact post-merge steps firstmate must run to update the primary and work-PC second mate. Do not update the primary copy, push directly to fork main, restart mates, or run fleet update scripts from this isolated worker. Firstmate owns the guarded merge and update after the PR passes.

Budget: finish within 120 tool calls and about 140k tokens. If the reconciliation exceeds that budget or any history question remains ambiguous, write partial state, the exact graph, conflicts, and next commands to /tmp/fm-upstream-fork-ci-sync/partial.md, append one blocked: line pointing there, and stop. Return a structured summary under Session Intent, Files Modified, Files Read, Decisions Made, Current State, and Next Steps; keep it under 500 words plus file references.

Later accepted Firstmate clarification

Adopt upstream creation-ordered, ID-addressed run selection, because the newest verified run is the authoritative attempt and a newer failure must not be hidden by an older live record.
Retain the fork strict ownership boundary: an unfetched head needs submitted-head or active-custody proof, and adjacent ledger rows never prove ownership.
Preserve the fork intent-provenance contract.
Add behavioral regressions proving: newer verified failure outranks older live; newer verified live outranks older terminal; unrelated or unproved terminal records cannot win; adjacent rows confer no ownership.
Resolve the merge conflicts against those rules and continue the brief.

Later accepted Firstmate clarification on watcher reconciliation

An elapsed declared wait rechecks with the declared clearing time has passed before generic wedge escalation.
Update that regression to assert the expired-wait recheck while retaining separate coverage for verified-working suppression and non-paused wedge escalation.
Fix the fake sleep fixture so its child is always reaped. Run the focused watcher regression, the complete watcher-triage and turn-end guard suites, bin/fm-doc-audience-check.sh, bin/fm-lint.sh, the focused workflow suite, bin/fm-test-run.sh --check-coverage, and git diff --check.

What Changed

  • Merges upstream main into the fork base 308bdfab in two steps. 425591d brings in upstream through b182d0f (145 commits). 0f30207 brings in upstream through 888871de (44 commits). Both merges keep the fork-only work: curated memory, quality gates, local landing, agy support, strict ownership proofs, and intent provenance. From upstream, this adds the OMP primary extensions, the Claude firstmate-calm mod, per-harness adapter references (including Gemini and Rovo), and new bin/ subsystems: fm-mail, fm-extension bindings, fm-secondmate-restart, fm-contributions, fm-dispatch-resolve, fm-quota-choose, and the Claude and agy trust helpers.
  • Takes upstream's CI savings in .github/workflows/ci.yml: per-PR concurrency that cancels only superseded pull_request runs, timeouts on every job, a fifth portable serial shard, and a Pi package install that fails the lane if the typecheck gate skips. tests/fm-ci-workflow.test.sh now uses a full expression parser and adds four regressions. Push, workflow_dispatch, schedule, and release runs keep independent groups that are never cancelled. no-mistakes-required body events keep their own groups. Each PR head reports the complete result set. The CI suite does not wait for lint.
  • Adds ten no-mistakes review commits on top of the merges. They fix teardown ordering and ownership, restart safety and acknowledgments, mail polling, lifecycle races, durable delivery recovery, and the refresh and polling rules for closed and merged contributions. They also strip reply correlation metadata, sync docs and comments with these fixes, and clear lint findings.

Risk Assessment

⚠️ Medium: The latest fix round is small and correct, but the whole change is a 199-commit, 438-file upstream plus fork merge touching teardown abort ownership, process-event locking, and secondmate restart gates, so it carries a broad, only partly verifiable behavioral surface.

Testing

I ran 16 focused suites through bin/fm-test-run.sh: CI workflow, watch-triage, turn-end guard, crew-state, teardown, secondmate restart, remote secondmate lifecycle, remote reply, procevent, extension binding, captain hold, contributions, mail, PR reviewers, task inbox, and Pi branch extension. I also ran --check-coverage. All passed except one pre-existing watcher test. That test failed once under heavy host load. I reproduced the cause, fixed the test, and the full watcher suite then passed alone. Live drives covered the publication graph and private-path audit against the real fork and upstream refs, and the real watcher with real tmux on an elapsed declared wait. They also covered a real Bearings board in a real Lavish session: I answered it in headless Chromium, and the answer went through the armed poller into tasks-axi. The other live drives were captain re-hold replay before and after the fix, the live doorbell test under a UTF-8 path with a real pi agent (Claude stopped at its folder-trust dialog), reviewer discovery on real PR kunchenguid#80, and seven contribution polls against 13 real upstream PRs. Seven scenarios stay untested live for credential or authority reasons, and their automated suites passed. I did not run the linters, doc-audience check, or git diff --check, because this step does not allow linters.

  • Live validation: ✅ go - 7 of 14 scenarios driven live against the product
Scenario Result Live Evidence
Fork publication contains fork main and upstream 888871d, drops no fork commits, and tracks no private or ignored paths ✅ pass live publication-audit.txt: fork main and upstream 888871d are ancestors of HEAD, 199 commits ahead of fork main with 0 behind, no .env/data/state/config/projects/.no-mistakes/credential paths, no ignored…
A new push to one PR cancels only that PR's older CI run; other PRs, main pushes, non-PR events, release work, and no-mistakes body-event groups keep their runs; CI does not wait for lint ⏸️ untested no Only GitHub Actions evaluates concurrency groups. Proving cancellation needs pushes to a fork PR, and this step has no authority to push. The PR phase's CI runs will show it.
Firstmate reports the authoritative no-mistakes run: newer failure beats older live, newer live beats older terminal, unproved terminal records and adjacent ledger rows cannot win ⏸️ untested no Needs real competing no-mistakes runs on one branch. This phase may not init or start no-mistakes runs. This worktree's repo is not initialized, and the initialized checkout is outside the worktree.
The watcher rechecks an elapsed declared wait with 'the declared clearing time has passed' before any generic wedge escalation ✅ pass live watcher-expired-declared-wait-live.txt: run 2 surfaced 'stale: wedgelab:fm-wedge (paused 3s, awaiting external - the declared clearing time has passed ...)', no 'possible wedge'
Captain answers a board card with Yes plus a 505-character note containing 'café' under a 41-character title; the full note is stored as valid UTF-8 and the call closes ✅ pass live board-answer-queued.png, board-answer-live-transcript.txt: state done, the complete 506-byte note is in the persisted decision, and the wake queued as 'check: procevent lavish ... 1'
Re-holding a released task and repeating the same release answer closes the second parent decision ✅ pass live captain-hold-rehold-replay.txt: at 48b9e5e the captain-hold-X-2 decision stays open; at HEAD 'resolved [key=captain-hold-X-2]: captain hold X: released'
Steering a worker whose home path contains 'café' still rings the doorbell, and a real agent acts and acknowledges; a path with a terminal control byte gets no doorbell line ✅ pass live doorbell-utf8-home-live.txt: 'ok - pi (0.85.1): the doorbell reached a real worker, which acted and acked with the mv'; the ESC-byte path returns rc=1 with no line
Reviewer discovery on a PR with a renamed file queries the file's authorship under its previous name at the base commit ✅ pass live pr-reviewers-rename-live.txt: PR kunchenguid#80 base has 4 commits for the old path and 0 for the new one; HEAD queries tests/fm-secondmate.test.sh, 48b9e5e queried tests/fm-secondmate-safety.test.sh
Settled merged contributions do not starve open ones: merged PRs are never re-read and the open PR is reached within the poll budget ✅ pass live contributions-poll-live.txt: each merged PR is read only in the poll that settled it; open kunchenguid#4804 is reached in poll 5 and re-read in polls 6-7; unacked pending signals cause no forge reads
A second mate's persistence reply through fm-secondmate-report.sh --doc, the plain helper, or plain correlated text releases the restart; negative or incomplete replies get a nudge ⏸️ untested no Needs a running second mate agent that fm-control.sh can stop and relaunch. This worker may not restart mates or touch the fleet.
A remote second mate's --doc report pointer is fetched and rewritten for the parent ⏸️ untested no Needs an SSH-reachable remote second mate. ssh localhost fails host-key verification, and the work-PC mate is fleet state this worker may not touch.
Teardown refuses to remove a worker while its no-mistakes run is parked or owned by another run, and publishes the backlog-close marker only after cancel succeeds ⏸️ untested no Needs a real parked no-mistakes run. This phase may not start or abort runs.
Mail polling eventually retries recovered headers under slow FETCH responses ⏸️ untested no Needs an IMAP mailbox with controllable latency. Using the user's real mailbox is out of scope.
Extension-source registration, publication, interrupted capture, and announcement recovery keep one lock order and stay recoverable ⏸️ untested no The race and mid-write interruption cases need fault injection, so they cannot be driven through the product surface.
Evidence: Publication graph and private-path audit

Source: Publication graph and private-path audit

# Publication audit of HEAD 82b3f928e9f1d0610ba9c6ccbff08b0408bc89b7 vs fork main 308bdfab525e28d11edc16953d872cb8e5d3bf4c
## graph
fork main ancestor of HEAD: yes
intake upstream 888871de ancestor of HEAD: yes
commits on HEAD not in fork main: 199
commits in fork main not in HEAD: 0
merge commits in range: 2
## tracked private-path candidates at HEAD (paths ending state/ config/ data/ projects/ .no-mistakes/ .env, credentials, keys)
(none)
## tracked files under top-level state/
## files added in range that .gitignore would ignore
(none)
## upstream drift since the merge (fresh git ls-remote of kunchenguid/firstmate at test time)
upstream main now: 9bc051ff43c6e4d23c163ee8f1d87551a11050c0 (3 commits after merged 888871de: #4775, #4799, #4800)
Evidence: CI workflow supersession model test output

Source: CI workflow supersession model test output

ok - a newer push to one PR supersedes that PR's in-flight CI
ok - distinct PRs get distinct concurrency groups
ok - every main push keeps its own group and is never cancelled
ok - non-PR events, including release work, keep independent non-cancelling runs
ok - compliance body events retain independent per-event groups
ok - every PR head publishes the complete CI result set
ok - the CI suite starts independently of lint
ok - every ci.yml job carries a finite timeout
ok - the incident's unbounded jobs keep their recommended caps
ok - the already-measured lane bounds are unchanged
![Board with Yes and café note queued (real Lavish, headless Chromium)](https://github.com/user-attachments/assets/feb57b32-f3f8-48b7-9556-dc740f1ae635) - Evidence: [Board with Yes and café note queued (real Lavish, headless Chromium)](https://github.com/BohnBawerick/firstmate/blob/9efa959027e7d1de10ccbc3eead354f255dabd18/.no-mistakes/evidence/fm/fm-upstream-fork-ci-sync/board-answer-queued.png) ![Board with the note typed, before queueing](https://github.com/user-attachments/assets/b9cfc5be-ef46-4ba9-aca3-18255d171565) - Evidence: [Board with the note typed, before queueing](https://github.com/BohnBawerick/firstmate/blob/9efa959027e7d1de10ccbc3eead354f255dabd18/.no-mistakes/evidence/fm/fm-upstream-fork-ci-sync/board-answer-filled.png)
Evidence: Board answer persisted through poller and keyed-answer intake

Source: Board answer persisted through poller and keyed-answer intake

# Live board answer: real fm-bearings-board.sh build -> real Lavish session -> headless Chromium click "Yes" + 505-char note starting "café" -> real armed fm-procevent lavish poller -> fm-captain-hold.sh keyed-answer intake -> tasks-axi.
# Title is 41 chars, so "title -> yes - note" is > 512 bytes (review finding r14); note has a 2-byte é (finding r19).
$ wake queue:
1789705869	1	check	procevent:lavish-1d2865b88a5db3af:1	check: procevent lavish lavish-1d2865b88a5db3af 1
$ tasks-axi show board-call --full (body decoded):
task:
  id: board-call
  title: Choose the publication route for the fork
  state: done
  blocked: no
  blocked_by: none
  held: no
  hold_reason: route choice pending
  hold_kind: captain
  hold_until: "-"
  kind: captain
  repo: sample
  priority: "-"
  created: "-"
  closed: 2026-09-18
  deps: none
  links: none
  body (decoded):
Resolution recorded by fm-captain-hold.
Decision digest: e3ee7b113758c814c224be6bae37ebd5d3104b19c4eada632d7119266b48b067
Resolution mode: answered

Captain decision:
Captain answered this call through the captured result lavish-1d2865b88a5db3af sequence 1.
Task: board-call
Answer: yes - café nnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnTAIL
Answer as shown to the captain: Choose the publication route for the fork -> yes - café nnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnn

CHECK state done: True
CHECK complete 506-byte note (café...TAIL) present in persisted decision: True
Evidence: Captain re-hold replay, before and after fix

Source: Captain re-hold replay, before and after fix

===== BEFORE fix (48b9e5e) =====
$ fm-captain-hold.sh hold X --title "Pick route" --reason "first choice"
X
rc=0
$ fm-captain-hold.sh answer X --decision-file decision.txt --release   # decision: Proceed
released: X
rc=0
$ fm-captain-hold.sh hold X --reason "second choice"
X
rc=0
$ fm-captain-hold.sh answer X --decision-file decision.txt --release   # same words again
released: X
rc=0
--- parent channel (parent/state/mate1.status):
needs-decision [key=captain-hold-X-1]: captain hold X: first choice
resolved [key=captain-hold-X-1]: captain hold X: released
needs-decision [key=captain-hold-X-2]: captain hold X: second choice
--- tasks-axi show X:
  id: X
  state: queued
  held: no
--- verdict: second decision opened=1 resolved=0

===== AFTER fix (HEAD 82b3f92) =====
$ fm-captain-hold.sh hold X --title "Pick route" --reason "first choice"
X
rc=0
$ fm-captain-hold.sh answer X --decision-file decision.txt --release   # decision: Proceed
released: X
rc=0
$ fm-captain-hold.sh hold X --reason "second choice"
X
rc=0
$ fm-captain-hold.sh answer X --decision-file decision.txt --release   # same words again
released: X
rc=0
--- parent channel (parent/state/mate1.status):
needs-decision [key=captain-hold-X-1]: captain hold X: first choice
resolved [key=captain-hold-X-1]: captain hold X: released
needs-decision [key=captain-hold-X-2]: captain hold X: second choice
resolved [key=captain-hold-X-2]: captain hold X: released
--- tasks-axi show X:
  id: X
  state: queued
  held: no
--- verdict: second decision opened=1 resolved=1
Evidence: Doorbell under a UTF-8 home path with a real pi agent

Source: Doorbell under a UTF-8 home path with a real pi agent

# Live doorbell under a UTF-8 home path. The test lab is mktemp -d "$TMPDIR/fm-inbox-live.XXXXXX", so the inbox path contains "café".
$ TMPDIR="/tmp/nmtest-café" FM_SEND_INBOX_LIVE_E2E=1 FM_SEND_INBOX_LIVE_HARNESSES=claude bash tests/fm-send-inbox-doorbell-live-e2e.test.sh
not ok - claude (2.1.273 (Claude Code)): composer stayed visibly pending; the pane is not steerable
#    Quick safety check: Is this a project you created or one you trust? (Like your own code, a well-known open source project, or work from your team). If not, take a moment to review what's in this folder first.
#    Claude Code'll be able to read, edit, and execute files here.
#    Security guide
#    ❯ No, exit
#      Yes, I trust this folder
#    Enter to confirm · Esc to cancel
not ok - live steering-inbox doorbell guard found failures above
exit=1
# claude blocked on its folder-trust dialog for the untrusted worktree path (environment, not product).

$ TMPDIR="/tmp/nmtest-café" FM_SEND_INBOX_LIVE_E2E=1 FM_SEND_INBOX_LIVE_HARNESSES=pi bash tests/fm-send-inbox-doorbell-live-e2e.test.sh
ok - pi (0.85.1): the doorbell reached a real worker, which acted and acked with the mv
ok - live steering-inbox doorbell guard: 1 harness(es) honored the doorbell contract
exit=0

# Doorbell line built by the lib for a UTF-8 path and for a path with an ESC control byte:
UTF-8 path -> : Firstmate instruction waiting: list '/tmp/nmtest-café~/t1.inbox'/*.msg and, in numeric order, read and act on each, then mv each handled file to '/tmp/nmtest-café~/t1.inbox'/handled/. [rc=0]
ESC-byte path -> (no line) [rc=1]
Evidence: Watcher expired declared-wait recheck, live

Source: Watcher expired declared-wait recheck, live

# Live watcher drive: real bin/fm-watch.sh + real bin/fm-crew-state.sh + real tmux (private socket) + real bin/fm-wake-drain.sh acks.
# The pane runs a quiet script that prints one line and then stays silent. The task status declares a wait whose clearing time passed 2 hours ago.
# Driver: /tmp/nmtest-run/watch-live.sh (FM_POLL=1 FM_SIGNAL_GRACE=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999)
pane command: bash
status: paused: waiting on the build queue until 2026-09-18T02:41:21Z
=== watcher run 1
  exit=0
  watch: signal: /tmp/nmtest-watch/state/wedge.status
  drain: 1789706483	2	signal	wedge.status	signal: /tmp/nmtest-watch/state/wedge.status
  drain: wake annotation: latest wake-EVENT observed at drain, not current state: wedge.status: paused: waiting on the build queue until 2026-09-18T02:41:21Z
=== watcher run 2
  exit=0
  watch: stale: wedgelab:fm-wedge (paused 3s, awaiting external - the declared clearing time has passed, rechecked on a long cadence not a wedge; confirm the wait cleared)
=== verdict
expired-wait recheck surfaced: yes
generic wedge escalation first: no
Evidence: Watcher live driver script

Source: Watcher live driver script

#!/usr/bin/env bash
# Live drive: real fm-watch.sh + real fm-crew-state.sh + real tmux (private socket).
set -u
ROOT=$1; LAB=/tmp/nmtest-watch; SOCK=nmtest-watch-$$
rm -rf "$LAB"; mkdir -p "$LAB/state" "$LAB/shim" "$LAB/agent"
REAL_TMUX=$(command -v tmux)
printf '#!/usr/bin/env bash\nexec %s -L %s "$@"\n' "$REAL_TMUX" "$SOCK" > "$LAB/shim/tmux"; chmod +x "$LAB/shim/tmux"
# A quiet "grok" agent: prints one line and then produces no output.
printf '#!/usr/bin/env bash\necho "waiting at the gate"\nwhile :; do read -r -t 3600 _ || :; done\n' > "$LAB/agent/grok"; chmod +x "$LAB/agent/grok"
trap '"$REAL_TMUX" -L "$SOCK" kill-server 2>/dev/null' EXIT
"$REAL_TMUX" -L "$SOCK" new-session -d -s wedgelab -n fm-wedge -x 160 -y 40 "$LAB/agent/grok"
sleep 1
echo "pane command: $("$REAL_TMUX" -L "$SOCK" display-message -p -t wedgelab:fm-wedge '#{pane_current_command}')"
past=$(date -u -d "@$(( $(date +%s) - 7200 ))" +%Y-%m-%dT%H:%M:%SZ)
printf 'window=wedgelab:fm-wedge\nkind=ship\nharness=grok\nbackend=tmux\n' > "$LAB/state/wedge.meta"
printf 'paused: waiting on the build queue until %s\n' "$past" > "$LAB/state/wedge.status"
echo "status: $(cat "$LAB/state/wedge.status")"
ack() {
  local err=$LAB/drain.err seq gen
  FM_STATE_OVERRIDE="$LAB/state" "$ROOT/bin/fm-wake-drain.sh" 2> "$err" | sed 's/^/  drain: /'
  seq=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation .*/\1/p' "$err")
  gen=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--recovery-generation \([A-Za-z0-9._-]*\)$/\1/p' "$err")
  [ -n "$seq" ] && FM_STATE_OVERRIDE="$LAB/state" "$ROOT/bin/fm-wake-drain.sh" --ack-through "$seq" --recovery-generation "$gen" >/dev/null 2>&1
}
for round in 1 2 3 4 5 6; do
  echo "=== watcher run $round"
  PATH="$LAB/shim:$PATH" FM_STATE_OVERRIDE="$LAB/state" FM_WATCH_HANDLING_SUCCESSOR=1 \
    FM_PAUSE_RESURFACE_SECS=999 FM_STALE_ESCALATE_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=1 \
    FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 timeout 90 "$ROOT/bin/fm-watch.sh" > "$LAB/watch.out" 2>&1
  echo "  exit=$?"; sed 's/^/  watch: /' "$LAB/watch.out"
  if grep -qF 'declared clearing time has passed' "$LAB/watch.out" || grep -qF 'possible wedge' "$LAB/watch.out"; then break; fi
  ack
done
echo "=== verdict"
grep -F 'declared clearing time has passed' "$LAB/watch.out" >/dev/null && echo "expired-wait recheck surfaced: yes" || echo "expired-wait recheck surfaced: no"
grep -F 'possible wedge' "$LAB/watch.out" >/dev/null && echo "generic wedge escalation first: yes" || echo "generic wedge escalation first: no"
Evidence: Reviewer discovery on real renamed-file PR kunchenguid#80

Source: Reviewer discovery on real renamed-file PR #80

# Live read-only GitHub check of fm-pr-reviewers.sh on https://github.com/kunchenguid/firstmate/pull/80 (file renamed: tests/fm-secondmate.test.sh -> tests/fm-secondmate-safety.test.sh; base 9e21160a)
commits at base for tests/fm-secondmate.test.sh: 4
commits at base for tests/fm-secondmate-safety.test.sh: 0

$ bin/fm-pr-reviewers.sh https://github.com/kunchenguid/firstmate/pull/80   # BEFORE fix (48b9e5e)
e-jung	2 recent commits
rc=0

$ bin/fm-pr-reviewers.sh https://github.com/kunchenguid/firstmate/pull/80   # AFTER fix (HEAD)
e-jung	2 recent commits
rc=0

# Pass-through gh spy (real API calls, arguments logged): which path the renamed file's authorship query used
old: path=tests/fm-secondmate-safety.test.sh 
new: path=tests/fm-secondmate.test.sh 
Evidence: Contribution polling against real upstream PRs

Source: Contribution polling against real upstream PRs

# Live drive: real bin/fm-contributions.sh poll in a scratch home against 12 merged + 1 open real upstream PRs (kunchenguid/firstmate), read-only GitHub API.
# gh is a pass-through spy that logs each real call. FM_CONTRIBUTIONS_BUDGET=25 seconds per poll (about 3 PR observations).
# Expected: once a PR settles as merged it is never re-read; unacked pending signals cause no forge reads; the open PR is reached and then re-read every poll.
=== poll 1
  out: contribution-wake: check: contributions merged-4689 17f131b96c3425bca45a2b34e578a701b76d2a60698d62594a3cc0487d196854
  out: contribution-wake: check: contributions merged-4689 e2d6ac5ea2991b81380be994279e8e3d3e6622336954fb7fcfd2404addf473f2
  out: contribution-wake: check: contributions merged-4710 40948fb3c274c9af57c97b7ae30bb9effb749ad7e50f23bd95cc47bc22109ccb
  exit=0
  records:
    merged-4689	merged	2026-09-18T04:42:19Z	-	2 pending
    merged-4692	merged	2026-09-18T04:42:19Z	-	0 pending
    merged-4710	merged	2026-09-18T04:42:19Z	-	1 pending
  forge reads this poll per PR:
          3 pulls/4689
          3 pulls/4692
          3 pulls/4710
          1 pulls/4738
  total gh calls: 25
=== poll 2
  out: contribution-wake: check: contributions merged-4775 9582c58f8603f31423ae8173a12f8cd021e377fde1d2836f4fc484194d57bf9e
  out: contribution-wake: check: contributions merged-4775 736a8443709ce057b94ff2fdf23a4d50fc47ab41e7549439a41d9cabaf05a8cf
  out: contribution-wake: check: contributions merged-4775 930226fdf60f00f963a32088253f72f5088ddae3dfd7ede94be4120185f42b4c
  exit=0
  records:
    merged-4689	merged	2026-09-18T04:42:19Z	-	2 pending
    merged-4692	merged	2026-09-18T04:42:19Z	-	0 pending
    merged-4710	merged	2026-09-18T04:42:19Z	-	1 pending
    merged-4738	merged	2026-09-18T04:42:44Z	-	0 pending
    merged-4775	merged	2026-09-18T04:42:44Z	-	3 pending
  forge reads this poll per PR:
          3 pulls/4738
          3 pulls/4775
          3 pulls/4777
  total gh calls: 23
=== poll 3
  out: contribution-wake: check: contributions merged-4778 5d3459ed8559627302ce64f38a0b25198b582e5d3b5c7806464a11af23d30330
  out: contribution-wake: check: contributions merged-4778 9d530ad04369794d77de4d114b5a46687d3a1ca536714abda0471f339c6825fa
  exit=0
  records:
    merged-4689	merged	2026-09-18T04:42:19Z	-	2 pending
    merged-4692	merged	2026-09-18T04:42:19Z	-	0 pending
    merged-4710	merged	2026-09-18T04:42:19Z	-	1 pending
    merged-4738	merged	2026-09-18T04:42:44Z	-	0 pending
    merged-4775	merged	2026-09-18T04:42:44Z	-	3 pending
    merged-4777	merged	2026-09-18T04:43:09Z	-	0 pending
    merged-4778	merged	2026-09-18T04:43:09Z	-	2 pending
  forge reads this poll per PR:
          3 pulls/4777
          3 pulls/4778
          3 pulls/4779
  total gh calls: 24
=== poll 4
  forge reads this poll per PR:
          3 pulls/4779
          3 pulls/4783
          3 pulls/4788
          1 pulls/4799
  total gh calls: 25
=== poll 5
  out: contribution-wake: check: contributions merged-4799 2f7346a225c758a24821da32a5beba31af763a01c6665521bb84b7d6ed273efe
  out: contribution-wake: check: contributions merged-4799 2d52804d6fb0b2faa4eac6ed2a6f6592124db4f1f2a59f9bbce208888e9c3ad0
  out: contribution-wake: check: contributions merged-4799 094ecb0ec1024b25895a1586d4f89906bb3e933d5e417fa8725c7775aa3a988b
  forge reads this poll per PR:
          3 pulls/4799
          3 pulls/4800
          3 pulls/4804
  total gh calls: 24
=== poll 6
  forge reads this poll per PR:
          3 pulls/4804
  total gh calls: 8
=== poll 7
  forge reads this poll per PR:
          3 pulls/4804
  total gh calls: 8
merged-4689	merged	2026-09-18T04:42:19Z	-	2 pending
merged-4692	merged	2026-09-18T04:42:19Z	-	0 pending
merged-4710	merged	2026-09-18T04:42:19Z	-	1 pending
merged-4738	merged	2026-09-18T04:42:44Z	-	0 pending
merged-4775	merged	2026-09-18T04:42:44Z	-	3 pending
merged-4777	merged	2026-09-18T04:43:09Z	-	0 pending
merged-4778	merged	2026-09-18T04:43:09Z	-	2 pending
merged-4779	merged	2026-09-18T04:43:44Z	-	0 pending
merged-4783	merged	2026-09-18T04:43:44Z	-	0 pending
merged-4788	merged	2026-09-18T04:43:44Z	-	0 pending
merged-4799	merged	2026-09-18T04:44:10Z	-	3 pending
merged-4800	merged	2026-09-18T04:44:10Z	-	0 pending
open-4804	open	2026-09-18T04:44:44Z	-	0 pending
Evidence: Contribution polling driver script

Source: Contribution polling driver script

#!/usr/bin/env bash
# Live drive: real fm-contributions.sh poll against real GitHub PRs (read-only), gh calls logged by a pass-through spy.
set -u
ROOT=$1; H=/tmp/nmtest-contrib
rm -rf "$H"; mkdir -p "$H/data" "$H/state" "$H/config" "$H/projects" "$H/fakebin"
printf '#!/bin/sh\nexit 1\n' > "$H/fakebin/tmux"; printf '#!/bin/sh\nexit 0\n' > "$H/fakebin/no-mistakes"
REAL=$(command -v gh)
printf '#!/bin/sh\nprintf "%%s\\n" "$*" >> "%s/gh-calls.log"\nexec %s "$@"\n' "$H" "$REAL" > "$H/fakebin/gh"; chmod +x "$H/fakebin/"*
printf '# Backlog\n\n## Queued\n' > "$H/data/backlog.md"
for n in 4800 4799 4788 4783 4779 4778 4777 4775 4738 4710 4692 4689; do
  printf -- '- [ ] merged-%s - Upstream contribution https://github.com/kunchenguid/firstmate/pull/%s (repo: firstmate) (kind: ship)\n' "$n" "$n" >> "$H/data/backlog.md"
done
printf -- '- [ ] open-4804 - Upstream contribution https://github.com/kunchenguid/firstmate/pull/4804 (repo: firstmate) (kind: ship)\n' >> "$H/data/backlog.md"
cp "$ROOT/.tasks.toml" "$H/.tasks.toml" 2>/dev/null || true
poll() { PATH="$H/fakebin:$PATH" FM_HOME="$H" FM_ROOT_OVERRIDE="$ROOT" FM_STATE_OVERRIDE="$H/state" FM_DATA_OVERRIDE="$H/data" FM_CONFIG_OVERRIDE="$H/config" FM_CONTRIBUTIONS_BUDGET=25 "$ROOT/bin/fm-contributions.sh" poll; }
summ() { for f in "$H"/data/*/contributions.json; do jq -r '.task as $t | .records[] | [$t, (.observation.state // "none"), .checked_at, (.error // "-"), ((.pending|length)|tostring)+" pending"] | @tsv' "$f"; done; }
for round in 1 2 3; do
  : > "$H/gh-calls.log"
  echo "=== poll $round"; poll 2>&1 | cut -c1-200 | sed 's/^/  out: /'; echo "  exit=${PIPESTATUS[0]}"
  echo "  records:"; summ | sed 's/^/    /'
  echo "  forge reads this poll per PR:"; grep -oE 'pulls/[0-9]+' "$H/gh-calls.log" | sort | uniq -c | sed 's/^/    /'
  echo "  total gh calls: $(wc -l < "$H/gh-calls.log")"
done
Evidence: Flaky watcher test diagnosis and fix

Source: Flaky watcher test diagnosis and fix

# Flaky test found during the focused run: tests/fm-watch-triage.test.sh
# test_live_declared_pause_ticking_footer_keeps_the_bounded_cadence (pre-existing on fork main, commit 406d850).
# Under heavy host load the full suite took 1192 s (hint 263 s) and failed once:
#   not ok - footer tick 2 printed a wake reason for a standing declared pause: stale: test:fm-paused-ticking (paused 1000s, ...)
# Alone it passed 3 of 3 times, and the full suite passed alone.
# Cause: rounds 1-6 ran with FM_PAUSE_RESURFACE_SECS=999. When 999 s of wall time pass between the
# first-sight surface and round 2, the watcher correctly re-surfaces the declared pause, so the test failed.
# Reproduction (temporary copy, declaration and reminder marker aged 1000 s before round 2):
#   old cadence 999   -> not ok - footer tick 2 re-surfaced a standing declared pause: ... (paused 1001s, ...)   rc=1
#   new cadence 86400 -> ok - a live declared pause keeps its bounded cadence ...                                  rc=0
# Fix: rounds 1-6 use a one-day cadence. The final recheck round still sets FM_PAUSE_RESURFACE_SECS=240 and ages the pause by 500 s.
- Outcome: ⚠️ 2 issues (1 warning, 1 info) across 1 run (57m27s)

Pipeline

Updates from git push no-mistakes

... (2 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Rebase** - 1 warning

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

🔧 **Review** - 4 issues found → auto-fixed (8) ✅
  • 🚨 bin/fm-secondmate-restart.sh:338 - A correlated reply saying blocked: open work could not be saved satisfies fm_pending_reply_try_resolve and immediately triggers restart, discarding the secondmate's conversation despite failed persistence. The timeout recheck has the same problem. Require a successful persistence acknowledgment at both restart paths; negative or incomplete replies must retain the conversation and use the existing fallback.
  • ⚠️ bin/fm-classify-lib.sh:1085 - The snapshot reader selects the last nonblank line rather than the latest recognized status event. For done: report complete followed by continuation prose, main's recovery backstop rejects the prose. If the branch already acknowledged the signal without publishing an outcome, subsequent main drains miss the completion. Reuse the existing event-recognition semantics while preserving the event's byte endpoint.
  • ⚠️ bin/fm-mail-check.sh:208 - The default 15-second watchdog can prevent mail polling from making progress. fm-mail.py collects up to 20 headers and logs out before emitting rows; twenty one-second FETCH responses exceed that budget. The failed batch advances no cursor, so every poll repeats it. Its 20-second socket timeout also exceeds the watchdog. Give the poll engine an inner deadline and return completed or degraded rows before the watchdog expires, using the existing cursor protocol.
  • ⚠️ tests/fm-task-delivery.test.sh:1328 - The new assertions grep AGENTS.md for two sentences and compare their line numbers to claim worker-role disambiguation. This is a source-content-only assertion prohibited by the test-quality rule. Remove this assertion block; retain the executable spawn checks against the generated launch brief and unchanged instruction files.

🔧 Fix applied.
6 issues (1 error, 5 warnings) still open:

  • 🚨 bin/fm-secondmate-restart.sh:171 - Successful replies sent through fm-secondmate-report.sh include '(via-helper)', so the new exact payload comparison rejects them and skips the restart. Accept the existing helper serialization when checking the affirmative acknowledgment, while retaining rejection of negative and incomplete replies.
  • ⚠️ bin/fm-mail.py:379 - The default inner deadline is nine seconds, but all new-mail candidates precede retries. With 15 new messages per poll and 0.65-second FETCH responses, every poll expires before reaching retries, permanently starving recovered headers. Preserve the existing retry reservation in execution order or time allocation, including when returning partial batches.
  • ⚠️ bin/fm-contributions.sh:286 - If polling stops after recording a merged PR for owner A but before updating existing owner B, subsequent polls enter settle_final. B's open observation is never replaced because this function only creates missing records or clears errors. Its stale fleet-action row can then override A's settled row indefinitely. Reconcile existing owners to the final observation while preserving their individual acknowledgment state.
  • ⚠️ bin/fm-contributions.sh:315 - A poll interrupted between write_record and publish_pending can leave a terminal observation containing an unnotified maintainer signal. Every subsequent poll takes this shortcut and bypasses publish_pending permanently. Replay already-persisted, unnotified signals through the existing publisher before skipping the forge read.
  • ⚠️ bin/fm-contributions.sh:314 - Permanent retirement on 'closed' misses ordinary reopening: after an issue or unmerged PR is closed, observed, and reopened, polling never reads it again and projection continues reporting a fresh closed result with nobody responsible. Remove 'closed' from the permanent-final shortcut and freshness exemption. This changes the deliberately introduced retirement policy, so it needs user review.
  • ⚠️ bin/fm-pr-reviewers.sh:63 - For a renamed file, this selects the new filename but queries its history at the base commit, where that path does not yet exist. A rename-only PR therefore loses its actual authors and can report no candidates. Use previous_filename for renamed entries when querying base-commit authorship.

🔧 Fix applied.
7 issues (1 error, 6 warnings) still open:

  • 🚨 bin/fm-procevent.sh:559 - Re-registering an extension source while its result publishes can deadlock indefinitely. Registration holds the lifecycle lock before requesting the source lock; publish_result holds the source lock before requesting the lifecycle lock. Enforce one lock order across both operations.
  • ⚠️ bin/fm-procevent.sh:2019 - External operations hold the home-wide lifecycle lock throughout invocation, including supported hour-long polls. One source therefore blocks unrelated sources and can delay completed results' classification and publication. Limit locking to ownership transitions, using the existing invocation records to protect active work.
  • ⚠️ bin/fm-procevent-extension-capture.pl:258 - An interruption after publishing .adapter but before .result permanently blocks subsequent captures. Sequence selection checks only .result, so retries reuse the sequence, but publish_new refuses the surviving .adapter. Make partial publication recoverable under the existing source and binding ownership checks.
  • ⚠️ bin/fm-procevent-lavish.sh:484 - A selected answer's note survives only in the label, which is truncated to 512 characters. The board accepts selection 'yes' plus a 500-character note, but adding a 40-character title makes that label exceed the limit. The answer intake then closes the call with part of the accepted note missing. Preserve the complete note in ordinary-answer serialization.
  • ⚠️ bin/fm-captain-hold.sh:1070 - Hold task X, answer --release with 'Proceed', hold X again, then repeat that answer. The second hold opens parent decision captain-hold-X-2, but the historical digest triggers replay and resolves captain-hold-X-1. X is released while the parent's new decision remains open. Scope replay detection to the current hold while preserving historical records and interrupted-answer retries.
  • ⚠️ .pi/extensions/lib/fm-branch-dispatch.ts:294 - A captain-held status followed by continuation prose makes a later stale wake branch-eligible. The held event closes the keyed decision, while this last-line check reads the prose and misses main's ownership. Select the latest recognized status event when classifying stale wakes, matching the watcher's event semantics.
  • ⚠️ bin/fm-task-inbox-lib.sh:264 - With a UTF-8 home path such as ~/firstmate, LC_ALL=C makes the printable-character check reject the inbox pathname. Steering is durably enqueued, but no doorbell is typed, and watcher retries repeat the failure. Reject terminal controls while permitting valid Unicode paths.

🔧 Fix applied.
4 warnings still open:

  • ⚠️ bin/fm-teardown.sh:1837 - The captured CLI response in tests/captures/no-mistakes-v1.70.1/parked.toon has top-level status 'running' alongside awaiting_agent and an awaiting-approval gate. This early return skips the gate checks even when branch and head match, allowing teardown to remove the worker while its run remains parked. Distinguish parked gates from actively executing steps before rejecting 'running'.
  • ⚠️ bin/fm-procevent-lavish.sh:484 - For selection 'yes' and note 'café', decode_json produces decoded characters, but the unencoded Perl stdout emits é as a lone Latin-1 byte. Joining the note into the answer therefore sends invalid UTF-8 into captain-hold's persisted decision. Encode decoded fields as UTF-8 at the TSV output boundary without double-encoding the byte-oriented label.
  • ⚠️ bin/fm-mail.py:379 - Retries still starve when each poll has a new message whose first FETCH consumes the nine-second deadline. That message is emitted as degraded, then the loop exits before attempting any recovered retry. Sustained slow new arrivals repeat this indefinitely because first-class alternation applies only when cap equals one. Extend the existing alternation to deadline-limited polls so retries eventually receive the first fetch opportunity.
  • ⚠️ bin/fm-secondmate-restart.sh:170 - The documented --doc report helper emits an affirmative reply as 'open records written down (data/report.md via-helper)'. This suffix removal handles only '(via-helper)', so the valid acknowledgment resolves the pending request but incorrectly falls back to a nudge instead of restarting. Separate documented helper metadata from the acknowledgment before comparing the exact affirmative payload, retaining rejection of negative or incomplete answers.

🔧 Fix applied.
4 warnings still open:

  • ⚠️ bin/fm-teardown.sh:3382 - The backlog-close marker is written before confirming cancellation of the parked run. If cancellation fails, teardown refuses, but bootstrap later replays the marker, deletes the task metadata, and closes its backlog row while the run remains parked. Confirm cancellation before publishing the marker.
  • ⚠️ bin/fm-procevent.sh:277 - When sources A and B share an extension binding, B's active long poll makes this binding-wide cleanup refuse recovery of A. Retirement uses the same cleanup and also fails. Scope source cleanup to its existing source_id, including handshake ownership; reserve binding-wide cleanup for binding retirement.
  • ⚠️ bin/fm-procevent-extension-capture.pl:183 - Interrupting capture after sidecar publication leaves the fixed <id>.runner file. Neither runner EXIT cleanup nor stale-claim reclamation removes it, so subsequent starts fail this exclusive creation before reaching the repaired sequence loop. Reclaim the marker under verified source ownership. The regression plants only a sidecar and misses this interruption artifact.
  • ⚠️ bin/fm-procevent-extension-capture.pl:242 - Partial A.1.adapter evidence now makes capture publish A.2.result, but next_result_sequence in fm-procevent.sh still checks only .result files and repeatedly selects missing sequence 1. Subsequent polls therefore reuse one request identity, causing supported idempotent extensions to replay the same event indefinitely. Make request identity and capture allocation use the same sequence rule.

🔧 Fix applied.
1 warning still open:

  • ⚠️ bin/fm-contributions.sh:308 - Merged URLs retain their old checked_at values and remain at the front of every poll. When processing this settled history consumes the configured budget, already-observed open PRs are never reached: subsequent polls repeat the same prefix and exit successfully. Exclude fully reconciled, already-notified merged records from the budgeted queue, while preserving reconciliation and pending-signal replay.

🔧 Fix applied.
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-teardown.sh:1845 - With two runs on the same branch, axi status can return parked run A while axi sync --check returns run B. If A's head is unfetched and B's submitted head matches local HEAD, teardown accepts B's proof and aborts A without establishing ownership. Pass run_id as fm_nm_submitted_head's third argument so missing or mismatched run identities cannot authorize cancellation.
  • ⚠️ bin/fm-procevent.sh:1299 - Termination after writing this marker but before enqueueing the wake leaves the generation permanently marked as reported. Subsequent reconciliation skips publication, and the watcher discards reconciliation stdout, leaving the stranded source silent. Record successful announcement only after durable wake publication; preserve retry through the existing queue and key.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ bin/fm-secondmate-restart.sh:186 - A supported plain reply such as 'done: open records written down [corr=0123456789abcdef]' resolves the pending request, but the correlation suffix remains in payload and fails this comparison. The mate receives a nudge instead of restarting. Remove the matched correlation metadata before comparing the affirmative payload, while retaining rejection of negative or incomplete replies.
  • ⚠️ bin/fm-procevent-remote-reply.sh:315 - The documented fm-secondmate-report.sh --doc command still emits '(data/x/report.md via-helper)', but this matcher now fetches only report=data/... pointers. A remote helper reply therefore resolves its request while leaving its report unavailable to the parent, without a transfer warning. Update the --doc producer to emit the structured report pointer.

🔧 Fix applied.
✅ Re-checked - no issues remain.

⚠️ **Test** - 2 issues (1 warning, 1 info)
Scenario Result Live Evidence
Fork publication contains fork main and upstream 888871d, drops no fork commits, and tracks no private or ignored paths ✅ pass live publication-audit.txt: fork main and upstream 888871d are ancestors of HEAD, 199 commits ahead of fork main with 0 behind, no .env/data/state/config/projects/.no-mistakes/credential paths, no ignored…
A new push to one PR cancels only that PR's older CI run; other PRs, main pushes, non-PR events, release work, and no-mistakes body-event groups keep their runs; CI does not wait for lint ⏸️ untested no Only GitHub Actions evaluates concurrency groups. Proving cancellation needs pushes to a fork PR, and this step has no authority to push. The PR phase's CI runs will show it.
Firstmate reports the authoritative no-mistakes run: newer failure beats older live, newer live beats older terminal, unproved terminal records and adjacent ledger rows cannot win ⏸️ untested no Needs real competing no-mistakes runs on one branch. This phase may not init or start no-mistakes runs. This worktree's repo is not initialized, and the initialized checkout is outside the worktree.
The watcher rechecks an elapsed declared wait with 'the declared clearing time has passed' before any generic wedge escalation ✅ pass live watcher-expired-declared-wait-live.txt: run 2 surfaced 'stale: wedgelab:fm-wedge (paused 3s, awaiting external - the declared clearing time has passed ...)', no 'possible wedge'
Captain answers a board card with Yes plus a 505-character note containing 'café' under a 41-character title; the full note is stored as valid UTF-8 and the call closes ✅ pass live board-answer-queued.png, board-answer-live-transcript.txt: state done, the complete 506-byte note is in the persisted decision, and the wake queued as 'check: procevent lavish ... 1'
Re-holding a released task and repeating the same release answer closes the second parent decision ✅ pass live captain-hold-rehold-replay.txt: at 48b9e5e the captain-hold-X-2 decision stays open; at HEAD 'resolved [key=captain-hold-X-2]: captain hold X: released'
Steering a worker whose home path contains 'café' still rings the doorbell, and a real agent acts and acknowledges; a path with a terminal control byte gets no doorbell line ✅ pass live doorbell-utf8-home-live.txt: 'ok - pi (0.85.1): the doorbell reached a real worker, which acted and acked with the mv'; the ESC-byte path returns rc=1 with no line
Reviewer discovery on a PR with a renamed file queries the file's authorship under its previous name at the base commit ✅ pass live pr-reviewers-rename-live.txt: PR kunchenguid#80 base has 4 commits for the old path and 0 for the new one; HEAD queries tests/fm-secondmate.test.sh, 48b9e5e queried tests/fm-secondmate-safety.test.sh
Settled merged contributions do not starve open ones: merged PRs are never re-read and the open PR is reached within the poll budget ✅ pass live contributions-poll-live.txt: each merged PR is read only in the poll that settled it; open kunchenguid#4804 is reached in poll 5 and re-read in polls 6-7; unacked pending signals cause no forge reads
A second mate's persistence reply through fm-secondmate-report.sh --doc, the plain helper, or plain correlated text releases the restart; negative or incomplete replies get a nudge ⏸️ untested no Needs a running second mate agent that fm-control.sh can stop and relaunch. This worker may not restart mates or touch the fleet.
A remote second mate's --doc report pointer is fetched and rewritten for the parent ⏸️ untested no Needs an SSH-reachable remote second mate. ssh localhost fails host-key verification, and the work-PC mate is fleet state this worker may not touch.
Teardown refuses to remove a worker while its no-mistakes run is parked or owned by another run, and publishes the backlog-close marker only after cancel succeeds ⏸️ untested no Needs a real parked no-mistakes run. This phase may not start or abort runs.
Mail polling eventually retries recovered headers under slow FETCH responses ⏸️ untested no Needs an IMAP mailbox with controllable latency. Using the user's real mailbox is out of scope.
Extension-source registration, publication, interrupted capture, and announcement recovery keep one lock order and stay recoverable ⏸️ untested no The race and mid-write interruption cases need fault injection, so they cannot be driven through the product surface.
  • git ls-remote fork and upstream main; git merge-base --is-ancestor for fork main and upstream 888871de; git ls-tree -r HEAD private-path grep; git check-ignore on added files
  • bin/fm-test-run.sh --json ... tests/fm-ci-workflow.test.sh tests/fm-watch-triage.test.sh tests/fm-turnend-guard.test.sh tests/fm-crew-state.test.sh tests/fm-teardown.test.sh tests/fm-secondmate-restart.test.sh tests/fm-procevent.test.sh tests/fm-extension-binding.test.sh tests/fm-captain-hold-lifecycle.test.sh tests/fm-contributions.test.sh tests/fm-mail.test.sh tests/fm-pr-reviewers.test.sh tests/fm-task-inbox.test.sh tests/fm-pi-branch-extension.test.sh tests/fm-remote-secondmate-lifecycle-e2e.test.sh
  • bin/fm-test-run.sh tests/fm-watch-triage.test.sh tests/fm-remote-reply.test.sh (isolated rerun, both pass)
  • bin/fm-test-run.sh --check-coverage
  • FM_TEST_ONLY=test_live_declared_pause_ticking_footer_keeps_the_bounded_cadence bash tests/fm-watch-triage.test.sh x3 before the fix, then again after it; slow-host reproduction in a temporary copy (old fails, new passes)
  • Live watcher: real bin/fm-watch.sh + bin/fm-crew-state.sh + tmux private socket + bin/fm-wake-drain.sh acks on paused: ... until &lt;2h ago&gt;
  • Live board: bin/fm-captain-hold.sh hold board-call, bin/fm-bearings-board.sh build, headless Chromium via chrome-devtools-axi selects Yes + 505-char 'café…TAIL' note, Send to Agent, armed fm-procevent.sh lavish poller captures and feeds answers, tasks-axi show board-call --full
  • Live captain hold: fm-captain-hold.sh hold X / answer --release / re-hold / same answer, at 48b9e5e and HEAD, with real tasks-axi and the parent channel
  • TMPDIR=/tmp/nmtest-café FM_SEND_INBOX_LIVE_E2E=1 FM_SEND_INBOX_LIVE_HARNESSES=pi bash tests/fm-send-inbox-doorbell-live-e2e.test.sh (claude attempt blocked by folder trust)
  • bin/fm-pr-reviewers.sh https://github.com/kunchenguid/firstmate/pull/80 at 48b9e5e and HEAD, with a pass-through gh argument log
  • Live bin/fm-contributions.sh poll x7 in a scratch home over 12 merged + 1 open real upstream PRs, with a pass-through gh call log
⚠️ **Document** - 2 issues (1 warning, 1 info)
  • ⚠️ bin/fm-secondmate-restart.sh:188 - bin/fm-lint.sh exits 1 at HEAD, and the doc edits did not cause this. ShellCheck reports SC1010 on the unquoted done in [ &#34;$verb&#34; = done ] at bin/fm-secondmate-restart.sh:188 and in the --doc done argument at tests/fm-remote-reply.test.sh:398. It also reports SC2016 (info) at tests/fm-extension-binding.test.sh:1639. Quoting the word ('done') keeps behavior unchanged. This phase may only edit docs, so the code and test fix belongs to the lint or review phase.
  • ℹ️ docs/scripts.md - The upstream merge added bin scripts that docs/scripts.md does not list: fm-agy-trust.sh, fm-claude-trust.sh, fm-dispatch-resolve.sh, fm-env-lib.sh, fm-gemini-lib.sh, fm-landed-lib.sh, fm-remote-herdr-guard.sh, fm-remote-herdr-owner-lib.sh, fm-contributions.jq, and fm-procevent-extension-capture.pl. Upstream's own docs/scripts.md also lacks these rows. Adding them only in the fork would cause merge conflicts on each future sync, so the better follow-up is an upstream contribution.
🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits August 29, 2026 10:34
…nchenguid#3273)

* fix(bin): keep home-summary publication bounded and off the watcher beat

A home whose tasks had accumulated ordinary status history could not publish
state/home-summary.json at all, and every attempt starved the watcher's
liveness beacon while it failed silently.

The producer's per-task open-decision fold spent tens of milliseconds per
status line on a bash 3.2 global bracket-class substitution used only as a
blank-line guard. On a real home that made the whole ledger producer take
minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every
attempt and never completed. Replace that guard with an equivalent case glob
in the one fold owner, which both the whole-file and cursor-backed folds use.

Bound each per-task current-state read in the snapshot with
FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh,
whose dead-peer detection deliberately never kills a slow-but-alive remote
command, so nothing else bounded it.

Detach the watcher's two publication triggers from the poll loop. The loop
owns the beacon that fm-guard.sh reads as proof supervision is alive, and an
inline publication put up to a full publication deadline between two beacon
touches. A single in-flight publication is tracked so a slow one cannot
accumulate clones.

Report a repeatedly failing publication at session start. Publication stays
deliberately non-fatal to its caller, so the existing bounded home-local
failure record is now surfaced as a HOME_SUMMARY bootstrap line once the
ledger is absent or stale and failures have been recorded since.

* no-mistakes(review): Preserve home-summary failure attempt ordering

* no-mistakes(review): Enforce durable home-summary single-flight and ordering

* no-mistakes(review): Derive failure ordering from publication boundaries

* no-mistakes(review): Restore best-effort failure logging and publication scoping

* no-mistakes(review): Make ordering regression sensitive to one failure

* no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance
…henguid#3268)

* fix(supervision): classify the appended status span, not the last line

An actionable project update could be classified as routine and absorbed, so
a worker that raised a decision, hit a blocker, failed, or finished stalled
silently with the captain never told.

Trigger, mask, symptom. A worker appends a captain-relevant event
(`needs-decision`, `blocked`, `failed`, `done`). Any later routine append -
a `working:` progress note - lands before the supervisor classifies the
batch; the watcher's 30s signal-grace linger exists precisely to coalesce a
status write with the same turn's turn-end, so this window is ordinary
rather than rare. Both supervisors then asked "is the LAST line
captain-relevant?", read the routine line, and absorbed the wake. The
`.seen-*` suppressor advanced either way, so nothing ever re-read the event.
When the crew was also provably working, the no-verb fallback absorbed it
too, which is why the event disappeared completely instead of surfacing late.

Reproduced end to end against a real watcher before any change: with the
trailing `working:` append the watcher never exits and the wake queue stays
empty; with that one line removed - the smallest counterfactual - the same
`needs-decision` surfaces and queues. The away-mode daemon's `classify_signal`
returns `self|routine signal` for a `blocked:` event under the same mask,
which is the worse case because no captain is present to notice.

The proven path was already in the tree: `status_open_decisions` fixed this
exact masking for the durable decision fold, and its header states the rule -
reading an append-only event log last-event-wins cannot represent an earlier
event that a later unrelated line moved past. The classification path was
never migrated to that read model. That is the earliest divergence, and the
fix is to migrate it rather than to special-case the symptom.

`status_span_first_actionable` in bin/fm-classify-lib.sh is the new single
owner: it reads the bytes at or after a caller-supplied position and returns
the first still-live captain-relevant event. Each supervisor supplies its own
position, because the always-on watcher and the away-mode daemon classify the
same stream independently and must not share one cursor: the watcher reads
the size already recorded in its `.seen-*` signature (no new state) and its
`.hb-surfaced-<task>` backstop marker, and the daemon its
`.subsuper-seen-status-<task>` marker. Those two markers held the escalated
line and now hold the escalated-through byte offset, which also removes a
second defect in the same code - content dedup silently swallowed a genuinely
new event whose text repeated an older one. An absent, malformed, or
past-the-end position reads the whole log, so uncertainty surfaces events
rather than losing them, and a marker an older build wrote as a status line
reads that way too. Status logs are only ever appended to, including across a
reused task id, so a recorded position keeps its meaning.

A `needs-decision`/`blocked` event in the span is retired only when the
whole-file fold proves its key closed; `status_open_decisions` stays the sole
owner of that rule, so same-key reopening and reserved-key namespaces need no
second implementation here. Every other captain-relevant event is terminal
and always actionable.

Both backstops now walk every status log instead of only those whose last
line looks captain-relevant, because the event a backstop most needs to catch
is exactly one a later append has moved past. That leaves
`scan_captain_relevant_statuses` with no callers, and it is removed rather
than left as a working copy of the defective read model.

Regression coverage exercises the classifier and both supervisors through
their own interfaces: the masked decision, the captain-reported
release/install completion followed by cleanup chatter, and the away-mode
blocker all surface; a routine append after an already-classified event stays
absorbed, so the fix does not convert ordinary progress into wakes; and the
heartbeat backstop catches a masked event the per-wake path missed. The
end-to-end watcher tests drive a real fm-watch.sh with the crew reported as
provably working, which is the configuration that made the original stall
silent.

Two further claims in the supplied RCA are deliberately not patched here.
"Repeated operational recoveries produced all-clear replies despite known
actions" is downstream of this same cause, not an independent contributor: an
all-clear reply is the documented response when the specific event needs no
action, so a classification that wrongly reported "no action" produces it, and
correcting the classification removes it. "The project was subjected to
validation requirements outside its accepted path" is delivery-mode selection,
which AGENTS.md section 7 owns; no code changed here touches it, so it is out
of scope.

Harness and backend axes were inspected rather than assumed: nothing in this
path reads a vendor-emitted signal. The status log's format and append
protocol are Firstmate's own and identical for every harness, and no runtime
backend reads or writes `.status` files (`bin/backends/*` contain no reference
to them). The surrounding triage's only backend touchpoints - pane capture and
the authoritative crew-state read - are unchanged. No live-harness guard
applies and no per-harness verification record changes.

Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and
`bin/fm-test-run.sh --changed --base origin/main`.

* no-mistakes(review): Prevent status races and surface classification failures

* no-mistakes(review): Surface unreadable signals and preserve AFK endpoints

* no-mistakes(review): Route stale wakes through captured span verdicts

* no-mistakes(review): Retire supervision offsets with reused task state

* no-mistakes(review): Bind status offsets and preserve live decision origins

* no-mistakes(review): Strengthen status identity with verified birth time

* no-mistakes(review): Skip turn-end markers during status classification

* no-mistakes(review): Preserve status presentation with platform-strength identities

* no-mistakes(review): Retain failed wakes and advance routine checkpoints

* no-mistakes(review): Surface all events and retain unreadable wakes

* no-mistakes(review): Treat absent status logs as successful empty spans

* no-mistakes(review): Bound repeated classification failures with durable receipts

* revert(supervision): drop the failure-receipt and durable-retry machinery

Captain-authorized revert to the minimal fix. Review rounds added a durable
failure-receipt store and wake-retention-on-failure to bound repeated
classification failures. That machinery grew larger than the fix it protected
and kept producing its own defects: an unreadable log still looped forever
because the always-on watcher never consulted the receipt, and the receipt was
persisted before its diagnostic was durably queued, so a crash in between
swallowed the alarm outright. Those two defects go away with the code that
contained them rather than being repaired.

Removed: the failure-receipt path, fingerprint, record and clear helpers and
their retirement bookkeeping; the retention of a durable wake when
classification fails; and the error-propagation plumbing in both supervisors
that existed only to drive them.

Kept, because it is the accepted fix rather than the declined machinery: span
classification of the events appended since a supervisor last looked, in both
supervisors and both backstops; reporting every actionable event in a span and
committing a position only through what was reported; naming the live opening of
a reopened decision; treating an absent log as ordinary and an unreadable one as
worth reporting; the non-.status filter; and the platform-strength identity that
guards a position commit without failing a read.

Replacement behavior for a log that cannot be classified: report it once, do NOT
advance the classification position so the content is classified from where it
stopped once readable, and DO advance the wake signature so the report is
bounded to one per distinct file state. Reporting and reading are different acts:
telling the captain about a log is not the same as having read it, and only the
latter may move a classification position.

The residual risk is explicit and accepted: there is no guaranteed automatic
retry inside a crash-mid-read window, and the locked session-start replay of the
durable queue covers it. That rationale is recorded at mark_escalated_seen so a
future reader does not reintroduce the retry as a "missing" guarantee.

Also fixes lint failures that arrived with the review-fix commits and were never
caught because the run never reached its lint step: an unfollowable conditional
source directive, a second unquoted-expansion site left after a call was split
across lines, cleanup of the file being read inside its own read loop (restructured
to one post-loop teardown rather than three in-loop copies), stub functions in
tests that are invoked indirectly, and a test local left unused when its
assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch,
so these were introduced here.

Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode
reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue,
wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures).

* no-mistakes(review): Correct classification failure contract documentation

* no-mistakes(review): Bound unreadable status reports without skipping classification

* no-mistakes(review): Preserve escalation markers when buffering fails

* no-mistakes(review): Detect permission recovery without advancing classification

* no-mistakes(document): Document status span classification contract

* no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed

* no-mistakes(review): Escalate blockers while preserving declared-wait cadence

* no-mistakes(review): Clarify actionable events override wait self-handling

* no-mistakes(review): Surface rejected decisions and dangling status links

* no-mistakes(document): Document reserved-key reconciliation classification

* no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check

* no-mistakes(document): Correct away-mode classification documentation
…#3289)

* docs: split harness adapter operations reference

* no-mistakes(review): Fix harness adapter routing and ownership contracts

* no-mistakes(review): Prune duplicate harness adapter ownership prose

* no-mistakes(review): Fix default effort routing and Grok max semantics

* no-mistakes(review): Remove source-only routing test and duplicate semantics

* no-mistakes(review): Add local harness adapter instruction evaluation

* no-mistakes(review): Fix harness evaluation gating and change mapping

* no-mistakes(test): Captain, require explicit harness instruction evaluator model

* no-mistakes(document): Fix harness adapter documentation references
* test(fixtures): share fake-toolchain and spawn-world builders

Future tests can start from tests/fixtures.sh instead of copying stubs, and a
no-mistakes version-floor bump is one constant rather than a multi-file edit.

Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen,
fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile.
Left for opportunistic migration: remaining make_spawn_fakebin copies
(trace-context, kimi, muse, backend), the make_stubs send cluster, and the
fake no-mistakes version banners in bootstrap/session-start/secondmate suites.
Did not touch tests/fm-pr-check-security.test.sh.

* no-mistakes(review): Prevent fake SSH test from blocking on stdin

* no-mistakes(document): Clarify shared fixture documentation

* no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks

* no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head

* no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head
* feat(bin): retire completed PR-check migration machinery

Every registered home already carried both completion markers, and no
installer still creates pre-migration checks. Remove the one-time migrate
script, its bootstrap/watch/teardown/docs surface, and migration-path tests
without weakening live check-trust or PR-poll authentication.

* no-mistakes(review): Restore live PR-check security coverage

* no-mistakes(document): Refresh retired PR-check documentation

* no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks
…3247)

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* no-mistakes(review): Split extension conformance and forward remote transfer input

* no-mistakes(review): Forward malformed remote payloads through fm-on

* no-mistakes(review): Bound extension coordinator failure cleanup

* no-mistakes(test): Skip repeated orphan sweep in coordinator children

* no-mistakes(test): Queue isolated extension sections through bounded workers

* no-mistakes(test): Bound extension coordinator lane cleanup

* no-mistakes(test): Split remote lifecycle coordinator sections

* no-mistakes(test): Coordinator probes pass; aggregate deadline remains

* no-mistakes(test): Launch extension sections concurrently

* no-mistakes(test): Fix coordinator marker publication

* no-mistakes(test): Stabilize extension binding coordinator timing

* no-mistakes(lint): Fix extension binding ShellCheck warnings

* fix(extensions): prove invocation cleanup before retirement

* no-mistakes(review): Harden process-event inbox confinement

* no-mistakes(review): Preserve legacy capture parity

* no-mistakes(review): Protect external registry staging

* no-mistakes(test): Stabilize bounded extension conformance aggregate

* no-mistakes(document): Document external evidence confinement

* no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh`

* no-mistakes(review): Harden extension staging and lifecycle reservation

* no-mistakes(review): Harden external staging and lifecycle reservations

* no-mistakes(review): Wire capture helper into remote conformance

* no-mistakes(review): Pin external capture handoff and signal failures

* no-mistakes(review): Bind pinned capture authority to inherited descriptor

* no-mistakes(review): Harden descriptor-bound capture authority

* no-mistakes(review): Harden core capture reservation authority

* no-mistakes(review): Harden capture reservation boundaries

* no-mistakes(review): Harden capture reservations and cleanup

* no-mistakes(review): Harden capture handoff and reservation cleanup

* no-mistakes(review): Bind capture handoff to claim descriptors

* no-mistakes(review): Release lifecycle locks after host crashes

* no-mistakes(review): Pin reservation recovery to recorded state roots

* no-mistakes(review): Reject control bytes in claim state roots

* no-mistakes(test): Stabilize extension capture descriptor handoff

* no-mistakes(document): Document extension capture authority boundary

* no-mistakes(lint): Fix ShellCheck extension binding warnings

* no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks

* no-mistakes(document): Correct extension namespace creation timing

* no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes

A promoted scout used to receive a free-form placeholder instead of the
mode-specific Definition of done a briefed ship worker gets, so it never
saw the ask-user escalation rule or the --yes prohibition. That gap is the
concrete reason one incident's worker drove validation with --yes and
answered its own ask-user findings.

- Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific
  Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh
  so the two contracts cannot drift.
- bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the
  scratch inventory, clean base, ship branch, and that Definition of done,
  and prints the fm-send.sh command that delivers it.
- State the --yes ban as a prohibition rather than a preference, without
  claiming an enforcement the tool does not provide.
- Cover both through the real promotion and brief paths in
  tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh.

* no-mistakes(review): Publish promotion instructions before committing task state

* no-mistakes(review): Supersede conflicting scout delivery rules after promotion

* no-mistakes(review): Reject invalid promotion instruction destinations

* no-mistakes(document): Align documentation with promotion delivery contracts

* no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check

* no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean

* no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head

* no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
)

* fix(bin): present complete Lavish board feedback as structured output

Give the Lavish adapter a read-only presentation so a handler sees every
annotation and the session-ending tag=message as its own field, instead of
grepping a truncated raw capture.

* no-mistakes(review): Preserve unquoted messages and prioritize captain prose

* no-mistakes(document): Document structured Lavish result reads

* no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves

Dispatch and completion each moved a task's physical record and its
backlog row as two independently timed steps, so a crash or a forgotten
follow-up could leave the two disagreeing: a record with no in-flight
row, an in-flight row with no owner, or a finished task still shown in
flight.

Fold each backlog transition into the script that performs the physical
change, under the per-task lock it already holds and before it reports
success. Dispatch moves the item to In flight after publishing the task
record and fails loudly, removing its provisional record, when that
transition cannot land. Completion records an authoritative close and
performs it before removing the record, so an interrupted cleanup can be
finished later, and its closing message now confirms what already
happened rather than instructing a future step.

Add a same-home reconciliation sweep to session start so a home that was
interrupted mid-transition settles its own books on restart, replaying a
recorded close and restoring an in-flight row it already owns a worker
for. It never reads or writes another home; the fleet snapshot and the
cross-home nudge stay as backstops.

Close records are validated before they are trusted: the file is read as
raw bytes and rejected outright when it carries a NUL or other control
byte, every field must be well formed and non-duplicated, the id must
match the record it was found under, the data location must resolve
inside this home, and each close argument must carry a permitted,
well-formed value. Writer and reader share one validator so a record
this home publishes always remains replayable, independent of locale.

Homes configured for a manual backlog, and homes with no backlog at all,
stay exempt and are unaffected.

* no-mistakes(review): Remove stale bootstrap migration helper invocation

* no-mistakes(review): Preserve pending closes and narrow signal deferral

* no-mistakes(review): Record close before destructive teardown

* no-mistakes(review): Refuse pending closes before creating resources

* no-mistakes(review): Guard relaunches and preserve cleanup warnings

* no-mistakes(review): Reject symlinked records and clarify cleanup guidance

* no-mistakes(review): Align dispatch eligibility and protect close replay

* no-mistakes(review): Unify exact task incarnation parsing

* no-mistakes(review): Render resolved configured backlog path

* no-mistakes(review): Harden transition path boundaries against symlinks

* no-mistakes(review): Validate lifecycle state before resource actions

* no-mistakes(review): Enforce transition tooling and continuous state locks

* no-mistakes(review): Consolidate same-home lifecycle file boundaries

* no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts

* no-mistakes(review): Reject final-component lifecycle record symlinks

* no-mistakes(document): Document lifecycle record path boundaries

* no-mistakes(lint): Quote literal done tokens in atomicity tests

* no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched

* no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check`

* no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks

* fix(records): validate record bytes without an uncurated tool

The byte validation added for close records and directory paths shelled
out to od. The spawn and teardown lifecycle runs under a curated command
set that deliberately excludes it, so on any restricted PATH the check
could not run, the data directory read as unresolvable, and dispatch and
cleanup refused - wedging the lifecycle rather than protecting it.

An earlier attempt made the failing test pass by adding od to that
curated set. That fixed the test to agree with the defect and quietly
widened the contract the fixture exists to pin, so it is reverted here.

Inspect the bytes with perl instead, which is already in the curated set
and already used in this repo for the same portability reason. The
emitted values are identical to od's, so the rejection semantics are
unchanged: NUL and other control bytes are still refused, legitimate
paths containing spaces or non-ASCII characters still round-trip, and
the check stays independent of the process locale.

The restricted-PATH teardown case now passes because the validator no
longer needs od, not because the fixture was loosened.

* no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication

* no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342)

* fix: publish promote and Relay meta rewrites through contained replace

Bare mv still rewrote live task records in place, so a symlink meta could
be followed to a target outside state/. Route those field rewrites through
the shared publisher and drop the unused library aliases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse dangling symlinks during X metadata clear

* no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects

* no-mistakes(review): Exercise dangling symlink refusal through clear helper

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877)

* fix(watch): absorb a turn-end whose pane churned since the previous poll

The watcher's "absorb a benign turn-end when the crew is provably working"
triage was structurally unreachable for any harness whose semantic busy state
has no verified source. crew_absorb_class only reports working for an actively
running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can
only answer unknown for such an adapter, so codex crewmates surfaced a signal
wake at every turn boundary with nothing to act on - a full supervisor drain,
inspect and acknowledge turn per worker turn, scaling with the number of workers
in flight and drowning the wakes that matter in identical noise.

Widen the proof rather than bound the wake rate. A wake carrying only bare
turn-ended markers is now also benign when the task's pane content changed since
the previous poll, compared against the same state/.hash-* marker the staleness
backbone already records and already trusts as liveness. That evidence claims no
harness semantics, so it fabricates no busy verdict an adapter has not earned,
and it needs no adapter cooperation.

Absorb stays evidence-driven in both directions. A wake naming any status file
keeps the strict proof, every captain-relevant verb still surfaces immediately,
and an unresolvable task, a missing prior hash, a failed or empty capture, or an
unchanged pane all surface exactly as before. The absorb defers rather than
swallows: a crew that has stopped renders nothing further, so its now-static pane
surfaces through the staleness backbone within a poll or two. Bounding the
surfacing rate instead would have suppressed genuinely stopped workers.

The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which
owns it, and costs one bounded capture reached only for a no-verb turn-end whose
crew is not already provably working.

* no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Captain, reject malformed pane-churn hashes

* no-mistakes(document): Document pane-churn turn-end evidence

* no-mistakes: apply CI fixes

* fix(watch): gate and bound the pane-churn turn-end absorb

Make the pane-churn form of positive work evidence opt-in per home and
bound how long it may defer one endpoint's bare turn-ends.

Absorbing a bare turn-end on pane churn is now reached only when the home
creates config/turnend-churn-absorb. The other two proofs read a verdict
the harness itself vouches for, while this one infers execution from
rendered bytes, so widening the absorb is a home's choice rather than a
default every fleet inherits. With the flag absent the predicate returns
on its first line and triage is unchanged.

Churn and pane staleness read the same pane, so neither can be the
other's only backstop. A pane that renders continuously never presents
the two consecutive identical hashes the staleness backbone needs, so an
unbounded churn absorb left a worker that had genuinely stopped behind
such a renderer with no path to surface at all. One endpoint's turn-ends
may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS,
tracked in state/.churn-since-*, after which the wake surfaces and the
window restarts. The bound is evaluated before any .stale- state is
touched, so a wake that surfaces there leaves the staleness backbone's
own classification alone.

Covers both with behavioral tests: the same churning fixture that absorbs
with the flag surfaces and queues without it, and a spent deferral window
surfaces and restarts. The four existing safety guards now run with the
flag enabled so they keep proving their specific guard.

* no-mistakes(review): Fail closed on invalid churn deferral state

* no-mistakes(review): Validate persisted churn deadlines before arithmetic

* no-mistakes(review): Make churn deadlines transactional and bounds safe

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Clarify pane-churn supervision documentation

* no-mistakes(lint): Fix watcher arithmetic lint issues

* no-mistakes: apply CI fixes

* no-mistakes(document): Clarify pane-churn fail-closed documentation

* fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194)

* fix(bin): bind the live pipeline-owned run instead of a superseded failed row

fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead
of the LIVE replacement run: the live run's pipeline-owned lane head is not a
git object in the task worktree, so head-equality attribution rejected it and
the coarse runs-list fallback silently continued past the RUNNING row onto an
older failed row whose head equalled the stale worktree HEAD. The home summary
then flipped invalid and Bearings hid the home's live work (F10).

Attribution precedence now follows the daemon's own identity:
- An ACTIVE run for the task's branch binds without head equality while
  branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active);
  the pipeline owning the branch is itself the attribution.
- A genuinely failed run with no later run on the branch still reports failed
  through the unchanged head-equality path - real failures are not hidden.
- In the coarse runs scan, an unresolvable head is unknown attribution and
  stops the scan (fm_nm_head_resolvable) instead of falling through to an
  older row; a resolvable-but-mismatched head keeps the historical
  reused-branch skip.

The exemption never applies to a terminal run and requires pipeline_owned
specifically, both pinned by negative-control tests. Fixture shape verified
against the live incident run's real axi status output.

* no-mistakes(document): Updated run-attribution documentation ownership

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Align pane-churn watcher documentation

* no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement

Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse explicitly empty custom-check state overrides

* no-mistakes(document): Document custom-check retirement safety contract

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose
)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract
)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics
kunchenguid and others added 28 commits September 16, 2026 11:32
* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
…kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
…nchenguid#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by kunchenguid#4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: kunchenguid#3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and
kunchenguid#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
kunchenguid#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation
…ing is committed yet; the pipeline commits these changes. Each failure below was reproduced at HEAD and passes after the fix. ci-1, serial shard 4, two failures: - tests/fm-agy-harness.test.sh: the earlier upstream merge sends the shell-only verdict through the fork's strict idle-shell proof (bin/backends/herdr.sh). That proof needs foreground_process_group_id and a real pane-shell process named like the shell. Kept the earlier agent's fixture fix. It matches the shape tests/fm-backend-herdr.test.sh already uses (sleep symlinked as the shell name, pgid = pid = shell). Full suite passes. - tests/fm-session-lock-ownership.test.sh: upstream kunchenguid#4777 changed --claude mode. A non-owner Claude Stop now exits 0 with a diagnostic. The merged guard header documents this. The fork's report-once decline now covers only non-Claude harnesses, so the test now drives that mode. The earlier agent had deleted the block-budget assertion. I restored it: one --claude turn in the same home must exit 0, print "OWNED BY ANOTHER LIVE SESSION", and not write state/.turnend-claude-blocks. All 16 tests pass. ci-2, Herdr: tests/fm-afk-inject-self-deadlock-e2e.test.sh had a race. The native daemon's shutdown flush outlived stop_fixture. Its cleanup then deleted the shared pidfile after the new daemon wrote it. Kept the earlier agent's wait for "daemon shutting down", which the daemon logs after that delete. HEAD failed locally with the same message; the fix passed 3 of 3 runs against real herdr 0.9.0. ci-3, macOS: the 19th test is the fork-only test_large_backlog_does_not_hit_jq_argument_limit. It runs the real snapshot, bearings, and home-summary scripts. The CI log shows it passed under GNU bash 3.2.57. Kept the count change to 19; it also counts 19 locally. ci-4, parallel shard 1: no test hung. All 11 scripts finished (586 s of script time). The job was cancelled 0.5 s after the last one passed. The job log also hid a real failure: tests/fm-pi-primary-types.test.sh failed because review fix r16 used Array.findLast, which is ES2023, and the typecheck targets ES2022. Changed it to filter(...).at(-1), with the same behavior. The typecheck passes against Pi 0.85.1, and all 42 fm-pi-branch-extension checks pass, including the r16 regression. The shards grew after the merge: shard 1 took 3m44s on fork main and now takes about 10 min; the fork adds tests on top of upstream's shard and 28 more shell files to lint. Every subtest was also about 30% slower on that runner. Raised both parallel caps from 10 to 15 minutes, the same 1.5x margin the serial shards use, and updated the pinned caps in tests/fm-ci-workflow.test.sh. Shard 2 also changed because it shares the cap rationale and took 8.5 min. The commit message must say: raised the portable parallel caps because the shards grew after the upstream merge and fork additions; no test hangs. Checks run: bin/fm-lint.sh passes. bin/fm-test-run.sh --check-coverage passes. tests/fm-ci-workflow.test.sh passes. git diff --check is clean
@BohnBawerick
BohnBawerick merged commit d000503 into main Sep 18, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.