Skip to content

Merge upstream kunchenguid/firstmate main into the fork - #1

Merged
dsantosg1103 merged 100 commits into
mainfrom
fm/firstmate-sync-upstream-kun
Sep 24, 2026
Merged

dsantosg1103 merged 100 commits into
mainfrom
fm/firstmate-sync-upstream-kun

Conversation

@dsantosg1103

Copy link
Copy Markdown
Owner

Intent

Update this fork to the latest upstream kunchenguid/firstmate main.
The fork was 98 commits behind upstream (31c47af5) and 12 ahead.
This brings in the upstream work the update was wanted for, including Lavish feedback routed to the owning worker (bd65e4a), dispatch confidence floors and typed dispatch (795e4b5, 69d660a), the quota-axi schema 6 fix (259a669), and the remote-secondmate fixes (804394e, 36c9814).

How

A single merge commit of upstream/main into the fork's main, with no rebase, because the fork's history is published.
Git merged every file cleanly except five, all in no-mistakes run attribution: bin/fm-crew-state.sh, bin/fm-nm-run-lib.sh, docs/architecture.md, docs/configuration.md, and tests/fm-crew-state.test.sh.
The fork's 12 commits are two fixes plus their review and doc follow-ups, and both fixes collide with upstream's rewrite of how crew state picks a validation run.

Upstream replaced the approach both fork fixes patched.
It now reads the no-mistakes axi run overview, which carries run ids, and picks the newest run on the crew's branch by creation order (fm_nm_select_run).
It then reads that one run by id with axi status --run <id>, which returns its full step and gate detail.
Upstream also dropped the old rule that let an older live ledger row win over a newer terminal one.
Its new rule is that an older live row never hides a newer failure.
Both sides of every hunk were read, and each fork change turned out to belong to one of the two mechanisms upstream replaced.
So the conflicting code and docs now match upstream, and the two fork tests that still describe current behavior are kept.

The 12 fork commits

Commit Subject Outcome
d77216b fix(crew-state): let a rebased live run outrank the previous failed one Superseded by f5d7f5f (kunchenguid#4476) and dd9b2ef (kunchenguid#4973)
5ba4e9f review: bind rebased live sibling rows through the exact-head anchor Superseded with d77216b
0bedfd8 review: scope header no-widening claim, cover unanchored live sibling Superseded with d77216b
d796106 review: cover the live-class guard against diverged terminal rows Superseded with d77216b
62039d0 review: revert sibling widening, reject strict-ancestor live heads Superseded with d77216b
5290f64 document: document superseded-ancestor run heads in architecture Superseded with d77216b (the text it edits was rewritten upstream)
878de07 document: scope unbindable-run-head claim to the newest ledger row Superseded with d77216b (the text it edits was rewritten upstream)
643f551 fix(bin): report a live run parked at its gate as parked, not working Superseded by f5d7f5f (kunchenguid#4476), except the fleet-view test, which is kept
79d560f review: reject headless home-view rows when naming the live run Superseded with 643f551
19c2d2f review: match parked-run fixtures to the CLI's real shapes Partly kept: the steps-row-only parked fixture and its test are kept
4c232d7 review: bound the per-run inspection as enrichment, document its horizon Superseded with 643f551
3e033b8 document: retarget ledger-rule doc pointers, scope crew-state query timeout Superseded with 643f551 (the renamed function and the 3-second inspection bound no longer exist)

d77216b series: a rebased live run read as failed

  • Fork problem: crew state reported failed for a task whose live run had been rebased onto an advanced upstream.
    The older failed run at the worktree's exact commit bound and answered.
  • Fork fix: a live newest ledger row whose head diverged from the worktree could be recognized through the exact-head anchor row behind it.
  • Upstream: the selected run is the newest one on the branch by creation order, so the older failed run can no longer answer on a current CLI.
    An executing live run binds whatever its head (dd9b2ef).
    test_live_rebased_run_beats_older_failed_run_at_local_head pins this.
  • Conflict: upstream deliberately pins the opposite rule for the coarse ledger in test_coarse_live_rebased_row_is_not_attributed.
    That test sets up the fork's exact anchored shape and requires the row not to be attributed.
    Keeping the fork's widening would mean overriding that upstream contract, so this resolves to upstream.
  • What is different from the fork now (flagged, not changed):
    1. A live run that is parked at a gate, sits at a diverged but locally resolvable head, and shows no pipeline custody now reads unknown ("selected run code identity unverified", naming both run ids).
      The fork read it as parked.
      It is no longer read as failed.
    2. On a no-mistakes CLI too old to print the overview's runs table, the fallback path can still read the rebased-run shape as failed.
      The installed CLI (v1.72.0) prints the table, so that path is not reached here.
    3. The fork's upstream PR fix(bin): let a live rebased run outrank the previous failed one in crew state kunchenguid/firstmate#4590 is still open, so upstream can still choose the fork's rule.

643f551 series: a parked gate read as working

  • Fork problem: when the direct status read could not be bound, the ledger fallback reported a gate-parked run as working - validating (background run).
  • Fork fix: it took the run id from the home view and inspected that run for its gate.
  • Upstream: the selected route always reads the chosen run by id, so its gate detail is there directly.
    Upstream also reads the state database when the overview is capped, which removes the fork's ten-most-recent-runs limit.
  • Resolution: nm_inspect_live_run, fm_nm_home_view_run_id, the fm_nm_runs_decision_for_worktree rename, and their tests are dropped.
    In the new structure the coarse ledger word is only reached when the overview is missing or has no row for the branch, so that code could never run.
    The fork's fixture tests that still hold were also run against the merged code.
    The two gate-reporting cases still report the gate, now in upstream's detail format, which includes the run id.
  • Kept: tests/fm-fleet-snapshot-view.test.sh::test_parked_live_run_is_a_hold_not_active_work, which checks that the fleet view lists a gate-parked crew as a hold rather than active work.
    Its no-mistakes fake now also answers bare no-mistakes axi with the home view, as the real CLI does, because upstream's selection reads it.
  • Kept: tests/fm-crew-state.test.sh::test_awaiting_agent_without_gate_block_names_the_step_row_gate and its run_parked_awaiting_only fixture.
    They cover the real v1.72.0 shape where a parked run has no gate: block, which none of upstream's fixtures do.
    The fixture comment was reworded because the sibling fixture it pointed to is dropped.

Diff against upstream

After the merge the tree matches upstream/main except for those two kept tests (git diff upstream/main --stat: 2 test files, 152 insertions, no code or doc changes).
No new features, and no private or gitignored files touched.

Checks

The captain asked for no broad local checks and no wait for CI.
Only the tests this resolution changes were run:

  • test_parked_live_run_is_a_hold_not_active_work: ok - a live run parked at its gate is a hold, not active work
  • test_awaiting_agent_without_gate_block_names_the_step_row_gate: ok - an awaiting_agent run with no gate block reports its step row's gate
  • bash -n on both edited test files: clean.

Shellcheck and the full suites were not run locally; CI covers them.

aminry and others added 30 commits September 15, 2026 19:21
…kunchenguid#4586)

* fix(watch): honour a declared wait before wedge-escalating a quiet pane

wedge_timer_check escalated on elapsed idle time alone. Nothing asked
whether the worker had already said why its pane was quiet, so a lane
that declared a bounded external wait climbed the escalation ladder for
as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every
repeat carried demand-deep-inspection - which by its own wording forbids
re-absorbing on the run-step or pane state, so the supervisor could not
use the evidence that was there either.

The generated brief promises that declaring `paused:` buys the long
recheck cadence instead of a wedge, but the timer was still reachable
while that declaration stood: a crew that declares a wait and then has an
active run or busy pane attributed to it is handed to the timer as
provably-working. The declaration is what the worker said about its own
silence, so it now outranks a liveness verdict that only says something
is running.

The consult runs in the at-threshold branch that was about to escalate,
beside the worktree walk already there, and costs one status-line read.
Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS
recheck the declared-wait absorber already uses, so the wait is still
rechecked and cannot rot invisibly. Which verb declared it decides the
wording, because the two block on different people: a `paused:` wait is
owed by an external dependency and asks the reader to confirm it still
holds, while a `captain-held:` transfer is owed by the captain reading
the recheck and asks them to answer or release the hold. A hold is not
rechecked at all while the away-posture record exists, as on every other
captain-held path, and that absorb arms no throttle so the recheck is
owed in full on return.

A declared clearing time that has already passed stops counting, and a
lane that never declared one keeps the identical escalation schedule,
reason, count and demand-deep-inspection wording, so detection and its
worst-case time are unchanged. The deferral restarts the idle timer
rather than cancelling it, so a lane that stops waiting escalates again
within one threshold.

A lane quiet because its own validation run is parked at a gate awaiting
a human decision is deliberately out of scope: reading that state needs a
signal carrying who the wait is on and what clears it, rather than one
inferred from a parked verdict that also covers gates awaiting the
crewmate itself.

Tests pin both directions for each case and were each confirmed to fail
with the consult removed.

* no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record

A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists.

For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim.

Fixes kunchenguid#4607

* no-mistakes(review): Add bounded GitLab merge-request state reads

* no-mistakes(review): Preserve network-free inactive crew-state scans

* no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627)

* fix: restore published contribution follow-up (Fixes kunchenguid#4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
…kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
…nchenguid#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by kunchenguid#4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: kunchenguid#3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and
kunchenguid#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
kunchenguid#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
Lakescape and others added 29 commits September 23, 2026 00:07
…nguid#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
…kunchenguid#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859.
…enguid#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing
…unchenguid#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg
kunchenguid#5391)

The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.
…#5390)

* fix(bin): prune a torn-down task's wake rows at teardown

Prune pending durable wake rows (.wake-queue) for a task when it is torn
down, clearing stale wakes for its target window, signal wakes for its status
or turn-ended files, and task-specific check wakes.

Fixes kunchenguid#3419.
Adjacent to kunchenguid#5252.

- bin/fm-wake-lib.sh: add fm_wake_queue_prune_task
- bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown
- tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task

* no-mistakes(document): docs: note teardown prunes a task's wake rows

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
…ixture readiness (kunchenguid#5392)

* Make portable tests match resolved host paths

Summary:
- Match macOS full Node command paths by basename in the Gemini behavior test.
- Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom.

Testing:
- bin/fm-lint.sh
- bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh

Related:
- None

* no-mistakes(review): Mirror production PATH helper rules per directory group in tests

* no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep

* no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm

---------

Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
…ns (kunchenguid#5389)

The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation.

Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427)

Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358)

* feat(bin): add an opt-in per-home worker account pin

A home that mixes work and personal accounts for one runner had no way to
say which account its workers launch on: Claude workers inherited whatever
CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient
root, and an ambient API key outranked both, with no signal at launch.

config/claude-account and config/pi-account now pin that choice per home.
With neither file every launch is unchanged. With one, every launch of that
runner from the home (ship, scout, local secondmate, raw Claude command, and
relaunch) runs under the declared root, and the spawn refuses before any
endpoint exists when the file is malformed or the runner's own check
(claude auth status, pi auth check with a model-listing fallback) says the
pinned account is not signed in. The check runs in a cleared environment so
an ambient credential cannot answer for an empty root. A pinned Claude launch
sheds the environment credentials Claude ranks above a stored login; a pinned
Pi launch needs an explicit <provider>/<id> model for a declared provider and
also carries --provider. The chosen account is printed on the spawned line and
recorded in the task record, and relaunch checks the pin before stopping the
running agent.

* test(secondmate): give the concurrent config-push wait room for a slow host

test_config_reread_serializes_concurrent_pushes waited about two seconds for
the first fm-config-push.sh to reach its first send-keys. On a slower host
that push takes four to five seconds, so the test failed on main before the
push ever got there. The loop still leaves as soon as the marker appears, so
the larger bound costs nothing where the push is fast.

* no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470)

* fix(bin): keep the Herdr lab session option before a -- delimiter

fm-herdr-lab.sh run appended --session <lab> after every argument, so a
command with a passthrough delimiter such as agent start ... -- <agent args>
handed the session flag to the agent and Herdr routed the call by the
caller's ambient socket instead of the lab.
The helper now inserts --session <lab> immediately before the first --
delimiter and keeps the trailing form otherwise.

* no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host

Lease liveness is now the pure record test in every calling context, so an
unmarked main honors a live branch lease held by a separate process, and a
lease file engages the guard's claim serialization for any caller; a home
with no lease files still takes no lock.

bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while
FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another
harness resolves own, crew, and secondmate to the primary's harness.

fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with
a separate grace: the perl watchdog is preferred, runs the command in its own
process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps
the group, so a descendant holding captured output can no longer keep the
caller waiting past the bound on a host without timeout.

The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook
it terminated at the configured timeout, re-measured on Claude Code 2.1.281.

* fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group

Live runs of real Claude and Pi engine turns under the bound showed both CLIs
start every tool command in a process group of its own, so those processes end
through the engine's own TERM handling rather than the group signal or reap.
Also clears the new timeout test's ShellCheck findings.

* no-mistakes(document): Clarify cross-harness lease documentation
…henguid#2648)

* feat(bin): make the ship-branch prefix configurable per project

fm-brief.sh hardcoded every generated ship branch to fm/<task-id>, which
leaks that firstmate produced the branch/PR - unwanted for a third-party
public repo that does not use this tooling.

Add an optional --branch-prefix flag to fm-brief.sh (default "fm/", so
existing installs are unaffected) and teach fm-project-mode.sh - the
registry's single-owner parser - to resolve a project's optional
"branch=<prefix>" data/projects.md annotation via a new --branch-prefix
query, order-independent with the existing mode/+yolo tokens. Firstmate
resolves the override at task intake and passes it explicitly, mirroring
how --mode already works; fm-brief.sh itself never reads the registry.

An empty override resolves to a bare "<task-id>" branch rather than a
leading slash. All five previously hardcoded fm/$ID sites (branch
creation, never-push rule text, definition-of-done text, and the status
message) now render the resolved prefix consistently.

* no-mistakes(review): Wire branch-prefix intake in AGENTS.md; fix fm-merge-local.sh hardcoded fm/ prefix

* no-mistakes(document): docs: document configurable ship-branch prefix in architecture.md

* no-mistakes(review): Persist immutable branch contracts

* no-mistakes(document): Document configurable ship branch prefixes

* no-mistakes(lint): Captain: fix ShellCheck test warnings

* fix(bin): map bearings PR rows to their recorded ship branch (kunchenguid#1887)

fm-bearings-snapshot.sh keyed a PR back to its task by string-matching
the headRefName against the fm/ prefix, so any project whose branch
prefix was overridden (e.g. via kunchenguid#2648's branch=<prefix> registry
annotation) had its PRs silently drop to task "-" in the bearings
view, exactly the third fm/-assumption issue kunchenguid#1887 named alongside
fm-merge-local.sh and fm-bearings-snapshot.sh itself.

fm-fleet-snapshot.sh now surfaces each task's recorded branch=
metadata field in its JSON task rows, and fm-bearings-snapshot.sh
cross-references a PR's headRefName against those recorded branches
before falling back to the legacy fm/ prefix heuristic, so a custom
branch prefix maps a PR back to its real task.

Adds a regression test proving a PR opened against a fix/<task-id>
branch resolves to that task instead of "-"; confirmed it fails on
the prior startswith("fm/") logic and passes with this change.
ShellCheck clean; full fm-bearings-snapshot.test.sh and
fm-fleet-snapshot-view.test.sh suites pass.

* fix(ci): align lint arithmetic-looking assignment and stale Bearings snapshot count

- Quote the --branch-prefix want_value assignment in fm-brief.sh, fm-promote.sh,
  and fm-spawn.sh so ShellCheck SC2100 no longer misreads the plain string
  'branch-prefix' as arithmetic shorthand.
- Bump the Stock macOS Bash snapshot job's hardcoded Bearings test-count
  assertion from 59 to 60: this PR added a Bearings test, so the count was
  stale, not the feature.

* no-mistakes(review): fix(bin): honor recorded ship branch in relaunch and review-diff

* no-mistakes(document): docs: complete branch-prefix flag in brief and promote headers

* fix(lint): quote branch-prefix parser token; drop unused BRANCH_Q after rebase

* no-mistakes(review): Restore %q branch escaping in promotion instructions with regression test

* no-mistakes(document): document recorded ship branch and prefix flag

fm-review-diff.sh's header is the owner of its branch-resolution
contract; it still described only the legacy local-branch behavior
after the change made review-diff honor state/<id>.meta's recorded
ship branch. README's feature bullet enumerates the registry's
optional flags and was missing the new branch=<prefix> override.

* no-mistakes(lint): Silence SC2016 on intentional single-quoted sed expression

* no-mistakes(review): address branch-prefix review findings in DoD and project-mode

* no-mistakes(test): branch-prefix suites pass under tasks-axi 0.2.6; environment-only failure

* no-mistakes(document): purge stale fm/ branch naming from docs and headers

* fix(test): assert the merged epoch status wording in the branch-prefix override test

The rebase resolution of tests/fm-brief.test.sh kept the branch's
pre-merge \`done: ready in branch ...\` assertion while the merged
fm-dod-lib.sh (carrying main's epoch-stamped status line) renders
\`done [at=<epoch>]: ready in branch ...\`. Align the assertion so the
override-consistency test matches the behavior it verifies.

* no-mistakes(review): Address remaining branch-prefix findings in four bin scripts

* no-mistakes(test): skip real-tasks-axi tests below the repo's 0.2.6 floor

* no-mistakes(document): document spawn's branch-prefix registry deviation notice
…d per-rule confidence floors (kunchenguid#5478)

* feat(bin): send dispatch resolver only the brief's task sections

* Sent Jev only the scaffolded Captain's intent and Firstmate spec
  sections, falling back to the whole brief when neither heading is
  present, so the identical setup, rules, and definition-of-done
  boilerplate no longer reads as a signal about the task
* Added an optional per-rule min_confidence that replaces the global 0.6
  floor for that rule; a picked rule below its own floor falls to the most
  probable other option that clears its floor, or returns ambiguous
* Kept files with no declared floor on the exact previous behavior and
  kept the model blind to the new field
* Recorded the live old-versus-new comparison over scaffolded fixtures

* no-mistakes(review): share brief heading parser, add kind line, fix floors

* no-mistakes(test): stop sending ship delivery mode to jev, keep scout tag

* no-mistakes(document): docs: list shared brief heading lib in scripts inventory
* feat(bin): supervision host core behind config/supervision-host

Add the supervision host (bin/fm-supervision-host.sh): beside a Claude
primary it owns the watcher cycle for the Stop auto-arm and, while the
away-posture record exists, hands each wake to a bounded headless Claude
engine session that runs the supervision branch's contract - the same
generated prompt, row eligibility, wake grant, per-actor drain, outcome
store, leases, and away relocation the Pi branch uses. Attended wakes pass
straight to main. Every path that cannot finish a wake hands it to main
with a supervision-host line; the park ends itself before the Stop hook
timeout with a cycle-boundary wake.

- bin/fm-supervision-engine-lib.sh: opt-in parse, verified engines
  (claude, default sonnet), one bounded engine turn, and a reap of engine
  tool processes that sit in their own process groups.
- bin/fm-branch-report.sh: the command twin of fm_branch_report, scoped
  to the tasks the current host turn claimed.
- bin/fm-branch-dispatch.mjs: command entry to the Pi dispatch module, so
  eligibility and the wake prompt have one owner.
- bin/fm-claude-stop-autoarm.sh runs the host in the arm's place when
  config/supervision-host exists; nothing changes without the file.
- bin/fm-watch-arm.sh --stop: home-scoped stop without a re-arm.
- bin/fm-lease-lib.sh: an opted-in home takes the lease-command lock for
  unmarked main too, closing the first-claim race; the refusal tells the
  caller to leave the lease alone and retry.
- /afk launches no away daemon on an opted-in Claude home; /quiet still
  does. Session start renders the host's main-side protocol there.

* fix(bin): relay a host turn's outcomes when the captain returns mid-turn, and log per-turn engine cost

Live validation found two supervision host gaps. A captain who returns while
an engine turn is running gets a return brief rendered before that turn's
outcomes exist, so the host now hands the close to main with those outcomes.
Claude reports a resumed conversation's running cost, so the engine lib now
derives each turn's cost from the total the host records, and the host log
records every close's destination.

* docs(verification): record the supervision host's live evidence

The dated live results behind docs/supervision-host.md: the Claude engine's
live guard, the away-wake cases against real workers, the engine's cost
reporting, and the flag-off before-and-after regression.

* docs: describe the supervision host ledger as covering every close

* no-mistakes(review): Harden supervision host ownership, boundary, ack, and late outcomes

* no-mistakes(review): Recheck park boundary just before starting an engine turn

* no-mistakes(review): Cap park boundary, deliver all host lines, reject incomplete results

* no-mistakes(document): Correct supervision host documentation and stale pointers
…t-in, and delivery (kunchenguid#5506)

Attestation MATCH; contract-class restore; CI/NM green. Squash-merged by Kun's firstmate.
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
…dth (kunchenguid#5517)

* fix(bin): keep the ps fallback identity independent of terminal width

fm_pid_identity's portable fallback read the command column at the
ambient COLUMNS width, so an identity recorded from a wide shell never
matched the one recomputed inside a narrow hook and the continuity guard
denied every fleet command. Pass -ww so the column is never cut.

Fixes kunchenguid#799

* no-mistakes(ci): Fixed CI failure in Behavior portable serial 4. Root cause: the -ww flag added in commit ac7ab5d to fm_pid_identity (bin/fm-wake-lib.sh) shifted the ps argv so $1 became -ww instead of -p, breaking the positional fake-ps fixtures in tests/fm-procevent.test.sh (lines 2736, 3571) which then fell through to real ps and failed the fm-procevent test. Fix (already applied in the worktree, matching the authoritative user instruction exactly): replaced -ww with a COLUMNS=10000 environment pin so the call is COLUMNS=10000 LC_ALL=C ps -p "$pid" -o lstart= -o command=, mirroring fm_pending_reply_pid_identity in bin/fm-pending-reply-lib.sh:982. argv is back to -p PID -o lstart= -o command=, so the fixtures match again with no fixture edits. Comments above the call in bin/fm-wake-lib.sh and in test_pid_identity_is_terminal_width_invariant (tests/fm-watcher-lock.test.sh) now describe the COLUMNS pin instead of -ww; the regression test still asserts narrow-vs-wide byte equality and the full command. Verified: the terminal-width-invariant regression test passes. The only local not-ok results were flaky, run-varying timing tests (procevent launch/claim confirmation, listener reparenting) that differ each run and are unrelated to the ps argv change
…thorized intent (kunchenguid#5526)

Fixes kunchenguid#3608

When a scout is promoted to a ship, the captain's authorized intent is
extracted from a legacy `# Task` body by matching `Captain:` and `[captain]`
lines anywhere in the body, including inside fenced code blocks and indented
examples, while the heading reader already tracks fences. A fenced `Captain:`
example therefore passed the provenance gate and became the ship contract's
intent while the real ask was dropped.

Make the captain-words extractor fence-aware like the heading reader: a line
inside a ``` or ~~~ fenced block, or indented four spaces or a tab, is never a
marked line. The promotion and spawn callers need no change. The regression
test covers both the extractor and the promotion provenance gate refusing a
brief whose only Captain lines are fenced or indented examples.
…riefs (kunchenguid#2868)

* fix(bin): forbid administering the shared worktree pool in crewmate briefs

A crewmate ran a `git worktree remove` loop over the treehouse pool its own
worktree came from, destroying five worktrees - four belonging to tasks that
were running mid-pipeline. The generated brief's rule 2, "stay inside this
worktree; modify nothing outside it", is a rule about files: removing a
worktree is administration of shared state, not an edit outside a directory,
so the sentence never reached the act. The worker satisfied its brief
completely.

Rule 7 already named one piece of shared infrastructure - the no-mistakes
daemon, one instance serving every lane - with the reason stated plainly. The
worktree pool is the same class of thing and was unnamed.

Fold the pool into that existing rule rather than adding a second warning:
state the constraint around the act (create, remove, return, prune, move,
reassign a worktree or pool slot; write into a sibling slot), keep concrete
commands as examples rather than as the definition so no single provider is
pinned, and give the prohibition a real exit through `blocked:`.

The rule is emitted from one shared string interpolated into both crewmate
scaffolds, so the ship and scout copies cannot drift apart. The secondmate
charter deliberately omits it: that home runs its own fleet and legitimately
allocates and returns slots for its own crewmates.

Contract text only; no runtime enforcement layer.

* no-mistakes(document): Distill pool-safety comment rationale

* no-mistakes(review): align pool-rule test grep patterns with emitted [at=<epoch>] text
… a dispatch record (kunchenguid#5524)

* fix(bin): refuse tasks-axi add --start so In flight always has a dispatch record

Fixes kunchenguid#4753

Dispatch (bin/fm-spawn.sh) is the only path that moves a backlog row to
In flight, because it creates the task record, status file, and inbox
that go with the row. A row hand-placed there through the wrapper's
`add --start` had none of those, and nothing later noticed, so the
live-task count included work nobody was doing. The wrapper now refuses
`add --start` (exit 2) and names the dispatch path; plain `add` and
`start <id>` pass through unchanged, and the lifecycle transitions
address tasks-axi directly so dispatch is unaffected.

The issue's other half, a reconcile sweep in bin/fm-inactive-reconcile.sh
that notices an In flight row with no task record, is left as is; this
change closes the only path that creates such a row.

* no-mistakes(review): refuse create --start alias, not just add --start

* no-mistakes(review): reword add --start guard docs to drop only-path overclaim

* no-mistakes(review): scope add/create --start guard docs, drop universal claim
…#5503)

* feat(bin): run the supervision host beside the other non-Pi primaries while away

Cursor's stop-hook park, the OpenCode plugin, the omp watch extension, Grok's
model-owned background arm, and Codex's foreground checkpoint now run
bin/fm-supervision-host.sh in the watcher arm's place when the home opted in
with config/supervision-host, so the host's Claude engine takes away-posture
wakes beside those primaries exactly as it does beside Claude. Without the
file nothing changes.

- The host streams its first cycle's status line, accepts --restart and the
  owner's predecessor arm for its first cycle, and prints each exit in one
  write, so owners that wait for arm readiness and restart their own
  successor (OpenCode, omp) keep their handling handoff.
- Codex's checkpoint passes its bound to the host as the park boundary,
  raises it to FM_CODEX_WATCH_CHECKPOINT_AWAY (3600 s) while the away record
  exists, and lets an engine turn that starts before the bound finish after
  it (FM_SUPERVISION_HOST_PARK_LIMIT).
- /afk launches no away daemon on an opted-in home of those harnesses and
  says so at entry when the file selects no engine for that primary.
- Session start renders the host protocol for each arm owner, and Grok's
  arm command becomes the host.

* fix(bin): keep the watcher-down banner away from the supervision branch actor

A supervision host's engine turn runs guarded commands after its successor
watcher cycle may already have closed on a newer wake, so the guard showed it
the watcher-down banner with the primary's repair line. Under a Codex primary
pin that line is the checkpoint, and a live Codex lab run showed the away
session running it mid-turn (the nested host stood down on its ownership
check). The branch actor never owns watcher continuity, so the banner, its
reminder, and the episode state now leave that actor out, as the queued-wake
warning already does.

The lint telemetry fixture counts bin/fm-afk-launch.sh's source directives,
which the host engine note raised from four to five.

* fix(bin): queue away-session outcomes recorded after the return for main

A Cursor park superseded by the captain's return stops its host as the
engine turn ends, so the host's own handoff of that turn's outcomes was
never printed and the outcomes never reached main. The report surface now
queues every outcome it records after the away record is gone as a durable
check wake; the return owner archives the record before it reads the store,
so each outcome is in the return brief, queued, or both. A host stopped
mid-turn also removes its turn's result and error files.

The stream test now acknowledges its first close and accepts a restarted
cycle that closes on its resurface before the arm confirms it.

* fix(bin): clear a hard-killed host's turn at the next activation

A Cursor park superseded mid-turn can kill its host outright, which runs no
cleanup, so the turn's result, error, and descendant files stayed behind and
any tool process the engine started was left running. The next host's
activation now reaps the descendants that turn recorded and removes its
files. The host suite also registers its homes in a file, because
make_home runs in a command substitution, so its cleanup now stops every
host a case leaves running.

* fix(bin): leave rows that arrive after main's drain unclaimed at its acknowledgement

Main's acknowledgement re-claimed every unreserved queued row, including one
that arrived after the drain above the acknowledged cutoff. That row stayed
main's without ever being shown to it, so while away the supervision host
refused every later wake that included it and handed each back to main until
main drained again. The acknowledgement now claims only unreserved rows at or
below its cutoff.

* docs: name the killed turn's engine and files in the host's failure direction

* docs: record live supervision host runs on the non-Pi primaries

* no-mistakes(review): Replay host-only supervision boundaries across omp session replacement

* no-mistakes(review): Deliver omp supervision-host wakes only at the host's close

* no-mistakes(document): Correct supervision host documentation for non-Pi primaries

* no-mistakes(ci): Fixed the CI failure by naming FM_CODEX_WATCH_CHECKPOINT_AWAY in the rendered Codex host instructions. The focused instruction and checkpoint suites pass
)

* feat(bin): auto-relaunch dead persistent secondmates during ordinary supervision

A persistent secondmate whose primary agent exits mid-session previously
stayed down until the next session-start liveness sweep. Extract the
sweep's probe/classify/relaunch mechanics into a shared library and drive
the same contract from a cadence-gated watcher tick, so a positively dead
or missing endpoint is relaunched through the guarded spawn path within a
poll cycle instead of an hour later.

Only the recovery-grade `dead` and `missing` verdicts authorize relaunch;
ambiguous, unreadable, unverified, and unreachable-remote reads stay
fail-closed and a remote route is never replaced by a local endpoint.
Each relaunch emits exactly one `check` wake and appends to a durable
per-mate ledger; a mate exceeding the bounded attempt budget is parked
behind a marker until a live probe rearms it. A per-mate liveness lock
serializes the tick against a concurrent session-start sweep.

* no-mistakes(review): Fail closed on relaunch ledger errors; clear state on remote teardown

* no-mistakes(review): Share ledger read guard; retire relaunch state under liveness lock

* no-mistakes(review): Lazy-load wake lib; live rearm restores full relaunch budget

* no-mistakes(review): Finish liveness tick for every mate before waking once

* no-mistakes(review): Keep liveness tick scanning past per-mate errors, then wake

* no-mistakes(review): Wake only on queued rows; teardown holds liveness lock

* no-mistakes(review): Queue liveness outcome wake before releasing mate lock

* no-mistakes(document): Update secondmate liveness documentation for mid-session recovery

* no-mistakes(lint): Fix empty assignments flagged by ShellCheck

* no-mistakes(ci): Added ShellCheck analysis boundaries for the shared liveness library in both callers and marked its result globals as intentional library outputs. Changed-file lint passed; full CI partitions were not run locally

* no-mistakes(ci): Fixed Lint 2 by removing an unused test variable in tests/fm-wake-queue.test.sh. ShellCheck, bash syntax, and the full wake-queue test script pass

* no-mistakes(ci): Fixed the CI wake-queue fixture: stall-only watcher legs now seed the liveness cadence marker, preventing the new endpoint probe from interfering with their assertions. The full wake-queue test, ShellCheck, and diff checks pass locally
…cycle ends (kunchenguid#5550)

* fix(bin): start a successor when the Claude Stop-hook arm's attached cycle ends

Fixes kunchenguid#2381

When the Claude Stop hook's foreground arm attached to a peer watcher cycle
and that cycle ended, the arm reported the delivered wake and the hook exited
2 without starting a successor, so the handling turn ran with no watcher.
Pi, omp, and OpenCode start the next arm before delivering the wake and pass
the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID; the Claude hook never
passed that predecessor identity.

The hook now runs its arm as a tracked child it waits on, so it holds that
arm's pid, and after any actionable close starts one handling-successor
bin/fm-watch-arm.sh with the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID.
The successor is launched the one way a process outlives a Claude hook's
exit-2 rewake (nohup, detached stdio, own process group, the shape
bin/fm-startup-network.sh already uses); the hook waits for its status line
and adds one banner line when no live watcher was confirmed, never withholding
the wake. The supervision-host path is unchanged, as is the arm wrapper.

The regression test drives the real hook against an arm fixture whose attached
peer cycle ends: it fails on the previous tip because no successor starts, and
now asserts the successor names the closed arm as its predecessor and outlives
the rewake. A second case pins the unconfirmed-successor banner line.
docs/watcher-continuity.md no longer records the Claude asymmetry.

* no-mistakes(ci): Serial-4 failure was a real regression: tests/fm-session-lock-ancestry.test.sh asserts exact cumulative arm-invocation counts while driving the real fm-claude-stop-autoarm.sh hook against a stubbed fm-watch-arm.sh. This PR makes the hook start a handling successor after an actionable close, so every owned actionable phase now records TWO arm invocations (foreground arm + successor) instead of one, breaking "healthy chain: expected 1 arm(s), got 2". Fixed by updating the cumulative expectations to match the new behavior: owned phases 1/2/6 -> 2/4/6, foreign carry phases 3/4/5 -> 4, plus a comment explaining the +2-per-owned-phase model. Verified: phase-1 (the CI failure point) now passes on every run, syntax checks clean, and sibling arm-count tests (fm-claude-stop-autoarm.test.sh, fm-cursor-primary, fm-turnend-guard) pass unchanged. The only remaining local not-ok is a WSL-only environmental artifact (orphan reparents to a subreaper, not PID 1) that passes on the CI runner. Parallel-1 failure is an unrelated flake: its 11 tests (fm-lint, fm-pr-merge, fm-test-run, fm-cd-pretool-check, fm-pi-primary-types, fm-grok-harness, fm-composer-lib, fm-review-diff, fm-tmux-submit-busy, fm-composer-ghost, fm-brief) do not include fm-session-lock-ancestry and none reads any file this PR touches; all pass locally. It should clear on CI re-run. Made the smallest root-cause fix (one test file, 6 count updates + a clarifying comment). Validation of the branch continues through the no-mistakes pipeline, which owns re-running CI
Brings the fork's main up to upstream 31c47af (98 upstream commits).

The fork's two crew-state fixes conflicted with upstream's rewrite of
no-mistakes run attribution. Upstream's identity-aware run selection
(f5d7f5f, kunchenguid#4476) and its rebased-run handling (dd9b2ef, kunchenguid#4973) solve
the same problems by a different mechanism, so the conflicting code and
docs resolve to upstream's version:

- d77216b and its review commits (live rebased run read as failed):
  the newest same-branch run is now selected by id, so the older failed
  run no longer answers; upstream pins the opposite coarse-ledger rule
  in test_coarse_live_rebased_row_is_not_attributed.
- 643f551 and its review commits (parked gate read as working): the
  selected run is read by id with its full gate detail, which replaces
  the fork's home-view inspection.

Kept from the fork: the steps-row-only parked gate test and the
fleet-view hold test, with the fleet fake now printing the home view
upstream's selection reads.
The merge keeps the fork's fleet-view test that a ship crew parked at a
no-mistakes gate is listed as a hold, which upstream does not cover.
@dsantosg1103
dsantosg1103 merged commit aa78b7f into main Sep 24, 2026
18 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.