Skip to content

feat(bin): sync upstream supervision and runtime updates - #15

Merged
knowttl merged 32 commits into
mainfrom
fm/fm-upstream-sync-5
Aug 2, 2026
Merged

knowttl merged 32 commits into
mainfrom
fm/fm-upstream-sync-5

Conversation

@knowttl

@knowttl knowttl commented Aug 2, 2026 •

Copy link
Copy Markdown
Owner

Syncs the fork with upstream kunchenguid/firstmate main at cd73e75, covering the 27 commits the fork was behind since 99533c5.

This is a true --no-ff merge, so upstream ancestry is preserved. Upstream was fetched by URL only; no upstream remote was added, and git remote -v shows only origin (knowttl) plus the pipeline's own no-mistakes remote before and after.

Upstream commit review

Each of the 27 commits, in merge order.

6ec5e08 fix(bin): normalize relative durable paths (kunchenguid#1256) - Resolves relative home, data, and state inputs to absolute paths before generating durable charters, and uses absolute paths at the spawn, AFK daemon, and X-mode cross-process handoffs so a later process cannot reinterpret them from a different working directory. Also handles dash-leading harness process names. Relevant to this fork because it hardens the same ancestry and spawn surfaces the fork's own Claude spawn work touches.

c21bf54 refactor(skills): make Bearings chat-only by default (kunchenguid#1136) - /bearings now reports in chat by default; writing the dated data/status-report-<date>.md artifact requires the explicit file argument, and live PR enrichment stays opt-in. Behavioral change to a captain-facing command: plain /bearings no longer leaves a file behind.

96e027e Clarify follow-up routing during validation (kunchenguid#1277) - One line in AGENTS.md clarifying how follow-ups route while a validation run is active. Instruction-only.

a24eac1 fix: honor concrete approval for project operations (kunchenguid#1272) - Adds an exception to hard rule 1: firstmate may directly edit, create, move, or delete project files when the captain concretely approves that specific operation or scope in the moment. The approval is never inferred, broadened, or standing, and the force, discard, unlanded-work, and merge-authority boundaries stay independently in force. This is a real widening of firstmate's own authority over projects/ and worth the captain's attention.

0bbb27b fix(skills): route new project intake through secondmate scopes (kunchenguid#1275) - New-project intake is matched against registered secondmate scopes rather than defaulting to the main home. Affects routing for any new project added to this fork's fleet.

daf6dce fix: scope validation corrections by accepted behavior (kunchenguid#1281) - Narrows what counts as an autonomous validation correction to what accepted behavior already covers, and classifies stale delivery evidence as autonomous. Tightens ask-user-authority; interacts with this fork's yolo-off posture by reducing what gets auto-decided.

a2d5f26 test: replace source assertions with behavioral coverage (kunchenguid#1282) - Large test cleanup removing ~3300 lines of assertions that grepped source text rather than exercising behavior. Reduces false failures when an implementation is rewritten but behavior is preserved, which is directly why several of this merge's conflicts were small.

56a7ac6 fix(watch): escalate busy workers with no completed turn (kunchenguid#1286) - Adds FM_BUSY_TURN_MAX_SECS (default 3600s) bounding how long a busy pane may run with no completed turn, routing an over-age pane through the existing wedge_timer_check for human inspection only. Motivated by a hung foreground call sitting behind an unchanging "Working..." footer for 25h. Directly overlaps the fork's PR 14 - see the drop note below.

e595611 fix(gitignore): ignore config/ as a directory (kunchenguid#1261) - Replaces a name-by-name config/ ignore list with a directory ignore, so a new home-local config file no longer makes the tree read dirty and block guarded sync paths. Practical benefit for this fork's own sync paths.

a53ffc1 fix(tests): replace source-content .gitignore assertion (kunchenguid#1304) - Swaps the .gitignore spelling assertion for a real git check-ignore behavioral test. Test-only.

79e62b8 feat: bound and consolidate startup memory during stow (kunchenguid#1303) - Adds a per-home startup-memory budget (config/startup-memory-budget, materialized at 7,500 estimated tokens) that /stow curates against, inherited into secondmate homes. New bootstrap diagnostic STARTUP_MEMORY_BUDGET. This fork's homes will start materializing that file.

f0d7cbe fix(herdr): place workers in the launching workspace (kunchenguid#1328) - Herdr spawn previously resolved its container by taking the first workspace whose label matched, so with two same-labeled workspaces a worker landed in the wrong one. Placement now binds to the launching process's own Herdr pane identity, resolved live, and refuses rather than degrading to a label search. Meaningful for this fork if it runs Herdr-backed work.

3a112a1 fix(calm): refine Calm working boat animation (kunchenguid#1339) - Replaces Pi's working row with an animated boat while Calm is on, with directional sail and animated water. Presentation-only, gated on config/calm.

28b02d2 fix(dispatch): preflight candidate auth before quota escalation (kunchenguid#1349) - Added bin/fm-auth-preflight.sh to resolve a dispatch candidate's authentication surface from quota-axi's own emitted auth sources rather than from harness or model names, so one vendor CLI's expired token could not gate an unrelated candidate. Gates quota-axi at 0.1.16. Largely superseded two commits later by f7d0d0a.

5fca47f feat(x-mode): reconcile promised public replies deterministically (kunchenguid#1350) - Makes a promised final public reply durable state rather than something the primary must remember: typed terminal results are reported into the owning home's inbox, commitments reconcile through tasks-axi public-followup, and teardown refuses while a home still owes a public reply. Inert unless X mode is opted into via .env.

96542a4 feat(bin): replace busy heuristics with semantic lifecycle state (kunchenguid#1327) - The largest behavioral change in this sync. Replaces pane-text busy heuristics with a semantic per-task busy record (bin/fm-busy-lib.sh) written only by bin/fm-busy-event.sh, with per-harness trusted-source classification and busy/idle/unknown/dead semantics where missing, malformed, stale, or untrusted data is unknown and never idle. Pi, OpenCode, and Claude are converted to real lifecycle hooks; Codex classifies unknown behind an explicit probe; a Grok-only rendered-tail fallback remains. This subsumes both fork changes dropped below.

621299a fix: preserve Calm boat continuity across working periods (kunchenguid#1356) - Calm's boat resumes from its frozen column across runs instead of restarting at the left edge. Presentation-only.

f7d0d0a fix: restore evidence-based dispatch eligibility (kunchenguid#1358) - Retires deterministic shell dispatch eligibility. fm-auth-preflight.sh built a pi:<model-prefix> source id, which wrongly rejected supported Pi candidates in the Codex family; it is replaced by bin/fm-vendor-auth-probe.sh, which holds no routing knowledge. The dispatching firstmate now establishes support and provider family from each harness's own catalog and only concrete contradictory evidence blocks a candidate. Also fixes a zero FM_*_TIMEOUT silently removing the hard bound.

3772964 docs: define captain instruction precedence (kunchenguid#1362) - Adds an always-loaded rule that a current explicit concrete captain instruction overrides a conflicting firstmate-written standing rule within its exact scope, never above platform/system/developer instructions. Keeps the red-PR default and the yolo boundary.

9fdef64 docs: define validation supersession sequence (kunchenguid#1407) - Supplies the missing sanctioned path when a captain instruction invalidates work mid-validation: abort through no-mistakes, confirm through status, recover branch custody through sync, then replace the obsolete work from the correct pre-invalidation base and validate once against the final head.

000c1db fix: bind backend overrides to exact-task authority (kunchenguid#1413) - An explicit --backend is authorized only for the exact task it was given for, after a Herdr secondmate carried a prior one-task --backend tmux exception forward by analogy and put its child in the wrong runtime.

66b0f77 fix(herdr): prevent focus flashes during projected workspace cleanup (kunchenguid#1229) - Removes workspace-emptying panes through Herdr's focus-preserving pane-death path instead of an explicit close, which was stealing the attached client's focus and misrouting in-flight keystrokes. Also fixes a case where a refused close could erase a task's records while its pane survived.

a805766 fix: prioritize completion runway in quota-aware dispatch (kunchenguid#1431) - Dispatch selection weights usable completion runway rather than raw remaining percentage, so a candidate that cannot finish the work is not preferred on headroom alone.

68641a3 fix(bin): preserve full task contract in no-mistakes intent (kunchenguid#1447) - fm-brief.sh carries the complete current task contract into the no-mistakes intent instead of a truncated form.

1e24757 fix(bin): parse punctuated secondmate registry entries safely (kunchenguid#1452) - Centralizes secondmate registry parsing into bin/fm-secondmate-registry-lib.sh with EOF, symlink, and unreadable-registry validation, replacing ad-hoc parsing that mishandled punctuated entries.

8c21b10 feat(bin): add durable process-event supervision (kunchenguid#1483) - Adds a domain-neutral process-to-event runner (bin/fm-procevent.sh) letting firstmate wait on a blocking external source without holding a turn, with one machine-wide owner per source, durable 0600 result capture before publication, and an explicit handled acknowledgement as the only thing that stops bounded re-announcement. Introduces the process-event-sources skill.

cd73e75 fix(bin): retire terminal process events and surface queued wakes (kunchenguid#1500) - Two fixes on the above: the runner now asks a source's own adapter whether it has terminated and retires it, ending a loop where one "Send & End" produced recurring empty results; and a captured result queued as a wake is now proactively surfaced by a healthy watcher instead of waiting for a manual drain.

Conflicts and how they were resolved

Six files conflicted.

  • bin/fm-push-transition-lib.sh - combined. The fork publishes a delivery receipt inside wake(); upstream wrapped the same function's echo in an output-status/FM_WAKE_POST_OUTPUT_ACTION hook. The two are orthogonal, so both are kept, with the receipt published before the wake line is written (the fork's original ordering).
  • bin/fm-session-lock-lib.sh - kept the fork's version. Upstream's only change here was basename to basename -- dash-hardening. The fork's fm_process_row walk uses ${FM_ROW_COMM##*/} and never invokes basename, so it already satisfies that intent, and the surrounding fork code references FM_ROW_*. Upstream's second basename -- site merged cleanly and is retained.
  • bin/fm-watch.sh - resolved to upstream exactly. See the two drops below. This file now matches cd73e75 byte for byte.
  • docs/configuration.md - resolved to upstream (FM_BUSY_TURN_MAX_SECS documented, PR 14's FM_WEDGE_WORKING_ESCALATE_SECS removed).
  • tests/fm-crew-state.test.sh - resolved to upstream, with one fork test restored (see retained work below).
  • tests/fm-daemon.test.sh - resolved to upstream.

Fork work dropped

PR 14 ("fix: defer stale wedge alarms for active crews") is dropped in full. Its purpose was suppressing false stale-wedge alarms for workers sitting in long in-contract calls, and upstream's semantic-lifecycle rework subsumes that purpose at the source. Under 96542a4, a worker inside a long blocking call reports semantic busy for the whole turn, and bin/fm-watch.sh gates the stale path on busy_now -ne 0 - so such a worker never reaches the 240s stale threshold that produced the false alarms in the first place. PR 14's deferral machinery is therefore unreachable for its own motivating case. Where it would still be reachable, it is actively harmful: 56a7ac6 routes a busy pane with no completed turn into wedge_timer_check at FM_BUSY_TURN_MAX_SECS precisely to bound a hung call, and PR 14 would defer that escalation for up to another hour, re-widening the exact bound upstream shipped to close. Dropped: wedge_escalation_deferred and FM_WEDGE_WORKING_ESCALATE_SECS from bin/fm-classify-lib.sh, the deferral in bin/fm-supervise-daemon.sh housekeeping and in bin/fm-watch.sh's wedge_timer_check, the .wedge-verified-* / .subsuper-wedge-verified-* markers, and the associated documentation.

The bin/fm-watch.sh half of the fork's residual-busy-frame fix (16c9e67) is dropped. That fix added an agent-presence override so a dead agent's last painted frame could not keep matching the harness busy signature forever. Upstream's rewritten window_is_busy does no rendered-frame matching at all outside the Grok-scoped fallback, so there is no frame branch left to guard. The same dead-agent condition is still bounded upstream: the stale record keeps reporting busy, and busy_turn_over_age escalates it at FM_BUSY_TURN_MAX_SECS. This is a detection-latency tradeoff, not a silence - the fork caught it within a poll, upstream catches it within an hour - and fork policy is to match upstream.

Fork-only tests dropped with the implementations they asserted:

  • tests/fm-watch-triage.test.sh test_window_is_busy_requires_a_present_agent - asserted the dropped window_is_busy presence override.
  • tests/fm-daemon.test.sh test_housekeeping_in_contract_stale_defers_then_escalates and test_housekeeping_in_contract_stale_escalates_past_allowance - asserted PR 14's deferral.

Fork work retained

The bin/fm-crew-state.sh and bin/fm-backend.sh half of 16c9e67 is kept, because upstream has not subsumed it. window_is_busy and fm-crew-state.sh are different consumers. Upstream's fm_busy_classify_meta - the entry point fm-crew-state.sh reaches - performs no endpoint-liveness read; only the separate fm_busy_classify_live does. And pane_readable tests structural pane existence, which an agent-less but still-alive pane passes. So on this path the fork's fm_backend_agent_confirmed_absent gate is still the only thing that settles an agent-less endpoint, and dropping it would reintroduce working · pane forever for a dead agent on a targeted crew-state read, which has no busy_turn_over_age backstop. Its regression test test_no_run_herdr_dead_agent_busy_frame_reads_gone was restored alongside it.

All other earlier fork work merged cleanly and is retained: the coalesced wake drain, Herdr Pi queued-input steering confirmation, U+00A0 composer classification, the Claude spawn and process-ancestry work, and the codev-session skill.

Pipeline document-step edits: per-file parity outcome

The validation pipeline's document step rewrote explanatory comments in five files. The captain applied a per-file parity test - a file that was byte-identical to upstream before the document step gets its edits reverted, because a comment improvement is still drift that re-conflicts on every future sync for zero behavioral value, and better comments belong upstream. Files already divergent in this merge keep their updated comments. Outcome per file:

File Upstream parity before the document step Outcome
bin/fm-watch.sh byte-identical Reverted. Restored to byte-identical with cd73e75, verified by an empty git diff cd73e75 -- bin/fm-watch.sh
bin/fm-backend.sh already divergent (the retained agent-presence block) Kept - the edits sit inside the fork's own block and added no new drift
bin/fm-crew-state.sh already divergent (the retained presence gate) Kept
docs/architecture.md already divergent Kept
docs/herdr-backend.md already divergent Kept

Note for future syncs: the reverted comments in bin/fm-watch.sh are genuinely stale upstream - they still describe the rendered-frame matching that 96542a4 removed. That is an upstream documentation defect, recorded here rather than fixed in the fork.

Accepted limitation: Codex and standalone Kimi workers can still be falsely wedge-alarmed

Dropping PR 14 leaves one class of false alarm open, and the captain accepted it as a documented limitation rather than fixing it in this PR.

The rationale for the drop - that a worker in a long in-contract call reports semantic busy and so never reaches the 240s stale threshold - holds only for the adapters upstream actually converted: Pi, pi-signed, OpenCode, and Claude. Upstream 96542a4 deliberately classifies Codex as unknown codex-unverified and standalone Kimi as unknown kimi-unverified, because neither lifecycle source was verifiable on the installed binaries. window_is_busy returns busy only on an exact busy verdict, so such a worker stays on the stale path. Meanwhile fm-crew-state.sh classifies an active no-mistakes run as authoritative working with source: run-step, so the pane is first absorbed as working and then surfaced as possible wedge at 240s. The away-mode daemon has the same reachable path because it checks only semantic busy state.

If a Codex- or Kimi-backed crewmate is ever dispatched, it re-enters the 240s false-alarm class that PR 14 was suppressing. The durable fix belongs upstream, in the semantic busy contract, not in fork-only watcher behavior.

Why it is accepted here rather than fixed: this is a sync PR under match-upstream-exactly, and the proposed remedy - having timed escalation re-check authoritative source: run-step state - is new fork-only watcher behavior, the same drift shape PR 14 was and that upstream just subsumed. Building its successor inside the sync would restart that cycle. Current exposure is zero: every dispatch route and the secondmate pin run claude today and Codex is dormant, so no affected worker exists. This section is the evidence base for an upstream request if that fleet configuration returns.

Surfaced by the pipeline's own review step as ask-user finding codex-active-run-stale-escalation (bin/fm-watch.sh:273, severity error), and approved by the captain with this documentation condition.

Defects found in upstream

Recorded as evidence, not fixed here, and no upstream PR was opened.

  1. Five upstream suites are red on upstream itself. Each was reproduced against a pristine git archive extraction of cd73e75 containing no fork code, compared by exit status: tests/fm-calm-pi-extension.test.sh, tests/fm-pi-watch-extension.test.sh, tests/fm-public-followup.test.sh, tests/fm-test-run.test.sh (whose assertions all pass but whose trailing coverage guard errors comm: file 2 is not in sorted order and exits 1), and tests/fm-session-start.test.sh (concurrent session-lock acquisition produced 40 winners). All five fail identically on the merged tree, so this sync neither introduced nor worsened them.
  2. fm_busy_record_read has no time-based expiry. "Stale" there means a generation mismatch only, so an agent hard-killed mid-turn leaves a gen-matching state=busy record that fm_busy_classify_meta reports as busy indefinitely. This is bounded in the watcher by busy_turn_over_age and on the crew-state path by the fork presence gate retained above, so it is not an unbounded hole.

Verification

Baseline method: every failing suite was re-run against a pristine extraction of upstream cd73e75 and compared by exit status, not by counting not ok lines - tests/fm-test-run.test.sh fails with zero not ok lines, so a text-based comparison misclassifies it.

Merged tree, bin/fm-test-run.sh --all: 104 suites, 7 failing.

Suite Verdict
fm-brief Pre-existing fork defect, fixed in this PR (c0a8c1f)
fm-secondmate-harness Merge-induced (upstream=0, merged=1), fixed in this PR (e31c463) - see drift note below
fm-calm-pi-extension Inherited (upstream=1, merged=1)
fm-pi-watch-extension Inherited (upstream=1, merged=1)
fm-public-followup Inherited (upstream=1, merged=1)
fm-test-run Inherited (upstream=1, merged=1) - all assertions pass; the trailing coverage guard errors comm: file 2 is not in sorted order and exits 1
fm-session-start Inherited (upstream=1, merged=1) - concurrent session-lock acquisition produced 40 winners

A post-fix --all run then completed: 104 suites, 6 failing. Both fixes are confirmed - fm-brief and fm-secondmate-harness moved to passing.

This suite is not deterministic under full-suite load, and that is stated here rather than smoothed over. Across the pre-fix and post-fix --all runs the failing set was not stable:

  • fm-pi-watch-extension failed in the pre-fix run and on pristine upstream, then passed in the post-fix run.
  • fm-kimi-harness and fm-watcher-lock failed in the post-fix run having passed earlier; fm-watcher-lock had been run individually to completion at exit 0 during the ancestry verification above.

A suite that flips in both directions across runs of an unchanged tree is load- or order-dependent, not a stable signal. fm-pi-watch-extension flipping is direct proof of that. So the honest statement is not "exactly five inherited failures": it is that four suites failed identically on pristine upstream when compared head to head (fm-calm-pi-extension, fm-public-followup, fm-test-run, fm-session-start), that the two suites this PR set out to fix now pass, and that fm-kimi-harness and fm-watcher-lock are unclassified pending a stable comparison against pristine upstream under the same conditions.

Neither unclassified suite is reached by the retained fork behavior in an obvious way, but that is a hypothesis, not a verdict, and it is recorded as unresolved rather than asserted. The pipeline's own test step is the authoritative check for this PR and its result governs.

bin/fm-lint.sh passes clean (exit 0) against its pinned ShellCheck 0.11.0, run against the exact head being shipped. bin/fm-doc-audience-check.sh passes (62 surfaces, 173 local links).

The validation pipeline's own lint step could not run: it exited 127, ShellCheck not found, because the binary is absent from the pipeline's environment. That is a tooling gap, not a clean lint result, so it was verified out of band instead of taken on trust.

This PR has no CI. GitHub Actions reports 0 passed, 0 failed - this PR has no CI checks configured, and the repository has no workflow-run history. The pipeline's checks-passed outcome therefore rests on an empty check list, not on green CI. Nothing here has been verified by CI; the evidence above is local.

Deliberate drift in an upstream test (captain-approved)

tests/fm-secondmate-harness.test.sh carries one fork-only line: FM_PROC_ROOT_OVERRIDE="$dir/no-proc" on the session-lock ancestry assertion. This edit to upstream-authored content will conflict on every future sync that touches this test, and must be re-applied until upstream adopts an equivalent /proc fast path. The line carries an inline comment naming the fork behavior it accommodates so the next sync's conflict is self-explaining.

Why it is needed: upstream 6ec5e08 added test_dash_leading_process_names_are_basename_operands, which stubs ps. This fork's fm_process_row reads /proc before ps, so on Linux it walks the real process tree and never observes the stub. Diverting the proc root forces the portable ps fallback, which is the form the fixture stubs. Upstream needs no override because it only ever calls ps.

The fork fast path is retained rather than dropped because upstream has not subsumed it: upstream still forks ps, basename, and grep per hop, measured at 2.5-3.3s, which pushed Claude auto-arm past FM_CLAUDE_AUTOARM_SYNC_WAIT_MS and made the turn-end guard block turns whose auto-arm was working normally.

One supporting change in fork code, bin/fm-session-lock-lib.sh: a pid the ps -A snapshot does not list now falls through to a per-pid ps read instead of failing. This is purely additive - the /proc fast path and the one-snapshot-per-shell behavior are untouched - and it is what lets the fixture be satisfied by a single-line test override rather than by teaching upstream's stub about this fork's snapshot form.

Proof the kept behavior still holds, per the captain's condition: tests/fm-claude-stop-autoarm.test.sh's test_ancestry_resolution_is_fork_frugal_and_bounded passes on the merged tree. It fails outright if ps is called at all on the /proc path (its stub exits 97) and if resolution exceeds 1000ms, so the fast path is proven intact, not merely present. tests/fm-watcher-lock.test.sh and tests/fm-procevent.test.sh, the other ancestry consumers, also pass.

One note on why the fork's code was never actually unsafe for dash-leading names: /proc's comm field never carries a login shell's leading dash, and the ps fallback uses ${FM_ROW_COMM##*/} rather than basename. The fixture's intent held all along; only its ps-coupled mechanism did not.

sparkus and others added 30 commits July 29, 2026 22:40
* fix(bin): handle dash-leading harness process names (#2)

* fix: handle dash-leading harness process names

* no-mistakes(review): Make dash-leading harness regression hermetic

* fix: preserve secondmate reply routes across relative homes

Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved.

Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory.

* no-mistakes(review): Preserve absolute overrides and normalize relative durable paths

* no-mistakes(review): Normalize relative home before deriving durable paths

* no-mistakes(document): Document relative durable-path normalization

* no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization

* no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck
* Add internal status skill

* no-mistakes(document): register /status skill in documentation-audiences inventory

* no-mistakes(lint): replace grep|wc -l with grep -c in status skill test

* test: silence literal status skill patterns

* Refactor bearings default to chat-only

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* no-mistakes(review): Align project removal preflight with approved exception

* no-mistakes(document): Align project removal documentation with approved exception

* fix: restore removal test byte-for-byte and preserve the default sentence

tests/fm-instruction-owners.test.sh had been changed to assert different
text; restore it byte-for-byte to origin/main. project-management SKILL.md's
Remove section now keeps the exact default "Never issue a raw removal
command from Firstmate." sentence that test still asserts, immediately
followed by the already-approved captain-operation-or-scope exception, so
the default and the exception both stay explicit and consistent.

* no-mistakes(document): Align project-write boundary documentation
…henguid#1275)

* Route project intake through secondmate scopes

* no-mistakes(test): Guard all main-home project registry mutations

* no-mistakes(document): Consolidate secondmate routing documentation

* no-mistakes: apply CI fixes

* Restore new-project routing scope

* no-mistakes(document): Clarify secondmate routing for new-project intake

* no-mistakes: apply CI fixes
)

* fix: scope validation corrections by accepted behavior

* no-mistakes(review): Classify stale delivery evidence as an autonomous correction
…#1282)

* test: remove source-content assertions

* no-mistakes(review): Replace source assertions with runtime behavior coverage

* no-mistakes(review): Isolate Kimi task temp runtime coverage

* no-mistakes(document): Refresh test cleanup documentation

* no-mistakes: apply CI fixes
…#1286)

* fix(watch): bound how long a busy pane may run with no completed turn

A busy pane (backend busy state or the harness's rendered footer) was
unconditional, unbounded proof of liveness in every escalation path, so a
hung foreground tool call behind a busy signature could run for hours
undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic-
backtracking regex hung one bash call for 25h behind an unchanging
"Working..." footer).

FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may
run with no completed turn (state/<id>.turn-ended, or its spawn record
before any turn has completed). Past the bound, busy_turn_over_age routes
the pane through the existing wedge_timer_check, reusing the identical
stale reason, escalation counter, and demand-deep-inspection marker for
human inspection only - never an automatic interrupt, signal, or restart
of the worker or its tool process. A completed turn resets the age.

Reproduced end-to-end against the real installed Pi TUI: a foreground
`sleep 999999` bash call with no timeout renders the actual busy footer,
and two captures ~15s apart show the elapsed counter changing the pane
hash while the same turn stays unfinished. Running the pre-fix watcher
against the real captures showed it never starts a wedge timer no matter
how long the pane stays busy; the fixed watcher starts and escalates the
timer through the same mechanism, while the real hung process remained
untouched and alive throughout.

* no-mistakes(review): fix: parse enriched AFK stale reasons

* no-mistakes(review): fix: preserve enriched wedges during AFK supervision

* no-mistakes(review): fix: route all enriched AFK wedges

* no-mistakes(document): Clarify busy-turn age supervision documentation
…unchenguid#1261)

A name-by-name list of config/ entries silently stops ignoring any new or
home-local file placed there, which makes the working tree read as dirty and
blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md
already documents config/ as captain-private and gitignored as a category;
this makes .gitignore match that contract.
…al coverage (kunchenguid#1304)

The second assertion in fm-gitignore-config.test.sh (added by kunchenguid#1261) greps
.gitignore for a specific spelling of the config/ ignore pattern. It fails
on a semantically equivalent pattern like config/** and does not prove Git
actually ignores anything, per the completed source-content-test audit.

Replace it with a real git check-ignore control test on a generated
unrelated path, and strengthen the existing directory-coverage test with
generated unpredictable direct and nested config/ paths.
)

* Add bounded startup memory curation

* no-mistakes(review): Record reproducible stow verification evidence

* no-mistakes(review): Validate inherited secondmate stow evidence

* no-mistakes(document): Document editable startup-memory budget propagation
* fix(herdr): place workers in the launching agent's exact workspace

Herdr enforces no workspace-label uniqueness, and spawn resolved its
container by taking the FIRST workspace whose label matched the home
label. With two workspaces both labeled "firstmate", a worker launched
from the second one was created in the first, so it appeared in a
different space than the Firstmate the captain was watching.

Reproduced end to end on Herdr 0.7.5 protocol 17 by running the real
bin/fm-spawn.sh inside a launcher pane in the second "firstmate"
workspace: the worker landed in w1 while its launcher was in w2, with an
unrelated third workspace focused throughout, which also rules out any
dependence on the focused workspace.

Placement now binds to the launching process's own Herdr identity. Herdr
injects HERDR_PANE_ID, HERDR_SESSION, and HERDR_SOCKET_PATH into every
process it manages a pane for, and fm_backend_herdr_launcher_identity
resolves that pane's current owning tab and workspace live from Herdr,
cross-checking the pane against its tab and confirming the workspace
exists exactly once in the session. The injected HERDR_TAB_ID and
HERDR_WORKSPACE_ID are creation-time snapshots and are deliberately not
read as current identity. Labels are no longer placement authority.

A claimed parent identity that is unreadable, contradictory, stale, or
from another named session or Herdr server stops the spawn before any
worker endpoint exists, rather than degrading to a label search. A
launcher with no Herdr ancestry has no workspace to inherit and keeps
the per-home labeled container, which must now resolve to exactly one
workspace; two same-labeled candidates refuse instead of adopting
either. A --secondmate launch keeps standing up that home's own
workspace by design.

With presentation spaces enabled, the projected child is created and
bound under that same exact parent and anchors its ordering on it, so a
duplicated home label no longer makes the layout ambiguous. Projection,
focus restoration, restart binding, and quarantine rules are unchanged,
and children are never collapsed into the parent. tmux, Zellij, cmux,
Orca, and the away-mode daemon terminal were each inspected and are not
affected: none resolves a container by searching mutable labels.

tests/fm-backend-herdr-launcher-workspace-e2e.test.sh drives the real
spawn and teardown against an isolated Herdr lab, with its headline case
running fm-spawn.sh inside a real Herdr pane so the identity comes from
Herdr's own injection. The refusal matrix and the ordering anchor are
covered deterministically in tests/fm-backend-herdr.test.sh.

Eight existing real-Herdr suites inherited the developer terminal's own
Herdr pane into their isolated lab sessions, which the new cross-session
check correctly refuses. tests/herdr-test-safety.sh now owns
herdr_forget_inherited_pane and those suites call it, so what they assert
no longer depends on where they were launched from.

Two unrelated fixes found along the way. tests/fm-secondmate-harness.test.sh
had the same class of environment leak through CLAUDECODE, which outranks
PI_CODING_AGENT in bin/fm-harness.sh and made its pi-signed ancestry case
resolve "claude" whenever the suite ran inside Claude Code. And
fm-spawn.sh's usage() printed a fixed line range that had already been
truncating its own help mid-sentence.

* no-mistakes(review): Enforce exact Herdr launcher and projection identity

* no-mistakes(document): Document exact Herdr launcher workspace placement
* feat(calm): replace Pi's working row with an animated ship while Calm is on

While Calm is active and one logical agent run is under way, Calm now hides
Pi's built-in working row and renders a small two-row SSHHIP-derived boat in
its place. When Calm is off, Pi's stock working row is left untouched.

The presentation uses only public Pi extension API: setWorkingVisible(false)
plus a temporary setWidget() component whose render(width) owns the responsive
geometry and whose timer requests a TUI render. Visibility follows agent_start
through agent_settled, so the boat does not flicker between tool calls,
automatic continuations, retries, or compaction inside the same run, and
settle, abort, and failure all reach the same cleanup.

fm-calm.ts stays the sole owner of the presentation choice and the only caller
of setWorkingVisible(); the new lib owns the sprite geometry and widget.

* no-mistakes(review): Guarded Calm-off lifecycle visibility writes; focused tests pass

* no-mistakes(test): Fixed Calm E2E wait to include tmux scrollback

* no-mistakes(document): Document Calm working boat behavior

* no-mistakes: apply CI fixes

* feat(calm): slow the Calm boat, animate blue water, and make the sail directional

The boat now moves one column every 880ms while a bounded fixed-cell water phase
advances every 220ms, so the water ripples several times between boat steps and
the presentation reads as calm. One scheduler drives both clocks and disposing
the widget stops them together; ticks rather than wall-clock timestamps drive
every state change, so tests seek animation time exactly.

Colors are standard ANSI foreground codes instead of theme lookups: blue for
every water cell and yellow for the complete boat, each run closed with a
default-foreground reset so nothing bleeds into padding or later frames. ANSI
bytes never enter geometry, so visible width stays exact.

The mainsail is directional and trails aft of the mast: <| travelling right and
|> travelling left. Direction reverses the moment the boat lands on an endpoint,
so the endpoint frame already shows the new heading and no frame at or after a
bounce shows the previous sail.

* test(calm): wait for the Ctrl+O expansion redraw this block asserts

* docs(calm): record the revised working-presentation verification evidence

* no-mistakes(document): Fix Calm feasibility document EOF whitespace
…henguid#1349)

* fix(dispatch): scope candidate authentication to its own surface

A locally expired timestamp in one credential store was reported to the
captain as a sign-out, including for dispatch candidates that never read
that store. A `harness=pi, model=xai/grok-*` candidate authenticates
through Pi's own xAI credential, but the only Grok quota reading
available was gated on the standalone Grok CLI's separate token, whose
expiry clock drifts independently. The always-loaded intake rule then
turned that unreadable quota into a mandatory captain escalation.

Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts
that must not depend on agent memory: it resolves a tuple's
authentication surface from quota-axi's own emitted auth sources rather
than from a harness or model name, so another harness's CLI can never
gate a candidate that does not use it. A vendor CLI is launched only
when the tuple's own harness owns the credential store under test and a
non-destructive discovery command is registered for it, which today is
`grok models` alone. That probe runs at most once with stdin closed and
a hard timeout, reads its verdict from the first stdout line because the
command exits 0 either way, treats unrecognized output as indeterminate,
and never invokes login, logout, or the interactive TUI. Quota is read
at most twice, and unknown headroom never makes a candidate ineligible
on its own.

Update the dispatch procedure to match: usable authentication with
unmeasurable headroom stays eligible at lower preference with the
unknown disclosed, and stop-and-report is reserved for unresolved
authentication, an unresolved relationship, or malformed configuration.
Record that Grok's `credits.remaining` is a prepaid balance rather than
window headroom.

Gate quota-axi at 0.1.16 in bootstrap, the first build reporting
per-credential auth sources. A stale install previously passed the
presence check silently, which is why a fix published two days earlier
was still not in effect.

Replace the orphaned quota-array-dispatch fixtures, which encoded a
`provider: "xai"` shape the tool never emits and had no consumer, with
fixtures shaped like real 0.1.16 output that the new suite drives the
script against. The suite asserts the verdict and, separately, which
vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI
fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture
change selects the right tests instead of refusing.

* refactor(bootstrap): give the quota-axi floor one owner

The floor was stated twice - once in bootstrap's gate and once inline in
the auth preflight - so bumping it needed two edits that could drift.
Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the
existing tasks-axi library, and derive the comparison from the constant
so the number appears exactly once. Bootstrap turns a failing check into
the operator diagnostic; the preflight refuses to emit an unscoped
verdict. Map the new library to both consuming suites so a bump re-runs
them, and record that any usable source means the surface authenticates.

* no-mistakes(review): Captain: bound quota checks and removed Python dependency

* no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry

* no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing

* no-mistakes(review): Captain: reject malformed OpenCode model relationships

* no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation

* no-mistakes(document): Updated dispatch authentication documentation

* no-mistakes: apply CI fixes
…nchenguid#1350)

* feat(x-mode): reconcile promised public replies deterministically

A promised final reply in an X or Discord thread was only kept while the
primary remembered it. Compaction or restart erased that memory, so a typed
public-followup obligation could sit at pending-work after its PR merged and
the original thread never got its reply.

Make the promise durable state instead:

- bin/fm-public-followup-emit.sh reports a typed terminal work result (source
  home, work id, generation, outcome, safe deliverables, bounded public-safe
  text) into the owning home's private inbox. The event id is derived from
  that identity tuple, so duplicate reports and restart replay converge with
  no coordination, and nothing ever parses a free-form done: sentence.
- bin/fm-public-followup.sh registers a commitment, reconciles events through
  tasks-axi public-followup, and runs the idempotent delivery sequence
  (begin-delivery with the payload hash, post, record the posted receipt or a
  typed error) against the stored platform and opaque thread binding. A
  delivery interrupted between post and receipt refuses rather than risk a
  second public reply.
- Session start surfaces unresolved commitments from disk, the existing relay
  poll surfaces a new terminal-result set once, and teardown refuses while
  this home still owes a public reply for that exact work.

tasks-axi public-followup remains the only owner of the obligation state
machine, state/x-context/ the only owner of the private request context, and
fm-x-reply.sh the only thing that posts. Its new optional --receipt-file is
the one addition there, so a caller can record how many messages were sent.

A home that never opted into the myfirstmate relay gates out on a single
[ -f "$FM_HOME/.env" ] test: no tasks-axi call, no backlog or context scan,
no output, and no artifact. Evidence in docs/verification/public-followup.md.

* no-mistakes(review): Hardened public-followup reconciliation and ownership guards

* no-mistakes(review): Hardened typed terminal cleanup and receipt reconciliation

* no-mistakes(review): Automated typed-delivery cleanup and strict backlog validation

* no-mistakes(review): Fail-closed parent resolution and registration-safe delivery

* no-mistakes(review): Harden relay gating and validate secondmate bindings

* no-mistakes(review): Use owner-aware single-gate teardown protection

* no-mistakes(document): Correct public-followup documentation drift

* no-mistakes(lint): Quote done literals to fix ShellCheck warnings

* no-mistakes: apply CI fixes
…chenguid#1327)

* feat: add semantic busy-state contract owner and event writer

One owner (bin/fm-busy-lib.sh) for the captain-approved semantic
busy-state redesign: a per-task gen-bound record written only by
bin/fm-busy-event.sh, per-harness trusted-source classification with
explicit source attribution, busy/idle/unknown/dead semantics where
missing, malformed, stale, or untrusted semantic data is unknown -
never idle - and endpoint death is the only process-level override.
The Grok-only rendered-tail fallback and the standalone-Kimi
verification gate live behind the same classifier.

* feat: arm busy-state at spawn and convert Pi to the semantic extension path

fm-spawn arms the busy-state contract for converted adapters and seeds
busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed
per-task extension now reports agent_start -> busy and agent_settled ->
idle confirmed by ctx.isIdle(), covering auto-retries, compaction
retries, tool loops, and queued continuations, while turn_end stays a
wake notification touch. Teardown removes the new record, gen sidecar,
and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy ->
agent-settled idle with the marker still touched.

* feat: convert OpenCode to the semantic session.status plugin path

The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now
classifies from OpenCode's semantic session.status events - busy and
retry are active, idle is inactive - latched to the worker's own
session so a subagent child session can never clear the worker's busy
state. The session.idle marker touch stays a wake notification.
Teardown removes both the new and the legacy plugin filenames.
Live-verified on OpenCode 1.17.18 in a real TUI pane: seed ->
session-busy -> session-status-idle.

* feat: convert Claude to the full lifecycle hooks path

The per-task settings.local.json now wires UserPromptSubmit -> busy
and Stop, StopFailure, and SessionEnd -> idle, so API-error and
shutdown turn ends can never strand a busy record; Stop keeps the
turn-ended notification touch. A refused (stale-gen) event exits 0 and
stays silent so Claude's own lifecycle is never broken. Live-verified
on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch
prompt, Stop closes each turn, a mid-stream Escape interrupt fires no
closing hook, and the firstmate-controlled idle/fm-interrupt clear
resolves it.

* feat: gate Codex busy state behind verified semantic sources

The approved contract prefers Codex's app-server turn lifecycle with
capability negotiation and sanctions its lifecycle hooks as the
intermediate. Live probes on codex-cli 0.145.0 show neither is usable
for a pane worker: the app-server daemon is unreachable for a TUI
thread and refuses to start outside the managed standalone install,
and firstmate-written project hooks never fired (interactive with
directory trust granted, and exec, both with
--dangerously-bypass-hook-trust) while global hooks fired in the same
runs. Codex therefore classifies unknown codex-unverified behind an
explicit probe rather than falling back to idle or footer text, and
fm-spawn installs no unverified Codex wiring.

* feat: gate standalone Kimi busy state on live verification

Standalone Kimi has no installed binary here, so per the approved
contract its semantic path stays guarded and it classifies unknown
kimi-unverified rather than idle - and never from its locale-sensitive
moon-phase spinner, which the redesign forbids inventing as a state
source. The gate records the preferred source order (Wire prompt
request lifetime, which brackets a turn and reports cancellation, then
the documented hooks including Interrupt because Stop does not fire on
interrupts) and the exact evidence required to open it. Arming without
wiring would seed a busy record nothing could clear, so both land
together behind the same gate.

* feat: route busy consumers through the contract and drop the global OR

The watcher, crew-state reader, and away-mode daemon now decide busy
state through bin/fm-busy-lib.sh: only an exact busy verdict counts as
working, and unknown never becomes working or a silent idle, so a crew
whose semantic state is missing, malformed, stale, or unverified
surfaces instead of being absorbed. Crew-state reports the producing
source in its detail. The watcher's global OR regex default is gone;
Grok keeps its isolated fallback inside the contract. The daemon's
supervisor-pane reader stays rendered-text - that pane is not a
recorded task - but is now scoped to firstmate's own detected harness
instead of every vendor signature. Secondmate pending-reply
observation is deliberately unchanged and documented as a
delivery-confirmation signal, not task state.

* docs: point busy-state documentation at the single contract owner

Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as
the owner of what busy means, with per-adapter sources, the
unknown-never-idle rule, the endpoint-death override, and the two
rendered-text readers that deliberately stay outside the contract.
Replaces the stale regex-first prose in architecture, tmux-backend,
herdr-backend, and configuration; converts the harness-adapters
per-harness rows from UI signatures to the semantic source each
harness uses; and records the live verification evidence, including
why Codex and standalone Kimi stay unknown.

* fix: arm away-launch signal handlers before acquiring the lifecycle lock

fm_afk_launch_main acquired its lock and only then installed the EXIT,
INT, and TERM traps. A signal arriving in that window terminated the
process by default action and left the lock directory behind, which
blocks the next away-mode launch until the stale-owner reclaim path
clears it. The release helper only removes a lock this process owns,
so the handlers are now armed first. The accompanying test also killed
the child whether or not the lock had appeared and sampled cleanup the
instant wait returned; it now requires the lock, then allows a bounded
settle, so it proves the guarantee instead of racing it.

* test: align fleet, Kimi, lifecycle, and detection suites with the contract

The fleet snapshot and wake-daemon lifecycle fixtures now prove a
working crew through its own semantic busy-state record instead of
rendered pane text, which is what those consumers read. The Kimi
watcher test asserts the approved contract directly: a standalone Kimi
task classifies unknown rather than matching its moon-phase spinner,
while Grok's isolated fallback still classifies only Grok. The
pi-signed detection cases clear ambient harness markers, fixing a
pre-existing failure where the running session's own CLAUDECODE
outranked the fixture's marker.

* fix: stop teardown from deleting a project's own Codex hooks file

An intermediate revision wired Codex through a firstmate-written
<worktree>/.codex/hooks.json, and teardown removed it alongside the
other generated wiring. The Codex wiring was dropped when its probes
came back unverified, so that removal now targets a file firstmate
never creates - and a project may legitimately track its own
.codex/hooks.json, which teardown would then delete from a pooled
worktree.

* fix: keep busy-record parsing from disturbing its sourcing caller

The record parser split fields with set -- under a temporary noglob,
which clobbers a sourcing caller's positional parameters and restores
glob expansion even when the caller had disabled it. The watcher, the
daemon, and the crew-state reader all source this library, so it now
reads fields with read -a, which never globs and never touches caller
state.

* docs: state exactly which Claude hook paths were reproduced live

The busy-state record listed all four wired Claude hooks in the source
column, which could read as a claim that every one fired during the
pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are
wired from hook names confirmed present in the installed binary, but
the abnormal turn ends they cover were not reproduced.

* test: let reset_fakes own the crew-state busy-text fixture lifecycle

The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so
the variable's lifetime was owned by one test rather than by the
shared reset that every other fake already uses.

* no-mistakes(review): Fix semantic busy-state lifecycle races

* no-mistakes(review): Make busy-state retirement idempotent

* no-mistakes(review): Enforce semantic state boundaries for status and injection

* no-mistakes(review): Restore harness-scoped away-mode busy guard

* no-mistakes(document): Refresh semantic busy-state documentation

* no-mistakes: apply CI fixes
…d#1356)

* fix(calm): resume working boat from frozen column across runs

Keep one extension-owned boat animation for the Pi session so settling
freezes column and direction, the next working period resumes there
without hidden-time jumps, and only a fresh session resets to the left edge.

* no-mistakes(review): Freeze Calm boat from last rendered state

* no-mistakes(document): Document Calm boat continuity contract
* fix(dispatch): judge candidate provider relations instead of rejecting them

Firstmate deterministically dropped supported Pi candidates in the
openai-codex family. bin/fm-auth-preflight.sh resolved a harness=pi tuple's
credential surface by constructing the source id `pi:<model-prefix>`, so
`pi + openai-codex/gpt-5.6-terra` looked for a `pi:openai-codex` source. That
source does not exist, because Pi's Codex family authenticates through the
Codex store quota-axi already lists as `auth-json`/`cli-rpc`. The tuple
returned `eligible=no reason=surface-unresolved` while the Pi catalog listed
the model and the Codex provider reported fresh, usable credentials with 64
effective percent remaining on its all-model scope.

The prefix construction was only ever valid where Pi holds its own credential
(`pi:xai`, `pi:kimi-coding`), which is why every previously configured Pi tuple
resolved and the defect stayed hidden until a Codex-family Pi model was
configured.

Retire dispatch eligibility from deterministic shell. The dispatching first
mate now establishes model support and provider family from each harness's
authoritative catalog, applies quota at the granularity the vendor supplies,
and shows that reasoning. Provider-level and all-model evidence bounds every
model established in that family; a named-model window bounds only its own
model. Missing model-level quota, a missing auth source, unmeasurable headroom,
and unmodeled authentication are disclosed uncertainty. Only concrete
contradictory evidence blocks a candidate.

Replace the preflight with bin/fm-vendor-auth-probe.sh, which keeps the
captain's approved bounded probe envelope without any routing knowledge: it
takes no harness, model, or provider, reads no quota, renders no verdict, and
holds only a fixed-argv safety allowlist. Its behavior suite proves the absent
identity surface, the untouched quota, the uniform exit status, the fixed argv
with stdin closed, and a real bound even when the configured bound is zero.

Also fixed along the way: a zero FM_*_TIMEOUT silently removed the hard bound,
the pinned Grok version had drifted to 0.2.117, and --changed selection refused
outright on any deleted bin/ script.

AGENTS.md section 4 and quota-array-dispatch own the corrected policy,
harness-adapters gets the catalog-responsibility correction, and
docs/verification/dispatch-auth.md records the 2026-07-30 evidence on
Pi 0.82.0, quota-axi 0.1.16, and grok 0.2.117.

* no-mistakes(review): Reject all-zero vendor probe timeouts
* docs: add captain-authorized inherent red-check merge exception

Keep the default red-PR ban and own one always-loaded exception in the
merge-authority section: captain-explicit PR or bounded batch plus exact
check, only when the failure is inherent to the selected delivery path.
Yolo cannot activate it; final head and the full current check suite must
be verified; other substantive failures remain non-waivable.

* docs: replace narrow red-check exception with captain precedence

Supersede the inherent failing-check merge exception with one always-loaded
Firstmate-local rule: a current explicit concrete captain instruction
overrides a conflicting Firstmate-written standing rule only within exact
scope, never above platform/system/developer instructions. Keep the ordinary
red-PR default and yolo boundary; point section 7 at the section 1 owner.
* fix: give validation-time captain overrides a supersession sequence

The Validate section let a captain instruction that completely
invalidates the work being validated keep the same task and worker, but
never said how: the adjacent rule flatly bans hand-editing, committing,
aborting, or restarting during an active run with no carve-out, so a
worker facing full invalidation had no sanctioned path forward.

Add the missing sequence: cancel through no-mistakes axi's abort
command, confirm the run has stopped through axi status, recover branch
ownership through axi sync's guarded recovery, only then replace the
obsolete work, and validate once against the final head. The existing
ban on hand-editing an active run now cross-references this sequence
instead of contradicting it.

* no-mistakes(review): Make validation custody recovery conditional

* no-mistakes(document): Clarify validation supersession abort exception

* fix: keep obsolete pipeline commits out of the superseded deliverable

The review-applied fix made custody recovery conditional on
branch_sync.next_action.code, but left an open gap: recovering custody
settles who owns the branch, not what content ships. As written, a
worker could recover an obsolete run's branch and build the
replacement on top of its now-irrelevant commits instead of from the
correct pre-invalidation base, carrying obsolete content into the
final deliverable.

Make that explicit: custody recovery settles ownership, not content,
so the worker replaces obsolete work from the correct base and keeps
the obsolete run's commits out of what gets validated and shipped.

* no-mistakes(test): Restore minimal pre-invalidation replacement instruction

* fix: dedupe redundant "replace the obsolete work" restatement

Line 309 already says the worker replaces the obsolete work from the
correct pre-invalidation base, excluding the obsolete commits. The
closing sentence restated "replace the obsolete work" again before
gating the final validation run, layering the same fact twice instead
of stating it once.

Trim the closing sentence to just the ownership gate and the
single-run-against-final-head requirement it uniquely adds.
* fix: bind explicit --backend to exact-task authority

A Herdr-backed second mate carried a prior one-task --backend tmux
exception forward by analogy, so its child landed in tmux and never
appeared under the second mate in Herdr. Runtime detection was correct;
the authority surface was not.

docs/configuration.md now owns that an explicit --backend is authorized
only for that exact task. AGENTS.md and fm-spawn help point there.

* no-mistakes(document): Consolidate backend selection authorization documentation
…unchenguid#1229)

* fix: remove projected workspaces through Herdr's focus-preserving pane-death path

Herdr 0.7.5's explicit close of a workspace-emptying last pane moves the
attached client's focus to a neighbor workspace, flashing the captain's
whole window and routing in-flight keystrokes to the wrong pane until
Firstmate's exact-tab restore masks it 56-197 ms later.

Teardown and cleanup now plan a workspace-emptying close as a focus-safe
removal: verify the close empties the workspace, reposition the doomed
workspace behind the focused one through the verified workspace.move
transport when it sits before a non-last focused workspace, prove the pane
holds one lone idle shell, and end that shell so Herdr removes the emptied
workspace through its focus-preserving pane-death path. Any ambiguity or
failure falls back to the plain close behind the existing restore backstop,
and fm_backend_herdr_kill applies the same plan for non-projected removals.

Two conditions proven on real hardware are encoded in the adapter: BSD ps
reports a login shell's comm as "-zsh", and an idle shell transiently
hosts a prompt helper right after a workspace.move relayout, absorbed by a
bounded strict-sample settle window in the idle-shell proof, now the single
owner shared with session-start cleanup.

An isolated-lab regression reproduces the raw steal on 0.7.5 and proves the
plan removes a doomed workspace with zero wrong-focus samples and no
corrective focus; unit fixtures cover the position, edge, ambiguity, move
and kill failure, escalation, and transient-helper cases. Upstream fixes
(kunchenguid#1877 explicit close, kunchenguid#1912 pane death) are merged but unreleased; once
released the plan degrades to a harmless reorder-then-remove.

* no-mistakes(review): Confirm pane death from structured not-found responses

* no-mistakes(review): Serialize Herdr kills and sample focus continuously

* no-mistakes(review): Synchronize Herdr focus evidence output

* no-mistakes(review): Refuse unlocked Herdr pane closes

* no-mistakes(document): Correct Herdr focus-safety documentation

* no-mistakes: apply CI fixes

* fix: never erase a Herdr task's records while its pane survives a refused close

A transient presentation-lock contention could produce a completed teardown
while the exact Herdr pane stayed alive as an unowned restored shell: the
kill refused the unlocked close (correctly), returned success, the warning
was suppressed, and cleanup erased the task's status, turn-end, and
metadata records after the isolated copy had already been returned.

Teardown now acquires the named-session presentation lock before anything
destructive: a contended lock refuses up front while the isolated copy, the
task branch, every durable record, and the endpoint are all intact for a
plain rerun, and the projected and flat close paths both run under that one
held lock instead of acquiring their own. Durable records are erased only
once the exact pane is confirmed gone through its structured presence; a
refused, skipped, or failed close retains every record with a visible,
retryable error, and after a skipped close (unresolvable lock path) only a
structured pane_not_found counts as gone - unknown never does.

The teardown regression drives a live contending lock holder end to end:
the refusal touches nothing (no worktree return, no branch drop, no close
attempt), and the retry after release returns the copy, closes the pane
under the lock, and removes the records. The unconfirmed projected close
now refuses with records retained, and the structured-presence gate has a
strict/default unit matrix.

* no-mistakes(review): Require structured pane-not-found before Herdr record removal

* no-mistakes(document): Correct Herdr record-retention verification date

* fix: refuse ambiguity, revalidate SIGKILL ownership, and roll back failed removals

Three accepted-contract corrections from the post-CI personal review of the
Herdr keep-spaces focus-flash mitigation.

Ambiguous endpoint identity no longer counts as a confirmed-gone pane: a
missing or malformed target refuses record removal in the structured
presence gate, and teardown treats missing confirmation machinery as a
refusal instead of skipping the gate, so only an exact structured
pane_not_found ever erases durable task records.

The pane-death SIGKILL escalation re-reads the exact pane's process
information and refuses to signal unless the same shell pid still passes
the strict bare-idle ownership proof, so a pid that exited and was reused
by an unrelated process is never signaled; the refused escalation falls
back to the plain close with the unrelated process untouched.

A reposition whose removal is not confirmed no longer outlives the attempt:
the emptying-close plan records the verified pre-move order and original
index whenever it invokes the mover, and both close owners restore the
exact original workspace order through a second verified move, under the
same held session lock, before reporting the close as failed.

Each defect was reproduced first: the unit matrix documented malformed
identity as gone, the PID-reuse regression showed SIGKILL reaching a
disowned pid, and the rollback regression showed a single unrestored move.
Teardown-level regressions cover unparseable presence retention alongside
the strict identity matrix.

* no-mistakes(review): Require confirmed Herdr removal and resolvable teardown locks

* no-mistakes(review): Enforce structured Herdr closes and teardown preflight

* no-mistakes(review): Preflight explicit Herdr close confirmation helper

* no-mistakes(document): Document Herdr rollback failure semantics

* no-mistakes(review): Captain, harden recursive Herdr teardown safety

* no-mistakes(document): Document recursive Herdr teardown evidence

* fix: retain nested secondmate home when a recursive child cleanup fails

Captain-decided Option A correction for nm-askuser-flash-r6, found during
complete-diff rereview of the merged head.

cleanup_firstmate_home_children's recursive secondmate branch called
itself for a nested child's home without checking the result, then
unconditionally removed that home right after. remove_firstmate_home
ends in an unconditional recursive delete with no check for leftover
records, so a nested secondmate whose own Herdr grandchild failed its
confirmed-gone check would have its entire home - retained grandchild
records included - erased by the very next line.

Guard the recursive call the same way every other fallible call in this
function already is: || return 1, skipping remove_firstmate_home and
leaving the nested home and its records for a safe rerun.

Empirically, fm-teardown.sh's set -eu already halted the script on the
prior unguarded call before reaching removal (verified by hand with the
guard reverted, under both this session's bash and stock macOS bash
3.2) - the reachable behavior was already correct. The explicit guard
is still applied exactly as decided: it matches every sibling call site
in the function, and it stops the correctness of this path depending on
errexit's well-known fragility under refactors (a wrapping if/&&, or a
future subshell) rather than on an explicit check.

Adds a teardown-level regression building on the existing direct-child
Herdr fixtures: a top-level secondmate contains a nested secondmate,
whose own Herdr child's close goes unconfirmed. Proves through the
public fm-teardown.sh interface that the nested home, the nested
secondmate's own record, and the grandchild's metadata and status all
survive, and that the top-level secondmate's record survives too.

* no-mistakes(document): Document nested Herdr teardown retention
…d#1431)

* fix(dispatch): prioritize quota completion runway

* no-mistakes(document): Document completion-aware quota runway selection
…uid#1447)

* Preserve task contract in no-mistakes intent

* no-mistakes(review): Preserve complete current task contract in no-mistakes intent
…nguid#1452)

* fix: centralize secondmate registry parsing

* no-mistakes(review): Centralize secondmate registry binding validation

* no-mistakes(review): Harden registry EOF and symlink validation

* no-mistakes(review): Reject unreadable registries before parsing

* no-mistakes(document): Document punctuation-safe secondmate registry validation

* no-mistakes: apply CI fixes
* feat(procevent): supervise long-polling sources into durable events

Firstmate had no way to wait on a blocking external process without holding
a conversational turn. Add a domain-neutral process-to-event runner plus a
thin adapter around the currently published `lavish-axi poll` interface:
canonical physical source identity, one machine-wide owner per source, direct
argv execution, and durable 0600 result capture before any event referencing
it is published on the existing wake queue. No second notifier, no polling
control plane, and no retry machinery.

A captured result with no durable handled acknowledgement stays eligible for
bounded re-announcement across any number of drains and restarts. Draining a
wake before acting on it and then starting a replacement session resurfaces
the same exact source and sequence, and never puts result payload text in an
event line. `fm-procevent.sh handled <source-id> <sequence>` is the only thing
that stops re-announcement: generation-keyed, private, path-safe, durable, and
atomically idempotent, so a paired external effect gated on its first-time
versus repeat report is never authorized twice.

An acknowledgement is refused unless matching captured result and adapter
records already exist, so a premature or mistyped call cannot suppress a
future result.

The source side is unchanged and still lossy: the published poll clears
feedback destructively before returning it, so a result lost in that window
is unrecoverable. This is never at-least-once, no-loss, or lossless, and the
handled acknowledgement is not a generic exactly-once effect either - a crash
between an external effect and its acknowledgement can still repeat that
effect on replay.

Integrate registered sources with watcher supervision, the guards, and
recoverable secondmate teardown across nested homes, and cover source
identity, lifecycle races, supervision, restart handling, and cleanup safety
with regressions.

* no-mistakes(review): Prevent Lavish prompt text from spoofing missing sessions

* no-mistakes(review): Serialize publication and secure handled acknowledgements

* no-mistakes(document): Document hardened process-event acknowledgement guarantees

* fix(procevent): never reclaim a source whose owned group still runs

A runner is its own process group leader and starts the blocking source in
that group, but the claim records only the leader PID and its identity. If the
leader died while the source child kept running, the missing PID was
classified stale: reconciliation released the claim and started a second
runner while the old blocking source was still consuming the same canonical
source. For the Lavish adapter that means two destructive long polls racing on
one review session, so it is not harmless process litter. It also contradicted
the documented promise that ownership is never released until the whole group
is gone.

Ownership state now distinguishes a generation that is really gone from one
whose leader crashed with its group still alive. Reconcile stops that
surviving group and releases its exact generation before starting any
replacement, and keeps the claim for a later cycle when it cannot prove the
group stopped or another home owns it. Acquisition and `start` treat the same
state as held rather than reclaimable.

Signalling that group is safe precisely because only an absent leader reaches
this state. A reused PID leaves the leader alive, so the identity comparison
still classifies it stale or uncertain and no group signal follows, which
keeps the existing PID-reuse refusal intact.

Add a public-interface regression for the exact crash cut - SIGKILL only the
leader, prove the child group survives, reconcile, and prove the old group is
gone with no second source running - plus its counterexample that a generation
with no leader and no surviving group is still reclaimed. Update the runner
help, operating documentation, skill, and verification record where they
described reclaim in terms of the leader alone.

* no-mistakes(review): Enforce runner group ownership and detect poller overlap

* no-mistakes(review): Isolate runner groups from unrelated caller processes

* no-mistakes(document): Document isolated process-event runner launch

* no-mistakes(lint): Suppress Perl literal ShellCheck false positive
…nchenguid#1500)

* fix(bin): deliver process-event results and retire ended sources

Two defects reproduced during a real Lavish adapter session.

One human `Send & End` produced four captured results: the real feedback,
then recurring empty ended sessions. The generic runner had no way to learn
a source was finished, so every reconcile restarted a poll that returned
immediately. The runner now asks the source's own adapter -
`fm-procevent-<adapter>.sh terminal <result-file>` - and on exit 0 alone
re-proves ownership, drops the registration, and releases its own claim
under one source boundary. Terminal knowledge stays adapter-owned: for
Lavish that is an ended session, a missing session, and the final feedback
delivery the published poll marks with `session_ended`. An adapter with no
terminal command keeps its source armed exactly as before. Capture before
publication, captured-result durability, queued wake durability, bounded
re-announcement, handled deduplication, one-owner ownership, and explicit
idempotent retirement are all unchanged.

A captured result queued its `check` wake durably, but a healthy watcher
with a fresh beacon never delivered it; the result surfaced only after a
manual drain. Publication happens outside the watcher (in the runner) or
unconditionally (in reconcile), so the watcher had no newly actionable
signal to report and never reached its rewake path. It now reports a
queued-but-unsurfaced process-event record through the same actionable exit
every other wake uses, deduplicated by the same `.seen-*` marker discipline
the signal scan uses, so the record is always durable before it is
suppressed. The durable queue remains the authority and no second notifier,
poller, timer, queue, or adapter-specific wake path is added.

Regressions cover both, driven end to end: an armed Lavish source against a
stand-in for the published poll polls once, captures once, publishes one
distinct event, and retires itself; two fixture adapters prove the terminal
decision follows the adapter alone; and a real capture plus a real watcher
prove one proactive wake before any drain, with no duplicate wake while the
record stays queued or after it is acknowledged.

* no-mistakes(review): Harden process-event retirement and proactive delivery

* no-mistakes(review): Route process-event delivery through shared wake owner

* no-mistakes(document): Clarify process-event delivery and retirement documentation

* no-mistakes(lint): Fix ShellCheck control-flow warnings

* no-mistakes(lint): Fix wake output status lint warning
Sync of the 27 upstream commits since 99533c5, resolved toward upstream
per the fork's match-upstream-exactly policy.

Conflicts resolved:
- bin/fm-push-transition-lib.sh: combined. The fork's delivery-receipt
  publication and upstream's FM_WAKE_POST_OUTPUT_ACTION output wrapper are
  orthogonal; both kept, receipt published before the wake line is written.
- bin/fm-session-lock-lib.sh: kept the fork's fork-free fm_process_row
  ancestry walk. Upstream's change here was `basename --` dash-hardening,
  which the fork's ${FM_ROW_COMM##*/} expansion already satisfies by not
  invoking basename at all. Upstream's second basename site merged cleanly.
- bin/fm-watch.sh, docs/configuration.md, tests/fm-crew-state.test.sh,
  tests/fm-daemon.test.sh: resolved to upstream.

Fork work dropped as subsumed by upstream:
- PR #14 (defer stale wedge alarms for active crews) is dropped in full.
  Upstream 96542a4 replaces pane-text busy heuristics with a semantic turn
  lifecycle, so a worker inside a long in-contract call now classifies busy
  and never reaches the 240s stale threshold that produced PR #14's false
  alarms. Upstream 56a7ac6 bounds that busy state at FM_BUSY_TURN_MAX_SECS
  (3600s) with no completed turn. PR #14's deferral would have re-widened
  exactly that bound.
- The bin/fm-watch.sh half of the fork's residual-busy-frame fix (16c9e67):
  upstream's window_is_busy no longer matches rendered frames outside the
  Grok-scoped fallback, so there is no frame branch to guard. The
  dead-agent case is bounded upstream by busy_turn_over_age instead.

Fork work retained (not subsumed):
- The bin/fm-crew-state.sh and bin/fm-backend.sh half of 16c9e67. Upstream's
  fm_busy_classify_meta performs no endpoint-liveness read, and pane_readable
  cannot see an agent-less but structurally alive pane, so the agent-presence
  gate is still the only thing settling that case on this path.
- All other earlier fork work: the wake-drain coalescing, Herdr Pi
  queued-input steering confirmation, the U+00A0 composer classification,
  the Claude spawn/ancestry work, and the codev-session skill.

Fork-only tests dropped with their implementations:
- tests/fm-watch-triage.test.sh test_window_is_busy_requires_a_present_agent
- tests/fm-daemon.test.sh in-contract wedge deferral tests
The fork's shared CONTEXT_RULE (a45035b) built its text with
$(cat <<'EOF' ... EOF), which tests/fm-brief.test.sh forbids because
Bash 3.2 mis-parses a heredoc nested in a command substitution. That
test predates this fork's merge base and upstream passes it, so the
fork's default branch was already red here before the upstream sync.

Use the same 'IFS= read -r -d ...' form the adjacent HERDR_SECTION
already uses, stripping the single trailing newline so the rendered
brief stays byte-identical.
…d ps

Upstream 6ec5e08 added a dash-leading ancestry assertion that stubs ps.
This fork's fm_process_row reads /proc before ps, so on Linux it walked
the real process tree and never observed the stub, and the assertion
failed on the merged tree while passing on upstream.

The fast path is kept: upstream has not subsumed it, still forking ps,
basename, and grep per hop at a measured 2.5-3.3s, which pushed Claude
auto-arm past FM_CLAUDE_AUTOARM_SYNC_WAIT_MS.

fm_process_row now falls through to a per-pid ps read when the ps -A
snapshot does not list the pid, instead of failing. Purely additive: the
/proc path and the one-snapshot-per-shell behavior are unchanged.

The upstream assertion gains one line, FM_PROC_ROOT_OVERRIDE pointing at
a nonexistent root, so it exercises the portable fallback the fixture
actually stubs. It carries a comment naming the fork behavior it
accommodates, because it will conflict on every future sync of this test
until upstream adopts an equivalent fast path.

tests/fm-claude-stop-autoarm.test.sh's fork-frugal ancestry case still
passes; it fails if ps is called at all on the /proc path or if
resolution exceeds 1000ms, so the fast path is proven intact.
@knowttl
knowttl merged commit 10a4241 into main Aug 2, 2026
@knowttl
knowttl deleted the fm/fm-upstream-sync-5 branch August 26, 2026 20:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants