Skip to content

fix(bin): sync upstream batch 15 — teardown, spawn, supervision, and bearings fixes - #49

Merged
zeeshaanahmad merged 14 commits into
mainfrom
fm/upstream-batch-15-to-tip
Sep 7, 2026
Merged

zeeshaanahmad merged 14 commits into
mainfrom
fm/upstream-batch-15-to-tip

Conversation

@zeeshaanahmad

@zeeshaanahmad zeeshaanahmad commented Sep 7, 2026 •

Copy link
Copy Markdown
Owner

Intent

Captain's standing policy for this fork (2026-09-05, verbatim): "the main goal for firstmate is to become sync with upstream. once we reach that state, we avoid introducing more changes and stay aligned with the upstream. until then you can keep adding work that helps aligning with the upstream easier and reaches the intended state quickly."

Captain's conflict rule (2026-09-05, verbatim): "if there is something upstream and our code both worked on and fixed but approaches differ, take the upstream change, however, if our change is better than upstream, file a PR. for other items, take the upstream changes."

Captain's direction on 2026-09-06 (paraphrased from "We need to close this sooner so we can focus on the real work"): sync batches are now the only firstmate work; keep them tight, no new fork-side improvements, no new upstream PR candidates unless a measured defect forces one.

The ask this task serves: upstream sync batch 15 - bring the fork's origin/main (now 8e41bd1: batch 14 landed as #48 with merge commit 3b20f6c, waypoint f91a950) up to upstream's current tip, waypoint 5592cb6 ("feat(tests): run live harness guards by default when available (kunchenguid#3889)"). Nine upstream commits in range f91a950..5592cb6, 116 files, +7863/-517:

What Changed

  • Merges nine upstream commits (f91a950..5592cb6) plus two fork-specific fixups: verifies Treehouse slot ownership before teardown, keeps supervision armed for registered custom checks, repairs bearings board listening and decision reconciliation in bin/fm-captain-hold.sh/bin/fm-procevent.sh, allows pooled spawns without a git origin, resolves Treehouse locks for remote secondmate homes, fixes native-Windows Bash helper invocation in the Pi adapter, and restores/rewires the live harness guard source gating in the fork's own live-guard tests.
  • Adds a new verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary roles: .omp/extensions/fm-primary-omp-watch.ts, .omp/extensions/fm-primary-turnend-guard.ts, .omp/fm-worker-overlay.yml, plus matching harness-adapters skill reference, docs/supervision-protocols/omp.md, and CI wiring.
  • Makes live harness guards run by default when available (bin/fm-test-run.sh, tests/lib.sh) and substantially expands test coverage and docs to match — new/updated suites for teardown endpoint safety, captain-hold lifecycle, omp harness/primary live e2e, live-gate, Pi Windows shell invocation, spawn pool base freshening, and updated architecture/configuration/captain-hold-lifecycle/verification docs.

Risk Assessment

✅ Low: This is a well-executed, true (non-rebase) merge of nine trusted upstream commits plus two small, precisely-scoped fork-side follow-up fixes; independent verification (silent-splice line-set diffs, byte-level trio diffs, family-enumeration cross-checks, and ancestry checks) confirms every claim in the unusually detailed merge commit message and both fix commits, with no dropped upstream content, no missed live-guard files, and no scope creep beyond what the captain's sync-batch policy allows.

Testing

Both targeted regression suites (fm-test-fixture-cleanup.test.sh and fm-live-gate.test.sh) pass cleanly, and manual CLI runs of the four rewired guards confirm they still emit their original per-guard opt-in skip messages under the new shared gate — the batch's two fork-specific fixes (shared-gate wiring and source-guard restoration) work end-to-end exactly as described in the commit messages, with no regressions found.

Evidence: fm-test-fixture-cleanup.test.sh full pass (source-guard regression coverage)
ok - fm_test_tmproot cleans up its fixture root on normal exit
ok - fm_test_tmproot cleans up its fixture root on SIGTERM
ok - the cleanup registry cannot be injected through path precreation
ok - failed fixture registration rolls back the new root
ok - the orphan sweep reaps only old fixtures without a live owner
ok - fm_test_rmtree refuses an empty fixture path
ok - fm_test_rmtree refuses a non-empty path outside the fixture temp root
ok - fm_test_rmtree refuses the working directory an unset TMP_ROOT resolves to
ok - every tests/lib.sh-dependent file refuses to run outside tests/ (176 files)
ok - fm_test_own_fixture permits teardown of a declared root outside the temp root
ok - fm_test_own_fixture refuses the working directory, its ancestors, and the repo root
ok - a fixture rooted outside $TMPDIR is removable only once declared
ok - the orphan sweep reaps read-only package fixtures
Evidence: fm-live-gate.test.sh full pass (shared-gate wiring coverage, 29 guards)
ok - a default-on guard runs wherever its tools are installed
ok - an absent tool is a named capability skip, not a silent pass
ok - a prompt-submitting guard stays off until it is asked for
ok - a guard's own variable turns it on
ok - a demanded run refuses to pass as a skip
ok - FM_LIVE switches the whole family
ok - a guard's own setting wins over FM_LIVE
ok - any entry point of a multi-mode guard turns it on
ok - the shared gate carries the gate-refusal bypass into every live guard
ok - all 29 live guards refuse together on FM_LIVE=0
Evidence: Manual CLI run of the four rewired fork guards with no opt-in env set
=== tests/fm-away-delivery-bound-live-e2e.test.sh ===
skip: live: opt-in; set FM_AWAY_BOUND_LIVE=1 to run
exit=0
=== tests/fm-claude-attribution-live-e2e.test.sh ===
skip: live: opt-in; set FM_CLAUDE_LIVE_E2E=1 to run
exit=0
=== tests/fm-nm-status-shape-live-e2e.test.sh ===
skip: live: opt-in; set FM_NM_STATUS_SHAPE_DRIFT=1 to run
exit=0
=== tests/fm-send-agent-pane-live-e2e.test.sh ===
skip: live: opt-in; set FM_SEND_AGENT_PANE_LIVE=1 to run
exit=0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Step was skipped.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-test-fixture-cleanup.test.sh — all 13 checks pass, including test_lib_dependent_files_refuse_to_run_outside_tests which copies every lib.sh-dependent test file (176 found) out of tests/ into a sandbox and asserts each aborts at its source line rather than running on with undefined helpers
  • bash tests/fm-live-gate.test.sh — all 10 checks pass, including test_every_live_guard_is_wired_to_the_shared_gate which drives bin/fm-test-run.sh --family live-harness-optin --list and asserts every one of the 29 listed guards (including the fork's 4) prints the shared skip: live: disabled by FM_LIVE=0 refusal
  • bash bin/fm-test-run.sh --family live-harness-optin --list — confirms the fork's four guards (fm-away-delivery-bound-live-e2e, fm-claude-attribution-live-e2e, fm-nm-status-shape-live-e2e, fm-send-agent-pane-live-e2e) are enumerated in the family alongside upstream's guards
  • Manually ran each of the four rewired fork guards directly with a clean environment (no opt-in vars set) to confirm the end-user-visible skip message still names each guard's original opt-in variable
  • git status --short before and after testing to confirm no transient artifacts were left in the worktree
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Firstmate merge reasoning

  • Upstream sync batch 15: fork main 8e41bd1 to upstream waypoint 5592cb6 (nine upstream commits) as one merge commit 77f3336; the waypoint is recorded on the task record and asserted by the merge script. Every contested file resolved upstream-first; the only fork hunk restored in bin/fm-captain-hold.sh is proven load-bearing by four fork tests with no upstream producer of the string they assert. Silent-splice sweep: zero upstream lines missing.
  • Two fork follow-up commits (cb3043b, 870000d) make the fork's four extra live guards obey upstream's new shared live gate and restore the fork's || exit 1 source guard on the guards feat(tests): run live harness guards by default when available kunchenguid/firstmate#3889 rewrote. Two pipeline ci-fix commits: e5d495e fixes a real pre-existing fork bug (an intact operational envelope with prose in front was classified as machine input), and 239a017 corrects a test fixture that omitted the envelope terminator every genuine envelope carries. All read in full.
  • Validation: run 01M1Y9GFH03TEY1R72APM3GVWT, rebase skipped, CI enabled; 14 of 14 named checks green on head 239a017. The one red check-run is a stale attestation evaluation from before the second fix commit; the same check passes on the same head in run 34147774702, and a workflow re-run cannot go green because it replays the original event payload.
  • Deviation recorded: pipeline commit 239a017 carries an agent co-author trailer added by the ci fix agent, against this repo's rule. Accepted rather than rewritten: rewriting a validated gate-fix commit would cost a fresh validation, and this history is fork-only and never goes upstream.
  • Merge authority: standing yolo for firstmate sync batches; merged with --merge to preserve upstream commit identities.

kunchenguid and others added 14 commits September 6, 2026 14:57
* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout
…ndmate, and primary (kunchenguid#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>
…unchenguid#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>
…nguid#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes kunchenguid#3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner
…nguid#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership
…nchenguid#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks
* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior
…enguid#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass
Brings the fork from waypoint f91a950 (batch 14) to upstream tip
5592cb6 "feat(tests): run live harness guards by default when available
(kunchenguid#3889)". Nine upstream commits, one merge commit, no rebase or squash, so
every upstream commit identity stays in ancestry.

Nine files conflicted; every resolution is recorded below.

Group 1 - captain-hold trio (upstream kunchenguid#3872 cf7e2fa)

bin/fm-captain-hold.sh, docs/captain-hold-lifecycle.md and
.agents/skills/captain-hold-lifecycle/SKILL.md were first reset to upstream's
version wholesale, leaving all three byte-identical to upstream.

One fork-only hunk was then restored on top of upstream's body: the freed-work
precondition reminder. It is justified, not discretionary. Four fork-only cases
in tests/fm-captain-hold-lifecycle.test.sh assert the exact string
"recheck freed task preconditions per .agents/skills/captain-hold-lifecycle/SKILL.md:",
and upstream's bin/ cannot produce that string anywhere, so those cases cannot
pass without the hunk. The restored code is additive: 74 added lines plus five
printf call sites rewritten to route through print_answer_outcome, which
preserves upstream's own output words on all five exits. Nothing else of the
fork's prior divergence on these files survived.

Its documentation was restored with it, because the feature it documents
survived: four lines on the reminder mechanism and one on the corrupted-binding
diagnostic in docs/captain-hold-lifecycle.md (purely additive, 15/0), the three
fork verification dates and the paragraph naming the four fork-only regressions,
and in SKILL.md the freed-work policy sentence plus recheck steps 7-10.
Upstream's own step 7 is preserved verbatim as step 11.

bin/fm-procevent.sh needed no conflict resolution and its fork-only
corrupted-binding forwarding auto-merged intact; upstream's read_binding still
fails loudly on a corrupted record, so the two remain consistent.

Group 2 - bin/fm-test-run.sh (upstream kunchenguid#3889 5592cb6)

Both conflicts were alphabetical-insertion collisions where each side added a
neighbouring entry. Both sides were kept in sorted order:
fm-bearings-board-lavish-live-e2e (upstream) before fm-claude-attribution-live-e2e
(fork) in the live-harness-optin family, and tests/fm-live-gate.test.sh
(upstream) before tests/fm-liveness-source.test.sh (fork) in the duration hints.

Every fork-only test file resolves to a family under upstream's scheme, with no
fork-side family list retained: bin/fm-test-run.sh --check-coverage reports ok
total=201, and --list --all selects all 201 tests/*.test.sh present on disk with
no file unlisted and none listed but absent.

Group 3 - tests/fm-watch-arm.test.sh

The fork's fixture-cleanup guard (ce640e3) was kept over upstream's simpler
shape. Both compute the same status and fail with an identical message; the fork
side additionally ends the fixture through a status transition so a failing path
leaves no child process behind, which upstream's shape does not provide. PR 47's
work survives: both of its fixtures are present, bin/fm-wake-lib.sh still carries
its single "the guard is progress, not depth" marker, and bin/fm-watch-arm.sh
still records stderr= before successor=.

Group 4 - four live-e2e headers (upstream kunchenguid#3889)

Upstream's new default-on guard from tests/lib.sh was taken in all four.

The fork's "|| exit 1" source guard was kept in
tests/fm-afk-pi-herdr-return-e2e.test.sh and tests/fm-pi-branch-live-e2e.test.sh,
because upstream's new header does not provide an equivalent: fm_live_gate is
itself a lib.sh function, so a failed source leaves it command-not-found and,
with no set -e, the script would run on. This is the same shape git auto-merged
without conflict in tests/fm-grok-stop-live-e2e.test.sh.

tests/fm-claude-stop-autoarm-live-e2e.test.sh: upstream's header was already
auto-merged; the conflict was the fork's command -v checks against upstream
deleting them. The command -v claude check was dropped because upstream's gate
provides an equivalent guard for exactly that tool. The fork's surviving
concurrency case still needs tmux and jq, which are now declared through
upstream's own gate tool list rather than as duplicate command -v scaffolding,
so an absent tool skips rather than fails. The fork's CONCURRENCY_MODES_CHECKED,
SOCKET and CONC_ROOT behaviour case is kept in full.

tests/fm-harness-liveness-drift-live-e2e.test.sh: upstream added a local
resolve_harness_binary; the fork had extracted the identical resolver into
tests/lib.sh as fm_test_resolve_harness_binary, which is what the merged call
site uses and what two fork-only live guards also depend on. Upstream's local
copy was dropped as dead code; the behaviour is unchanged, only its location
differs.

Silent-splice sweep

For bin/fm-spawn.sh, bin/fm-teardown.sh, bin/fm-supervision-instructions.sh,
bin/fm-turnend-guard.sh, bin/backends/tmux.sh, bin/fm-composer-lib.sh,
bin/fm-procevent.sh, bin/fm-wake-lib.sh and AGENTS.md, every line upstream added
since the merge base is present in the merged tree: 211, 245, 11, 4, 4, 48, 30,
103 and 2 added lines respectively, none missing.

AGENTS.md carries both upstream additions (the reconcile-requests/ record and
the omp harness row) and retains all twelve fork-only lines documenting
fork-only features, matching batch 14's restored baseline exactly at 12 added
and 1 replaced against upstream.

The omp adapter (kunchenguid#3867) was taken whole, along with every other new upstream
file in the range.
…e gate

Upstream kunchenguid#3889 adds tests/fm-live-gate.test.sh, which enumerates the whole
live-harness-optin family and requires every member to honour FM_LIVE=0 through
the shared fm_live_gate helper in tests/lib.sh. The fork carries four live
guards upstream does not have, each written before that helper existed and each
gating itself with its own env check, so the new conformance test failed on
them and the batch was red.

Each guard now opens with the shared gate and keeps its own opt-in variable and
live tools, so nothing silently starts running and nothing stops being checked:

- fm-away-delivery-bound-live-e2e: FM_AWAY_BOUND_LIVE, tmux
- fm-claude-attribution-live-e2e:  FM_CLAUDE_LIVE_E2E, claude
- fm-nm-status-shape-live-e2e:     FM_NM_STATUS_SHAPE_DRIFT, no-mistakes
- fm-send-agent-pane-live-e2e:     FM_SEND_AGENT_PANE_LIVE, tmux

The two that did not already source tests/lib.sh now do. They keep their own
EXIT traps, which is the same shape upstream's own live guards use.

The three tool checks the gate now owns were dropped rather than left as dead
code beneath it. An absent tool is now a named capability skip instead of a
hard failure, which is the gate's documented behaviour.

tests/fm-live-gate.test.sh passes: "all 29 live guards refuse together on
FM_LIVE=0". Each guard still prints its own opt-in skip naming its original
variable when run with no environment set.
 rewrote

tests/fm-test-fixture-cleanup.test.sh pins a fork guarantee with real evidence
behind it: a test file copied out of tests/ cannot resolve `. lib.sh`, bash
treats that failed source as an ordinary non-zero return, and a file that keeps
running with fm_test_tmproot undefined once `rm -rf`d a live task worktree. Every
lib-dependent file must therefore abort at its source line.

Upstream kunchenguid#3889 rewrote the live-e2e headers to source tests/lib.sh and then call
fm_live_gate. Both are lib.sh functions, so on a failed source the gate is merely
command-not-found and the file runs on instead of refusing. Copied out of tests/,
tests/fm-bearings-board-lavish-live-e2e.test.sh reached its real body and invoked
lavish-axi for real; the autoarm and cmux guards did the same.

This restores the fork's `|| exit 1` on the 19 live-family files that gained or
kept an unguarded lib.sh source in this batch. It is the same one-line guard
commit 806826c added for exactly this reason, and it is load-bearing rather than
cosmetic: upstream's new header provides no equivalent, because the thing that
would catch the failure is itself defined by the file that failed to load.

tests/fm-test-fixture-cleanup.test.sh now passes with no failures, and
tests/fm-live-gate.test.sh still reports "all 29 live guards refuse together on
FM_LIVE=0".
…operational-input.sh has a fallback for front-truncated envelopes (a delivery that lost its leading header but kept its trailing FIRSTMATE_OP_END terminator). That fallback (`fm_operational_terminator_kind`) only checks for the terminator via a substring/"contains" case pattern, without verifying the header is actually absent. So a message like "Captain quote: <full watcher envelope with header+terminator>" — a complete, intact envelope with human prose glued to the front, not a truncated fragment — was wrongly classified as genuine operational input of kind "watcher" and hidden by the Pi Calm extension's operational-user-layout patch. This is exactly the failure in tests/fm-calm-pi-extension.test.sh: "Calm hid an operational near miss: Captain quote:". Neither bin/fm-operational-input.sh nor tests/fm-calm-pi-extension.test.sh were touched by this sync-batch-15 diff (8e41bd1..870000d) — this is a pre-existing bug in fork-owned code, newly exposed as a real, deterministic (non-flaky) failure, not an infra/attestation artifact, so it needed a real fix. Fix: in fm_operational_input_kind, before falling back to the terminator-only match, bail out (return 1, i.e. not operational) if the message contains the full FM_OPERATIONAL_HEADER_PREFIX anywhere — a genuinely truncated fragment can never contain that literal prefix text (it was cut off), so its presence means the envelope is intact and merely has extra text in front, which must stay classified as ordinary (visible) captain speech. Verified: - Reproduced the exact bug directly against bin/fm-operational-input.sh before the fix (classified the near-miss as "watcher"), and confirmed it now returns nothing (not classified) after the fix, while an intact envelope with no leading prose still classifies correctly as "watcher". - Ran tests/fm-operational-input.test.sh (the dedicated unit suite for this script, including the front-truncated/tail-truncated/near-miss contract tests) — all 9 tests pass. - Ran tests/fm-turnend-guard.test.sh (another heavy consumer of this classifier) — all tests pass, no regressions. - Could not run tests/fm-calm-pi-extension.test.sh itself end-to-end locally because it requires a globally-installed @earendil-works/pi-coding-agent package not present in this environment/network-restricted worktree; verified instead at the exact boundary (bin/fm-operational-input.sh via the "kind" CLI command) that the JS layer (classifyFirstmateCurrentOperationalText) calls into, which is the sole place all Calm operational-input classification data flows through (confirmed no parallel JS/TS reimplementation exists). - shellcheck on the modified file reports no warnings. Change is a single, minimal, root-cause fix (7 lines added, 1 removed) confined to bin/fm-operational-input.sh; no new subsystems, no unrelated refactoring
…-extension.test.sh

The generic MONITOR_${label}_${suffix} fixture in the Pi follow-up E2E
expected a watcher envelope without its FIRSTMATE_OP_END terminator, while
every genuine watcher envelope (and the fixture's own exact_watcher case)
carries one. This mismatch was latent until e5d495e made
fm_operational_input_kind() require the terminator to be genuinely present
for the terminator-only fallback to fire, at which point the encode/decode
round-trip started persisting the terminator in the stored message text,
and the outdated fixture without it failed a real, deterministic assertion:
"not ok - Pi follow-up loaded_on persisted the wrong turn or input
semantics" (reproduced by reverting this fix).

Verified: tests/fm-calm-pi-extension.test.sh (12/12, exit 0),
tests/fm-operational-input.test.sh (9/9), tests/fm-turnend-guard.test.sh
(87/87), bin/fm-lint.sh clean on the modified file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@zeeshaanahmad
zeeshaanahmad merged commit bb429c7 into main Sep 7, 2026
15 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants