Skip to content

merge: bring the fork up to date with upstream (150 commits), preserving both fork commits - #4

Merged
elixlabssolutions merged 155 commits into
mainfrom
fm/fm-upstream-merge
Aug 6, 2026
Merged

elixlabssolutions merged 155 commits into
mainfrom
fm/fm-upstream-merge

Conversation

@elixlabssolutions

@elixlabssolutions elixlabssolutions commented Aug 6, 2026 •

Copy link
Copy Markdown

Intent

Bring this firstmate fork up to date with its upstream project (kunchenguid/firstmate) by MERGING upstream/main into our history, preserving both our local commits and upstream's 150. The captain chose merge over rebase explicitly, to avoid rewriting history and to guarantee our work is not lost. A merge commit with two parents is therefore the intended shape; the large diff is upstream's work arriving, not new authoring.

Our two commits that had to survive semantically, not just textually:

  1. The scout-implementation-contract skill plus its AGENTS.md section 13 trigger.
  2. The vault guard, which blocks crewmate secret-value exposure. This is a security control written after a real 2026-07-30 incident where a crewmate printed secret VALUES into its transcript. A merge resolution that weakened, bypassed, or silently dropped it was explicitly forbidden.

Deliberate conflict-resolution policy the captain set: prefer upstream's version of THEIR work and ours of OURS, and where the two genuinely interleave keep BOTH behaviors rather than picking a side. Nothing was resolved by deleting anyone's behavior.

Decisions a reviewer reading only the diff would not know:

  • bin/fm-spawn.sh codex launch template keeps BOTH upstream's new OPINPUT brief encoding AND our --dangerously-bypass-hook-trust flag. That flag is load-bearing, not redundant: without it the worktree .codex/hooks.json never loads and the vault guard is present but inert.
  • bin/fm-spawn.sh claude settings.local.json keeps upstream's full semantic busy-state hook set (UserPromptSubmit/Stop/StopFailure/SessionEnd) AND re-adds our PreToolUse vault guard, which upstream's rewrite had dropped. Upstream's Stop hook already carries the turn-end touch, so it supersedes the simpler one it replaced. A deliberate comment marks PreToolUse as a security control so a future merge does not silently drop it again.
  • bin/fm-spawn.sh pi extension keeps upstream's busyEvent helper and restores spawn alongside execFile in the import, because our tool_call vault handler needs it.
  • bin/fm-spawn.sh opencode branch takes upstream's renamed fm-busy-state.js in exclude_path while leaving our vault plugin install unchanged.
  • bin/fm-spawn.sh codex branch keeps upstream's busy-state negotiation comment and our vault hooks.json install side by side; they are independent concerns.
  • tests/fm-busy-adapter-wiring.test.sh: upstream's NEW test asserted the codex worktree .codex/hooks.json must be ABSENT, using file absence as a proxy for "no busy wiring installed". Our vault guard legitimately writes that same file for an unrelated security purpose, so their proxy forbade a control it was never aimed at. Rather than delete our security control, the assertion was narrowed to its actual intent: if the file exists it must contain no fm-busy-event.sh wiring. This was verified with a negative control to still trip on real busy wiring, the sibling assertion that codex arms no busy contract is untouched and still passes, and fm-busy-lib.sh never reads that file when classifying, so a vault-only hooks.json cannot make codex look verified. This is the single most reviewer-relevant change in the diff and is intentional.
  • docs/documentation-audiences.json: upstream added a rule requiring every tracked prose surface to be classified exactly once. Our two files predate it. They are classified by sibling precedent - the scout contract as agent-runtime like every other SKILL.md, the vault guard as maintainer-architecture like the sibling PreToolUse guard docs arm-pretool-check.md, cd-guard.md and subagent-guard.md.
  • tests/fm-vault-guard.test.sh: two test-harness adaptations to upstream's new contracts, no behavior change. Ship spawns now require explicit --mode and --yolo, so the helper passes them or the per-harness install assertions never execute; and the assertion guarding upstream's turn-end plugin now names their renamed fm-busy-state.js.
  • One small follow-up commit removes a duplicated codex turn-end sentence that the merge left stated twice, per this repo's one-owner rule.

Deliberately NOT included: docs/rfcs/ is the captain's unlanded work, untracked and left exactly as-is by explicit instruction.

Verification already done locally: lint clean on pinned ShellCheck 0.11.0; the vault guard's 26 checks and the scout contract's 5 both pass and match their pre-merge baselines; the guard's live allow/deny behavior is unchanged including adversarial glued-flag, case-variant and run-child dump cases; and the merged tree differs from upstream/main only in the files our own commits touch.

What Changed

  • Add remote secondmate homes, job-worker execution, trace propagation, readiness checks, and durable parent/cleanup bindings.
  • Expand lifecycle supervision with semantic busy state, process-event handling, decision holds, wake/follow-up recovery, and improved multi-harness reliability.
  • Preserve the fork’s scout-contract and vault secret-exposure guard through the upstream merge, while updating documentation, CI sharding, and behavioral coverage for the expanded runtime.

Risk Assessment

⚠️ Medium: Captain, this is a broad 150-commit upstream merge, but the two-parent shape, preserved fork behavior, conflict-resolution integrations, and stated acceptance criteria were inspected without a substantiated changed-code defect.

Testing

Completed 1 recorded test check.

  • Outcome: ⚠️ 1 error across 1 run (37m20s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

⚠️ **Test** - 1 error
  • 🚨 tests failed with exit code 1
  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Test results a human reviewer should read before merging

The 2 real failures this merge caused, both found and fixed

These are the only test failures attributable to this change. Both are our two fork commits meeting contracts upstream added during the 150 commits.

  1. tests/fm-busy-adapter-wiring.test.sh - upstream's new test asserted the codex worktree .codex/hooks.json must be absent, using file absence as a proxy for "no busy wiring installed". The vault guard legitimately writes that same file for an unrelated security purpose, so upstream's proxy forbade a control it was never aimed at. Fixed by narrowing the assertion to its actual intent: if the file exists it must contain no fm-busy-event.sh wiring. Verified with a negative control that it still trips on real busy wiring; the sibling assertion that codex arms no busy contract is untouched and still passes; and fm-busy-lib.sh never reads that file when classifying, so a vault-only hooks.json cannot make codex look verified. This is the single most reviewer-relevant change in the diff.
  2. tests/fm-documentation-audiences.test.sh - upstream added a rule requiring every tracked prose surface to be classified exactly once, and this fork's two files predate it. Fixed by classifying them on sibling precedent: the scout contract as agent-runtime like every other SKILL.md, the vault guard as maintainer-architecture like the sibling PreToolUse guard docs (arm-pretool-check.md, cd-guard.md, subagent-guard.md).

Two further changes to tests/fm-vault-guard.test.sh are test-harness adaptations with no behavior change: ship spawns now require explicit --mode and --yolo (without them the per-harness install assertions never execute at all), and the assertion guarding upstream's turn-end plugin now names their renamed fm-busy-state.js.

The pre-existing upstream failures on macOS

Running the full suite on this macOS machine surfaces failures that are not caused by this merge. Each was reproduced identically - same assertion text - against a clean extraction of pure upstream/main, and upstream's own CI is 13/13 green on the exact commit merged here (6f0f87f).

Why these stay invisible upstream: every behavior-test lane (tests-portable-parallel-1/2, tests-portable-serial, tests-herdr) runs on ubuntu-latest. There is one macOS job, macos-stock-bash on macos-latest, but it only exercises snapshot consumers under stock Bash 3.2.57 rather than the behavior suite, so macOS-specific failures in the wider suite are never executed in CI.

In this PR's own pipeline run, 6 remain:

Test Cause
fm-backend-orca macOS bash 3.2: an EXIT trap masks a failed metadata write, so fm-spawn can exit 0 having recorded no state/<id>.meta. Real platform behavior, and upstream's bug, not ours.
fm-kimi-harness this machine's python3 is 3.9.6; tomllib needs 3.11+
fm-teardown herdr is not installed
fm-remote-secondmate-lifecycle-e2e needs an SSH-reachable remote host
fm-remote-secondmate-trace-context needs an SSH-reachable remote host
fm-gotmp pre-existing upstream, deterministic, reproduced on pure upstream and idle

A further four failures seen during investigation (fm-backlog-handoff, fm-secondmate-lifecycle-e2e, fm-remote-backlog-handoff, fm-decision-hold-lifecycle) shared one root cause: tasks-axi 0.2.3 against upstream's new required 0.2.4. That mattered operationally, not just in tests - fm-teardown.sh gates every non-forced scout teardown on fm-decision-hold.sh verify, which refused outright on 0.2.3 even when the gate was already attested. Upgrading to tasks-axi 0.2.4 (plus gh-axi 0.1.29 and lavish-axi 0.1.45, both also below the floors this merge enforces) resolved all four and restored scout teardown. Anyone else merging this should upgrade those three tools first.

fm-watcher-lock: load-induced, NOT a supervision gap

Stated plainly because supervision is load-bearing for a macOS fleet: this is not a real macOS supervision gap.

During the first full-suite run it failed with arm did not exit with HUP status (got 124) - a timeout waiting for the watcher-arm to exit on SIGHUP. I initially flagged that as a possible real supervision defect. That was wrong. Re-running the same test on an idle machine passes 30 of 30 checks, exit 0, and it also passes in this PR's own pipeline run. The failure was contention from running the full suite in parallel with other work, not a defect in SIGHUP handling. Watcher-arm supervision behaves correctly on macOS.

Every other failure above was re-checked idle and remained deterministic, so only this one was load-sensitive.

kunchenguid and others added 30 commits July 15, 2026 19:24
…kunchenguid#626)

* docs: compress firstmate operating contract

* docs: make delivery rigor single-owner

PR B already removed personal and stacked review requirements, but it did not explicitly assign rigor to the selected delivery path or forbid risk-based manual clean gates. That gap still permitted the Hi Bit inversion.

* no-mistakes(review): Honor configured merge authority across faster delivery paths

* no-mistakes(document): Align docs with compressed operating contract
* Add durable captain decision holds

* no-mistakes(review): Validate decision hold retries and origin paths

* no-mistakes(review): Enforce durable decision lifecycle boundaries

* no-mistakes(review): Harden decision display and partial retry recovery

* no-mistakes(test): Update scout teardown fixtures for decision inventory

* no-mistakes(document): Align decision lifecycle and scout teardown documentation

* no-mistakes: apply CI fixes

* no-mistakes(review): Reconcile terminal decision holds

* no-mistakes(document): Align captain decision-hold documentation
* fix: harden PR check artifacts

* fix: close PR check migration gaps

* fix: close PR check retry gaps

* fix: clarify migration outcomes and ESM boundary

* fix: keep failed migrations authoritative

* test: use inert PR validation fixtures

* no-mistakes(review): Reserve noncanonical PR quarantine namespace

* no-mistakes(review): Prevalidate final PR-check teardown artifacts

* no-mistakes(review): Preserve X metadata and validate teardown IDs

* no-mistakes(review): Initialize migration state before watcher exclusion

* no-mistakes(review): Isolate failed poll migrations from bootstrap recovery

* no-mistakes(review): Allow safe polling during incomplete private repairs

* no-mistakes(review): Authenticate watcher checks at execution time

* no-mistakes(review): Preserve custom checks with hash-bound registration

* no-mistakes(review): Clean custom check snapshots on watcher signals

* no-mistakes(review): Stop watcher checks promptly on signals

* no-mistakes(review): Terminate watcher check groups before cleanup

* no-mistakes(document): Correct stale X-mode watcher documentation

* fix: drain returned watcher check groups

* no-mistakes(review): Harden quarantine links and recover validated replacement polls

* no-mistakes(review): Preserve X mode across shim version transitions

* no-mistakes(review): Refresh legacy X shims before marker short-circuits

* no-mistakes(document): Correct persisted PR-check artifact documentation

* no-mistakes(document): Correct stale PR-check documentation

* fix: bind PR poll repair provenance

* no-mistakes(review): Enforce single-link ownership for custom check artifacts

* no-mistakes(review): Preserve private checks, X polling, and lifecycle IDs

* no-mistakes(review): Separate task creation and legacy teardown validation

* no-mistakes(review): Restore safe legacy operations and teardown validation

* no-mistakes(review): Disambiguate migration obligations and preserve legacy retries

* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence

* no-mistakes(review): Reconcile legacy migration retries and teardown collisions

* no-mistakes(review): Force legacy namespace reconciliation before marker short-circuits

* no-mistakes(document): Document private poll artifact safety contracts

* no-mistakes(lint): Suppress intentional literal-dollar lint finding

* fix: migrate historical X poll identity

* fix: harden PR check artifacts

* no-mistakes(review): Preserve fail-closed diagnostics and legacy quarantine evidence

* no-mistakes(review): Reconcile legacy migration retries and teardown collisions

* fix: migrate historical X poll identity

* no-mistakes(review): Harden X-mode artifact publication against symlink corruption

* no-mistakes(review): Guard X artifact publication

* no-mistakes(review): Enforce private X artifact reads

* no-mistakes(test): Fix backend compatibility fixture dependencies

* no-mistakes(document): Refresh PR-check documentation

* no-mistakes(lint): Remove unused x-mode test locals

* no-mistakes: apply CI fixes
* fix: compact session-start backlog digest

* no-mistakes(test): Fix legacy backend fixture helper

* no-mistakes(test): Fix watcher exit wait helper

* no-mistakes(document): Document compact backlog digest
* fix: dedupe stale watcher guard banner

* no-mistakes(review): Keep read-only guard state nonmutating

* no-mistakes(document): Clarify stale watcher docs
* fix: balance bearings landed defaults

* no-mistakes(document): Document balanced landed baseline

* no-mistakes: apply CI fixes
* docs: clarify captain-facing translation contract

* no-mistakes(review): Restore runtime fallback mandate

* no-mistakes(document): Align Bearings translation wording
…#646)

* fix: make bootstrap nudges deterministic

* no-mistakes(review): Honor state override for bootstrap nudges

* no-mistakes(review): Update benign bootstrap documentation labels

* no-mistakes(review): Validate bootstrap nudge retry markers

* no-mistakes(document): Align bootstrap nudge documentation

* no-mistakes: apply CI fixes
…nchenguid#649)

* Clarify concise secondmate registry contract

* no-mistakes(review): Expand secondmate registry boilerplate coverage

* no-mistakes(document): Point route docs to owner
…kunchenguid#654)

* fix(bin): strip quotes on blocked_by in decision-hold resolve

tasks-axi quotes multi-entry blocked_by as "a,b,c", so the comma-boundary
membership test only matched middle elements. Strip surrounding quotes
before matching so first and last hold ids resolve correctly.

* no-mistakes(document): Refresh decision-hold regression evidence
* feat(secondmate): inherit shared captain preferences

* no-mistakes(review): Honor shared captain data overrides

* no-mistakes(review): Honor bootstrap data override registry

* no-mistakes(document): Refresh shared inheritance docs

* no-mistakes(document): Clarify inherited local-material docs
* feat(spawn): gate local agent secret injection

* fix(spawn): align final Keychain slot

* no-mistakes: apply CI fixes
* test: isolate herdr autodetect smoke session

* no-mistakes(review): Restored autodetect smoke gate bypass

* no-mistakes(test): Harden Herdr lab provisioning

* no-mistakes(document): Refresh Herdr lab docs
* fix(pi): distinguish stale locks when arming watcher

* no-mistakes(test): Stabilize watcher extension async waits

* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate as house vocabulary

* no-mistakes(test): Update captain vocabulary contract test

* no-mistakes(document): Align secondmate documentation vocabulary
…uid#686)

* fix: parse secondmate home after pre-field parentheses

Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.

* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges

* no-mistakes(document): Document nudge script inventory
…nguid#688)

Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
…nsion (kunchenguid#205)

* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn

Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(brief): harden fm-brief regression coverage for parse and scaffolds

Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the

Co-authored-by: Cursor <cursoragent@cursor.com>
kunchenguid#166 apostrophe regression cannot return unnoticed.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#693)

* fix: make watcher supervision continuous

* no-mistakes(review): Bound watcher retries and log attached signals

* no-mistakes(review): Add bounded successor-recovery wake fallbacks

* no-mistakes(review): Prevent overlapping successor-arm retries

* no-mistakes(review): Resume supervision after late arm closes

* no-mistakes(review): Bind OpenCode recovery to attempted arm

* no-mistakes(test): Synchronize peer beacon regression fixture

* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures

* no-mistakes(document): Captain: document watcher successor protocol behavior

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix: always fetch PR head for review diffs

Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.

* no-mistakes(test): Isolate session-start nudge tests from gate state

* no-mistakes(document): Correct review-diff documentation
…and skills (kunchenguid#736)

* docs: resolve five contract contradictions

* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): reverify grok exit command

* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for exited paused crew

* no-mistakes(review): Gate pause suppression on confirmed agent death

* no-mistakes(test): Fixed stale pause cadence

* no-mistakes(document): Document dead-agent hold cadence
…id#744)

* fix(supervision): distinguish ordinary wakes from repair

* no-mistakes(review): Make passive guard follow-ups recovery-only

* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes

* no-mistakes(review): fix x-poll claim error deduplication

* no-mistakes(review): separate claim diagnostics from relay recovery

* no-mistakes(document): Document X-mode once-only mention wakes
…enguid#747)

* feat(wake): enrich drained signal context

* no-mistakes(review): Bound wake enrichment reads

* no-mistakes(document): Document wake-drain annotations

* docs(wake): explain at-least-once drain boundary

* no-mistakes(review): Prevent symlink races in wake annotations

* no-mistakes(review): Exercise wake symlink race regression

* test: document intentional AFK marker subprocesses

* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces

* no-mistakes(review): Harden Herdr projection creation and spawn serialization

* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission

* no-mistakes(test): Correct stale Orca metadata failure fixture

* no-mistakes(document): Document Herdr presentation projection accurately
kunchenguid and others added 27 commits August 3, 2026 09:38
* Add generic remote secondmate transport

* Add routed remote secondmate replies

* Add remote outbox backlog handoff

* Integrate remote secondmate lifecycle

* no-mistakes(review): Fix remote snapshot and handoff races

* no-mistakes(review): Serialize remote home provisioning transactions

* no-mistakes(review): Harden remote lifecycle transaction boundaries

* no-mistakes(review): Serialize remote lifecycle mutations and fail closed

* no-mistakes(review): Close remote lifecycle and file race windows

* no-mistakes(review): Serialize remote reply retirement and inheritance

* no-mistakes(review): Harden remote transfer integrity and recovery

* no-mistakes(review): Serialize remote respawn with registry retirement

* no-mistakes(document): Document remote bootstrap convergence accurately

* no-mistakes(document): Clarify skipped remote secondmate mutations

* no-mistakes(lint): Resolve remote script ShellCheck warnings

* no-mistakes: apply CI fixes
* feat(spawn): propagate a native W3C traceparent to spawned agents

Add a default-off capability that resolves one W3C traceparent for a task,
injects it into the agent's pane shell as the TRACEPARENT environment
variable immediately before launch, and records the identical value as
traceparent= in state/<id>.meta, so an external observer that explicitly
reads that env value or meta field can correlate a worker, a Secondmate, and
their nested children into one trace with no collector, storage, UI, or
vendor coupling.

TRACEPARENT as an environment variable is a firstmate convention carrying a
W3C-formatted value: W3C Trace Context standardizes the header, not an env
var, and OpenTelemetry SDKs do not read it automatically, so a downstream
must consume it deliberately; this feature parents no SDK span by itself.

Identity is per task, not per spawn: the carrier is minted with random ids on
the first spawn, adopted as a child (fresh span, same trace) for a nested
spawn whose parent already holds one, and reused verbatim from the meta on
relaunch, so a task keeps one stable logical identity across restarts. A
malformed or all-zero inherited value is treated as absent and roots a fresh
trace. A new root is sampled (01) - a sampling decision a downstream
parent-based sampler honors, not a guarantee that any collector stores a
span, and firstmate emits no spans; a child preserves the inherited flag.

Trust boundary: a firstmate-minted root is random and reads no prompt, path,
task prose, credential, or arbitrary environment key. An inherited
TRACEPARENT is opaque caller-controlled data - up to 24 bytes of id passed
through after syntax validation - so whoever set it controls those bytes, a
bounded fixed-width channel rather than a general content or secret channel.
The feature adds no OTEL_* variable, no tracestate, and no arbitrary
environment injection; it runs no configurable or arbitrary command, only the
fixed local od and tr (resolved from PATH) to read a few bytes of entropy - a
small local pipeline with no network or watchdog and no hard latency
guarantee. Any entropy or validation failure that returns omits the carrier
without aborting the spawn. A default-off spawn leaves the generated meta and
launch environment unchanged.

Enablement is default-off (config/trace-context, or FM_TRACE_CONTEXT where a
non-empty value overrides and unset or empty defers to the file) and is
propagated into secondmate homes, taking effect at each agent's next launch:
a Secondmate launched or relaunched after enablement carries the primary
trace into its nested workers, while an already-running Secondmate roots new
traces for its own workers until relaunched. Injection reuses the existing
GOTMPDIR channel, so all spawn backends and harnesses and the ship, scout,
and secondmate paths are covered.

Covered by a pure-library suite and a spawn-path integration test (fake tmux
plus a real worktree, hermetic against ambient FM_TRACE_CONTEXT) proving the
recorded and injected carriers are identical and sent before launch, that
default-off writes and injects neither, that a relaunch reuses the recorded
carrier, and that an explicit FM_TRACE_CONTEXT overrides the file both ways;
plus a source-owner inheritance test proving trace-context propagates and
absence-mirrors through propagate_inheritable_config.

Documentation follows the repository documentation-audiences contract:
docs/trace-context.md is maintainer-architecture rationale, the configuration
schema lives in docs/configuration.md, and the repeatable test evidence is
separated into docs/verification/trace-context.md (maintainer-verification),
registered in docs/documentation-audiences.json.

* fix(spawn): propagate the effective trace-context decision to secondmates

FM_TRACE_CONTEXT overrode trace context only in the process that read it. A
newly launched secondmate decided enablement from the inherited
config/trace-context file alone, so the override did not cross the
primary-to-secondmate boundary: FM_TRACE_CONTEXT=off with the file present left
the secondmate's nested workers traced (a broken kill switch), and
FM_TRACE_CONTEXT=on with the file absent left them untraced despite the
inherited carrier.

Deliver the primary's effective decision to a newly launched secondmate as a
normalized on/off FM_TRACE_CONTEXT in the launch prefix, so a FM_TRACE_CONTEXT
override governs the nested primary -> secondmate -> worker chain both ways, not
just the copied file. The value is bounded to the literal on/off and does not
broaden environment injection; the already-running secondmate boundary is
unchanged.

Add a genuine two-level spawn regression that drives fm-spawn twice with the
exact environment the primary injects into the secondmate and proves both
divergent directions end to end. Correct the documentation that implied
secondmate coverage on every backend, since orca and cmux reject secondmate
spawns, and refresh the verification evidence for the new assertion count.

* no-mistakes(review): Clarify Secondmate trace-context launch snapshots

* no-mistakes(document): Correct trace-context documentation ownership and relaunch semantics

* fix(spawn): resolve the trace-context decision once for carrier and snapshot

The effective trace-context decision was read twice per spawn: once inside
fm_trace_context_resolve for the recorded carrier, and again for the secondmate
FM_TRACE_CONTEXT launch snapshot. A config-file change between the two reads
could pair a carrier with the opposite enable state - an injected carrier with
an off snapshot, or no carrier with an on snapshot.

Freeze the effective on/off decision once, drive the carrier resolution under
that frozen FM_TRACE_CONTEXT so it cannot independently re-read the file, and
reuse the same frozen decision for the secondmate launch snapshot. Add a
spawn-path regression that drives the file-decided path and proves the recorded
carrier and the delivered snapshot always agree, and refresh the verification
evidence for the new assertion count.

* no-mistakes(review): Preserve legacy Secondmate trace boundary

* no-mistakes(document): Correct trace-context verification comparison base

* no-mistakes(review): Captain, prevent failed trace delivery metadata claims

* no-mistakes(review): Captain, align trace-context tests and verification evidence

* no-mistakes(document): Correct trace-context verification evidence

* no-mistakes(lint): Suppress intentional ShellCheck literal-dollar warnings

* no-mistakes(review): Captain: freeze trace context at session start

* no-mistakes(test): Captain: stabilize scheduler test and document Kimi trace coverage

* no-mistakes(document): Document trace-context safety boundaries

* fix(trace): fail off on stale session snapshots

Publish each home session decision atomically through a same-directory temporary file and bind it to the current session lock. A replacement failure can no longer leave an earlier on decision active in a later session; missing, stale, malformed, or unpublishable state defaults safely to off.

Add regressions for read-only replacement and failed publication, update spawn and session-start fixtures for the lock-bound format, and refresh the architecture and verification records.

* no-mistakes(review): Fix trace spawn failure independence and duplicate safety

* no-mistakes(document): Refresh trace-context documentation and verification

* no-mistakes(review): Clear partial backend input after failed trace submission

* no-mistakes(review): Stop unsafe trace delivery before launch append

* no-mistakes(document): Document unsafe trace delivery handling

* fix(trace): bound each trace to one routed task, never the routing agent

A persistent Secondmate holds its launch-time TRACEPARENT in the process
environment for its whole life, and routed requests never replace it, so
resolving new-task carriers from the ambient environment chained every
routed task into one ever-growing trace per Secondmate with distinct
parent ids. Resolve now reuses the task's recorded carrier or mints a
fresh sampled root, never reading ambient TRACEPARENT, so each routed
task is its own trace boundary while relaunch, recovery, and
scout-to-ship promotion keep one stable per-task identity.

The spawn regression models the reviewed scenario exactly: two unrelated
tasks spawned sequentially through one persistent Secondmate environment
record and inject distinct trace ids, adopt nothing from the Secondmate's
carrier, and a relaunch of the first task reuses its original carrier
verbatim.

* docs(trace): define the per-task trace boundary

The design contract is one task per trace: a persistent Secondmate is
routing infrastructure with its own agent identity, never a shared trace
root for the unrelated tasks routed through it. Root/recovery semantics
replace the removed child-inheritance path, the sampling and safety
sections drop inherited-carrier language because ambient TRACEPARENT is
never read, and the verification page records the refreshed suite
inventories including the two-task Secondmate boundary regression.

* test(trace): adopt the explicit per-task delivery contract in spawn fixtures

Rebasing onto current main brings the explicit per-task delivery contract:
ship spawns now require --mode and --yolo instead of resolving them from the
project registry. The trace spawn fixtures pass the same explicit contract
canonical spawn tests use, preserving the per-task trace boundary coverage
unchanged, and the verification page records the refreshed comparison base.
* fix(bin): classify tmux agent liveness independent of process titles

`fm_backend_tmux_agent_state` attributed a pane solely from
`#{pane_current_command}`, which is a process TITLE a harness can rewrite,
not a structural fact. Claude Code 2.1.220 reports its version string there,
so a live Claude endpoint classified `ambiguous`: the session-start secondmate
liveness sweep could no longer see it, and any consumer that gates on a
positive classification refuses outright.

Read a second, independent name source: the kernel `comm` of every process in
the pane tty's foreground process group. Either source naming a verified
harness yields `alive`, because a false `dead` is the one verdict that can
start a duplicate agent on a live worktree. Scoping to the foreground process
group rather than the pane's descendants keeps a harness-named background
process from faking an agent, and covers multi-process launchers (the Pi
Launcher path) without a special case.

Verified on 2026-08-03 against all seven adapters running for real on tmux
3.6a / macOS 26.5.2 arm64: claude 2.1.220, codex-cli 0.146.0, opencode
1.18.11, pi 0.82.0, pi-signed 0.82.0, grok 0.2.118, kimi 0.31.1 all classify
`alive`, each attributed by a source independent of its title.

Two tests, because they fail for different reasons:
- tests/fm-tmux-agent-liveness.test.sh pins the logic with real processes and
  no harness, so it runs everywhere CI runs tmux. It drives the two name
  sources apart on purpose and asserts the divergence, so no case can go
  quietly vacuous.
- tests/fm-harness-liveness-drift-live-e2e.test.sh relaunches every installed
  harness and fails naming the harness and version when one stops being
  attributed by a title-independent source.

AGENTS.md section 4 carries the resulting standing rule, and
firstmate-coding-guidelines owns how to satisfy it.

* no-mistakes: apply CI fixes

* docs: move the harness-dependent-check policy out of AGENTS.md

The standing rule was stated in AGENTS.md section 4 with the mechanics in
firstmate-coding-guidelines, which split one contract across two owners and
charged every session for a rule that only fires when firstmate's own
harness-dependent code is being changed.

firstmate-coding-guidelines is now the single owner of both the rule and how
to satisfy it: real-harness proof required, that proof authorized to spend
tokens, structural signals preferred over vendor-rendered surfaces, and a
guard that fails loudly naming the harness and version where a surface signal
is unavoidable. No inline stub is left behind, because AGENTS.md already
carries the load trigger for that skill in sections 7 and 13, so it is read
before any change to firstmate's shared tracked material.

Also records the cross-platform lesson the pipeline caught in the portable
regression, and corrects that file's header: the divergence assertion lives
on the version-string case, which diverges on both supported platforms,
rather than on every case.

* no-mistakes(review): Harden tmux liveness identity and drift validation

* no-mistakes(document): Clarify cross-platform tmux liveness documentation
…#1609)

* feat(bin): trace remote secondmate routes and unify the inherit allowlist

Per-task W3C trace context (kunchenguid#995) resolved and injected its carrier only at
the local spawn path. A remote secondmate is routed through
spawn_remote_secondmate, which returns long before that site and wrote its own
metadata block, so a remote secondmate stayed silently untraced even with the
capability enabled.

The parent home still owns that task's identity, because it holds the metadata
an observer reads. It now resolves the carrier against the task's own meta
under its own frozen decision - reused verbatim on relaunch, freshly rooted
otherwise, never adopting the parent process's ambient TRACEPARENT - and hands
it to the configured host through a new fm-spawn --traceparent argument,
accepted only for a secondmate launch and only as a strict W3C value. The
remote host exports it at the same unconditional pre-launch site and reports
back the carrier its endpoint actually holds, which the parent records, so an
already-alive endpoint reports the identity its agent really received rather
than one the parent merely intended. Disabled remains byte-identical and off.

The remote inherit path also carried its own hardcoded copy of the inheritable
config set, already drifted from FM_INHERITABLE_CONFIG by trace-context. Both
remote ends now derive from that one declaration, so a future item cannot be
sent by one side and refused by the other, and session-scoped enablement items
are skipped on live convergence exactly as the local path skips them.

Also fixes a latent stderr leak: an absent session lock printed a raw redirect
failure, which the new remote resolve site made visible.

Adds tests/fm-remote-secondmate-trace-context.test.sh, driving the real
parent -> fm-on -> remote entrypoint -> control -> remote fm-spawn chain over
the deterministic SSH boundary and reading the carrier back from the remote
pane's own log.

* no-mistakes(document): Clarify remote trace and allowlist contracts
* feat(bin): widen the remote runtime PATH and add a remote doctor preflight

The fixed remote entrypoint hard-coded a four-directory PATH, so a remote
account whose tools live under nix or a per-user profile could not run basic
Firstmate work without a login shell. The entrypoint now composes its child
PATH from the code root's bin, the account's ~/.local/bin, the common
package-manager directories that actually exist on the host, and the portable
system tail, deduplicated and in a fixed order, still under env -i with the
same variable allowlist and no shell command string.

fm-remote-doctor.sh reports that exact PATH by inheriting it from its own
entrypoint launch rather than recomposing it, so the ordering keeps one owner.
It is read-only, reports where each required and optional tool resolved, and
exits non-zero naming every required tool that did not. Remote seeding runs it
as a preflight before anything is created on the host and restores the registry
when it fails.

* no-mistakes(review): Harden remote git authorization and missing-tool diagnostics

* no-mistakes(document): Document remote PATH doctor and safe shims

* no-mistakes(lint): Fix ShellCheck findings in remote path tests

* no-mistakes(lint): Suppress exported fixture's false-positive ShellCheck warning
* feat(bin): gate remote second mates on herdr readiness

A remote second mate now always runs on the Herdr backend, whose server
belongs to the host's GUI login session and therefore outlives the SSH
connections that supervise it. fm-spawn's remote route forces that backend
and the host-local control script refuses any other, so the requirement
cannot be dropped from either side.

fm-remote-doctor.sh becomes the single owner of what "ready" means. It keeps
its PATH and tool reporting from kunchenguid#1623 and adds the Herdr, Aqua LaunchAgent,
GUI-session, server-reachability, and entrypoint-symlink checks, tagging each
gap fixable: or human: with the exact operator step. --fix closes only the
automatable gaps - writing and loading the Aqua-scoped dev.firstmate.herdr
launch agent, starting the server where no launch agent applies, and
recreating the entrypoint symlink - then re-derives every check from the host,
so a human gap is never presented as fixed. It never creates a login session,
writes an auto-login password, or touches FileVault.

Remote seed, remote spawn, and the startup liveness relaunch all run the same
check, repair, re-check sequence through one shared library and fail closed
with the doctor's own gap text. Recovery inherits the gate because it respawns
through the same route.

Tests drive the real doctor against a controlled account fixture with a
private HOME, a state-backed launchctl, and a fake herdr, and prove the
dangerous actions are never attempted. The remote lifecycle suites gain a
stateful Herdr CLI fixture and answer the readiness gate at the SSH boundary,
so they never inspect or repair the runner's own account.

* no-mistakes(review): Validate launch-agent contract and confirm Herdr startup

* no-mistakes(review): Validate loaded launch-agent contract before readiness

* no-mistakes(review): Refuse legacy remote backends without altering routes

* no-mistakes(review): Clarify conditional remote readiness repair sequence

* no-mistakes(review): Repair remote readiness before liveness probing

* no-mistakes(review): Preserve unknown seeds and reject legacy liveness

* no-mistakes(document): docs: clarify remote Herdr backend ownership
…1659)

* Pin remote secondmates to fm-remote

* no-mistakes(review): Fail closed on legacy remote Herdr endpoints

* no-mistakes(review): Isolate fm-remote launch agent from interactive default

* no-mistakes(document): Document shared remote Herdr retirement safety
)

* feat: run remote commands through Aqua job worker

* no-mistakes(review): Enforce remote job deadlines and safe worker shutdown

* no-mistakes(review): Refresh stale workers and harden dependency-free supervision

* no-mistakes(review): Harden worker ownership recovery and shutdown quarantine

* no-mistakes(review): Fix doctor bootstrap, harness repair, and output draining

* no-mistakes(review): Probe doctor tools through authenticated worker bootstrap

* no-mistakes(review): Refresh stale workers before doctor tool probes

* no-mistakes(review): Recover stopped quarantines and extend job deadlines

* no-mistakes(review): Separate queue and execution timeout windows

* no-mistakes(review): Supervise Linux worker crashes and bind root identity

* no-mistakes(review): Resolve authorized Nix profile bin links

* no-mistakes(review): Clarify Nix path resolution documentation

* no-mistakes(review): Harden PATH safety and nvm selection

* no-mistakes(review): Honor nvm system defaults and refresh doctor digest

* no-mistakes(review): Keep workers ready during active jobs

* no-mistakes(review): Bound pre-execution validation by job timeout

* no-mistakes(document): Clarify remote worker documentation

* no-mistakes(lint): Fix remote worker ShellCheck diagnostics

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(remote): arm SSH dead-peer detection in fm-on.sh

A vanished remote host mid-poll (a reboot, a dropped link) left ssh
blocked indefinitely on a half-open TCP connection, because fm-on.sh's
ssh invocation had no ServerAliveInterval/ServerAliveCountMax. This
wedged the remote-reply ferry: fm-procevent.sh's runner blocked inside
the ssh child and never reached its own no-result -> claim-release ->
reconcile re-arm self-healing path, which otherwise already handles a
nonzero exit with empty output correctly. Recovery required a manual
retire and re-arm.

Arm ServerAliveInterval=15 and ServerAliveCountMax=3 by default
(bounded ~45s detection window), both overridable via
FM_SSH_ALIVE_INTERVAL and FM_SSH_ALIVE_COUNT_MAX. This is a transport-
level fix in fm-on.sh, so it covers every remote command routed
through it, not just the reply ferry. The remote sshd answers
keepalive probes independently of whatever the remote command is
doing, so a legitimately long-but-alive command (a 55s poll, a clone,
the doctor) is never falsely killed - only a truly vanished peer trips
it, turning that case into a bounded, detectable ssh failure (exit
255) instead of an indefinite hang.

Extends tests/fm-on.test.sh with a behavioral regression asserting a
bounded, positive ServerAliveInterval/ServerAliveCountMax on the real
ssh argv captured through the FM_SSH_BIN process seam, plus coverage
that both are env-overridable.

* no-mistakes(document): Document SSH dead-peer detection ownership
* feat(bootstrap): gate stale axi CLIs at the floors firstmate actually uses

Add gh-axi 0.1.29 floor so bare --squash PR merges stop failing quietly on
older builds. Raise tasks-axi to FM_TASKS_AXI_MIN=0.2.2 (multi-id mv) while
keeping feature probes. Keep quota-axi at 0.1.16 after verifying schema 3
and per-model availability already ship there; runway remains optional.

* no-mistakes(document): Clarify AXI compatibility documentation ownership
…d#1661)

* fix(guard): stop false watcher-down alarm mid-turn under Claude auto-arm

bin/fm-guard.sh derived its watcher-health verdict from fm_watcher_healthy,
which requires a live watcher process holding the home lock. Under the Claude
Stop-hook auto-arm supervision model the watcher is armed at each turn end and
exits on its wake, so it runs only between turns. Every guarded command run
mid-turn therefore found no live watcher and printed the "WATCHER DOWN -
SUPERVISION IS OFF" banner even though supervision was healthy. Because the
episode key was derived from the beacon mtime (which the between-turns watcher
advances every poll), the full banner re-printed on essentially every command,
and the message always blamed a "fresh beacon" that was in fact fresh.

Make the pull guard's health check model-aware via a new
fm_watcher_supervision_verdict in bin/fm-wake-lib.sh:

- Under the auto-arm model a beacon fresh within FM_GUARD_GRACE is healthy even
  with no live watcher process; only a beacon stale beyond grace (or absent) is
  a genuine lapse and alarms.
- Under every persistent-watcher harness (codex foreground checkpoint,
  opencode/pi/grok background arm, tmux, unknown) a live identity-matched
  watcher with a fresh beacon is still required, unchanged.

The banner now names the true failing condition, a missing live watcher process
versus a genuinely stale beacon, instead of always blaming the beacon, and the
once-per-episode dedup keys on that condition rather than the beacon mtime so a
genuine lapse announces once and does not re-print each turn.

The turn-end guard keeps the strict fm_watcher_healthy check because it fires at
the turn boundary, where the auto-arm brings a fresh watcher up and it
cooperates with that arm. fm_watcher_healthy itself is unchanged, so the arm
layer's start/attach/replace decisions are unaffected.

Tests in tests/fm-guard-stale-banner.test.sh cover the auto-arm healthy
fresh-beacon-without-a-watcher case, the auto-arm stale-beacon alarm and its
stable episode, the true-reason banner wording, and the reason-keyed episode
surviving a beacon mtime change; existing persistent-model cases are pinned to
that model.

* no-mistakes(review): Pin secondmate supervision model to launched harness

* no-mistakes(document): Align watcher documentation with model-aware supervision health
* fix(tests): stop fixture-tempdir helper from self-deleting under command substitution

fm_test_tmproot is almost always called as `TMP_ROOT=$(fm_test_tmproot prefix)`,
which forks a subshell to capture its stdout. The old implementation set its
EXIT cleanup trap inside that call, so the trap fired - and deleted the fixture
root - the instant the subshell exited, before the real caller's own EXIT trap
was ever installed. Every test using the documented call pattern leaked its
fixture root on every run; two suites had already independently discovered and
worked around this with ad-hoc mktemp calls.

Registration now goes through a $$-keyed registry file instead of in-process
state, since $$ resolves to the invoking shell's PID even inside the
subshell. The real cleanup trap is armed once at source time (always the real
caller, never a subshell) for EXIT, INT, and TERM. A best-effort orphan sweep
on next source reaps marked fixture roots old enough to be from a killed prior
run.

Simplifies the two existing ad-hoc workarounds (fm-procevent.test.sh,
wake-helpers.sh) back onto the shared helper now that it works correctly.

* no-mistakes(review): Preserve live fixtures during orphan reaping

* no-mistakes(review): Harden fixture ownership against PID reuse

* no-mistakes(review): Secure cleanup registry against path precreation

* no-mistakes(review): Make fixture registration transactional

* no-mistakes(document): Documentation already matches fixture cleanup behavior

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* feat(herdr): default presentation spaces on with an explicit opt-out

Herdr's disposable one-task presentation workspace was opt-in through the
presence of local config/herdr-presentation-spaces. It is now on by default,
and a home opts out by writing "off" into that same file.

Values are read with the whole-file whitespace-stripped convention the other
scalar config items already use, plus case folding. An absent file, an empty
file, and "on" all resolve on; only "off" opts out; an unrecognized value warns
and keeps the default rather than failing a spawn over a purely visual setting.
The empty file is exactly the historical opt-in form, so every home that had
already enabled the projection stays enabled with no migration step, and no
previously enabled home can be turned off by the flip.

Because absence now means on at both ends, secondmate inheritance needs no
item-specific convergence: mirroring an absent primary file converges a
secondmate to the same default-on rather than turning its projection off, and
only an explicit primary opt-out propagates the opt-out.

The gate itself moves into fm_backend_herdr_presentation_enabled in the Herdr
adapter so the semantics have one owner that regressions can exercise directly.

* no-mistakes(document): Document Herdr default-on presentation safety

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>
…henguid#1711)

* fix: surface consolidated open decisions on every wake-drain

A needs-decision or blocked event buried under later, unrelated status
appends was only ever shown via the last-line wake annotation, so a
still-open captain decision could go silently missed even though
status_open_decisions (fm-classify-lib.sh) already folds the whole
status stream correctly and fleet-snapshot/bearings already reuse it.

Wire that same fold into bin/fm-wake-drain.sh: a new fleet-wide
scan_open_decisions wrapper scans every state/<id>.status, and
fm-wake-drain.sh prints a separate, bounded OPEN DECISIONS section on
every drain (including the empty-queue fast path), so session-start
and every wake-handling turn surface it for free without duplicating
the open/resolved fold itself. Heartbeat wakes drain through the same
script, so this covers that surface too.

Also tighten status_open_decisions' file guard to skip an unreadable
status file instead of leaking a bash redirection error, now that a
fleet-wide directory scan can reach files a single targeted read
would not.

* no-mistakes(review): Prevent status symlinks leaking open decisions

* fix: drop unbounded perl subprocess from status symlink guard

The review step's own symlink-safety auto-fix (O_NOFOLLOW read via a
perl subprocess) forked one perl process per status file scanned by
the new fleet-wide open-decisions scan, with no cap - inflating
fm-wake-drain.sh's total external-read cost from 8 (the existing
annotation read_cap) to 18 in the enrichment-caps regression test.

The plain [ -L "$f" ] check already rejects any status file that is
itself a symlink before any read happens, which is exactly what the
new regression test exercises and is the same defense level the
sibling scan_captain_relevant_statuses/last_status_line already rely
on elsewhere in this file (no O_NOFOLLOW). Drop the subprocess-based
nofollow read and keep the cheap builtin guard.

* no-mistakes(document): Document actionable fleet-wide open decision drains
…kunchenguid#1710)

* fix(bin): abort orphaned no-mistakes runs and reap leaked processes at teardown

Teardown could remove a task's worker while its no-mistakes pipeline run was
still parked at a gate, leaving an orphaned run holding a fleet slot
indefinitely (observed 2026-08-03: runs parked 7h39m and parked at a
post-CI approval gate). It could also leave backgrounded/disowned
descendant processes rooted under the worktree or tasktmp surviving
reparented to init (observed: two `go test` binaries pinning CPU for
hours with no live task meta to attribute them to).

Add two coupled pre-teardown steps, both scoped to this task's exact
branch/head or worktree/tasktmp so they can never touch another task's
run or processes:
- conclude_task_no_mistakes_run aborts a run parked at a gate via
  `no-mistakes axi abort`, cd'd into the exact worktree so the daemon
  resolves the run itself rather than teardown naming a --run id.
- reap_task_worktree_processes sweeps for processes whose cwd is under
  the worktree or tasktmp (via `lsof -a -d cwd`) and TERM/KILLs them.

Both run before any worktree return, branch delete, or backend kill,
and are idempotent on a retried teardown. The branch+head attribution
logic is factored out of bin/fm-crew-state.sh into the new shared
bin/fm-nm-run-lib.sh so both scripts use the same ownership contract.

* no-mistakes(review): Fail closed on incomplete teardown cleanup

* no-mistakes(review): Bind teardown cleanup to verified run and process identities

* no-mistakes(review): Require confirmed aborts and convergent identity-safe process reaping

* no-mistakes(review): Handle process exits during teardown identity checks

* no-mistakes(review): Restore teardown library in hermetic gotmp fixtures

* no-mistakes(document): Document teardown run attribution and timeout

* no-mistakes(lint): Rename shell variable conflicting with done keyword

* no-mistakes: apply CI fixes
kunchenguid#1709)

The script installs as a symlink under ~/.local/bin. Taking dirname of
the symlink itself (instead of its real target) pointed SCRIPT_DIR at
~/.local/bin, breaking sourcing of the sibling fm-remote-job-lib.sh.
Resolve the real path first, preferring python3's os.path.realpath,
then realpath, falling back to the raw BASH_SOURCE on hosts with
neither.
…d#1724)

* fix(pi): stop Calm claiming a built-in tool name another extension owns

fm-calm.ts claimed bash/read/edit/write/grep/find/ls unconditionally at
extension load, regardless of whether Calm was on. Pi resolves two
extensions registering the same built-in name by first-registered-wins
with no merge and no unregister call, and Calm's project-local
.pi/extensions/ position beats any global or CLI-configured extension,
so a user who never even enabled Calm could have their own bash/read/etc
override silently replaced.

Captain-approved plan implemented:

- Registration is now gated on config/calm already being "on" at load
  time. A Calm-off session or reload registers nothing, so a non-Calm
  user never contests a name. This stays synchronous during the
  factory's own load, not deferred to session_start: /reload (and
  ctx.newSession/fork/switchSession) render the restored transcript from
  a pre-session_start snapshot of the tool registry, so a deferred claim
  would miss that render - confirmed by tests/fm-calm-pi-extension
  .test.sh's hidden-block-geometry E2E when trialed.
- The first time Calm turns on in a session that started off
  (activateBuiltInsIfNeeded, from the /calm command handler), Calm calls
  pi.getAllTools() - safe only once every extension has finished loading,
  unlike the load-time path above - to see whether a different extension
  already owns a name, and skips claiming only that one, leaving it and
  its owning extension fully intact and callable.
- A contested name found this way prints a prominent ctx.ui.notify()
  warning naming the tool, plus a console diagnostic.
- reportBuiltInLosses() remains the backstop for the one case neither of
  the above can reach: a session that starts or reloads with Calm already
  on, where the registry snapshot is taken before Calm gets any chance to
  check ownership. A symlink-safe realpath comparison avoids misreporting
  Calm's own registration as foreign when its path crosses a symlink
  (macOS /tmp, /var).

Confirmed, bounded trade-off: the very first time a session that started
Calm-off turns Calm on, tool-call rows already on screen from before that
toggle do not retroactively collapse, because Pi never lets an extension
re-point an already-rendered row at a definition registered later. Every
session after that first toggle starts with the preference already on and
takes the synchronous load-time path, so the guarantee is intact from
then on. docs/calm.md and the file's own header document this in full.

tests/fm-calm-pi-extension.test.sh gains test_builtin_gate_load_time
(config/calm off registers nothing, on registers all 7 synchronously at
load) and test_calm_activation_collision_and_regression_bound (first
activation claims every uncontested built-in, leaves a foreign bash tool
fully intact and callable, warns and logs the contested name, and locks
in the documented pre-activation bound against real ToolExecutionComponent
rendering). test_rendering_and_session_lifecycle and the live interactive
E2E are updated for the new gate-at-load and first-activation-bound
contract.

* no-mistakes(document): Document Calm tool collision boundaries

* no-mistakes: apply CI fixes
…#1727)

* fix(bin): give secondmate homes a durable parent binding record

Finished-worker cleanup on a remote second mate refused forever with
"cannot resolve the primary home ... durable parent binding". The
remote launch hands the child the remote code checkout as its parent
home (fm-spawn.sh's sole writer of FM_PUBLIC_FOLLOWUP_PRIMARY_HOME
receives FM_HOME=$FM_ROOT from fm-remote-secondmate-control.sh's
host-local launch), and that path can never carry the parent's real
records, so the guard refused unconditionally once relay looked active
anywhere on that host.

fm-home-seed.sh and fm-remote-home-provision.sh now write a durable
.fm-secondmate-parent record next to the .fm-secondmate-home identity
marker, naming the home's route to its parent as local (with the real
parent path) or remote (with the parent's SSH alias for diagnostics
only). fm-teardown.sh's cleanup gate reads it: a remote parent is out
of scope for the delegated-promise check (the whole promised-public-
reply subsystem is same-filesystem by construction, so a remote parent
can never hold one), while a token committed directly to the child's
own .env file - never the process environment - still refuses, so an
unrelated export in the remote host's login shell can no longer mask
in. For a local secondmate, the durable parent_home now also backs up
the launch-time env var, closing a silent fail-open where a restart
that dropped the launch prefix made the guard treat a genuinely active
parent relay as off.

Regression coverage drives the real remote route (SSH boundary + Herdr
fixture) and real fm-home-seed.sh seeding rather than hand-crafted
markers.

* no-mistakes(review): Captain: fail closed on unsafe durable parent records

* no-mistakes(review): Captain: enforce durable parent binding commit protocol

* no-mistakes(review): Captain: publish local parent binding before identity

* no-mistakes(review): Captain: refuse conflicting local parent bindings

* no-mistakes(review): Captain: reject non-regular secondmate seed leaves

* no-mistakes(review): Captain: enforce unique durable parent bindings

* no-mistakes(review): Captain: reject route-incompatible durable parent fields

* no-mistakes(document): Document durable secondmate parent bindings

* no-mistakes(lint): Fix secondmate parent parser ShellCheck warnings

* no-mistakes: apply CI fixes
* feat(bin): gate lavish-axi at its session_ended floor in bootstrap

bin/fm-procevent-lavish.sh decides that a human "Send & End" review is
terminal by reading session_ended from the poll response's leading session
block. That field first shipped in lavish-axi 0.1.35, so an older installed
build silently leaves every ended review source armed forever and captures
an empty ended result on each later cycle. The same release is what makes a
plain reopen refuse a session the human deliberately ended.

Add LAVISH_AXI_MIN=0.1.35 to the existing axi-family floor structure in
bin/fm-bootstrap.sh, reusing tool_version_at_least and the same MISSING
diagnostic gh-axi already emits, so an incompatible build is reported as an
upgrade request before any review surface is armed. Later lavish-axi
releases only add artifact-authoring surface the adapter never reads, so
the floor is the feature-introduction point rather than latest.

Fixtures that stubbed lavish-axi as a bare exit-0 tool would now be read as
unparseable builds, so tests/lib.sh gains fm_fake_version_tool and every
bootstrap-running suite uses it for lavish-axi.

* no-mistakes(review): Clarify lavish-axi version floor rationale

* no-mistakes: apply CI fixes

* feat(bin): set axi-family floors to current latest under the bump policy

The axi-family bootstrap floors are the CURRENT LATEST published version of
each tool, captain-bumped periodically to move the whole fleet onto the
newest axi tools. They are not the minimum feature-introduced version. The
earlier lavish-axi work set a feature-minimum floor, which is the opposite
of this policy, so replace it along with the older feature-minimum rationale
carried by tasks-axi and quota-axi.

State the policy explicitly in bin/fm-bootstrap.sh's header, which owns it,
and in each per-tool floor owner, so no future change argues a floor back
down to the earliest release that happens to satisfy some behavior. Remove
the lavish-axi session_ended and upstream-PR citation, the tasks-axi
multi-ID-mv minimum argument, and the quota-axi credential-source argument
as floor rationale; the tasks-axi feature probes remain as a separate
defense-in-depth concern.

Floors: lavish-axi 0.1.45 (was 0.1.35), tasks-axi 0.2.4 (was 0.2.2),
quota-axi 0.1.17 (was 0.1.16), gh-axi 0.1.29 unchanged and already latest.
Each was verified against the tool's current published version.

The mechanism is unchanged: the same shared version helper and the same
MISSING diagnostic path. The below-fires and at-or-above-silent regression
rows move to the new floors, keeping each boundary genuine by pinning the
patch immediately below each floor rather than a version that was only
below the old one. Fleet fixtures move to the new floors so a bootstrap-
running suite is not reported as an out-of-date build.

Three operator-facing backlog handoff and receipt errors named "0.2.2+"
while the enforced floor moved, so they now point at the floor's owner
instead of duplicating a version number that drifts.

* no-mistakes(review): Centralize AXI floor policy beside constants

* no-mistakes(review): Clarify bootstrap boundary test comment

* no-mistakes(document): Centralize AXI floor policy rationale
…guid#1737)

* fix(bin): bound OPEN DECISIONS scan cost with a per-status-file cursor

The fleet-wide OPEN DECISIONS scan added in kunchenguid#1711 re-reads and refolds
every task's entire lifetime status log on every drain, so its cost
grows unbounded with total log size. Add status_open_decisions_incremental
and scan_open_decisions_incremental to fm-classify-lib.sh: they persist a
per-status-file byte cursor plus the folded open-decision set, and fold
only newly appended bytes on each call, reusing status_open_decisions'
exact fold-line rule (extracted into _fm_decision_fold_line) so the two
strategies can never disagree on what is open. A missing or invalidated
cursor (new task, truncated/rewritten/shrunk log) falls back to a full
re-fold. bin/fm-wake-drain.sh now calls the incremental wrapper instead
of the whole-file scan.

* fix(bin): add O(1) rotation detection and read-failure guarding to the cursor fold

Add the two pieces the incremental open-decisions cursor was missing,
scoped to this repo's actual status-file usage (create-once, append-only,
never replaced or rewritten in place):

- An O(1) device+inode identity check (one stat call) alongside the
  existing size-shrink check, so a status file replaced/rotated/recreated
  at the same path is detected and falls back to a full re-fold, even
  when the replacement is the same size. A same-inode, same-size,
  in-place byte edit is a deliberately accepted gap: no code path in
  this repo ever does that to a status file.
- Checked reads: a stat/wc/tail failure is a genuine I/O error, not
  "the file is empty" - it now reports the already-trusted persisted
  open set unchanged instead of risking a silent invalidation.

Both stay O(1) plus new bytes per call, matching the cursor's bounded-
cost design; no content hashing or pending-fragment machinery.

* no-mistakes(review): Preserve cursor state across failed incremental reads

* no-mistakes(review): Refold status when cursor cache reads fail

* no-mistakes(document): Document cursor-backed open-decision scanning

* no-mistakes: apply CI fixes
…guid#1754)

* fix(bin): preempt remote reply long-polls for queued short jobs

Session start on a home with live remote second mates could stall silently
for many minutes: the single serial remote job worker ran each armed
fm-remote-delta-read.sh reply poll to its full 55s window while bootstrap's
short sync, inherit, state, and route commands sat queued behind it, and
non-FIFO queue pickup let re-armed polls keep winning the lane. Measured
end to end, a trivial short job took 31s behind one 30s poll window.

The worker now preempts a running preemptible job (the read-only, cursor-
anchored delta read is the only member of that class) as soon as a
non-preemptible job is queued, publishing exit 75 with emptied output -
byte-identical to the poll's own elapsed-window-with-no-data result - so
the parent runner takes its existing no-result path and the watcher re-arms
from the same cursor with nothing lost. The delta read translates SIGTERM
into that same exit after removing its staging directory. Sibling polls
never preempt each other, so two armed monitors cannot churn. The same
measured scenario now completes in 1s.

* no-mistakes(document): Clarify remote poll preemption documentation
Brings this fork up to date with 150 upstream commits while preserving both
local commits: the scout implementation contract (#1) and the crewmate
secret-value block (#3).

Conflicts resolved by keeping upstream's version of upstream's work and this
fork's version of its own, and both behaviors where they genuinely interleave:

- AGENTS.md section 13: upstream's project-management continuation line and
  this fork's scout-implementation-contract trigger were independent additions
  at the same point; both kept.
- bin/fm-spawn.sh codex launch template: upstream's new __OPINPUT__ brief
  encoding plus this fork's --dangerously-bypass-hook-trust flag, which the
  worktree vault hook needs to load at all.
- bin/fm-spawn.sh claude settings.local.json: upstream's semantic busy-state
  hook set (UserPromptSubmit/Stop/StopFailure/SessionEnd) plus this fork's
  PreToolUse vault guard. Upstream's Stop hook already carries the turn-end
  touch, so it supersedes the simpler one it replaced.
- bin/fm-spawn.sh opencode branch: upstream renamed its turn-end plugin to
  fm-busy-state.js; the exclude_path now names the renamed file and the vault
  plugin install is unchanged.
- bin/fm-spawn.sh pi extension: upstream's busyEvent helper plus this fork's
  tool_call vault handler, with spawn restored alongside execFile in the import
  the handler needs.
- bin/fm-spawn.sh codex branch: upstream's busy-state negotiation comment and
  this fork's worktree hooks.json vault install are independent; both kept.
- docs/scripts.md: upstream's fm-subagent-pretool-check.sh row and this fork's
  three vault rows are separate table additions; all four kept.
- docs/turnend-guard.md: upstream's rewritten intro, with vault-guard.md added
  to the related-guards list so this fork's cross-reference survives.

Test-harness adaptations to upstream's new contracts, no behavior change:

- tests/fm-vault-guard.test.sh passes the now-required --mode and --yolo on
  ship spawns, so the per-harness install assertions execute again.
- The same test asserts upstream's renamed fm-busy-state.js is still written.

Verification: bin/fm-lint.sh clean on pinned ShellCheck 0.11.0;
fm-test-run.sh --check-coverage ok at 125 scripts; the vault guard's 26 checks
and the scout contract's 5 both pass, matching their pre-merge baselines; and
the guard's live allow/deny behavior is unchanged, including the adversarial
glued-flag, case-variant, and run-child dump cases.
…olution

The merge kept upstream's busy-state paragraph and this fork's vault-install
comment side by side, and both stated the codex turn-end mechanism. Keep the
statement in upstream's paragraph only, so the two cannot drift.
Upstream's new busy-adapter test asserted the codex worktree .codex/hooks.json
is absent, using file absence as a proxy for 'no busy wiring installed'. This
fork's vault-guard seatbelt (docs/vault-guard.md) installs its PreToolUse hook
into that same file for every codex crewmate, so the proxy now forbids a
security control it was never aimed at.

Assert the intent directly: if the file exists it must contain no
fm-busy-event.sh wiring. Upstream's protection is unchanged - a hooks.json
carrying busy wiring still fails - and the sibling assertion that codex arms no
busy contract is untouched and still passes. fm-busy-lib.sh never reads this
file when classifying, so a vault-only hooks.json cannot make codex look
verified.
Upstream's documentation-audience inventory requires every tracked prose
surface to be classified exactly once, and this fork's two files predate that
rule, so the check reported them unclassified after the merge.

Classify them by their siblings' precedent: the scout contract as agent-runtime
like every other .agents/skills SKILL.md, and the vault guard as
maintainer-architecture like the sibling PreToolUse guard contracts
docs/arm-pretool-check.md, docs/cd-guard.md, and docs/subagent-guard.md.
@elixlabssolutions elixlabssolutions changed the title feat(bin): add remote secondmate orchestration and resilient supervision merge: bring the fork up to date with upstream (150 commits), preserving both fork commits Aug 6, 2026
@elixlabssolutions
elixlabssolutions merged commit 84de72c into main Aug 6, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Assignees

Couldn't load assignees.