Skip to content

feat(bin): sync upstream fleet runtime capabilities - #10

Merged
purple-phoenix merged 85 commits into
mainfrom
fm/fm-upstream-sync-u7
Jul 29, 2026
Merged

purple-phoenix merged 85 commits into
mainfrom
fm/fm-upstream-sync-u7

Conversation

@purple-phoenix

@purple-phoenix purple-phoenix commented Jul 28, 2026 •

Copy link
Copy Markdown
Owner

Intent

Captain's order: merge upstream kunchenguid/firstmate main (71 commits ahead, through fa0d85d) into this purple-phoenix fork, preserving the semantics of our six landed changes: #2 watcher hardening, #3 user-journey audits, #4 locale-invariant composer, #5 capacity dashboard, #6 dashboard redesign, #7 task-id contract. Hard constraints: MERGE, never rebase or force; never push to upstream (fetch only). Captain authorized merge-on-green for this sync PR.

Eleven files conflicted; each was reconciled semantically rather than by taking a side wholesale. Deliberate decisions a reviewer would not infer from the diff:

  • fm-watch-arm.sh: both sides independently fixed silent watcher arm-cycle death. Upstream's rewrite won structurally (cycle ledger, bounded successor wait, successor-following attach, typed failures) as the better maintenance baseline. Four things from our PR 2 were ported onto it because upstream lacks them: the durable .watcher-arm-dead marker on every FAILED path (fm-guard.sh reads its detail= field), the loud report plus marker when a signal ends supervision with tasks in flight, the child reap before any successor check, and the wake-sequence check that keeps a normal wake handoff from being reported as a death. PR 2's idle exemption was deliberately DROPPED in favor of upstream's unconditional loudness: that is the stronger reading of PR 2's own goal and closes a hole for X-only homes. Our two PR 2 tests were adapted to upstream's typed failure line while still asserting marker and non-zero exit.
  • Task-id minting converged with no code change: upstream's fm_task_id_creation_valid has the same name and contract as our existing quiet wrapper (path-safe, <= 64), so upstream's callers work unchanged against ours. Our captain-decision exemption in fm-decision-hold.sh (a kind=captain hold may exceed the length limit as non-dispatchable and never truncated, but must stay path-safe) is preserved with its tests.
  • fm-spawn.sh: kept upstream's per-task spawn lock plus our diagnostic-emitting validator.
  • fm-fleet-snapshot.sh: unioned upstream's multi-blocker resolution with our capacity evidence fields.
  • docs/supervision-protocols/claude.md and docs/herdr-backend.md: took upstream's rewrites; for herdr, PR 4's locale-invariance content was re-homed into the surviving sections rather than dropped.

Beyond the merge itself, four pre-existing upstream defects had to be fixed because they reproduce on a clean upstream/main tree and blocked a green suite. These are intentional additions, not scope creep:

  1. fm-brief.sh did not parse under stock macOS Bash 3.2 (apostrophe inside a heredoc nested in $( )), so no brief could be generated on macOS at all - upstream issue macOS bash 3.2: fm-brief.sh fails to parse at fa0d85d (unexpected EOF, heredoc-in-case); watcher cycles also exit 1 without reason since update kunchenguid/firstmate#1179. Fixed structurally by moving each Definition-of-done into a function body; upstream's wording restored verbatim.
  2. tests/fm-bearings-snapshot.test.sh used sed 'a' which drops the newline on BSD sed and corrupted its own fixture. Replaced with awk.
  3. fm-spawn.sh did not fail closed when task metadata could not be written - it printed "spawned ..." and exited 0, leaving a live worker nothing was tracking. Upstream's own orca test asserted this and was red.
  4. tests/fm-session-start.test.sh read $BASHPID, which does not exist in Bash 3.2, so under set -u every racer subshell aborted and the exactly-one-winner session-lock assertion always saw zero.

Two failures were genuinely caused by the merge and are fixed: our two task-id tests widened upstream's pinned ShellCheck source graph (switched to the non-following directive form), and our capacity and user-journey-audit skills were unclassified in upstream's new documentation-audience inventory (both classified agent-runtime).

One nuance to read carefully: two commits touch cleanup_child's TERM handling. The failure that prompted them was NOT a code defect - it was an artifact of running the suite under nohup, which sets SIGHUP to ignored and passes that to every descendant, so the test's kill -HUP was discarded. What survives on merit is the latency change: an intermediate commit introduced an unbounded blocking wait measured at 5-6s against an 8s budget, and the bounded 1s grace brings it to 1-2s. Treat it as a latency improvement, not a bug fix.

Verification: full suite 103/103 files with 3 failures, all environment-only and all skipping in CI (Pi 0.81.1+ vs local 0.80.10; Python 3.11+ for tomllib vs local 3.9.6; a harness-ancestry test that resolves the real Claude this runs under). fm-lint.sh clean. Every bin script parses under stock macOS Bash 3.2.

Mid-flight, PR #8 (persistent tailnet-only dashboard service) merged into origin/main. It was integrated as a second MERGE, never a rebase, because rebasing would flatten the upstream merge commit whose two parents are the record of this sync. That merge had three conflicts: .gitignore and AGENTS.md were additive on both sides, and bin/fm-fleet-snapshot.sh had #8 and upstream independently growing multi-blocker support. Both field families are emitted because both have live consumers (bin/fm-capacity.mjs reads #8's blocked_by_all, bin/fm-bearings-snapshot.sh reads upstream's blocked_by_ids/unresolved_blocker_ids). The one genuine contract collision was the scalar blocked_by on parsed backlog records: upstream captured only the last blocker while #8 joined all of them with commas. The join won because fm-capacity.mjs splits that scalar on commas, so it is the contract with a real consumer, whereas upstream's single-value form appeared only in test expectations and its information survives in blocked_by_ids; three upstream assertions were updated accordingly and still pin the structured fields. Captain-hold selection keeps upstream's stricter captain_actionable predicate so a blocked captain hold stays queued, with #8's origin and decision-key extraction layered on. docs/dashboard-service.md from #8 was classified maintainer-architecture in the documentation-audience inventory.

The PR body must be pipeline-generated. An earlier round of this work replaced the body by hand, which deleted the deterministic ## Pipeline section and its no-mistakes signature, failing the repository's PR-provenance check. That was corrected by moving the detailed conflict write-up into a PR comment and letting the pipeline own the body, per the guidance in CONTRIBUTING.md that the signature must never be pasted in by hand. Two sibling PRs, #8 and #9, merged into origin/main while this sync was in flight; both were integrated as merges rather than rebases for the same reason the upstream sync itself is a merge.

What Changed

  • Sync upstream fleet runtime capabilities, including Kimi and pi-signed adapters, quota-aware dispatch, Pi Calm mode, Herdr presentation spaces, session-start nudges, and GitLab merge watching.
  • Harden watcher, spawn, session, and cleanup lifecycles while preserving fork contracts for failure markers, task IDs, capacity evidence, and locale-invariant composer detection.
  • Add canonical timed and sharded test tooling, isolation coverage, runtime verification, and documentation-audience checks.

Risk Assessment

⚠️ Medium: No material defects were substantiated, but the large multi-merge change spans critical supervision, backend, dispatch, and snapshot paths, leaving meaningful integration risk despite careful reconciliation.

Testing

Completed 1 recorded test check.

  • Outcome: ⚠️ 1 error across 1 run (22m27s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

⚠️ **Test** - 1 error
  • 🚨 tests failed with exit code 1
  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits July 17, 2026 14:31
* fix(pi): distinguish stale locks when arming watcher

* no-mistakes(test): Stabilize watcher extension async waits

* no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate as house vocabulary

* no-mistakes(test): Update captain vocabulary contract test

* no-mistakes(document): Align secondmate documentation vocabulary
…uid#686)

* fix: parse secondmate home after pre-field parentheses

Registry summaries often include parentheticals before the structured
(home: ...) field. Match that field with a greedy prefix so handoff
no longer reports "has no home" for those entries.

* no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges

* no-mistakes(document): Document nudge script inventory
…nguid#688)

Relabel absent-captain and related domain defaults wording so it names
the firstmate repo rather than treating "template" as this domain's
identity label. Keep the design-tenet "shared template" statements and
unrelated launch/PR-poll template uses unchanged.
…nsion (kunchenguid#205)

* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn

Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(brief): harden fm-brief regression coverage for parse and scaffolds

Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the

Co-authored-by: Cursor <cursoragent@cursor.com>
kunchenguid#166 apostrophe regression cannot return unnoticed.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#693)

* fix: make watcher supervision continuous

* no-mistakes(review): Bound watcher retries and log attached signals

* no-mistakes(review): Add bounded successor-recovery wake fallbacks

* no-mistakes(review): Prevent overlapping successor-arm retries

* no-mistakes(review): Resume supervision after late arm closes

* no-mistakes(review): Bind OpenCode recovery to attempted arm

* no-mistakes(test): Synchronize peer beacon regression fixture

* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures

* no-mistakes(document): Captain: document watcher successor protocol behavior

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix: always fetch PR head for review diffs

Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.

* no-mistakes(test): Isolate session-start nudge tests from gate state

* no-mistakes(document): Correct review-diff documentation
…and skills (kunchenguid#736)

* docs: resolve five contract contradictions

* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): reverify grok exit command

* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for exited paused crew

* no-mistakes(review): Gate pause suppression on confirmed agent death

* no-mistakes(test): Fixed stale pause cadence

* no-mistakes(document): Document dead-agent hold cadence
…id#744)

* fix(supervision): distinguish ordinary wakes from repair

* no-mistakes(review): Make passive guard follow-ups recovery-only

* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes

* no-mistakes(review): fix x-poll claim error deduplication

* no-mistakes(review): separate claim diagnostics from relay recovery

* no-mistakes(document): Document X-mode once-only mention wakes
…enguid#747)

* feat(wake): enrich drained signal context

* no-mistakes(review): Bound wake enrichment reads

* no-mistakes(document): Document wake-drain annotations

* docs(wake): explain at-least-once drain boundary

* no-mistakes(review): Prevent symlink races in wake annotations

* no-mistakes(review): Exercise wake symlink race regression

* test: document intentional AFK marker subprocesses

* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces

* no-mistakes(review): Harden Herdr projection creation and spawn serialization

* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission

* no-mistakes(test): Correct stale Orca metadata failure fixture

* no-mistakes(document): Document Herdr presentation projection accurately
…nchenguid#775)

* fix(send): treat opencode busy-queued composer state as submitted

When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row.  The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.

Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy.  If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted).  On an idle pane, keep returning "pending"
(genuine swallow detection preserved).

Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pane + composer clears on first Enter -> empty (existing path)

* docs: document busy-queued Enter exception across backend docs and skills

Add explanatory comments and backend documentation for the
busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn
but keeps typed text in composer until the turn ends):

- bin/fm-tmux-lib.sh: document the busy-aware fallback in the
  file header and above fm_tmux_submit_enter_core
- .agents/skills/afk/SKILL.md: daemon-facing policy note
- .agents/skills/harness-adapters/SKILL.md: harness-specific fact
- docs/tmux-backend.md: submit-acknowledgement section with the
  busy-queue exception
- docs/herdr-backend.md: record the known gap
- docs/architecture.md: cross-reference in the daemon section

* test(tmux): fix SC2181 and make busy-submit test executable
…unchenguid#765)

* fix(spawn): require two stable reads before accepting worktree path

The treehouse-get worktree-detection loop in fm-spawn.sh accepted the
first pane_current_path read that differed from the project path, but
on some tmux/WSL setups a brand-new window transiently reports a
stale-but-real path before the pane actually settles into the
worktree. Since that stale path is itself a real, distinct git
checkout, it also passes validate_spawn_worktree's isolation check,
so the loop silently recorded the wrong worktree in state/<id>.meta
(and, for claude harness spawns, installed the turn-end hook there
too).

Require two consecutive polls to agree on the same non-project path
before accepting it, using the existing inter-poll sleep as the
confirmation gap so an already-settled pane isn't slowed down by an
extra cycle.

* fix(tests): drop unused CASE_DIR read in worktree-settle test

ShellCheck SC2034: CASE_DIR is split out of the case record but
never referenced; discard it with _ instead.

---------

Co-authored-by: Freudator86 <tim@allesknut.de>
…anges (kunchenguid#752)

* fix(watcher): stabilize Linux process identity

* no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state

Stop free-text tokens like "merged" from promoting nonterminal working: lines
to captain-relevant, so AFK no longer permanently suppresses idle recovery.
Defend wedge aging independently for nonterminal progress verbs, bind
no-mistakes current-state attribution to code identity (not branch alone),
and mark setup-complete as nonterminal in the ship brief scaffold.

* no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution

* no-mistakes(document): Document current-code-bound run attribution

* no-mistakes(test): Wait for stable Herdr shell readiness

* no-mistakes(test): Make Herdr and watcher readiness tests deterministic

* no-mistakes(test): Make tmux capture and watcher lifecycle deterministic

* no-mistakes(document): Document corrected supervision contracts
* fix: allow safe teardown during watcher recovery

* no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code

* no-mistakes(document): Sync continuity-gate docs to allow teardown recovery

* test: mark dynamic teardown fixture literal

* no-mistakes(document): docs: add teardown to continuity gate allow list
…nguid#790)

* feat(herdr): order presentation worker spaces

* fix(herdr): preserve focus during projected cleanup

* no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs

* no-mistakes(review): Serialize Herdr aborts with guarded focus regressions

* no-mistakes(review): Fall back flat when Herdr serialization is unavailable

* no-mistakes(test): Stabilize watcher startup and AFK handoff tests

* no-mistakes(document): Correct Herdr ordering and focus documentation
…#809)

* Send literal config reread after inherited config push

When declared inherited config changes under an already-running secondmate,
build a per-home instruction from validated destination post-write bytes and
deliver it on the routed secondmate path. Unchanged config sends nothing;
ABSENT represents removal; captain-shared is never inlined. Covers mid-session
config-push and the locked bootstrap convergence path without hardening spawn
against deliberate runtime choice.

* no-mistakes(review): Fix config reread framing, partial propagation, and respawn order

* no-mistakes(review): Send config rereads via durable single-line pointers

* no-mistakes(review): Make failed config rereads retryable

* no-mistakes(review): Make config reread retries generation-safe

* no-mistakes(review): Make config rereads durable and ordered

* no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode

* no-mistakes(review): Retain write retries and quarantine stale respawn generations

* no-mistakes(review): Preserve exact config reread retries and delivery order

* no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning

* no-mistakes(document): Consolidated config-reread documentation
* feat(watch): follow GitLab merge requests to merge

The merge watch only understood GitHub pull requests, so a task whose
deliverable is a GitLab merge request was never followed to merge.

Generalize the stored poll identity from owner/repository to a
provider-tagged provider/url/host/path/number record. GitLab runs mostly on
self-hosted instances and its projects nest under groups at no fixed depth,
so the host and the full project path are data in the record rather than
constants, and every consumer rebuilds the URL from those parts and refuses
any record that does not reconstruct it exactly.

The GitLab state is read with plain glab, matching the GitHub path's use of
plain gh, so an upstream checkout needs no extra tooling. Two things about
glab were established by running it rather than assumed, because a wrong
invocation here fails silently into a permanent "not merged":

- glab has no field selector, and its JSON would need a JSON processor that
  firstmate does not require, so the state is read from glab's own field
  output. Only an exact "merged" wakes firstmate, so a changed format stays
  silent instead of reporting a merge.
- glab cannot take a merge request URL the way gh can, because that form
  resolves through the current git repository and the watcher has none. It
  is addressed by project URL and merge request number instead.

An absent glab produces no wake rather than a false merge, and arming
refuses with a clear message since that is the one point where a missing
CLI can still be reported. A GitLab task records no pr_head, which both
consumers already treat as optional. The merge path still addresses GitHub
only and refuses a merge request URL rather than sending it to the wrong
forge.

The record version moves to v2, and the existing non-executing migration
rebuilds an already-armed watch from its recorded URL, so no watch is lost
by upgrading.

docs/gitlab-merge-watch.md records the evidence, taken against the public
fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture.

* no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation

* no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs
…uid#821)

* feat(herdr): correct all-home child presentation topology

Inherit the presentation opt-in to secondmate homes, label new projected
spaces with the approved corner format, insert each child under its owning
parent under one session-scoped lock, and keep flat non-destructive fallback.

* no-mistakes(review): Exclude secondmates from Herdr presentation projection

* no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering

* no-mistakes(review): Use adjacency-only Herdr child ownership

* no-mistakes(review): Reject foreign legacy projections safely

* no-mistakes(review): Validate Herdr session sockets before projection

* no-mistakes(test): Fix Herdr teardown fixture session socket metadata

* fix(herdr): canonicalize presentation lock socket paths

Always resolve the session socket parent directory so symlink parents
such as /tmp -> /private/tmp cannot split the shared cross-home lock
identity. Refuse relative socket paths. Clarify lock-unavailable warnings.

* no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing

* no-mistakes(document): Document all-home Herdr child topology

* no-mistakes(lint): Quote fallback provenance string for ShellCheck
* fix(no-mistakes): drop full-suite local Test override

Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad
tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a
focused contract test so the override cannot silently return.

* no-mistakes(lint): Make CI contract assertion ShellCheck-clean
* feat(test): add canonical timed suite runner and honest CI timeout

Introduce bin/fm-test-run.sh as the single serial owner for selecting
one script, a family, a conservative changed-file set, or the explicit
complete suite, with per-script timing markers and a JSON artifact.
Wire CI Behavior through the runner, raise the hang-tripwire timeout to
25 minutes, and document entry points without restoring a full-suite
local no-mistakes Test command.

* no-mistakes(review): Captain: fix changed selection and empty summaries

* no-mistakes(review): Captain: fail closed on unmapped changed sources

* no-mistakes(document): Document canonical timed test entry points
* fix: disclose main-home orphan and unstructured inventory gaps

Main Bearings could report an empty fleet while structured in-flight rows
lacked meta or current backlog rows were free-form. Emit main_inventory from
the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next
gate, and keep meta as the only live Underway source.

* no-mistakes(document): Document Bearings inventory-integrity projection

* no-mistakes: apply CI fixes
* feat: add concurrent test isolation proof for Phase 2

Prove an audited portable candidate set passes under concurrent
workers with private mode-0700 temp roots, without enabling
production CI sharding or fm-test-run --jobs.

* no-mistakes(review): Pin isolation proof to audited candidate manifest
* feat(secondmate): parent-owned guards for missed status reports

Marked parent-to-secondmate requests now create a durable pending-reply
expectation with a privacy-safe correlation id before delivery. Transport
success never resolves it; only a correlated parent status or document
pointer does. After a completed turn with no report, the parent sends one
recovery repost and escalates once if that turn is also missed, without
scraping the secondmate conversation or looping.

* no-mistakes(review): Deduplicate wrong-home pending-reply sightings

* no-mistakes(review): Harden pending-reply recovery and escalation guards

* no-mistakes(review): Bound pending-reply backend polling

* no-mistakes(review): Cache pending-reply status scans

* no-mistakes(review): Protect undelivered pending-reply records from scans

* no-mistakes(review): Close pending-reply delivery durability gaps

* no-mistakes(review): Separate pending-reply transport outcomes

* no-mistakes(review): Escalate stalled pending-reply deliveries once

* no-mistakes(review): Resolve attempted deliveries from correlated reports

* no-mistakes(review): Resolve late reports after delivery escalation

* no-mistakes(document): Document pending-reply grace and ownership

* no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings
* feat: add required pinned Herdr CI lane

Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and
SHA-256 pins, run the real-herdr-gated family serially through
fm-test-run with hard-fail on herdr-not-found, and keep portable
Behavior free of claimed Herdr coverage.

* no-mistakes(document): Consolidate real-Herdr CI documentation ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
kunchenguid and others added 26 commits July 24, 2026 14:45
…d#997)

* feat(claude): Stop-owned tokenless watcher continuity via asyncRewake auto-arm

Claude primaries (main home and marked secondmate homes) no longer depend
on the model remembering to re-arm the watcher after each wake. A tracked
Stop asyncRewake hook (bin/fm-claude-stop-autoarm.sh, timeout 28800s)
fires on every turn end, claims one home-scoped single-flight owner,
foregrounds bin/fm-watch-arm.sh inside the hook-owned process tree, and
translates an actionable close or typed watcher failure into exactly one
exit-2 rewake. The hook scopes to genuine primary checkouts, requires the
session lock to be held by its own harness ancestor, stays inert while
AFK owns triage or the home is idle, and hands AFK transitions mid-cycle
to the daemon without rewaking.

The synchronous turn-end guard gains a --claude cooperative mode: it
ignores stop_hook_active (true on every post-continuation stop, which is
what re-opened the 2026-07-21 blind window), waits briefly for a watcher
health proof, a live auto-arm owner claim, or a fresh rewake epoch, and
re-blocks only when the auto-arm genuinely failed to establish - bounded
to 3 consecutive blocks per session, safely below Claude Code's 8-block
override, then a degraded allow with a visible systemMessage. Codex
keeps the previous one-block loop guard byte-identically, and Pi,
OpenCode, and Grok adapters are untouched.

Continuity PreToolUse gate and durable wake queue are preserved; the
gate's recovery guidance now names the Stop-owned re-arm and reserves
manual background arms for auto-arm failure. Claude supervision protocol,
harness-adapters facts, architecture, configuration, and continuity docs
updated; docs/turnend-guard.md records the 2026-07-24 Claude 2.1.218
contract revalidation (tokenless multi-cycle rewake, no-dedup, timeout
process-group kill, 8-block cap, interactive non-stall) and the 2.1.219
product live E2Es.

Regression matrix: hermetic tests cover scope, identity, AFK, need,
single-flight, translation, guard cooperation, budget, and registration;
the new live E2E proves two full tokenless auto-arm rewake cycles with
zero model arm commands; Pi and OpenCode Option B live E2Es pass
unchanged.

* no-mistakes(review): Fix Claude X-mode auto-arm continuity backstop

* no-mistakes(review): Remove unsupported Claude contract-lab verification claims

* no-mistakes(document): Update Claude auto-arm continuity documentation
* Clean stale Herdr projections at session start

* no-mistakes(document): Document stale Herdr session-start projection cleanup

* no-mistakes(review): Enforce locked exact Herdr projection cleanup

* no-mistakes(review): Fail closed on unverified session lock ownership

* no-mistakes(review): Serialize session lock acquisition atomically

* no-mistakes(document): Align session-start and Herdr cleanup documentation

* no-mistakes(document): Generalize lock-refusal diagnostics

* no-mistakes(lint): Avoid reserved keyword in concurrency test

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…uid#1001)

* fix: recover Claude supervision at session start

* fix: remove Claude watcher-status command gate

* no-mistakes(document): docs: remove stale continuity gate references
* Replace quota dispatch selector instructions

* no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection
* remove vestigial dispatch selector

* no-mistakes(review): Synchronize isolation proof and portable shard evidence

* no-mistakes(review): Correct shard history and proof archive date

* no-mistakes(review): Remove reintroduced selector documentation reference

* no-mistakes(document): Remove stale dispatch strategy documentation
…1039)

quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per
provider, so the successor named in the interim rule has landed and the
rule's own removal condition is satisfied.

Keep the ownership clause so quota-axi remains the single owner of how
model or product windows relate to bounding account windows, and drop
the interim weakest-headroom instruction. The unknown-semantics case is
already covered by the existing requirement to stop and report a
candidate whose applicable quota data or interpretation cannot be
established.

Drop the matching assertion phrase from
tests/fm-instruction-owners.test.sh; the retained ownership phrase still
asserts.
…unchenguid#1049)

* fix(tmux): scope Claude busy detection by harness

* no-mistakes(review): Separate verified and fallback busy signatures

* no-mistakes(test): Scope busy signatures to supplied harnesses

* no-mistakes(document): Document harness-scoped busy detection
* Add verified Kimi crewmate harness adapter

* no-mistakes(review): Scope Kimi moon detection to spinner lines

* no-mistakes(review): Match only complete Kimi spinner rows

* no-mistakes(review): Resolve Kimi binary portably before pane creation

* no-mistakes(document): Align Kimi adapter documentation

* no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override

* Fix Kimi busy spinner detection

* no-mistakes(review): Recognize Kimi session-lock ancestry and holders

* no-mistakes(review): Scope pending-reply Kimi busy detection by harness

* no-mistakes(document): Correct Kimi spinner capture documentation

* no-mistakes(document): Clarify optional Kimi spinner whitespace

* no-mistakes(lint): Silence intentional pending-reply test stub warnings

* test: align rebased Kimi busy fixtures

* no-mistakes: apply CI fixes

* Reconcile Kimi busy detection after per-harness scoping

* no-mistakes(review): Clarify observed Kimi spinner whitespace contract

* no-mistakes(document): Clarify Kimi harness documentation
* fix kimi pointer submission and spinner conformance

* no-mistakes(review): Preserve Kimi submit target ownership guard
* Add guarded Kimi turn-end hook

* no-mistakes(review): Require jq before installing Kimi turn-end hook

* no-mistakes(review): Expose jq inside isolated Kimi test fixtures

* no-mistakes(review): Preserve Kimi config boundaries during hook removal

* no-mistakes(review): Document Kimi removal newline safeguard

* no-mistakes(document): Document Kimi shared-home preservation
)

* Fix structural tmux composer reading

* Verify Calm compatibility with Pi 0.82

* no-mistakes(review): Harden structural composer classification boundaries

* no-mistakes(review): Refresh composer and Kimi regression fixtures

* no-mistakes(review): Fail closed on unbounded composer edges

* no-mistakes(review): Enforce aligned composer geometry safely

* no-mistakes(review): Make composer ambiguity locale-safe

* no-mistakes(review): Preserve ambiguity through composer submission

* no-mistakes(review): Carry composer proof through retries

* no-mistakes(document): Document structural tmux composer delivery guarantees

* no-mistakes: apply CI fixes
* feat: add verified pi-signed adapter

* no-mistakes(review): Correct pi-signed maintainer verification date

* no-mistakes(review): Correct remaining pi-signed verification dates

* no-mistakes(review): Preserve authoritative pi-signed runtime identity

* no-mistakes(document): Document pi-signed shared adapter semantics

* no-mistakes: apply CI fixes
* fix(pi): rearm watcher across same-process session transitions

Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement
as well as terminal quit. The primary watcher extension latched a module-level
stopping flag on every shutdown, so a replacement session in the same process
could not arm monitoring until Pi restarted.

Own arm authority per session generation so only the active live generation
may start, stop, or rearm the child. Replacement sessions can arm again without
restarting Pi, stale prior-generation callbacks cannot mutate the active cycle,
and real quit still blocks late rearm.

* no-mistakes(review): Preserve Pi generation isolation and exit cleanup

* no-mistakes(document): Correct Pi watcher transition documentation
* Consume quota-axi pace signals in dispatch profile array selection.

Add quota-array-dispatch as the single owner of the pace-aware candidate
choice, keep AGENTS.md to the intake boundary and load trigger, and cover
the acceptance cases with sanitized schemaVersion 3 fixtures.

* no-mistakes(review): Stop and report genuine quota dispatch ties

* no-mistakes(document): Document quota pace freshness and uncertainty
Sync 71 upstream commits while preserving the semantics of the six changes
landed here (#2 watcher hardening, #3 user-journey audits, #4 locale-invariant
composer, #5 capacity dashboard, #6 dashboard redesign, #7 task-id contract).

Conflict resolutions, by file:

- bin/fm-watch-arm.sh: took upstream's rewrite (kunchenguid#693, kunchenguid#997) as the baseline for
  its cycle-exit ledger, bounded successor wait, successor-following attach, and
  typed nonzero failures, then ported this fork's PR 2 semantics onto it: the
  durable state/.watcher-arm-dead marker on every FAILED path (bin/fm-guard.sh
  reads its detail= field), the loud report plus marker when a signal ends
  supervision while tasks are in flight, the hardened child reap before any
  successor check, and the wake-sequence check that keeps a normal wake handoff
  from being reported as an unexplained death. Dropped PR 2's idle exemption:
  upstream fails loudly whenever an attached cycle ends without a successor, and
  that is the stronger reading of PR 2's own goal that supervision gaps must not
  be silent.
- bin/fm-spawn.sh: kept upstream's per-task spawn lock and this fork's
  diagnostic-emitting fm_task_id_creation_check, which names the charset or
  length reason instead of a bare "invalid task id".
- bin/fm-fleet-snapshot.sh: unioned upstream's multi-blocker resolution
  (unresolved_blocker_ids, blocked_by_ids, hold_reason/hold_kind selection) with
  this fork's capacity evidence fields (repo, kind, since, delivery_mode,
  project_resolved) on both blocked-record projections.
- docs/supervision-protocols/claude.md: took upstream's Stop-hook-owned
  supervision rewrite wholesale; it already carries PR 2's requirement to treat
  a watcher-failure wake as an alarm.
- docs/herdr-backend.md: took upstream's guidance-versus-evidence split (kunchenguid#994)
  and re-homed PR 4's content, with the locale-invariance rule moved into
  "Composer and injection safety" and its regression evidence into
  docs/verification/runtime-backends.md.
- docs/architecture.md, docs/configuration.md: unioned both sides' additions and
  repointed the herdr references at the sections that survived the split.
- tests/fm-watcher-lock.test.sh, tests/fm-brief.test.sh: kept both sides' new
  tests. PR 2's attached-death test now asserts upstream's typed failure line
  with this fork's marker detail, and the unconfirmable-watcher test uses
  upstream's looser stdout assertion plus this fork's marker assertions.
- .github/workflows/ci.yml: bearings count is 43, the merged total, not either
  side's 37 or 42.

Two upstream defects blocked a green suite on macOS and are fixed here:

- bin/fm-brief.sh did not parse under stock macOS Bash 3.2, so no brief could be
  generated on macOS at all. Bash 3.2 tracks quote state through a heredoc body
  while scanning for the closing paren of a command substitution, so the single
  apostrophe upstream added inside a Definition-of-done heredoc broke the whole
  file. Reworded that line, extended tests/fm-brief.test.sh to check the stock
  interpreter as well as the PATH one, and added a CI step that parses every
  shipped script under macOS Bash 3.2.
- tests/fm-bearings-snapshot.test.sh built a fixture with sed's "a\" append,
  which drops the following newline on BSD sed and glued a section header onto
  the appended row. Replaced it with an equivalent awk append.
…sh 3.2

Move each delivery mode's Definition-of-done text into its own function body
instead of an inline DOD=$(cat <<EOF ... EOF) assignment.

Stock macOS Bash 3.2 tracks quote state through a heredoc body while it scans
for the closing paren of a command substitution, so a single apostrophe in that
body makes the entire script unparseable - fm-brief.sh could not generate any
brief on macOS. A heredoc in a plain function body is parsed normally, so this
removes the bug class rather than avoiding apostrophes in the prose, and the
upstream wording that tests/fm-ask-user-authority.test.sh pins is restored
verbatim.
The metadata write used a bare redirect, and fm-spawn.sh runs under set -u
without set -e, so a failed write was ignored: the script went on to print
'spawned <id> ...' and exit 0 with no state/<id>.meta on disk.

That durable record is what makes a worker supervisable - the watcher, recovery,
and teardown all locate a task through it - so a successful-looking spawn left a
live worker that nothing was tracking. Stop instead; the existing abort cleanup
still owns backend teardown.

tests/fm-backend-orca.test.sh already asserted this and was failing.
…fixture

Three merge-integration fixes and one inherited upstream defect.

Merge integration:
- Upstream's source-graph boundary (kunchenguid#939) pins which tests may carry production
  shellcheck context. This fork's two task-id tests sourced bin/fm-pr-lib.sh with
  a followed directive purely for validator access; both already disabled SC1091,
  so the directive only widened the graph. Switched to the non-following form and
  said why inline.
- Upstream's documentation-audience inventory did not classify this fork's
  capacity and user-journey-audit skills, which upstream never had. Both are
  agent-runtime, matching every other .agents/skills entry.

Inherited defect:
- tests/fm-session-start.test.sh read $BASHPID in each racer subshell. BASHPID
  arrived in Bash 4.0, so under stock macOS Bash 3.2 with set -u every racer
  aborted and the exactly-one-winner assertion always saw zero. The fallback
  execs sh in the substitution subshell so its PPID is that racer's own pid; $$
  would report the shared parent and defeat the race the test exists to prove.
The signal path polled for up to 2s before escalating to SIGKILL, on top of the
ledger-lock wait and the successor check. The HUP regression has an 8s budget, so
under a loaded parallel suite the arm could miss it and the test saw a timeout
instead of exit 129.

Block on the owned child instead: it returns the moment the child dies, and
SIGKILL stays as the backstop for a child that ignores TERM outright. This keeps
the property the poll was added for - the exact child is fully reaped before any
successor check, so a still-fresh dying child cannot read as a healthy successor
and suppress the arm-death alarm - while removing the fixed delay.

Verified with six concurrent runs of tests/fm-watcher-lock.test.sh, the load that
exposed it: 29/29 green in every run.
…m exit

Measured HUP-to-exit latency was 5-6s against the regression's 8s budget, because
the blocking wait sat through a whole FM_POLL interval: Bash defers the watcher's
TERM trap until its poll sleep returns.

Give TERM a 1s grace, then SIGKILL, then block to reap. The blocking wait still
guarantees the exact owned child is fully reaped before any successor check, so a
still-fresh dying child cannot read as a healthy successor and suppress the
arm-death alarm; the bound just stops a deferred trap from dominating shutdown.

Measured after the change: 1-2s across five runs, a 4-8x margin on the budget.

This replaces the previous commit's unbounded wait, which traded the original
fixed 2s poll for a stall of up to one poll interval - slower, not faster.
PR #8 landed on origin/main while this sync was in flight. Integrated as a
MERGE, not a rebase: the captain's order for this branch is merge-only, and
rebasing would flatten the upstream merge commit whose two parents are the
record of the sync.

Three conflicts:

- .gitignore, AGENTS.md: additive on both sides; kept both entries.
- bin/fm-fleet-snapshot.sh: #8 and upstream independently grew multi-blocker
  support. #8 added blocked_by_all plus captain-hold origin/key extraction from
  the hold body; upstream added blocked_by_ids with resolution tracking through
  unresolved_blocker_ids. Both have live consumers - bin/fm-capacity.mjs reads
  #8's fields, bin/fm-bearings-snapshot.sh reads upstream's - so both are
  emitted rather than one replacing the other.

The one genuine contract collision was the scalar blocked_by on parsed backlog
records. Upstream's parser captured only the last blocked-by field; #8 joined
every blocker with commas. bin/fm-capacity.mjs splits that scalar on commas, so
the join is the contract with a real consumer, while upstream's single-value
form was asserted only in tests and its information survives in blocked_by_ids.
Kept the join and updated the three upstream assertions, which still pin
blocked_by_ids and unresolved_blocker_ids alongside.

Captain-hold selection keeps upstream's stricter captain_actionable predicate,
so a captain hold that is itself blocked stays queued instead of surfacing as
actionable, and adds #8's origin and decision-key extraction on top.
PR #8 added docs/dashboard-service.md, which upstream's documentation-audience
inventory did not know about, so the merged tree failed its own completeness
check. The doc states that it owns the architecture narrative, trust design, and
verification evidence while pointing at other owners for mechanics, which is the
maintainer-architecture audience rather than operator-current.
@purple-phoenix

Copy link
Copy Markdown
Owner Author

Merge resolution detail

Posted as a comment rather than in the PR body: the body is owned by the no-mistakes pipeline, which writes the deterministic ## Pipeline section.

Merges 71 commits from kunchenguid/firstmate main (through fa0d85d) into this fork, preserving the semantics of the six changes landed here: #2 watcher hardening, #3 user-journey audits, #4 locale-invariant composer, #5 capacity dashboard, #6 dashboard redesign, #7 task-id contract.

A real merge commit with both parents. Nothing was rebased or force-pushed, and nothing was pushed to upstream.

Conflict resolutions

Eleven files conflicted. For each, which side's approach won and why:

bin/fm-watch-arm.sh - upstream baseline, our semantics ported on

Both sides independently fixed silent watcher arm-cycle death. Upstream's rewrite (kunchenguid#693, kunchenguid#997) is the better maintenance baseline and won structurally: the bounded lifecycle ledger in state/.watch-cycle-exits.log, the bounded wait_for_healthy_successor, an attached arm that follows identity-matched successors instead of exiting, and typed nonzero failures.

Four things from our PR 2 were ported onto it, because upstream has no equivalent:

Ported Why it had to survive
state/.watcher-arm-dead marker on every FAILED path bin/fm-guard.sh reads its detail= field to name a supervision gap on the next fleet action; upstream only signals on stdout, which is lost if nothing is reading it
Loud report + marker when a signal ends supervision with tasks in flight Upstream's signal handlers only write a ledger row
Hardened child reap (TERM, wait, escalate to KILL, wait) before any successor check A still-fresh dying child otherwise reads as a healthy successor and suppresses the alarm on Linux
Wake-sequence handoff check in attach_and_wait An attached arm cannot read the watcher's output, so without it every normal wake handoff observed by a second attached arm is reported as an unexplained death

One PR 2 behavior was deliberately dropped: the idle exemption that exited 0 when no state/<id>.meta existed. Upstream fails loudly whenever an attached cycle ends without a successor, regardless of fleet size. That is the stronger reading of PR 2's own stated goal - supervision gaps must not be silent - and it closes a hole PR 2 left open, since an X-only home still needs a live cycle with no fleet work. Our two PR 2 tests were adapted to upstream's typed failure line while still asserting the marker and non-zero exit.

bin/fm-spawn.sh - both

Kept upstream's new per-task spawn lock and our diagnostic-emitting fm_task_id_creation_check, which names the charset or length reason instead of upstream's bare error: invalid task id.

Task-id minting - converged, no code change needed

The briefed expectation was that upstream had its own implementation to reconcile against. It does: fm_task_id_creation_valid (path-safe plus <= 64). Our fm-pr-lib.sh already exposes a function of that exact name and contract as a quiet wrapper over our diagnostic fm_task_id_creation_check, and FM_TASK_ID_MAX_LENGTH is 64, so upstream's callers (fm-herdr-session-cleanup.sh, tests/fm-pr-check-security.test.sh) work unchanged against ours. Upstream's name and contract are the baseline; our richer diagnostics sit underneath. Our captain-decision exemption in bin/fm-decision-hold.sh - a kind=captain hold may exceed the length limit as a non-dispatchable, never-truncated id, but must still be path-safe - is preserved intact, along with tests/fm-task-add.test.sh and tests/fm-decision-hold-lifecycle.test.sh.

bin/fm-fleet-snapshot.sh - union

Upstream added multi-blocker resolution (unresolved_blocker_ids, blocked_by_ids, hold_reason/hold_kind selection); PR 5/6 added capacity evidence (repo, kind, since, delivery_mode, project_resolved). Both are on both blocked-record projections now. Verified end-to-end that upstream's captain_actionable routing and our delivery_mode/order projection both survive.

docs/supervision-protocols/claude.md - upstream

Upstream's Stop-hook-owned supervision rewrite replaces manual background-task arming and already carries PR 2's requirement to treat a watcher-failure wake as an alarm, so nothing needed porting.

docs/herdr-backend.md - upstream restructure, our content re-homed

Upstream's kunchenguid#994 split current guidance from verification evidence, deleting ~650 lines of incident chronology in favor of a pointer. Taking that wholesale would have dropped PR 4's locale-invariance documentation, so it was re-homed: the behavior rule into "Composer and injection safety", its regression evidence into docs/verification/runtime-backends.md.

docs/architecture.md, docs/configuration.md - union

Both sides' additions kept; the herdr references in configuration.md were repointed at the sections that survived the split, so no reference dangles.

tests/fm-watcher-lock.test.sh, tests/fm-brief.test.sh, tests/fm-cd-pretool-check.test.sh - union

Both sides' new tests kept. Two of our watcher assertions were adapted: the attached-death test now expects upstream's typed failure line with our marker detail, and the unconfirmable-watcher test uses upstream's looser stdout assertion (upstream loosened it deliberately, because their refactor changed which FAILED line appears) plus our marker assertions.

.github/workflows/ci.yml

Bearings test count is 43 - the merged total (36 base + 1 ours + 6 upstream) - not either side's 37 or 42. Verified by running the suite, not by arithmetic alone.

Four upstream defects fixed here

The merged result could not go green without these. All four are pre-existing on upstream main and reproduce on a clean upstream/main tree; none was introduced by this merge. The first matches upstream issue kunchenguid#1179, which reports exactly this breakage at fa0d85d. Two of the four are the same root cause - code that assumes Bash 4+ while the repo ships a stock-macOS-Bash-3.2 CI lane.

1. bin/fm-brief.sh did not parse under stock macOS Bash 3.2 - no brief could be generated on macOS at all.

Bash 3.2 tracks quote state through a heredoc body while scanning for the closing paren of a command substitution, so a single apostrophe inside DOD=$(cat <<EOF ... EOF) makes the whole file unparseable. Upstream added firstmate's to one Definition-of-done body.

The repo already had a regression guard for this exact class (test_script_parses, from issue kunchenguid#166), but it ran bash -n against the PATH bash - Bash 4+ on most machines - so the regression sailed past it.

Fixed structurally rather than by avoiding apostrophes: each mode's Definition-of-done text moved into its own function body, where a heredoc parses normally. Upstream's wording is restored verbatim, so tests/fm-ask-user-authority.test.sh still passes. Also extended test_script_parses to check the stock interpreter when it differs from the PATH one, and added a CI step that parses every shipped script under macOS Bash 3.2.

2. tests/fm-bearings-snapshot.test.sh built a fixture with BSD-incompatible sed.

sed '/## In flight/a\' drops the newline after the appended text on BSD sed, gluing the ## Queued header onto the appended row. The corrupted backlog made the SSHHIP captain-hold assertion fail. Replaced with an equivalent awk append.

3. bin/fm-spawn.sh did not fail closed when it could not record task metadata.

The write used a bare redirect, and the script runs under set -u without set -e, so a failed write was ignored: it went on to print spawned <id> ... and exit 0 with no state/<id>.meta on disk. That record is what makes a worker supervisable - the watcher, recovery, and teardown all locate a task through it - so a successful-looking spawn could leave a live worker nothing was tracking. It now stops with a non-zero exit. tests/fm-backend-orca.test.sh already asserted this and was failing on upstream.

Endpoint teardown on that path remains owned only by the Orca and Herdr projection cleanup; see the pipeline-rounds section below for why per-backend rollback was deliberately left out of this change.

4. tests/fm-session-start.test.sh could never pass under stock macOS Bash 3.2.

Each of its 40 racer subshells read $BASHPID, which arrived in Bash 4.0. Under set -u that aborts every racer, so the exactly-one-winner session-lock assertion always saw zero winners. The fallback execs sh in the substitution subshell so its PPID is that racer's own pid; $$ would report the shared parent and defeat the race the test exists to prove.

Known environment prerequisite (not a code change)

tests/fm-calm-pi-extension.test.sh pins upstream's new Pi Calm work to Pi 0.81.1/0.82.0. This machine has 0.80.10, so two cases fail locally. They skip in CI, which does not install the Pi package, so this does not affect the PR's checks - but running Pi Calm mode from this sync needs a local Pi upgrade to 0.81.1 or newer. Left as-is rather than loosening upstream's deliberate compatibility guard.

Second merge: origin/main (#8) landed mid-flight

PR #8 (persistent tailnet-only dashboard service) merged into main while this sync was in flight. The validation pipeline's rebase step wanted to rebase onto it; that would have flattened the upstream merge commit whose two parents are the record of this sync, so it was integrated as a second merge instead. Re-running the pipeline afterwards, the rebase step passed as a no-op, confirming the integration.

Three conflicts. .gitignore and AGENTS.md were additive on both sides. bin/fm-fleet-snapshot.sh was substantive: #8 and upstream had independently grown multi-blocker support.

Side Added
#8 blocked_by_all, plus captain-hold origin / decision-key extraction from the hold body
upstream blocked_by_ids with resolution tracking via unresolved_blocker_ids

Both are emitted, because both have live consumers: bin/fm-capacity.mjs reads #8's fields, bin/fm-bearings-snapshot.sh reads upstream's.

The one genuine contract collision was the scalar blocked_by on parsed backlog records - upstream captured only the last blocker, #8 joined all of them with commas, and the merged test file asserted both against the same fixture shape. The join won on evidence rather than preference: fm-capacity.mjs:717 splits that scalar on commas, so it is the contract with real production code behind it, while upstream's single-value form appeared only in test expectations and its information survives intact in blocked_by_ids. Three upstream assertions were updated and still pin the structured fields alongside.

Captain-hold selection keeps upstream's stricter captain_actionable predicate, so a captain hold that is itself blocked stays queued rather than surfacing as actionable, with #8's origin and decision-key extraction layered on top. docs/dashboard-service.md from #8 was classified maintainer-architecture in the audience inventory, which the merged tree otherwise failed its own completeness check on.

A note on the watcher shutdown-latency commit

Two commits here touch cleanup_child's TERM handling, and the second one's message overstates its case. The failure that prompted them - tests/fm-watcher-lock.test.sh's HUP regression timing out - was not a code defect. It was an artifact of running the suite under nohup, which sets SIGHUP to ignored and passes that disposition to every descendant, so the test's kill -HUP was silently discarded and the arm never exited. Instrumenting the signal handler showed it receiving only TERM, never a single HUP; the same test through the same runner in the foreground is 29/29 green.

What survives on its own merit is the latency change. The intermediate commit replaced the original bounded poll with an unbounded blocking wait, which sat through a whole FM_POLL interval because Bash defers the watcher's TERM trap until its sleep returns - measured at 5-6s against the regression's 8s budget, via a direct probe unaffected by the nohup issue. The bounded 1s grace brings that to 1-2s. So the net effect against the merge baseline is a real improvement in shutdown latency and margin, but it should be read as that, not as a bug fix.

Verification

  • Full suite via bin/fm-test-run.sh --all: 103/103 files, 3 failures, all environment-only (see below). Run without nohup, so the HUP-dependent watcher tests genuinely exercise their signal.
  • bin/fm-lint.sh: clean (pinned ShellCheck 0.11.0).
  • Every bin/*.sh and bin/backends/*.sh parses under stock macOS Bash 3.2.
  • Every skill named in AGENTS.md section 13 resolves to a real SKILL.md; both sides' skills coexist.
  • No files lost from either side; upstream's deletion of the vestigial dispatch selector (fix(bin): remove vestigial dispatch selector kunchenguid/firstmate#1026) applied cleanly with no dangling references.

The three remaining failures are local toolchain shortfalls, not code, and all skip in CI:

Test Needs
fm-calm-pi-extension Pi 0.81.1+ (this machine: 0.80.10)
fm-kimi-harness Python 3.11+ for tomllib (this machine: 3.9.6)
fm-secondmate-harness No claude process ancestry; resolves the real host harness

Validation-pipeline rounds

The review step raised one finding against this branch's own spawn fix: the comment claimed the abort cleanup owns backend teardown, when spawn_abort_cleanup only removes Orca resources and Herdr presentation projections. It suggested tracking and rolling back endpoint creation for every backend plus atomic metadata publication.

That was scoped deliberately. The comment correction was taken. Per-backend endpoint rollback across tmux, flat Herdr, Zellij, and cmux was declined as a rider on an upstream sync - the residual gap predates this branch, and before the fail-closed fix the same failure reported a successful spawn and exited 0, so an untracked endpoint was strictly more likely, not less. It deserves its own task and tests.

Atomic publication was attempted and then backed out, because it broke two upstream structural guards that pin the expected source shape of bin/fm-spawn.sh:

  • tests/fm-gotmp.test.sh requires the literal echo "tasktmp=$TASK_TMP".
  • tests/fm-backend-herdr.test.sh requires the literal } > "$STATE/$ID.meta" to appear after the lock acquisition and after the backend case, proving the lock spans backend creation through metadata publication.

A temp-file-and-rename necessarily changes both. Rewriting upstream's invariant guards to accommodate optional hardening is a poor trade inside a large sync, and the fail-closed fix alone was already green across the full suite. The atomicity idea can be taken up separately, updating those guards deliberately.

A third failure in that round, tests/fm-backend-tmux-smoke.test.sh, was a 10-second live-tmux readiness wait failing under pipeline load; it passes 7/7 on an idle machine and was left alone. The two environment failures were explicitly excluded from the fix instructions so the pipeline could not "resolve" them by weakening upstream's Pi and Python version guards.

Second sibling PR to land on main while this sync was in flight. Integrated as a
MERGE for the same reason as #8: rebasing would flatten the upstream merge commit
whose two parents are the record of this sync.

No conflicts - #9 touches the capacity surface and its own new
bin/fm-wait-progress.mjs, which this branch does not modify. Verified the
consumers that the earlier blocker-model union touched are still green:
fm-capacity 24/24, fm-fleet-snapshot-view 15/15, fm-documentation-audiences 5/5,
and the new script is already classified in the toolbelt and audience
inventories that #9 updated.
# Conflicts:
#	bin/fm-fleet-snapshot.sh
@purple-phoenix
purple-phoenix merged commit c826482 into main Jul 29, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.