feat(bin): sync upstream fleet runtime capabilities - #10
Conversation
* fix(pi): distinguish stale locks when arming watcher * no-mistakes(test): Stabilize watcher extension async waits * no-mistakes(document): Document Pi lock recovery
* fix: accept secondmate as house vocabulary * no-mistakes(test): Update captain vocabulary contract test * no-mistakes(document): Align secondmate documentation vocabulary
…uid#686) * fix: parse secondmate home after pre-field parentheses Registry summaries often include parentheticals before the structured (home: ...) field. Match that field with a greedy prefix so handoff no longer reports "has no home" for those entries. * no-mistakes(document): Refresh handoff test comments
* feat: add native session-start nudges * no-mistakes(document): Document nudge script inventory
…nguid#688) Relabel absent-captain and related domain defaults wording so it names the firstmate repo rather than treating "template" as this domain's identity label. Keep the design-tenet "shared template" statements and unrelated launch/PR-poll template uses unchanged.
…nsion (kunchenguid#205) * fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4 (notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing. Co-authored-by: Cursor <cursoragent@cursor.com> * test(brief): harden fm-brief regression coverage for parse and scaffolds Tighten bash -n checking, pin literal backtick rendering in the no-mistakes DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the Co-authored-by: Cursor <cursoragent@cursor.com> kunchenguid#166 apostrophe regression cannot return unnoticed. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#693) * fix: make watcher supervision continuous * no-mistakes(review): Bound watcher retries and log attached signals * no-mistakes(review): Add bounded successor-recovery wake fallbacks * no-mistakes(review): Prevent overlapping successor-arm retries * no-mistakes(review): Resume supervision after late arm closes * no-mistakes(review): Bind OpenCode recovery to attempted arm * no-mistakes(test): Synchronize peer beacon regression fixture * no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures * no-mistakes(document): Captain: document watcher successor protocol behavior * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* fix: always fetch PR head for review diffs Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded pr_head= so reviewers never hold a merge over a "missing" fix that already landed on the remote PR. Recorded SHA is offline fallback only; local branch is last resort with a warning. Store the tip under refs/fm-review/ so a later base-branch fetch cannot clobber the compare tip via FETCH_HEAD. * no-mistakes(test): Isolate session-start nudge tests from gate state * no-mistakes(document): Correct review-diff documentation
…and skills (kunchenguid#736) * docs: resolve five contract contradictions * no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): reverify grok exit command * no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for exited paused crew * no-mistakes(review): Gate pause suppression on confirmed agent death * no-mistakes(test): Fixed stale pause cadence * no-mistakes(document): Document dead-agent hold cadence
…id#744) * fix(supervision): distinguish ordinary wakes from repair * no-mistakes(review): Make passive guard follow-ups recovery-only * no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes * no-mistakes(review): fix x-poll claim error deduplication * no-mistakes(review): separate claim diagnostics from relay recovery * no-mistakes(document): Document X-mode once-only mention wakes
…enguid#747) * feat(wake): enrich drained signal context * no-mistakes(review): Bound wake enrichment reads * no-mistakes(document): Document wake-drain annotations * docs(wake): explain at-least-once drain boundary * no-mistakes(review): Prevent symlink races in wake annotations * no-mistakes(review): Exercise wake symlink race regression * test: document intentional AFK marker subprocesses * fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces * no-mistakes(review): Harden Herdr projection creation and spawn serialization * no-mistakes(review): Captain, disarm Herdr cleanup before launch submission * no-mistakes(test): Correct stale Orca metadata failure fixture * no-mistakes(document): Document Herdr presentation projection accurately
…nchenguid#775) * fix(send): treat opencode busy-queued composer state as submitted When fm-send sends a message to a BUSY opencode crewmate on the tmux backend, opencode accepts the Enter and queues the message for the next turn, but leaves the typed text visible in the composer row. The submit-verification loop sees a pending composer, exhausts retries, and reports a false "Enter swallowed" failure while the message is actually delivered. Fix: after Enter retries are exhausted and the composer still shows pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn, footer shows "esc interrupt"), the harness queued the message, so return "empty" (accepted). On an idle pane, keep returning "pending" (genuine swallow detection preserved). Regression tests cover four scenarios: - busy pane + pending composer -> empty (message queued) - idle pane + pending composer -> pending (genuine swallow) - busy pane + composer clears on first Enter -> empty - idle pane + composer clears on first Enter -> empty (existing path) * docs: document busy-queued Enter exception across backend docs and skills Add explanatory comments and backend documentation for the busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn but keeps typed text in composer until the turn ends): - bin/fm-tmux-lib.sh: document the busy-aware fallback in the file header and above fm_tmux_submit_enter_core - .agents/skills/afk/SKILL.md: daemon-facing policy note - .agents/skills/harness-adapters/SKILL.md: harness-specific fact - docs/tmux-backend.md: submit-acknowledgement section with the busy-queue exception - docs/herdr-backend.md: record the known gap - docs/architecture.md: cross-reference in the daemon section * test(tmux): fix SC2181 and make busy-submit test executable
…unchenguid#765) * fix(spawn): require two stable reads before accepting worktree path The treehouse-get worktree-detection loop in fm-spawn.sh accepted the first pane_current_path read that differed from the project path, but on some tmux/WSL setups a brand-new window transiently reports a stale-but-real path before the pane actually settles into the worktree. Since that stale path is itself a real, distinct git checkout, it also passes validate_spawn_worktree's isolation check, so the loop silently recorded the wrong worktree in state/<id>.meta (and, for claude harness spawns, installed the turn-end hook there too). Require two consecutive polls to agree on the same non-project path before accepting it, using the existing inter-poll sleep as the confirmation gap so an already-settled pane isn't slowed down by an extra cycle. * fix(tests): drop unused CASE_DIR read in worktree-settle test ShellCheck SC2034: CASE_DIR is split out of the case record but never referenced; discard it with _ instead. --------- Co-authored-by: Freudator86 <tim@allesknut.de>
…anges (kunchenguid#752) * fix(watcher): stabilize Linux process identity * no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state Stop free-text tokens like "merged" from promoting nonterminal working: lines to captain-relevant, so AFK no longer permanently suppresses idle recovery. Defend wedge aging independently for nonterminal progress verbs, bind no-mistakes current-state attribution to code identity (not branch alone), and mark setup-complete as nonterminal in the ship brief scaffold. * no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution * no-mistakes(document): Document current-code-bound run attribution * no-mistakes(test): Wait for stable Herdr shell readiness * no-mistakes(test): Make Herdr and watcher readiness tests deterministic * no-mistakes(test): Make tmux capture and watcher lifecycle deterministic * no-mistakes(document): Document corrected supervision contracts
* fix: allow safe teardown during watcher recovery * no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code * no-mistakes(document): Sync continuity-gate docs to allow teardown recovery * test: mark dynamic teardown fixture literal * no-mistakes(document): docs: add teardown to continuity gate allow list
…nguid#790) * feat(herdr): order presentation worker spaces * fix(herdr): preserve focus during projected cleanup * no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs * no-mistakes(review): Serialize Herdr aborts with guarded focus regressions * no-mistakes(review): Fall back flat when Herdr serialization is unavailable * no-mistakes(test): Stabilize watcher startup and AFK handoff tests * no-mistakes(document): Correct Herdr ordering and focus documentation
…#809) * Send literal config reread after inherited config push When declared inherited config changes under an already-running secondmate, build a per-home instruction from validated destination post-write bytes and deliver it on the routed secondmate path. Unchanged config sends nothing; ABSENT represents removal; captain-shared is never inlined. Covers mid-session config-push and the locked bootstrap convergence path without hardening spawn against deliberate runtime choice. * no-mistakes(review): Fix config reread framing, partial propagation, and respawn order * no-mistakes(review): Send config rereads via durable single-line pointers * no-mistakes(review): Make failed config rereads retryable * no-mistakes(review): Make config reread retries generation-safe * no-mistakes(review): Make config rereads durable and ordered * no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode * no-mistakes(review): Retain write retries and quarantine stale respawn generations * no-mistakes(review): Preserve exact config reread retries and delivery order * no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning * no-mistakes(document): Consolidated config-reread documentation
* feat(watch): follow GitLab merge requests to merge The merge watch only understood GitHub pull requests, so a task whose deliverable is a GitLab merge request was never followed to merge. Generalize the stored poll identity from owner/repository to a provider-tagged provider/url/host/path/number record. GitLab runs mostly on self-hosted instances and its projects nest under groups at no fixed depth, so the host and the full project path are data in the record rather than constants, and every consumer rebuilds the URL from those parts and refuses any record that does not reconstruct it exactly. The GitLab state is read with plain glab, matching the GitHub path's use of plain gh, so an upstream checkout needs no extra tooling. Two things about glab were established by running it rather than assumed, because a wrong invocation here fails silently into a permanent "not merged": - glab has no field selector, and its JSON would need a JSON processor that firstmate does not require, so the state is read from glab's own field output. Only an exact "merged" wakes firstmate, so a changed format stays silent instead of reporting a merge. - glab cannot take a merge request URL the way gh can, because that form resolves through the current git repository and the watcher has none. It is addressed by project URL and merge request number instead. An absent glab produces no wake rather than a false merge, and arming refuses with a clear message since that is the one point where a missing CLI can still be reported. A GitLab task records no pr_head, which both consumers already treat as optional. The merge path still addresses GitHub only and refuses a merge request URL rather than sending it to the wrong forge. The record version moves to v2, and the existing non-executing migration rebuilds an already-armed watch from its recorded URL, so no watch is lost by upgrading. docs/gitlab-merge-watch.md records the evidence, taken against the public fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture. * no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation * no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs
…uid#821) * feat(herdr): correct all-home child presentation topology Inherit the presentation opt-in to secondmate homes, label new projected spaces with the approved corner format, insert each child under its owning parent under one session-scoped lock, and keep flat non-destructive fallback. * no-mistakes(review): Exclude secondmates from Herdr presentation projection * no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering * no-mistakes(review): Use adjacency-only Herdr child ownership * no-mistakes(review): Reject foreign legacy projections safely * no-mistakes(review): Validate Herdr session sockets before projection * no-mistakes(test): Fix Herdr teardown fixture session socket metadata * fix(herdr): canonicalize presentation lock socket paths Always resolve the session socket parent directory so symlink parents such as /tmp -> /private/tmp cannot split the shared cross-home lock identity. Refuse relative socket paths. Clarify lock-unavailable warnings. * no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing * no-mistakes(document): Document all-home Herdr child topology * no-mistakes(lint): Quote fallback provenance string for ShellCheck
* fix(no-mistakes): drop full-suite local Test override Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a focused contract test so the override cannot silently return. * no-mistakes(lint): Make CI contract assertion ShellCheck-clean
* feat(test): add canonical timed suite runner and honest CI timeout Introduce bin/fm-test-run.sh as the single serial owner for selecting one script, a family, a conservative changed-file set, or the explicit complete suite, with per-script timing markers and a JSON artifact. Wire CI Behavior through the runner, raise the hang-tripwire timeout to 25 minutes, and document entry points without restoring a full-suite local no-mistakes Test command. * no-mistakes(review): Captain: fix changed selection and empty summaries * no-mistakes(review): Captain: fail closed on unmapped changed sources * no-mistakes(document): Document canonical timed test entry points
* fix: disclose main-home orphan and unstructured inventory gaps Main Bearings could report an empty fleet while structured in-flight rows lacked meta or current backlog rows were free-form. Emit main_inventory from the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next gate, and keep meta as the only live Underway source. * no-mistakes(document): Document Bearings inventory-integrity projection * no-mistakes: apply CI fixes
* feat: add concurrent test isolation proof for Phase 2 Prove an audited portable candidate set passes under concurrent workers with private mode-0700 temp roots, without enabling production CI sharding or fm-test-run --jobs. * no-mistakes(review): Pin isolation proof to audited candidate manifest
* feat(secondmate): parent-owned guards for missed status reports Marked parent-to-secondmate requests now create a durable pending-reply expectation with a privacy-safe correlation id before delivery. Transport success never resolves it; only a correlated parent status or document pointer does. After a completed turn with no report, the parent sends one recovery repost and escalates once if that turn is also missed, without scraping the secondmate conversation or looping. * no-mistakes(review): Deduplicate wrong-home pending-reply sightings * no-mistakes(review): Harden pending-reply recovery and escalation guards * no-mistakes(review): Bound pending-reply backend polling * no-mistakes(review): Cache pending-reply status scans * no-mistakes(review): Protect undelivered pending-reply records from scans * no-mistakes(review): Close pending-reply delivery durability gaps * no-mistakes(review): Separate pending-reply transport outcomes * no-mistakes(review): Escalate stalled pending-reply deliveries once * no-mistakes(review): Resolve attempted deliveries from correlated reports * no-mistakes(review): Resolve late reports after delivery escalation * no-mistakes(document): Document pending-reply grace and ownership * no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings
* feat: add required pinned Herdr CI lane Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and SHA-256 pins, run the real-herdr-gated family serially through fm-test-run with hard-fail on herdr-not-found, and keep portable Behavior free of claimed Herdr coverage. * no-mistakes(document): Consolidate real-Herdr CI documentation ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…d#997) * feat(claude): Stop-owned tokenless watcher continuity via asyncRewake auto-arm Claude primaries (main home and marked secondmate homes) no longer depend on the model remembering to re-arm the watcher after each wake. A tracked Stop asyncRewake hook (bin/fm-claude-stop-autoarm.sh, timeout 28800s) fires on every turn end, claims one home-scoped single-flight owner, foregrounds bin/fm-watch-arm.sh inside the hook-owned process tree, and translates an actionable close or typed watcher failure into exactly one exit-2 rewake. The hook scopes to genuine primary checkouts, requires the session lock to be held by its own harness ancestor, stays inert while AFK owns triage or the home is idle, and hands AFK transitions mid-cycle to the daemon without rewaking. The synchronous turn-end guard gains a --claude cooperative mode: it ignores stop_hook_active (true on every post-continuation stop, which is what re-opened the 2026-07-21 blind window), waits briefly for a watcher health proof, a live auto-arm owner claim, or a fresh rewake epoch, and re-blocks only when the auto-arm genuinely failed to establish - bounded to 3 consecutive blocks per session, safely below Claude Code's 8-block override, then a degraded allow with a visible systemMessage. Codex keeps the previous one-block loop guard byte-identically, and Pi, OpenCode, and Grok adapters are untouched. Continuity PreToolUse gate and durable wake queue are preserved; the gate's recovery guidance now names the Stop-owned re-arm and reserves manual background arms for auto-arm failure. Claude supervision protocol, harness-adapters facts, architecture, configuration, and continuity docs updated; docs/turnend-guard.md records the 2026-07-24 Claude 2.1.218 contract revalidation (tokenless multi-cycle rewake, no-dedup, timeout process-group kill, 8-block cap, interactive non-stall) and the 2.1.219 product live E2Es. Regression matrix: hermetic tests cover scope, identity, AFK, need, single-flight, translation, guard cooperation, budget, and registration; the new live E2E proves two full tokenless auto-arm rewake cycles with zero model arm commands; Pi and OpenCode Option B live E2Es pass unchanged. * no-mistakes(review): Fix Claude X-mode auto-arm continuity backstop * no-mistakes(review): Remove unsupported Claude contract-lab verification claims * no-mistakes(document): Update Claude auto-arm continuity documentation
* Clean stale Herdr projections at session start * no-mistakes(document): Document stale Herdr session-start projection cleanup * no-mistakes(review): Enforce locked exact Herdr projection cleanup * no-mistakes(review): Fail closed on unverified session lock ownership * no-mistakes(review): Serialize session lock acquisition atomically * no-mistakes(document): Align session-start and Herdr cleanup documentation * no-mistakes(document): Generalize lock-refusal diagnostics * no-mistakes(lint): Avoid reserved keyword in concurrency test * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…uid#1001) * fix: recover Claude supervision at session start * fix: remove Claude watcher-status command gate * no-mistakes(document): docs: remove stale continuity gate references
* Replace quota dispatch selector instructions * no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection
* remove vestigial dispatch selector * no-mistakes(review): Synchronize isolation proof and portable shard evidence * no-mistakes(review): Correct shard history and proof archive date * no-mistakes(review): Remove reintroduced selector documentation reference * no-mistakes(document): Remove stale dispatch strategy documentation
…1039) quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per provider, so the successor named in the interim rule has landed and the rule's own removal condition is satisfied. Keep the ownership clause so quota-axi remains the single owner of how model or product windows relate to bounding account windows, and drop the interim weakest-headroom instruction. The unknown-semantics case is already covered by the existing requirement to stop and report a candidate whose applicable quota data or interpretation cannot be established. Drop the matching assertion phrase from tests/fm-instruction-owners.test.sh; the retained ownership phrase still asserts.
…unchenguid#1049) * fix(tmux): scope Claude busy detection by harness * no-mistakes(review): Separate verified and fallback busy signatures * no-mistakes(test): Scope busy signatures to supplied harnesses * no-mistakes(document): Document harness-scoped busy detection
* Add verified Kimi crewmate harness adapter * no-mistakes(review): Scope Kimi moon detection to spinner lines * no-mistakes(review): Match only complete Kimi spinner rows * no-mistakes(review): Resolve Kimi binary portably before pane creation * no-mistakes(document): Align Kimi adapter documentation * no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override * Fix Kimi busy spinner detection * no-mistakes(review): Recognize Kimi session-lock ancestry and holders * no-mistakes(review): Scope pending-reply Kimi busy detection by harness * no-mistakes(document): Correct Kimi spinner capture documentation * no-mistakes(document): Clarify optional Kimi spinner whitespace * no-mistakes(lint): Silence intentional pending-reply test stub warnings * test: align rebased Kimi busy fixtures * no-mistakes: apply CI fixes * Reconcile Kimi busy detection after per-harness scoping * no-mistakes(review): Clarify observed Kimi spinner whitespace contract * no-mistakes(document): Clarify Kimi harness documentation
* fix kimi pointer submission and spinner conformance * no-mistakes(review): Preserve Kimi submit target ownership guard
* Add guarded Kimi turn-end hook * no-mistakes(review): Require jq before installing Kimi turn-end hook * no-mistakes(review): Expose jq inside isolated Kimi test fixtures * no-mistakes(review): Preserve Kimi config boundaries during hook removal * no-mistakes(review): Document Kimi removal newline safeguard * no-mistakes(document): Document Kimi shared-home preservation
) * Fix structural tmux composer reading * Verify Calm compatibility with Pi 0.82 * no-mistakes(review): Harden structural composer classification boundaries * no-mistakes(review): Refresh composer and Kimi regression fixtures * no-mistakes(review): Fail closed on unbounded composer edges * no-mistakes(review): Enforce aligned composer geometry safely * no-mistakes(review): Make composer ambiguity locale-safe * no-mistakes(review): Preserve ambiguity through composer submission * no-mistakes(review): Carry composer proof through retries * no-mistakes(document): Document structural tmux composer delivery guarantees * no-mistakes: apply CI fixes
* feat: add verified pi-signed adapter * no-mistakes(review): Correct pi-signed maintainer verification date * no-mistakes(review): Correct remaining pi-signed verification dates * no-mistakes(review): Preserve authoritative pi-signed runtime identity * no-mistakes(document): Document pi-signed shared adapter semantics * no-mistakes: apply CI fixes
* fix(pi): rearm watcher across same-process session transitions Pi emits session_shutdown for ordinary /new, /resume, and /fork replacement as well as terminal quit. The primary watcher extension latched a module-level stopping flag on every shutdown, so a replacement session in the same process could not arm monitoring until Pi restarted. Own arm authority per session generation so only the active live generation may start, stop, or rearm the child. Replacement sessions can arm again without restarting Pi, stale prior-generation callbacks cannot mutate the active cycle, and real quit still blocks late rearm. * no-mistakes(review): Preserve Pi generation isolation and exit cleanup * no-mistakes(document): Correct Pi watcher transition documentation
* Consume quota-axi pace signals in dispatch profile array selection. Add quota-array-dispatch as the single owner of the pace-aware candidate choice, keep AGENTS.md to the intake boundary and load trigger, and cover the acceptance cases with sanitized schemaVersion 3 fixtures. * no-mistakes(review): Stop and report genuine quota dispatch ties * no-mistakes(document): Document quota pace freshness and uncertainty
Sync 71 upstream commits while preserving the semantics of the six changes landed here (#2 watcher hardening, #3 user-journey audits, #4 locale-invariant composer, #5 capacity dashboard, #6 dashboard redesign, #7 task-id contract). Conflict resolutions, by file: - bin/fm-watch-arm.sh: took upstream's rewrite (kunchenguid#693, kunchenguid#997) as the baseline for its cycle-exit ledger, bounded successor wait, successor-following attach, and typed nonzero failures, then ported this fork's PR 2 semantics onto it: the durable state/.watcher-arm-dead marker on every FAILED path (bin/fm-guard.sh reads its detail= field), the loud report plus marker when a signal ends supervision while tasks are in flight, the hardened child reap before any successor check, and the wake-sequence check that keeps a normal wake handoff from being reported as an unexplained death. Dropped PR 2's idle exemption: upstream fails loudly whenever an attached cycle ends without a successor, and that is the stronger reading of PR 2's own goal that supervision gaps must not be silent. - bin/fm-spawn.sh: kept upstream's per-task spawn lock and this fork's diagnostic-emitting fm_task_id_creation_check, which names the charset or length reason instead of a bare "invalid task id". - bin/fm-fleet-snapshot.sh: unioned upstream's multi-blocker resolution (unresolved_blocker_ids, blocked_by_ids, hold_reason/hold_kind selection) with this fork's capacity evidence fields (repo, kind, since, delivery_mode, project_resolved) on both blocked-record projections. - docs/supervision-protocols/claude.md: took upstream's Stop-hook-owned supervision rewrite wholesale; it already carries PR 2's requirement to treat a watcher-failure wake as an alarm. - docs/herdr-backend.md: took upstream's guidance-versus-evidence split (kunchenguid#994) and re-homed PR 4's content, with the locale-invariance rule moved into "Composer and injection safety" and its regression evidence into docs/verification/runtime-backends.md. - docs/architecture.md, docs/configuration.md: unioned both sides' additions and repointed the herdr references at the sections that survived the split. - tests/fm-watcher-lock.test.sh, tests/fm-brief.test.sh: kept both sides' new tests. PR 2's attached-death test now asserts upstream's typed failure line with this fork's marker detail, and the unconfirmable-watcher test uses upstream's looser stdout assertion plus this fork's marker assertions. - .github/workflows/ci.yml: bearings count is 43, the merged total, not either side's 37 or 42. Two upstream defects blocked a green suite on macOS and are fixed here: - bin/fm-brief.sh did not parse under stock macOS Bash 3.2, so no brief could be generated on macOS at all. Bash 3.2 tracks quote state through a heredoc body while scanning for the closing paren of a command substitution, so the single apostrophe upstream added inside a Definition-of-done heredoc broke the whole file. Reworded that line, extended tests/fm-brief.test.sh to check the stock interpreter as well as the PATH one, and added a CI step that parses every shipped script under macOS Bash 3.2. - tests/fm-bearings-snapshot.test.sh built a fixture with sed's "a\" append, which drops the following newline on BSD sed and glued a section header onto the appended row. Replaced it with an equivalent awk append.
…sh 3.2 Move each delivery mode's Definition-of-done text into its own function body instead of an inline DOD=$(cat <<EOF ... EOF) assignment. Stock macOS Bash 3.2 tracks quote state through a heredoc body while it scans for the closing paren of a command substitution, so a single apostrophe in that body makes the entire script unparseable - fm-brief.sh could not generate any brief on macOS. A heredoc in a plain function body is parsed normally, so this removes the bug class rather than avoiding apostrophes in the prose, and the upstream wording that tests/fm-ask-user-authority.test.sh pins is restored verbatim.
The metadata write used a bare redirect, and fm-spawn.sh runs under set -u without set -e, so a failed write was ignored: the script went on to print 'spawned <id> ...' and exit 0 with no state/<id>.meta on disk. That durable record is what makes a worker supervisable - the watcher, recovery, and teardown all locate a task through it - so a successful-looking spawn left a live worker that nothing was tracking. Stop instead; the existing abort cleanup still owns backend teardown. tests/fm-backend-orca.test.sh already asserted this and was failing.
…fixture Three merge-integration fixes and one inherited upstream defect. Merge integration: - Upstream's source-graph boundary (kunchenguid#939) pins which tests may carry production shellcheck context. This fork's two task-id tests sourced bin/fm-pr-lib.sh with a followed directive purely for validator access; both already disabled SC1091, so the directive only widened the graph. Switched to the non-following form and said why inline. - Upstream's documentation-audience inventory did not classify this fork's capacity and user-journey-audit skills, which upstream never had. Both are agent-runtime, matching every other .agents/skills entry. Inherited defect: - tests/fm-session-start.test.sh read $BASHPID in each racer subshell. BASHPID arrived in Bash 4.0, so under stock macOS Bash 3.2 with set -u every racer aborted and the exactly-one-winner assertion always saw zero. The fallback execs sh in the substitution subshell so its PPID is that racer's own pid; $$ would report the shared parent and defeat the race the test exists to prove.
The signal path polled for up to 2s before escalating to SIGKILL, on top of the ledger-lock wait and the successor check. The HUP regression has an 8s budget, so under a loaded parallel suite the arm could miss it and the test saw a timeout instead of exit 129. Block on the owned child instead: it returns the moment the child dies, and SIGKILL stays as the backstop for a child that ignores TERM outright. This keeps the property the poll was added for - the exact child is fully reaped before any successor check, so a still-fresh dying child cannot read as a healthy successor and suppress the arm-death alarm - while removing the fixed delay. Verified with six concurrent runs of tests/fm-watcher-lock.test.sh, the load that exposed it: 29/29 green in every run.
…m exit Measured HUP-to-exit latency was 5-6s against the regression's 8s budget, because the blocking wait sat through a whole FM_POLL interval: Bash defers the watcher's TERM trap until its poll sleep returns. Give TERM a 1s grace, then SIGKILL, then block to reap. The blocking wait still guarantees the exact owned child is fully reaped before any successor check, so a still-fresh dying child cannot read as a healthy successor and suppress the arm-death alarm; the bound just stops a deferred trap from dominating shutdown. Measured after the change: 1-2s across five runs, a 4-8x margin on the budget. This replaces the previous commit's unbounded wait, which traded the original fixed 2s poll for a stall of up to one poll interval - slower, not faster.
PR #8 landed on origin/main while this sync was in flight. Integrated as a MERGE, not a rebase: the captain's order for this branch is merge-only, and rebasing would flatten the upstream merge commit whose two parents are the record of the sync. Three conflicts: - .gitignore, AGENTS.md: additive on both sides; kept both entries. - bin/fm-fleet-snapshot.sh: #8 and upstream independently grew multi-blocker support. #8 added blocked_by_all plus captain-hold origin/key extraction from the hold body; upstream added blocked_by_ids with resolution tracking through unresolved_blocker_ids. Both have live consumers - bin/fm-capacity.mjs reads #8's fields, bin/fm-bearings-snapshot.sh reads upstream's - so both are emitted rather than one replacing the other. The one genuine contract collision was the scalar blocked_by on parsed backlog records. Upstream's parser captured only the last blocked-by field; #8 joined every blocker with commas. bin/fm-capacity.mjs splits that scalar on commas, so the join is the contract with a real consumer, while upstream's single-value form was asserted only in tests and its information survives in blocked_by_ids. Kept the join and updated the three upstream assertions, which still pin blocked_by_ids and unresolved_blocker_ids alongside. Captain-hold selection keeps upstream's stricter captain_actionable predicate, so a captain hold that is itself blocked stays queued instead of surfacing as actionable, and adds #8's origin and decision-key extraction on top.
PR #8 added docs/dashboard-service.md, which upstream's documentation-audience inventory did not know about, so the merged tree failed its own completeness check. The doc states that it owns the architecture narrative, trust design, and verification evidence while pointing at other owners for mechanics, which is the maintainer-architecture audience rather than operator-current.
Merge resolution detailPosted as a comment rather than in the PR body: the body is owned by the no-mistakes pipeline, which writes the deterministic Merges 71 commits from A real merge commit with both parents. Nothing was rebased or force-pushed, and nothing was pushed to upstream. Conflict resolutionsEleven files conflicted. For each, which side's approach won and why:
|
| Ported | Why it had to survive |
|---|---|
state/.watcher-arm-dead marker on every FAILED path |
bin/fm-guard.sh reads its detail= field to name a supervision gap on the next fleet action; upstream only signals on stdout, which is lost if nothing is reading it |
| Loud report + marker when a signal ends supervision with tasks in flight | Upstream's signal handlers only write a ledger row |
Hardened child reap (TERM, wait, escalate to KILL, wait) before any successor check |
A still-fresh dying child otherwise reads as a healthy successor and suppresses the alarm on Linux |
Wake-sequence handoff check in attach_and_wait |
An attached arm cannot read the watcher's output, so without it every normal wake handoff observed by a second attached arm is reported as an unexplained death |
One PR 2 behavior was deliberately dropped: the idle exemption that exited 0 when no state/<id>.meta existed. Upstream fails loudly whenever an attached cycle ends without a successor, regardless of fleet size. That is the stronger reading of PR 2's own stated goal - supervision gaps must not be silent - and it closes a hole PR 2 left open, since an X-only home still needs a live cycle with no fleet work. Our two PR 2 tests were adapted to upstream's typed failure line while still asserting the marker and non-zero exit.
bin/fm-spawn.sh - both
Kept upstream's new per-task spawn lock and our diagnostic-emitting fm_task_id_creation_check, which names the charset or length reason instead of upstream's bare error: invalid task id.
Task-id minting - converged, no code change needed
The briefed expectation was that upstream had its own implementation to reconcile against. It does: fm_task_id_creation_valid (path-safe plus <= 64). Our fm-pr-lib.sh already exposes a function of that exact name and contract as a quiet wrapper over our diagnostic fm_task_id_creation_check, and FM_TASK_ID_MAX_LENGTH is 64, so upstream's callers (fm-herdr-session-cleanup.sh, tests/fm-pr-check-security.test.sh) work unchanged against ours. Upstream's name and contract are the baseline; our richer diagnostics sit underneath. Our captain-decision exemption in bin/fm-decision-hold.sh - a kind=captain hold may exceed the length limit as a non-dispatchable, never-truncated id, but must still be path-safe - is preserved intact, along with tests/fm-task-add.test.sh and tests/fm-decision-hold-lifecycle.test.sh.
bin/fm-fleet-snapshot.sh - union
Upstream added multi-blocker resolution (unresolved_blocker_ids, blocked_by_ids, hold_reason/hold_kind selection); PR 5/6 added capacity evidence (repo, kind, since, delivery_mode, project_resolved). Both are on both blocked-record projections now. Verified end-to-end that upstream's captain_actionable routing and our delivery_mode/order projection both survive.
docs/supervision-protocols/claude.md - upstream
Upstream's Stop-hook-owned supervision rewrite replaces manual background-task arming and already carries PR 2's requirement to treat a watcher-failure wake as an alarm, so nothing needed porting.
docs/herdr-backend.md - upstream restructure, our content re-homed
Upstream's kunchenguid#994 split current guidance from verification evidence, deleting ~650 lines of incident chronology in favor of a pointer. Taking that wholesale would have dropped PR 4's locale-invariance documentation, so it was re-homed: the behavior rule into "Composer and injection safety", its regression evidence into docs/verification/runtime-backends.md.
docs/architecture.md, docs/configuration.md - union
Both sides' additions kept; the herdr references in configuration.md were repointed at the sections that survived the split, so no reference dangles.
tests/fm-watcher-lock.test.sh, tests/fm-brief.test.sh, tests/fm-cd-pretool-check.test.sh - union
Both sides' new tests kept. Two of our watcher assertions were adapted: the attached-death test now expects upstream's typed failure line with our marker detail, and the unconfirmable-watcher test uses upstream's looser stdout assertion (upstream loosened it deliberately, because their refactor changed which FAILED line appears) plus our marker assertions.
.github/workflows/ci.yml
Bearings test count is 43 - the merged total (36 base + 1 ours + 6 upstream) - not either side's 37 or 42. Verified by running the suite, not by arithmetic alone.
Four upstream defects fixed here
The merged result could not go green without these. All four are pre-existing on upstream main and reproduce on a clean upstream/main tree; none was introduced by this merge. The first matches upstream issue kunchenguid#1179, which reports exactly this breakage at fa0d85d. Two of the four are the same root cause - code that assumes Bash 4+ while the repo ships a stock-macOS-Bash-3.2 CI lane.
1. bin/fm-brief.sh did not parse under stock macOS Bash 3.2 - no brief could be generated on macOS at all.
Bash 3.2 tracks quote state through a heredoc body while scanning for the closing paren of a command substitution, so a single apostrophe inside DOD=$(cat <<EOF ... EOF) makes the whole file unparseable. Upstream added firstmate's to one Definition-of-done body.
The repo already had a regression guard for this exact class (test_script_parses, from issue kunchenguid#166), but it ran bash -n against the PATH bash - Bash 4+ on most machines - so the regression sailed past it.
Fixed structurally rather than by avoiding apostrophes: each mode's Definition-of-done text moved into its own function body, where a heredoc parses normally. Upstream's wording is restored verbatim, so tests/fm-ask-user-authority.test.sh still passes. Also extended test_script_parses to check the stock interpreter when it differs from the PATH one, and added a CI step that parses every shipped script under macOS Bash 3.2.
2. tests/fm-bearings-snapshot.test.sh built a fixture with BSD-incompatible sed.
sed '/## In flight/a\' drops the newline after the appended text on BSD sed, gluing the ## Queued header onto the appended row. The corrupted backlog made the SSHHIP captain-hold assertion fail. Replaced with an equivalent awk append.
3. bin/fm-spawn.sh did not fail closed when it could not record task metadata.
The write used a bare redirect, and the script runs under set -u without set -e, so a failed write was ignored: it went on to print spawned <id> ... and exit 0 with no state/<id>.meta on disk. That record is what makes a worker supervisable - the watcher, recovery, and teardown all locate a task through it - so a successful-looking spawn could leave a live worker nothing was tracking. It now stops with a non-zero exit. tests/fm-backend-orca.test.sh already asserted this and was failing on upstream.
Endpoint teardown on that path remains owned only by the Orca and Herdr projection cleanup; see the pipeline-rounds section below for why per-backend rollback was deliberately left out of this change.
4. tests/fm-session-start.test.sh could never pass under stock macOS Bash 3.2.
Each of its 40 racer subshells read $BASHPID, which arrived in Bash 4.0. Under set -u that aborts every racer, so the exactly-one-winner session-lock assertion always saw zero winners. The fallback execs sh in the substitution subshell so its PPID is that racer's own pid; $$ would report the shared parent and defeat the race the test exists to prove.
Known environment prerequisite (not a code change)
tests/fm-calm-pi-extension.test.sh pins upstream's new Pi Calm work to Pi 0.81.1/0.82.0. This machine has 0.80.10, so two cases fail locally. They skip in CI, which does not install the Pi package, so this does not affect the PR's checks - but running Pi Calm mode from this sync needs a local Pi upgrade to 0.81.1 or newer. Left as-is rather than loosening upstream's deliberate compatibility guard.
Second merge: origin/main (#8) landed mid-flight
PR #8 (persistent tailnet-only dashboard service) merged into main while this sync was in flight. The validation pipeline's rebase step wanted to rebase onto it; that would have flattened the upstream merge commit whose two parents are the record of this sync, so it was integrated as a second merge instead. Re-running the pipeline afterwards, the rebase step passed as a no-op, confirming the integration.
Three conflicts. .gitignore and AGENTS.md were additive on both sides. bin/fm-fleet-snapshot.sh was substantive: #8 and upstream had independently grown multi-blocker support.
| Side | Added |
|---|---|
| #8 | blocked_by_all, plus captain-hold origin / decision-key extraction from the hold body |
| upstream | blocked_by_ids with resolution tracking via unresolved_blocker_ids |
Both are emitted, because both have live consumers: bin/fm-capacity.mjs reads #8's fields, bin/fm-bearings-snapshot.sh reads upstream's.
The one genuine contract collision was the scalar blocked_by on parsed backlog records - upstream captured only the last blocker, #8 joined all of them with commas, and the merged test file asserted both against the same fixture shape. The join won on evidence rather than preference: fm-capacity.mjs:717 splits that scalar on commas, so it is the contract with real production code behind it, while upstream's single-value form appeared only in test expectations and its information survives intact in blocked_by_ids. Three upstream assertions were updated and still pin the structured fields alongside.
Captain-hold selection keeps upstream's stricter captain_actionable predicate, so a captain hold that is itself blocked stays queued rather than surfacing as actionable, with #8's origin and decision-key extraction layered on top. docs/dashboard-service.md from #8 was classified maintainer-architecture in the audience inventory, which the merged tree otherwise failed its own completeness check on.
A note on the watcher shutdown-latency commit
Two commits here touch cleanup_child's TERM handling, and the second one's message overstates its case. The failure that prompted them - tests/fm-watcher-lock.test.sh's HUP regression timing out - was not a code defect. It was an artifact of running the suite under nohup, which sets SIGHUP to ignored and passes that disposition to every descendant, so the test's kill -HUP was silently discarded and the arm never exited. Instrumenting the signal handler showed it receiving only TERM, never a single HUP; the same test through the same runner in the foreground is 29/29 green.
What survives on its own merit is the latency change. The intermediate commit replaced the original bounded poll with an unbounded blocking wait, which sat through a whole FM_POLL interval because Bash defers the watcher's TERM trap until its sleep returns - measured at 5-6s against the regression's 8s budget, via a direct probe unaffected by the nohup issue. The bounded 1s grace brings that to 1-2s. So the net effect against the merge baseline is a real improvement in shutdown latency and margin, but it should be read as that, not as a bug fix.
Verification
- Full suite via
bin/fm-test-run.sh --all: 103/103 files, 3 failures, all environment-only (see below). Run withoutnohup, so the HUP-dependent watcher tests genuinely exercise their signal. bin/fm-lint.sh: clean (pinned ShellCheck 0.11.0).- Every
bin/*.shandbin/backends/*.shparses under stock macOS Bash 3.2. - Every skill named in
AGENTS.mdsection 13 resolves to a realSKILL.md; both sides' skills coexist. - No files lost from either side; upstream's deletion of the vestigial dispatch selector (fix(bin): remove vestigial dispatch selector kunchenguid/firstmate#1026) applied cleanly with no dangling references.
The three remaining failures are local toolchain shortfalls, not code, and all skip in CI:
| Test | Needs |
|---|---|
fm-calm-pi-extension |
Pi 0.81.1+ (this machine: 0.80.10) |
fm-kimi-harness |
Python 3.11+ for tomllib (this machine: 3.9.6) |
fm-secondmate-harness |
No claude process ancestry; resolves the real host harness |
Validation-pipeline rounds
The review step raised one finding against this branch's own spawn fix: the comment claimed the abort cleanup owns backend teardown, when spawn_abort_cleanup only removes Orca resources and Herdr presentation projections. It suggested tracking and rolling back endpoint creation for every backend plus atomic metadata publication.
That was scoped deliberately. The comment correction was taken. Per-backend endpoint rollback across tmux, flat Herdr, Zellij, and cmux was declined as a rider on an upstream sync - the residual gap predates this branch, and before the fail-closed fix the same failure reported a successful spawn and exited 0, so an untracked endpoint was strictly more likely, not less. It deserves its own task and tests.
Atomic publication was attempted and then backed out, because it broke two upstream structural guards that pin the expected source shape of bin/fm-spawn.sh:
tests/fm-gotmp.test.shrequires the literalecho "tasktmp=$TASK_TMP".tests/fm-backend-herdr.test.shrequires the literal} > "$STATE/$ID.meta"to appear after the lock acquisition and after the backend case, proving the lock spans backend creation through metadata publication.
A temp-file-and-rename necessarily changes both. Rewriting upstream's invariant guards to accommodate optional hardening is a poor trade inside a large sync, and the fail-closed fix alone was already green across the full suite. The atomicity idea can be taken up separately, updating those guards deliberately.
A third failure in that round, tests/fm-backend-tmux-smoke.test.sh, was a 10-second live-tmux readiness wait failing under pipeline load; it passes 7/7 on an idle machine and was left alone. The two environment failures were explicitly excluded from the fix instructions so the pipeline could not "resolve" them by weakening upstream's Pi and Python version guards.
Second sibling PR to land on main while this sync was in flight. Integrated as a MERGE for the same reason as #8: rebasing would flatten the upstream merge commit whose two parents are the record of this sync. No conflicts - #9 touches the capacity surface and its own new bin/fm-wait-progress.mjs, which this branch does not modify. Verified the consumers that the earlier blocker-model union touched are still green: fm-capacity 24/24, fm-fleet-snapshot-view 15/15, fm-documentation-audiences 5/5, and the new script is already classified in the toolbelt and audience inventories that #9 updated.
# Conflicts: # bin/fm-fleet-snapshot.sh
Intent
Captain's order: merge upstream kunchenguid/firstmate main (71 commits ahead, through fa0d85d) into this purple-phoenix fork, preserving the semantics of our six landed changes: #2 watcher hardening, #3 user-journey audits, #4 locale-invariant composer, #5 capacity dashboard, #6 dashboard redesign, #7 task-id contract. Hard constraints: MERGE, never rebase or force; never push to upstream (fetch only). Captain authorized merge-on-green for this sync PR.
Eleven files conflicted; each was reconciled semantically rather than by taking a side wholesale. Deliberate decisions a reviewer would not infer from the diff:
Beyond the merge itself, four pre-existing upstream defects had to be fixed because they reproduce on a clean upstream/main tree and blocked a green suite. These are intentional additions, not scope creep:
Two failures were genuinely caused by the merge and are fixed: our two task-id tests widened upstream's pinned ShellCheck source graph (switched to the non-following directive form), and our capacity and user-journey-audit skills were unclassified in upstream's new documentation-audience inventory (both classified agent-runtime).
One nuance to read carefully: two commits touch cleanup_child's TERM handling. The failure that prompted them was NOT a code defect - it was an artifact of running the suite under nohup, which sets SIGHUP to ignored and passes that to every descendant, so the test's kill -HUP was discarded. What survives on merit is the latency change: an intermediate commit introduced an unbounded blocking wait measured at 5-6s against an 8s budget, and the bounded 1s grace brings it to 1-2s. Treat it as a latency improvement, not a bug fix.
Verification: full suite 103/103 files with 3 failures, all environment-only and all skipping in CI (Pi 0.81.1+ vs local 0.80.10; Python 3.11+ for tomllib vs local 3.9.6; a harness-ancestry test that resolves the real Claude this runs under). fm-lint.sh clean. Every bin script parses under stock macOS Bash 3.2.
Mid-flight, PR #8 (persistent tailnet-only dashboard service) merged into origin/main. It was integrated as a second MERGE, never a rebase, because rebasing would flatten the upstream merge commit whose two parents are the record of this sync. That merge had three conflicts: .gitignore and AGENTS.md were additive on both sides, and bin/fm-fleet-snapshot.sh had #8 and upstream independently growing multi-blocker support. Both field families are emitted because both have live consumers (bin/fm-capacity.mjs reads #8's blocked_by_all, bin/fm-bearings-snapshot.sh reads upstream's blocked_by_ids/unresolved_blocker_ids). The one genuine contract collision was the scalar blocked_by on parsed backlog records: upstream captured only the last blocker while #8 joined all of them with commas. The join won because fm-capacity.mjs splits that scalar on commas, so it is the contract with a real consumer, whereas upstream's single-value form appeared only in test expectations and its information survives in blocked_by_ids; three upstream assertions were updated accordingly and still pin the structured fields. Captain-hold selection keeps upstream's stricter captain_actionable predicate so a blocked captain hold stays queued, with #8's origin and decision-key extraction layered on. docs/dashboard-service.md from #8 was classified maintainer-architecture in the documentation-audience inventory.
The PR body must be pipeline-generated. An earlier round of this work replaced the body by hand, which deleted the deterministic ## Pipeline section and its no-mistakes signature, failing the repository's PR-provenance check. That was corrected by moving the detailed conflict write-up into a PR comment and letting the pipeline own the body, per the guidance in CONTRIBUTING.md that the signature must never be pasted in by hand. Two sibling PRs, #8 and #9, merged into origin/main while this sync was in flight; both were integrated as merges rather than rebases for the same reason the upstream sync itself is a merge.
What Changed
Risk Assessment
Testing
Completed 1 recorded test check.
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ No issues found.
command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.