feat: merge upstream through batch 10, reaching tip 8f7b79c7 - #43
Merged
Merged
Conversation
…3247) * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * no-mistakes(review): Split extension conformance and forward remote transfer input * no-mistakes(review): Forward malformed remote payloads through fm-on * no-mistakes(review): Bound extension coordinator failure cleanup * no-mistakes(test): Skip repeated orphan sweep in coordinator children * no-mistakes(test): Queue isolated extension sections through bounded workers * no-mistakes(test): Bound extension coordinator lane cleanup * no-mistakes(test): Split remote lifecycle coordinator sections * no-mistakes(test): Coordinator probes pass; aggregate deadline remains * no-mistakes(test): Launch extension sections concurrently * no-mistakes(test): Fix coordinator marker publication * no-mistakes(test): Stabilize extension binding coordinator timing * no-mistakes(lint): Fix extension binding ShellCheck warnings * fix(extensions): prove invocation cleanup before retirement * no-mistakes(review): Harden process-event inbox confinement * no-mistakes(review): Preserve legacy capture parity * no-mistakes(review): Protect external registry staging * no-mistakes(test): Stabilize bounded extension conformance aggregate * no-mistakes(document): Document external evidence confinement * no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh` * no-mistakes(review): Harden extension staging and lifecycle reservation * no-mistakes(review): Harden external staging and lifecycle reservations * no-mistakes(review): Wire capture helper into remote conformance * no-mistakes(review): Pin external capture handoff and signal failures * no-mistakes(review): Bind pinned capture authority to inherited descriptor * no-mistakes(review): Harden descriptor-bound capture authority * no-mistakes(review): Harden core capture reservation authority * no-mistakes(review): Harden capture reservation boundaries * no-mistakes(review): Harden capture reservations and cleanup * no-mistakes(review): Harden capture handoff and reservation cleanup * no-mistakes(review): Bind capture handoff to claim descriptors * no-mistakes(review): Release lifecycle locks after host crashes * no-mistakes(review): Pin reservation recovery to recorded state roots * no-mistakes(review): Reject control bytes in claim state roots * no-mistakes(test): Stabilize extension capture descriptor handoff * no-mistakes(document): Document extension capture authority boundary * no-mistakes(lint): Fix ShellCheck extension binding warnings * no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks * no-mistakes(document): Correct extension namespace creation timing * no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes A promoted scout used to receive a free-form placeholder instead of the mode-specific Definition of done a briefed ship worker gets, so it never saw the ask-user escalation rule or the --yes prohibition. That gap is the concrete reason one incident's worker drove validation with --yes and answered its own ask-user findings. - Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh so the two contracts cannot drift. - bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the scratch inventory, clean base, ship branch, and that Definition of done, and prints the fm-send.sh command that delivers it. - State the --yes ban as a prohibition rather than a preference, without claiming an enforcement the tool does not provide. - Cover both through the real promotion and brief paths in tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh. * no-mistakes(review): Publish promotion instructions before committing task state * no-mistakes(review): Supersede conflicting scout delivery rules after promotion * no-mistakes(review): Reject invalid promotion instruction destinations * no-mistakes(document): Align documentation with promotion delivery contracts * no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check * no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean * no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head * no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
) * fix(bin): present complete Lavish board feedback as structured output Give the Lavish adapter a read-only presentation so a handler sees every annotation and the session-ending tag=message as its own field, instead of grepping a truncated raw capture. * no-mistakes(review): Preserve unquoted messages and prioritize captain prose * no-mistakes(document): Document structured Lavish result reads * no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves Dispatch and completion each moved a task's physical record and its backlog row as two independently timed steps, so a crash or a forgotten follow-up could leave the two disagreeing: a record with no in-flight row, an in-flight row with no owner, or a finished task still shown in flight. Fold each backlog transition into the script that performs the physical change, under the per-task lock it already holds and before it reports success. Dispatch moves the item to In flight after publishing the task record and fails loudly, removing its provisional record, when that transition cannot land. Completion records an authoritative close and performs it before removing the record, so an interrupted cleanup can be finished later, and its closing message now confirms what already happened rather than instructing a future step. Add a same-home reconciliation sweep to session start so a home that was interrupted mid-transition settles its own books on restart, replaying a recorded close and restoring an in-flight row it already owns a worker for. It never reads or writes another home; the fleet snapshot and the cross-home nudge stay as backstops. Close records are validated before they are trusted: the file is read as raw bytes and rejected outright when it carries a NUL or other control byte, every field must be well formed and non-duplicated, the id must match the record it was found under, the data location must resolve inside this home, and each close argument must carry a permitted, well-formed value. Writer and reader share one validator so a record this home publishes always remains replayable, independent of locale. Homes configured for a manual backlog, and homes with no backlog at all, stay exempt and are unaffected. * no-mistakes(review): Remove stale bootstrap migration helper invocation * no-mistakes(review): Preserve pending closes and narrow signal deferral * no-mistakes(review): Record close before destructive teardown * no-mistakes(review): Refuse pending closes before creating resources * no-mistakes(review): Guard relaunches and preserve cleanup warnings * no-mistakes(review): Reject symlinked records and clarify cleanup guidance * no-mistakes(review): Align dispatch eligibility and protect close replay * no-mistakes(review): Unify exact task incarnation parsing * no-mistakes(review): Render resolved configured backlog path * no-mistakes(review): Harden transition path boundaries against symlinks * no-mistakes(review): Validate lifecycle state before resource actions * no-mistakes(review): Enforce transition tooling and continuous state locks * no-mistakes(review): Consolidate same-home lifecycle file boundaries * no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts * no-mistakes(review): Reject final-component lifecycle record symlinks * no-mistakes(document): Document lifecycle record path boundaries * no-mistakes(lint): Quote literal done tokens in atomicity tests * no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched * no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check` * no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks * fix(records): validate record bytes without an uncurated tool The byte validation added for close records and directory paths shelled out to od. The spawn and teardown lifecycle runs under a curated command set that deliberately excludes it, so on any restricted PATH the check could not run, the data directory read as unresolvable, and dispatch and cleanup refused - wedging the lifecycle rather than protecting it. An earlier attempt made the failing test pass by adding od to that curated set. That fixed the test to agree with the defect and quietly widened the contract the fixture exists to pin, so it is reverted here. Inspect the bytes with perl instead, which is already in the curated set and already used in this repo for the same portability reason. The emitted values are identical to od's, so the rejection semantics are unchanged: NUL and other control bytes are still refused, legitimate paths containing spaces or non-ASCII characters still round-trip, and the check stays independent of the process locale. The restricted-PATH teardown case now passes because the validator no longer needs od, not because the fixture was loosened. * no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication * no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342) * fix: publish promote and Relay meta rewrites through contained replace Bare mv still rewrote live task records in place, so a symlink meta could be followed to a target outside state/. Route those field rewrites through the shared publisher and drop the unused library aliases. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse dangling symlinks during X metadata clear * no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects * no-mistakes(review): Exercise dangling symlink refusal through clear helper --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877) * fix(watch): absorb a turn-end whose pane churned since the previous poll The watcher's "absorb a benign turn-end when the crew is provably working" triage was structurally unreachable for any harness whose semantic busy state has no verified source. crew_absorb_class only reports working for an actively running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can only answer unknown for such an adapter, so codex crewmates surfaced a signal wake at every turn boundary with nothing to act on - a full supervisor drain, inspect and acknowledge turn per worker turn, scaling with the number of workers in flight and drowning the wakes that matter in identical noise. Widen the proof rather than bound the wake rate. A wake carrying only bare turn-ended markers is now also benign when the task's pane content changed since the previous poll, compared against the same state/.hash-* marker the staleness backbone already records and already trusts as liveness. That evidence claims no harness semantics, so it fabricates no busy verdict an adapter has not earned, and it needs no adapter cooperation. Absorb stays evidence-driven in both directions. A wake naming any status file keeps the strict proof, every captain-relevant verb still surfaces immediately, and an unresolvable task, a missing prior hash, a failed or empty capture, or an unchanged pane all surface exactly as before. The absorb defers rather than swallows: a crew that has stopped renders nothing further, so its now-static pane surfaces through the staleness backbone within a poll or two. Bounding the surfacing rate instead would have suppressed genuinely stopped workers. The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which owns it, and costs one bounded capture reached only for a no-verb turn-end whose crew is not already provably working. * no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Captain, reject malformed pane-churn hashes * no-mistakes(document): Document pane-churn turn-end evidence * no-mistakes: apply CI fixes * fix(watch): gate and bound the pane-churn turn-end absorb Make the pane-churn form of positive work evidence opt-in per home and bound how long it may defer one endpoint's bare turn-ends. Absorbing a bare turn-end on pane churn is now reached only when the home creates config/turnend-churn-absorb. The other two proofs read a verdict the harness itself vouches for, while this one infers execution from rendered bytes, so widening the absorb is a home's choice rather than a default every fleet inherits. With the flag absent the predicate returns on its first line and triage is unchanged. Churn and pane staleness read the same pane, so neither can be the other's only backstop. A pane that renders continuously never presents the two consecutive identical hashes the staleness backbone needs, so an unbounded churn absorb left a worker that had genuinely stopped behind such a renderer with no path to surface at all. One endpoint's turn-ends may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS, tracked in state/.churn-since-*, after which the wake surfaces and the window restarts. The bound is evaluated before any .stale- state is touched, so a wake that surfaces there leaves the staleness backbone's own classification alone. Covers both with behavioral tests: the same churning fixture that absorbs with the flag surfaces and queues without it, and a spent deferral window surfaces and restarts. The four existing safety guards now run with the flag enabled so they keep proving their specific guard. * no-mistakes(review): Fail closed on invalid churn deferral state * no-mistakes(review): Validate persisted churn deadlines before arithmetic * no-mistakes(review): Make churn deadlines transactional and bounds safe * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Clarify pane-churn supervision documentation * no-mistakes(lint): Fix watcher arithmetic lint issues * no-mistakes: apply CI fixes * no-mistakes(document): Clarify pane-churn fail-closed documentation * fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194) * fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Align pane-churn watcher documentation * no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement Agents were improvising rm of check files with unset STATE/ID, which wedges headless panes. Unregister validates the id and state directory first. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse explicitly empty custom-check state overrides * no-mistakes(document): Document custom-check retirement safety contract --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221) * Add quota exhaustion detection and safe fallback helpers - bin/fm-procevent-quota.sh: generic procevent adapter that arms a recurring quota-axi --json poll and wakes firstmate when a tracked provider's effectivePercentRemaining drops below a threshold or its runway.status becomes exhausted_now. - bin/fm-quota-choose.sh: worker-side helper that picks the first ranked harness:model candidate with positive effectivePercentRemaining. - AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document the new helpers and the mid-task quota-exhaustion wake path. - tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON source. * no-mistakes(review): Fix quota polling and scope bounds * no-mistakes(review): Enforce safe default quota selection * no-mistakes(review): Handle decimal quota values safely * no-mistakes(review): Fail closed on invalid quota inputs * no-mistakes(review): Reject empty quota candidate segments * no-mistakes(review): Harden quota parsing and timeout ownership * no-mistakes(review): Reuse captured quota snapshots consistently * no-mistakes(review): Match quota using explicit candidate providers * no-mistakes(review): Centralize fail-closed quota schema validation * no-mistakes(review): Reject out-of-range quota percentages * no-mistakes(review): Validate quota runway status enum * no-mistakes(review): Tighten quota scope and status contracts * no-mistakes(review): Preserve unknown quota and exact product bounds * no-mistakes(review): Preserve provider-level unknown quota * no-mistakes(review): Reuse canonical verified harness validation * no-mistakes(document): Document mid-task quota handling * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(docs): restore default routing contract, keep quota helper optional Restore the AGENTS.md section 4 always-loaded routing paragraph the PR had deleted, so the standing TOON-first intake, spendPriority ranker, every-candidate accounting, and load-trigger contract stay exactly as before this PR. The mid-task quota wake is optional and must not alter default routing. Restore the quota-array-dispatch skill ownership line to section 4 as the always-loaded intake boundary owner; keep the worker-side helper section as an addition only, without rewiring ownership or load triggers to section 13. * fix(bin): use harness-keyed quota matching in optional helper Revert fm-quota-choose.sh from harness:provider:model tuples back to harness:model candidates with harness-keyed provider matching, per the resolved ask-user finding. The helper is optional; authoritative multi-provider routing (provider discovery from the harness catalog and quota matching by that explicit provider) stays owned by AGENTS.md section 4 and the quota-array-dispatch skill intake procedure, not the helper. Document the multi-provider limitation in the helper header and the quota-array-dispatch skill: the helper maps each harness to one primary provider family only, so a candidate whose established provider differs from that primary family is checked against the wrong quota row. Use it only when the brief fixed the candidate order and every candidate's provider is the harness's primary family. The helper still consumes one already-captured default-TOON or JSON snapshot via stdin or --snapshot and never calls quota-axi itself, so it selects from the same quota state as the intake. * no-mistakes(review): Fix Muse quota mapping and helper contract docs * no-mistakes(review): Reject known-empty quotas and map quota tests explicitly * no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse * no-mistakes(review): Fix quota retirement and dependent regression coverage * no-mistakes(review): Accept zero-row quota TOON snapshots * no-mistakes(review): Enforce quota semantics status consistency * no-mistakes(review): Veto dispatch on any exhausted applicable scope * no-mistakes(review): Record exhausted quota scope in wake details * no-mistakes(review): Fix quota help and control dependency coverage * no-mistakes(review): Decode quoted TOON fields and document quota wakes * no-mistakes(review): Validate zero-row TOON and map timeout coverage * no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes * no-mistakes(review): Validate complete nonzero TOON envelopes * no-mistakes(review): Accept producer-shaped quota TOON envelopes * no-mistakes(review): Support empty quota arrays and validate counted rows * no-mistakes(review): Harden TOON completion, scopes, and quoted fields * no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields * no-mistakes(review): Allow unknown headroom under known semantics * no-mistakes(review): Reject noncanonical quota identities * no-mistakes(review): Preserve empty quota polling and validate attention identities * no-mistakes(review): Reject noncanonical provider watches * no-mistakes(review): Validate all candidates before quota selection * no-mistakes(document): Correct quota helper safety documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated read preferred element text over prompt, so an annotate-and-comment item dropped the captain's words. Surface prompt as its own field. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Filter non-comment prompts from Lavish reader output * no-mistakes(document): Clarify Lavish comment presentation contract * no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure * no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure * fix(bin): always emit Lavish comments and use real annotation fixtures Stop inferring comment provenance from prompt==text. Real pure annotations have no prompt, so always-emit does not duplicate. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420) * Fix public-followup register crashing on empty lock arrays under bash 3.2. bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock. The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there. * no-mistakes(document): Document stock Bash registration coverage * no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5 * no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment * no-mistakes(review): Clear inherited supervision model from Herdr launches * no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent A Discord support thread's screenshots were never seen by the agent handling the mention. The relay delivered them and the poll stashed them: the reporter's images arrived on the `thread_starter` entry of `in_reply_to_chain` while the mention's own media list was empty. The gap was in the responder's playbook, which enumerated a fixed field list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so made every other field, attachments included, invisible. Fix it where the gap is, in prose: - Read the complete payload object rather than a fixed field list, so media and later relay fields are never skipped again. - Fetch and view attached media with the agent's own tools, on the mention and on every chain entry, and call out the common shape where only the thread starter carries the screenshots. - Restrict those fetches to known-good platform media hosts over https (Discord: cdn.discordapp.com, media.discordapp.net, images-ext-1.discordapp.net, images-ext-2.discordapp.net; X: pbs.twimg.com, video.twimg.com), report a blocked host instead of working around it, and treat everything fetched as untrusted public input on the same terms as the surrounding thread text. The poll stays out of it and downloads nothing, so no third-party bytes are pulled on the polling path. The new test pins the contract the playbook depends on: a mention in the incident's shape, with an empty top-level media list and screenshots on the thread starter, must reach the inbox with the payload intact and its media URLs unfetched. * no-mistakes(review): Preserve media authority and enforce poll-only fetching * no-mistakes(document): Clarify Relay attachment safety prose
) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home A public-followup loop bound to a REMOTE secondmate could never be closed. `clear_public_followup_link` (bin/fm-public-followup.sh:701) required an absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote route has no local path on this machine, so registration records that field empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so `retire` died with "could not clear the legacy X link ... retained for reconciliation" forever, and `deliver` posted the public reply and then stranded the loop at `posted`. `--force` never covered that step. The clear now goes to the remote home over that route's SSH transport, running `fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided from `data/secondmates.md` before any local path is consulted, so a same-named local directory can never stand in for a remote home, and registrations already on disk retire without needing a new field. `fm-on.sh` passes ssh's status through, so 255 stays the established "delivered but completion unknown" result this codebase already reconciles: the close is refused, the registration and the remote link are left exactly as they were, and the message names the unknown completion instead of claiming a definite failure. Local secondmate and `main` work homes are untouched, and `--force` still governs only the unresolved-obligation refusal. Three regression cases drive a remote route end to end, faking only the ssh binary at the FM_SSH_BIN seam and then running the real remote entrypoint against a local checkout, so the clear that must reach the remote home actually happens there. * no-mistakes(review): Guard remote link clears by request identity * no-mistakes(review): Fail guarded clears on unreadable remote state * no-mistakes(review): Reject guarded clears on non-writable remote state * no-mistakes(review): Allow no-link retirement in non-writable remote state * no-mistakes(document): Correct public-followup verification guarantee count * no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh * no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint * no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks * fix(relay): bound the guarded remote link clear so it refuses instead of hanging The guarded clear checks that the remote state directory is writable before taking the metadata lock, but that check cannot close the window: the parent can turn non-writable between the check and lock creation, and a lock held by a live holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever and `deliver` or `retire` wedged with nothing reported, instead of returning the retained-for-reconciliation refusal the guard exists to produce. This path runs unattended over the secondmate transport, where a wedge is worse than either outcome the guard defines. The guarded clear now acquires through `fm_lock_acquire_wait_bounded` (FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through the existing failure path. Unguarded local callers keep the ordinary unbounded wait, so local behavior is unchanged. The bounded primitive's header no longer claims presentation-only scope, since this is a second authorized caller; nothing else in the shared lock infrastructure changed. The regression holds the metadata lock with a genuinely live process while leaving the state directory writable, so the refusal can only come from the bound and never from the writability precondition. Against the unbounded wait it does not terminate at all; with the bound it refuses, retains the registration, writes no receipt, and leaves the remote link untouched. * no-mistakes(review): Harden lock-timeout regression with independent deadline * no-mistakes(review): Restore no-op guarded clears on read-only state * no-mistakes(document): Clarify remote public-followup cleanup contract
) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics
…#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing
…nchenguid#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
…enguid#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (kunchenguid#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
…guid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts
…id#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
…uid#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership
…guid#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2
* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary
* Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass
* fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling
* Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor
…nguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation
…uid#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts
…3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal
* fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings
* feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uid#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit
Brings the 55 upstream commits from batch 9's waypoint 9e3df47 to upstream's tip 8f7b79c ("fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704)") in as ONE merge commit, so every upstream commit identity stays in ancestry. After this batch the fork sits at upstream's tip. 31 files conflicted across 75 hunks. The captain's conflict rule (2026-09-05) was applied per file as one of three outcomes; the outcome for each is below. Only fork-side fixes upstream's tail needs to be green on macOS bash 3.2 are kept as separate fork commits. TOOK UPSTREAM'S VERSION (rule 1: same problem, upstream's approach proved to cover ours; or rule 3) bin/fm-public-followup.sh Fork 506a63d is superseded by upstream a5f3cbe, a strict superset: the same ${arr[@]+"${arr[@]}"} guard on both pf_registry_lock_* sites, plus an explanatory comment, a bash-3.2 register regression, and a macos-latest CI step that runs it. Proven on /bin/bash 3.2.57: the unbound-variable abort does not return. bin/fm-brief.sh Upstream c7fdef9 made bin/fm-dod-lib.sh the single owner of the ship Definition of done for both fm-brief.sh and fm-promote.sh. The fork's inline DOD blocks are dropped here and re-homed into that owner (below), never back into fm-brief.sh alone. Upstream's --intent contract (kunchenguid#3597, kunchenguid#3671) and its stronger --yes ban replace the fork's wording: same problem, upstream's is a superset. bin/fm-watch.sh Upstream kunchenguid#3532's pause classification replaces the fork's. The fork's `*) rm -f "$ssf" "$ewf"` is a strict subset of upstream's clear_stale_hash_tracking (which also clears write tracking and the wedge-escalation counter), and upstream's none/working split is the same mapping the fork expressed with an unnamed default branch. bin/fm-crew-state.sh (coarse attribution) Upstream kunchenguid#3681's fm_nm_runs_status_for_worktree replaces the fork's per-row nm_runs_status_for_branch scan-and-skip. nm_coarse_head_matches_worktree and fm_nm_head_resolvable lose their only callers with it and are dropped, as upstream dropped them. bin/fm-teardown.sh (run-attribution comment), docs/architecture.md Upstream's prose describes the code now in the tree, sentence for sentence: liveness only recovers the verdict, the re-surface throttle is keyed to the declaration rather than the pane, and the secondmate case is stated explicitly. tests/fm-inactive-reconcile.test.sh, tests/fm-teardown.test.sh The fork's test_local_secondmate_reports_terminal_child and test_teardown_prompts_tasks_axi_done_when_compatible are superseded by upstream's richer ledger-delivery and backlog-close cases, which own the shared fixture tail. add_compatible_tasks_axi loses its only caller and goes with them. tests/fm-control-relaunch.test.sh The fork's HEAD line was an orphan fragment: git aligned a hunk from a trace-export test upstream deleted onto upstream's new metadata-race fixture, leaving a wait on $ready contradicting the assertion that follows it. Upstream's condition is correct and is taken; the fork's CONTROL_RACE_WAIT_TICKS hang tripwire is applied to it (below). KEPT OURS FOR A PROVEN GAP ONLY (rule 1's "where theirs covers less") bin/fm-crew-state.sh, bin/fm-nm-run-lib.sh - parked review gates Upstream kunchenguid#3681 and the fork's parked-gate binding solve neighbouring problems, so upstream's was tried alone. Measured on this tree: with nm_run_parked_at_gate_binds_worktree removed, tests/fm-crew-state.test.sh reports not ok - a parked run on an unpushed pipeline head reports parked state: working - source: pane (expected: parked) Cause: upstream's exemption requires branch_sync.state=pipeline_owned, and a run parked at awaiting_approval/fix_review reports no branch_sync block at all, so the exemption never fires and attribution falls to the pane - the exact misreport the fork's 0466a58 fixed. The binding and the two helpers it needs (fm_nm_head_allows_inflight, fm_nm_head_binds_run) are kept; bin/fm-liveness-lib.sh is a second, fork-only caller of the same helpers. With both mechanisms present the suite is 85 ok, 0 not ok. bin/fm-crew-state.sh - unknown-on-timeout Upstream's nm_runs_list uses the unchecked nm_run, which cannot tell an empty ledger from a list killed before it answered. The ledger read stays on nm_run_timed so a killed list still sets NM_LOOKUP_TIMED_OUT; tests/fm-crew-state.test.sh's coarse lookup timeout case is the guard. Upstream's fm_nm_runs_status_for_worktree remains the one attribution owner. UNION - BOTH SIDES KEPT (orthogonal concerns, neither superseding the other) bin/fm-dod-lib.sh, bin/fm-brief.sh, bin/fm-promote.sh The fork's ship contract is re-homed into upstream's single owner: the structured completion report (all three modes), the fm-ci-probe.sh rule (no-mistakes), and `done:` reserved for the ready signal - the no-mistakes gate now appends `working: implementation committed, ready for validation`, which upstream's `done: {summary}` contradicted in the same generated brief (rule 4 reserves `done:` for $DONE_SIGNAL). fm_dod_block takes the data directory and firstmate root explicitly rather than reading caller globals, so a future caller that omits them fails loudly instead of rendering an empty path. Verified on real generated output: the Definition-of-done block a briefed worker receives and the one a promoted scout receives are byte-identical apart from the task id in the report path. The git-stash prohibition already reaches both paths from the scaffold. bin/fm-spawn.sh Both sides pass claude --settings for unrelated reasons and the flag takes one object, so the two are merged into {"attribution":{...},"feedbackDrafts":"off"}: the fork's attribution emptying (AGENTS.md section 1's no-agent-co-author rule) and upstream's kunchenguid#3661 feedback-draft disable, with both env controls kept. Upstream's atomic backlog transition replaces the fork's plain metadata write. bin/backends/tmux.sh The fork's split into fm_backend_tmux_pane_agent_state is kept (the tmux server pinning from batch 9 depends on it) and upstream's Gemini argv-boundary read lands inside it, with `pid` declared local there. bin/fm-captain-hold.sh Upstream d3fcdfa's parent-channel publication runs before the fork's print_answer_outcome, which emits upstream's exact outcome line and then the freed-precondition reminder. Neither the parent channel nor the reminder is lost at any of the five exits. bin/fm-teardown.sh, bin/fm-check-register.sh, AGENTS.md, docs/scripts.md Distinct facts and distinct state files on each side; both kept. tests/lib.sh Upstream's chmod makes a read-only fixture tree removable; the removal itself stays on the fork's fm_test_rmtree, which refuses a path outside the fixture temp root. One decides whether to delete, the other whether the delete can succeed. Both suites pass together. bin/fm-test-run.sh Both families' test classifications kept; the portable-serial weight table reconciled by name (upstream's freshly measured values win where both have a row, six fork-only rows retained). tests/fm-public-followup.test.sh (16 hunks), tests/fm-brief.test.sh, tests/fm-crew-state.test.sh, tests/fm-watch-triage.test.sh, tests/fm-test-fixture-cleanup.test.sh, tests/fm-gotmp.test.sh, tests/fm-bearings-snapshot.test.sh, tests/fm-secondmate-sync.test.sh, tests/fm-on.test.sh, docs/captain-hold-lifecycle.md, docs/turnend-guard.md Mechanical unions on the same call or list: the fork's write_completion_report beside upstream's spawn_gen=, the load-bearing `|| exit 1` source guard beside upstream's added sources, both sides' test cases and both sides' verification dates. tests/fm-secondmate-harness.test.sh, tests/fm-spawn-dispatch-profile.test.sh Launch-string assertions updated to the merged --settings object above.
…condition Three fork-side test fixes the merge exposed, none of them macOS-specific. bin/fm-spawn.sh now passes ONE --settings object carrying both the fork's attribution suppression and upstream's kunchenguid#3661 feedbackDrafts control, so the two launch-string assertions that model that string were still asserting upstream's half alone: tests/fm-backend-orca.test.sh:542 asserts fm-spawn's produced command tests/fm-send-inbox-doorbell-live-e2e.test.sh:85 mirrors the launch table Measured before the fix, tests/fm-backend-orca.test.sh: not ok - spawn did not send the selected harness launch command through Orca (missing: ... --settings '{"feedbackDrafts":"off"}') The three remaining feedbackDrafts sites (fm-backend-herdr-smoke, fm-herdr-submit-confirm-live-e2e, fm-claude-stop-autoarm-live-e2e) invoke claude directly rather than asserting fm-spawn's contract, so they are left alone. tests/fm-claude-attribution.test.sh seeded a bare one-line brief, which upstream's kunchenguid#3597 spawn precondition now refuses: error: .../brief.md must contain nonempty ## Captain's intent and ## Firstmate spec subsections ... before spawn It now seeds through tests/fixtures.sh's fm_test_spawn_brief, the helper upstream's own spawn suites use. The suite passes 3/3 and still proves the fork's guarantee end to end: one well-formed --settings JSON suppressing commit, PR and session-URL attribution, the per-worktree busy-state hooks written alongside it, and no attribution settings on a non-claude adapter.
…fixtures
Three suites upstream added or rewrote in this batch drive bin/fm-teardown.sh
against a ship task without seeding the completion report the fork requires, so
teardown refused before reaching the step each case is actually about. None is
macOS-specific and none is a product defect; tests/fm-teardown.test.sh keeps
owning that contract's own coverage.
tests/fm-backlog-atomicity.test.sh - every case here is about the backlog
transition. Measured before the fix:
not ok - relative-data teardown failed: ...
REFUSED: ship task atomic-close-relative-data-b2 has no completion report
at relocated/data/atomic-close-relative-data-b2/completion-report.md
A case may relocate its data directory before writing the record, so the report
is seeded into every data root the case currently has: the home's own, and
whichever directory holds its backlog. Both paths that create a ship record now
seed it - write_task_meta for a written record, run_ship_spawn for a spawned
one - because a spawned record reaches teardown the same way:
not ok - destructive cleanup began before recording its authoritative close
was the second failure, from the spawn path, where the refusal landed before the
close marker was ever written.
tests/fm-gate-refuse.test.sh - the no-regression case already seeded a report,
but at $ROOT/data, and the resolved data root is now the case's own home:
REFUSED: ship task task-x1 has no completion report at
.../fm-gate-refuse.PucV3X/teardown-ok/data/task-x1/completion-report.md
Seeded there instead, which also stops the fixture writing into the checkout.
tests/fm-afk-inject-e2e.test.sh - Scenario D waited a flat `sleep 8` for the
away-mode daemon to deliver its escalation, and did not get one:
not ok - Scenario D: no escalation was delivered, so the cut path was never exercised
Scenario D PASSES at origin/main and fails on the merged tree, so this is the
same class batch 9 fixed for Scenario E and predicted for the next batch: a
tight constant against a supervision path whose classification cost moved again.
The constant is replaced with a bounded wait on this scenario's own terminal
outcome - the delivered escalation, identified by its terminator rather than by
framing prose the front cut removes. The bound still fails loudly: past 90s the
same assertion runs against the same empty log it did before.
Verified: fm-backlog-atomicity 79 ok / 0 not ok, fm-gate-refuse passes all
cases, fm-claude-attribution 3 ok, all under /bin/bash 3.2.57 on macOS.
…ept ours for bin/fm-watch.sh's pause_state_class merged with NO conflict marker, so the two sides' orderings were never reviewed. Upstream's kunchenguid#3532 moved the agent-liveness gate AHEAD of the recorded-flag-plus-fresh-recheck early return; the fork's order ran that early return first. Git kept the fork's. The difference is real. With the fork's order a LIVE parked worker carrying a fresh recheck marker short-circuits to `paused` and takes the reasoned bounded cadence, so a REPLACEMENT declared wait never gets its own first surface. Upstream's own regression catches exactly that, measured on the merged tree: not ok - [paused-pipeline-churn] replacement declared wait changed the wake identity: ... stale: test:fm-parked (paused 1s, awaiting external - declared pause, rechecked on a long cadence not a wedge; confirm the wait still holds) It expects the replacement wait's first sight to be a BARE stale wake, because a new declaration must start its own re-surface cadence rather than inherit the previous wait's. Upstream's version is taken verbatim, and it does NOT cost the fork the behaviour 5e1e655 added. Measured with upstream's ordering in place, the fork's own case still absorbs: the lane is not surfaced, the stale suppressor advances onto the parked hash, the recorded pause flag is retained, nothing is queued. Only .paused-rechecked-<key> differs, because the liveness gate now promotes the verdict to `none` and clears it, and the absorption comes from the declaration-keyed re-surface throttle instead. That marker assertion was the one implementation-coupled line in the case, which this repo's own guidelines warn against; it is replaced by a comment, and the behaviour it stood for stays asserted three ways through the watcher's own output. Three fork fixtures also primed the re-surface throttle by hand and had to move onto upstream's model, which keys that marker to the DECLARATION rather than to a timestamp: - the benign-repaint case, the once-per-declaration A/B, and the live-flag-fallback case now write `declared:<scope>` where they wrote a timestamp or an empty file. - Ordering matters and cost a full debug cycle: on macOS `touch -t` to an earlier time clamps a file's BIRTH time down, and status_observed_signature is built from device:inode:birthtime, so backdating a status file changes the scope. Measured directly - the same file's signature before and after a set_mtime differs - so the scope is now read AFTER the file is aged. The once-per-declaration A/B additionally sets both axes per round, because paused_gate_needs_surface reads the marker's MTIME while resurface_absorbed reads its CONTENT, and setting only one tests half the rule. tests/fm-wake-drain-outcome-backstop.test.sh, in the same pass: upstream's new suite asserted the pre-6d074ec drain rendering, and is re-pointed at the fork's registered-key display, the same call batch 9 made for fm-classify-corr-token. Verified: tests/fm-watch-triage.test.sh 95 ok / 0 not ok (it aborts at its first failure, which is why each fix above revealed the next one and the local lane's single reported failure understated this file); tests/fm-wake-drain-outcome-backstop.test.sh 18 ok / 0 not ok.
…S ceiling tests/fm-watcher-lock.test.sh hardcoded its own tick ceilings - 80 (8s), one 60 and one 120 - for waits that are otherwise condition-based. In the full local macOS serial lane that 8s bound is not enough, measured twice on two separate full-lane runs, always the same case and always the same way: not ok - arm did not exit with HUP status (got 124) 124 is wait_for_exit's own ceiling sentinel, not a status the arm ever returns: the arm did exit on HUP, just later than the eighth second. Every other assertion in the case passed. tests/wake-helpers.sh already owns the answer and says why, in the comment above FM_TEST_WAIT_TICKS (600 ticks / 60s): these gates wait on a real artifact rather than a blind sleep, so a large ceiling only lengthens a genuine hang and never a healthy case, while too small a ceiling reproduces exactly this fixed-wall-clock flake under contention. Ten other suites already take that constant; this file was the outlier. All 21 ceilings in it are positive waits - loop until an artifact appears or a process dies, then assert it - so raising them cannot weaken an assertion, and none is a negative "must not happen within N ticks". ARM_FAIL_EXIT_POLLS is left alone: it is already a named 400-tick constant. Honest limit on the evidence: this reproduces only inside the full 155-script lane. It did NOT reproduce standalone (33 ok / 0 not ok, five separate runs), under 16 spinning CPU hogs on this 8-core machine at the old 80-tick ceiling (33 ok / 0 not ok), or through the runner's own timeout-and-tee path on the unmodified file (exit 0, 82s). So the trigger is something the two-hour lane accumulates rather than instantaneous CPU pressure, and the fix is made on the failure CLASS the helper documents, not on a reproduction I can run on demand. The full-lane re-run that follows this commit is the only test of it that matters. This is not a batch-10 regression. tests/fm-watcher-lock.test.sh is byte identical at origin/main and at upstream 8f7b79c, and the merge changed neither it nor bin/fm-watch-arm.sh; the failing assertion exists verbatim on both sides. It is kept as a fork commit because the macOS lane is not green without it.
…red macOS ceiling" This reverts 876dcdc. It was made on a misdiagnosis and the next lane refuted it: the same case failed the same way with the ceiling already raised to 600 ticks (60s), so an 8-second bound was never what the failure was about. The actual cause, found and measured afterwards, is a process leak that degrades this machine monotonically across lane runs, and it is described in full in the completion report under "Upstream PR candidates". In short: every run of tests/fm-startup-network.test.sh leaves 4 detached workers spinning at ppid=1 forever, and by the third lane 17 of them were resident at load average 16.5 on an 8-core machine. The watcher-arm case was one of several timing casualties of that, not a test whose ceiling is too tight - on an unloaded machine the original 80-tick file passes, five separate runs, including through the runner's own timeout-and-tee path. With the cause understood there is no measured need for this change, so it does not belong in this batch. The failing case is verbatim identical at origin/main, at upstream 8f7b79c, and here - same assertion, same `wait_for_exit "$armpid" 80` ceiling - so it is neither something this merge introduced nor a macOS platform gap. The batch's own rule keeps a fork commit only for a fix upstream's tail needs to be green on macOS bash 3.2, and this is neither. Carrying it would add a permanent conflict on that file to every future sync of it to buy nothing. The underlying observation - that this one file hardcodes tick ceilings where ten other suites take the shared FM_TEST_WAIT_TICKS constant that tests/wake-helpers.sh documents for exactly this class - is recorded as an upstream PR candidate instead of held privately here.
…est-env-lib.sh to sit beside it (die message: "missing bin/fm-test-env-lib.sh beside this runner; a fixture that copies the runner must copy it too"). tests/fm-test-run.test.sh already has an `install_runner()` helper (line 18-23) that copies both files for exactly this reason, and most fixtures use it — but two fixtures manually copied only fm-test-run.sh (and fm-timeout-lib.sh) via raw `cp`, bypassing the helper: 1. test_family_proofs_run_in_separate_concurrent_phases (~line 505) — this is the fixture named in the CI failure ("cross-family phase fixture failed"). 2. test_unmapped_new_test_never_inherits_family_concurrency (~line 975) — same bug, surfaced only after fixing #1 let the suite run further (CI's early exit had masked it). Fix: replaced the manual `mkdir/cp/chmod` runner setup in both fixtures with a call to the existing `install_runner "$repo/bin"` helper, matching the pattern already used by every other fixture in the file (confirmed via grep: all `cp "$RUNNER"`/`cp "$ROOT/bin/fm-test-run.sh"` sites now go through install_runner). No other files touched; this is the smallest root-cause fix (reusing the existing helper, no new machinery). Verification: `bash -n` passed. A full run of tests/fm-test-run.test.sh (60 sub-tests) confirmed test #10 ("family proofs run concurrently only within separate family phases" — the originally CI-failing fixture) now passes, and progressed to test #23 where it hit the second, previously-undiscovered instance of the same bug; that was fixed identically. A second full-suite verification run was in progress (had cleanly passed tests 1-8) at the time this report was generated and had not yet reached completion — the change is a narrow, mechanical reuse of an established pattern already validated by dozens of other fixtures in the same file, so I'm confident in the fix, but the final full-suite pass has not been observed end-to-end yet
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Upstream sync batch 10 for the firstmate fork - the FINAL batch, after which the fork sits at upstream's tip and a freeze begins.
GOAL: bring upstream kunchenguid/firstmate (remote
upstream, READ-ONLY, never pushed to) into this fork's main up to waypoint 8f7b79c ("fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704)"), measured against base origin/main = c159642 (batch 9, PR #42): 55 upstream commits, 31 conflicted files, all content conflicts.MANDATED SHAPE, deliberate and not to be rewritten: exactly ONE merge commit (git merge --no-ff 8f7b79c), never a rebase, squash or cherry-pick, because the file tree is not proof of a correct batch - a flattened branch can carry identical content while losing the 55 upstream commit identities from ancestry. The waypoint must remain an ancestor; bin/fm-pr-merge.sh asserts this at merge time and the batch is merged with --merge. This run is started with --skip=rebase on purpose: rebasing would destroy the merge shape the batch exists to preserve. Do not rebase, flatten, or linearise this branch. If origin/main moves it is folded in with a merge, never a rebase. git stash was never used and must not be (pooled worktrees share one object store, so refs/stash is a single fleet-wide stack).
THE CAPTAIN'S CONFLICT RULE (verbatim, 2026-09-05): "if there is something upstream and our code both worked on and fixed but approaches differ, take the upstream change, however, if our change is better than upstream, file a PR. for other items, take the upstream changes." Applied as three outcomes per conflicted file, each recorded in the merge commit body: (1) same problem/different approach -> take upstream's, but first PROVE upstream covers what ours covered by running the fork's own test against the merged file; where upstream covers less, keep ours FOR THAT GAP ONLY; (2) ours demonstrably better -> still take upstream's on our main, and list the file, fork commit and MEASURED evidence under "Upstream PR candidates" in the completion report (firstmate opens those PRs, not this worker; the fork does not keep a private better version); (3) anything else -> take upstream's. So a reviewer seeing upstream's approach replace a fork approach in a conflicted file is seeing the rule applied deliberately, not an accident, and reverting to the fork's version would violate it.
Only fork-side fixes that upstream's tail needs to be green on macOS Bash 3.2 are retained as fork commits. None proved necessary: the batch's parse sweep is clean under stock /bin/bash 3.2.57 over all 144 linted files, and the fork's one inherited macOS fix (506a63d, empty-array guards in bin/fm-public-followup.sh) was RETIRED because upstream's a5f3cbe is a strict superset - proven non-vacuously by stripping the guards and getting
PF_REGISTRY_LOCK_IDS[@]: unbound variable, then restoring.CARRIED-FORWARD REQUIREMENTS FROM BATCH 9, all delivered: (a) the fork's ship contract (completion report, CI probe, no-git-stash,${BASHPID:-$ $} cannot separate sibling subshells on Bash 3.2; upstream's kunchenguid#3696/kunchenguid#3697 change nothing. (d) Upstream's kunchenguid#3681 was taken and proven to cover less on the fork's own shape, so bin/fm-crew-state.sh keeps nm_run_parked_at_gate_binds_worktree for that one measured gap (without it: "not ok - a parked run on an unpushed pipeline head reports parked ... state: working"), because upstream's exemption needs branch_sync.state=pipeline_owned and a run parked at awaiting_approval/fix_review emits no branch_sync block. (e) The fork's tmux server pinning survives with all 8 fm-send-shell-pane-refusal cases passing.
done:reserved for the ready signal) is re-homed into bin/fm-dod-lib.sh, which upstream's c7fdef9 made the single owner for both fm-brief.sh and fm-promote.sh - never back into fm-brief.sh alone - and fm_dod_block deliberately takes four arguments instead of upstream's two so a new caller that omits them fails loudly rather than rendering an empty completion-report or CI-probe path; verified on REAL generated output that a briefed worker and a promoted scout receive a byte-identical contract apart from the task id. (b) The macOS lane counts were re-measured on the merged tree rather than assumed: 16 and 49, not the 45 predicted. (c) tests/fm-pending-reply.test.sh's concurrent-resolver case was investigated and NOT papered over with a skip: it fails 3/3 here and 3/3 at origin/main, a standalone repro shows "closes appended: 2" instead of 1, and a live-holder probe confirmsTHE BATCH'S MOST IMPORTANT RESOLUTION: bin/fm-watch.sh's pause_state_class merged with NO conflict marker - both sides had rewritten it in different regions, so git spliced them and silently kept the fork's ORDERING, discarding upstream kunchenguid#3532's deliberate move of the agent-liveness gate ahead of the recorded-flag/fresh-recheck early return. Upstream's own regression caught it. Upstream's version was taken verbatim and proven not to cost the fork the behaviour it protected: under upstream's ordering the fork's case still absorbs the lane, still advances the stale suppressor, still retains the recorded pause flag and still queues nothing; only the internal .paused-rechecked- marker differs, because the liveness gate now promotes to
noneand clears it while absorption comes from the declaration-keyed re-surface throttle. That one implementation-coupled assertion was replaced by a comment, which is deliberate and correct under this repo's guideline that tests must exercise behaviour through a public interface and never assert implementation internals.VERIFICATION ALREADY PERFORMED, so re-doing it is not a finding: bin/fm-lint.sh 0; bin/fm-doc-audience-check.sh 0 (surfaces=96 local_links=333); bin/fm-test-run.sh --check-coverage ok (total=192); /bin/bash 3.2.57 -n over all 144 linted files, parsed=144 failures=0, each file in its own invocation because
bash -n a b cparses only the first argument; git diff 8f7b79c HEAD -- .github/workflows/ empty (the fork makes zero CI changes and must not start now); progress-steal-guard grep = 1; BASHPID capability gate = 1; and four full local portable-serial lanes under /bin/bash on macOS with a 1200s per-script bound (CI does not run that lane on macOS). Final lane: 155 scripts, 2 failed, 29 gate-skipped. tests/fm-backend-herdr-focus-flash-e2e.test.sh was deliberately excluded and never run because upstream reclassified it into the real-herdr-gated family.OPERATIONAL FACT FOR THE TEST STEP, please read before choosing what to run: this repo's suite is extremely slow on this machine and CANNOT be run inside a bounded step. The portable-serial lane is 155 scripts and took 107, 122, 278 and 332 minutes on four separate runs. Individual scripts are long too: tests/fm-watch-triage.test.sh alone is 10-11 minutes, tests/fm-public-followup.test.sh about 5, tests/fm-remote-secondmate-lifecycle-e2e.test.sh about 7. A previous attempt at this step launched two of those in the background and then idled waiting for them, exhausting its 30-minute budget without finishing. Please do NOT start the full lane or those long scripts. Fast, high-value checks that finish in seconds to two minutes and actually discriminate:
git log --oneline --merges -1,git merge-base --is-ancestor 8f7b79c7 HEAD,git rev-list --count c1596426..8f7b79c7,git diff 8f7b79c7 HEAD -- .github/workflows/,bin/fm-lint.sh,bin/fm-doc-audience-check.sh,bin/fm-test-run.sh --check-coverage,bin/fm-ci-probe.sh, a/bin/bash -nsweep overbin/fm-lint.sh --list-files, and short suites such as tests/fm-send-shell-pane-refusal.test.sh, tests/fm-ci-probe.test.sh, tests/fm-crew-state.test.sh and tests/fm-dod-contract.test.sh if present. The four full lanes are already run and quoted above; re-running them is not needed and will not fit.THE TWO REMAINING LANE FAILURES ARE PRE-EXISTING AND NEITHER IS MERGE-EXPOSED - established by measurement, not assumed, and deliberately NOT fixed in this batch because firstmate scoped it that way ("fix only failures the merge exposed; a pre-existing process leak or any defect present on both sides of the merge goes into the completion report as a follow-up candidate, not into this batch"): (1) fm-pending-reply's concurrent-resolver case, which fails identically at origin/main; (2) fm-watcher-lock's "arm did not exit with HUP status (got 124)", which failed 3/3 inside a full 155-script lane and passed 6/6 outside one and could not be reproduced in a 6-script serial tail. Because bin/fm-watch.sh (which the arm spawns) IS heavily changed by this merge, byte-identity of the test alone was not accepted as proof; two probe roots differing only in bin/fm-watch.sh were built and the arm's HUP exit latency timed 10x each: merged watcher 0.44-1.08s, origin/main watcher 0.35-5.55s, all 20 runs returning the correct status 129, against the case's 8s budget. The merged watcher is strictly faster, so the merge is exonerated and the thin headroom is pre-existing.
FIVE UPSTREAM PR CANDIDATES are recorded with measured evidence in the completion report, which lives OUTSIDE this repository at ~/dev/firstmate/data/upstream-batch-10-to-tip/completion-report.md (data/ is gitignored fleet-private state, so its absence from the repo is expected and is not a finding). The most serious candidate: a detached fm-startup-network worker spins forever once its home directory is removed - fm_lock_acquire_wait in bin/fm-wake-lib.sh has no default ceiling, measured still spinning after 1015s - which leaked 4 orphan processes per suite run, accumulated to 17 and put this 8-core machine at load average 16.5. bin/fm-startup-network.sh is byte-identical to upstream, so under the captain's rule the fork keeps upstream's version and the fix belongs upstream; it is intentionally not patched here.
Commits 876dcdc and eebe02f cancel each other and leave no net change: the first raised tests/fm-watcher-lock.test.sh's hardcoded wait ceilings while those lane failures were still misdiagnosed as a too-tight bound, the second reverts it once measurement refuted that. They are left in history on purpose because the revert's reasoning is the record of how the cause was actually established; they are not accidental churn and should not be squashed away.
CONSTRAINTS THAT REMAIN IN FORCE: never push to the default branch; never merge a PR; ask-user findings are never this worker's to answer and must be escalated to firstmate; --yes is forbidden; the shared no-mistakes daemon must never be stopped, restarted or updated.
What Changed
kunchenguid/firstmatecommits intomainvia a single--no-ffmerge to waypoint8f7b79c7, resolving 31 conflicted files by taking upstream's approach per the fork's conflict rule, while keeping fork-specific logic only where upstream measurably covers less: the four-argumentfm_dod_blockcontract re-homed into the newbin/fm-dod-lib.sh(shared byfm-brief.shandfm-promote.sh), andnm_run_parked_at_gate_binds_worktreeretained inbin/fm-crew-state.shfor a parked-run state gap upstream's exemption doesn't cover.bin/fm-gemini-lib.sh,.agents/skills/harness-adapters/references/harness/gemini.md), an extension runtime (bin/fm-extension.mjs,bin/fm-extension-launch-barrier.mjs,docs/extension-bindings.md), quota-aware dispatch (bin/fm-quota-choose.sh,bin/fm-procevent-quota.sh), backlog transition handling (bin/fm-backlog-transition-lib.sh), secondmate restart and parent-channel support (bin/fm-secondmate-restart.sh,bin/fm-parent-channel-lib.sh,docs/secondmate-parent-channel.md), and Claude trust handling (bin/fm-claude-trust.sh), each with new test coverage (tests/fm-gemini-harness.test.sh,tests/fm-extension-binding.test.sh,tests/fm-backlog-atomicity.test.sh,tests/fm-quota-choose.test.sh,tests/fm-secondmate-restart.test.sh,tests/fm-claude-trust.test.sh, etc.).bin/fm-public-followup.shnow that upstream's version is a strict superset, and took upstream'sbin/fm-watch.shliveness-gate ordering inpause_state_class(an implementation-coupled test assertion was replaced with a comment describing the now-equivalent observable behavior).Risk Assessment
✅ Low: The sole finding from round 1 (a stale "Two"/"Three" rule-count in the DoD block text) was fixed with a minimal, correct one-word change verified against the current file; independent spot-checks of the merge shape (single --no-ff merge, waypoint ancestry, 55-commit count, zero CI diff), the most sensitive conflict resolution (fm-watch.sh's pause_state_class, confirmed byte-identical to upstream's actual tree), the retained fork-specific logic (fm-crew-state.sh's nm_run_parked_at_gate_binds_worktree, the four-arg fm_dod_block contract with both callers updated), and the honestly-investigated pending-reply test (not skipped) all corroborate the intent's claims with no contradictions found.
Testing
Confirmed the batch-10 merge retains its required single-merge-commit shape with the waypoint as an ancestor and no CI workflow changes, then proved the review-1 fix ("Two"→"Three" in bin/fm-dod-lib.sh) is correct by executing the real fm_dod_block function and a scoped, real end-to-end subset of tests/fm-brief.test.sh, plus ran three fast, unaffected regression suites (fm-ci-probe, fm-send-shell-pane-refusal, fm-crew-state) covering carried-forward batch-9 requirements — all passed with no failures or setup issues; the full 155-script portable-serial lane was intentionally not re-run per the user intent's explicit multi-hour runtime warning and its four prior completed runs.
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
Step was skipped.
🔧 **Review** - 1 issue found → auto-fixed ✅
bin/fm-dod-lib.sh:243- The merge's union of fork and upstream text in bin/fm-dod-lib.sh introduced a wrong count: the preamble at line 243 reads "Two firstmate-specific rules layer on top of that guidance:" but three bullets now follow (ask-user escalation, the --yes ban, and the fork's fm-ci-probe.sh rule appended at line 249-252). Verified directly (sed -n '243,252p' bin/fm-dod-lib.sh) and viagit diff 72ccd9e^1 72ccd9e -- bin/fm-brief.sh, which shows the fork's pre-merge version said "Three firstmate-specific rules" before this content was re-homed into fm-dod-lib.sh. No test asserts the count, so this generated text - shown verbatim to every briefed worker and promoted scout as part of the Definition-of-Done block - now miscounts its own rules. Confirmed independently by an opus-reviewer subagent and by my own direct read.🔧 Fix: Fix DoD block rule count: "Two" to "Three
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
git log --oneline --merges -1git merge-base --is-ancestor 8f7b79c77c2198a71a01082215227a64500015e3 HEADgit rev-list --count c15964260771856a9127941fe3f06ab49bb94006..8f7b79c77c2198a71a01082215227a64500015e3git diff 8f7b79c77c2198a71a01082215227a64500015e3 HEAD -- .github/workflows/Direct invocation of fm_dod_block "no-mistakes" from bin/fm-dod-lib.sh to confirm real generated output reads 'Three firstmate-specific rules' with exactly 3 bulletsScoped subset of tests/fm-brief.test.sh: test_no_mistakes_dod_wording, test_no_mistakes_ci_probe_guard_wording, test_ask_user_escalation_format (run end-to-end against bin/fm-brief.sh)bash tests/fm-ci-probe.test.shbash tests/fm-send-shell-pane-refusal.test.shbash tests/fm-crew-state.test.shgit status --short (clean worktree after cleanup)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
Firstmate merge reasoning
Merged autonomously under the standing yolo posture for this repository, through the upstream-batch path (
--merge, waypoint8f7b79c77c2198a71a01082215227a64500015e3asserted against the PR head at merge time).Why it is safe to merge:
Residuals accepted, not blockers: upstream has moved five commits past the waypoint, which become batch 11 immediately after this lands; the two known serial-lane failures are carried into that batch's scope check.