merge: sync upstream/main (9f8ad95a) into the fork, keeping every fork change - #34
Merged
Merged
Conversation
* fix(bin): keep typed Lavish comments when an element is also annotated read preferred element text over prompt, so an annotate-and-comment item dropped the captain's words. Surface prompt as its own field. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Filter non-comment prompts from Lavish reader output * no-mistakes(document): Clarify Lavish comment presentation contract * no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure * no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure * fix(bin): always emit Lavish comments and use real annotation fixtures Stop inferring comment provenance from prompt==text. Real pure annotations have no prompt, so always-emit does not duplicate. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420) * Fix public-followup register crashing on empty lock arrays under bash 3.2. bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock. The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there. * no-mistakes(document): Document stock Bash registration coverage * no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5 * no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment * no-mistakes(review): Clear inherited supervision model from Herdr launches * no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent A Discord support thread's screenshots were never seen by the agent handling the mention. The relay delivered them and the poll stashed them: the reporter's images arrived on the `thread_starter` entry of `in_reply_to_chain` while the mention's own media list was empty. The gap was in the responder's playbook, which enumerated a fixed field list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so made every other field, attachments included, invisible. Fix it where the gap is, in prose: - Read the complete payload object rather than a fixed field list, so media and later relay fields are never skipped again. - Fetch and view attached media with the agent's own tools, on the mention and on every chain entry, and call out the common shape where only the thread starter carries the screenshots. - Restrict those fetches to known-good platform media hosts over https (Discord: cdn.discordapp.com, media.discordapp.net, images-ext-1.discordapp.net, images-ext-2.discordapp.net; X: pbs.twimg.com, video.twimg.com), report a blocked host instead of working around it, and treat everything fetched as untrusted public input on the same terms as the surrounding thread text. The poll stays out of it and downloads nothing, so no third-party bytes are pulled on the polling path. The new test pins the contract the playbook depends on: a mention in the incident's shape, with an empty top-level media list and screenshots on the thread starter, must reach the inbox with the payload intact and its media URLs unfetched. * no-mistakes(review): Preserve media authority and enforce poll-only fetching * no-mistakes(document): Clarify Relay attachment safety prose
) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home A public-followup loop bound to a REMOTE secondmate could never be closed. `clear_public_followup_link` (bin/fm-public-followup.sh:701) required an absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote route has no local path on this machine, so registration records that field empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so `retire` died with "could not clear the legacy X link ... retained for reconciliation" forever, and `deliver` posted the public reply and then stranded the loop at `posted`. `--force` never covered that step. The clear now goes to the remote home over that route's SSH transport, running `fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided from `data/secondmates.md` before any local path is consulted, so a same-named local directory can never stand in for a remote home, and registrations already on disk retire without needing a new field. `fm-on.sh` passes ssh's status through, so 255 stays the established "delivered but completion unknown" result this codebase already reconciles: the close is refused, the registration and the remote link are left exactly as they were, and the message names the unknown completion instead of claiming a definite failure. Local secondmate and `main` work homes are untouched, and `--force` still governs only the unresolved-obligation refusal. Three regression cases drive a remote route end to end, faking only the ssh binary at the FM_SSH_BIN seam and then running the real remote entrypoint against a local checkout, so the clear that must reach the remote home actually happens there. * no-mistakes(review): Guard remote link clears by request identity * no-mistakes(review): Fail guarded clears on unreadable remote state * no-mistakes(review): Reject guarded clears on non-writable remote state * no-mistakes(review): Allow no-link retirement in non-writable remote state * no-mistakes(document): Correct public-followup verification guarantee count * no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh * no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint * no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks * fix(relay): bound the guarded remote link clear so it refuses instead of hanging The guarded clear checks that the remote state directory is writable before taking the metadata lock, but that check cannot close the window: the parent can turn non-writable between the check and lock creation, and a lock held by a live holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever and `deliver` or `retire` wedged with nothing reported, instead of returning the retained-for-reconciliation refusal the guard exists to produce. This path runs unattended over the secondmate transport, where a wedge is worse than either outcome the guard defines. The guarded clear now acquires through `fm_lock_acquire_wait_bounded` (FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through the existing failure path. Unguarded local callers keep the ordinary unbounded wait, so local behavior is unchanged. The bounded primitive's header no longer claims presentation-only scope, since this is a second authorized caller; nothing else in the shared lock infrastructure changed. The regression holds the metadata lock with a genuinely live process while leaving the state directory writable, so the refusal can only come from the bound and never from the writability precondition. Against the unbounded wait it does not terminate at all; with the bound it refuses, retains the registration, writes no receipt, and leaves the remote link untouched. * no-mistakes(review): Harden lock-timeout regression with independent deadline * no-mistakes(review): Restore no-op guarded clears on read-only state * no-mistakes(document): Clarify remote public-followup cleanup contract
) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics
…#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing
…nchenguid#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
…enguid#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (kunchenguid#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
…er (kunchenguid#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
…unchenguid#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
…chenguid#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy.
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.
bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.
Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.
The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.
Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.
Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np
* no-mistakes(review): Serialize captain holds and fix backend-aware listing
* no-mistakes(document): Update captain-call retention documentation
* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
…uid#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes kunchenguid#3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization
) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
…guid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained.
…chenguid#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one.
…chenguid#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes kunchenguid#3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task.
- fm-pr-merge: write pr_base= ahead of the pr= identity block, which upstream's metadata parser now requires to be the file's tail, and source fm-wake-lib after the state-directory check as upstream does. - fm-captain-hold: apply a retained --pr/--report artifact without the attribution note (tasks-axi update has no --note) and carry that note into the done close instead. - tests: fork pr-merge cases assert upstream's gh pr merge line; fork herdr classifier fixtures now answer upstream's process-table read; the brief test checks upstream's Task subsections in place of the retired Intent line.
…contracts - fm-free-lane-slim-prompt-live-e2e opens with the shared fm_live_gate (default-on: it spends no tokens) instead of a private env check. - fm-quota-pools and fm-usage-history build their board through a lavish-axi stub that reports an open session, which the board now requires. - fm-bearings-board-render's fork payloads carry the Underway name field the validator now requires.
…recheck default - fm-watch-triage: upstream's live-parked-worker churn case asserted that the first sight surfaces once; this fork absorbs a declared wait from first sight (PR #12), so the case now pins that rule, the replacement-declaration absorb, and the recheck at the window's end measured from the status line. - fm-watch-thresholds and docs/configuration.md follow upstream's four-hour FM_PAUSE_RESURFACE_SECS default.
…eclared-wait case The declaration reaches handle_paused_stale before the working-verdict deferral under fork PR #12, so the published age reads 'paused Ns' and an elapsed 'until' is rechecked as elapsed rather than wedge-escalated.
…asured Durations come from this sync's local serial-lane run; the shards doc says so and points at the first green fork CI run for the refresh.
… CPU Every spawn started a bridge-owner watcher even when the backend could not name a real pane pid (a tmux test double answers 'fakepane'), and each watcher walked the whole process table in a shell loop every five seconds for up to five minutes. Across a test suite that spawns dozens of tasks the watchers starved the work under test: upstream's fm-control-relaunch race case reached its two-second checkpoint budget 14 seconds late on a throttled Linux host, while upstream alone stayed inside it. The watcher now starts only for a numeric pane pid and pre-filters the table with one grep, so an idle watcher costs one ps and one grep per pass. tests: the Herdr presentation e2e's teardown helper captured its status after an if statement, which reads 0 when the condition fails, so a teardown refused on upstream's new project slot lock (contended by its concurrent sibling) was reported as success and its projected workspace survived into the next focus wave.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Merge upstream
kunchenguid/firstmatemain at9f8ad95a(163 commits since the shared base4ad8cba, 2026-08-31) into this fork's main ate59d14b6, as a real merge commit so later syncs fast-forward, while keeping every one of the fork's 34 commits (fork PRs #1 to #33).Captain ruling "merge" on 2026-09-16.
Delivery-contract deviation, recorded by firstmate at intake: this ships direct-PR rather than the repo's usual no-mistakes pipeline because a 163-commit upstream merge cannot be reviewed as a diff under the 500-line rule; upstream's own review already covered its commits, and the fork's own CI plus the verifications below are the gate.
The "PR must be raised via no-mistakes" check (
.github/workflows/no-mistakes-required.yml) is red on purpose: it verifies a no-mistakes pipeline attestation and has no waiver, and this delivery is direct-PR under the deviation above. Landing it therefore needs the attended--allow-red "PR must be raised via no-mistakes"onbin/fm-pr-merge.sh, which is the captain's call, not a standing one.Rollback: the pre-merge main commit is
e59d14b(e59d14b6).Do not merge while any worker is running: running homes fast-forward onto main.
Land with the "Create a merge commit" method only (
bin/fm-pr-merge.sh ... -- --merge, never--squashor--rebase): a squash would collapse upstream's 163 commits into one and the next sync would conflict on all of it again, while a merge commit keeps the shared history so future syncs fast-forward.Branch shape
484b64eamerge commit (43 conflicted files resolved, table below).3600435fport fork behaviours onto upstream's restructured contracts (pr_base placement, captain-hold attribution note, test fixture ports).41df23d6,628e88e7,be61f793,15d1be49, plus one helper-dedupe commit: fork and upstream test suites adapted to each other's contracts, and timing hints for the 37 serial scripts upstream's new coverage guard found unmeasured.757dcbe0fix found by the PR's own CI (both jobs reproduced in an Ubuntu container here, see below): the fork's browser-bridge owner watcher now starts only for a real numeric pane pid and pre-filters the process table, and the fork's Herdr presentation e2e teardown helper captures its exit status correctly.Acceptance
git merge-base HEAD upstream/main=9f8ad95adca24ed33134d3459cbb0acbbbb7a8d6=git rev-parse upstream/main.757dcbe0: Lint, Test coverage guard, portable parallel 1-2, portable serial 1-6, Behavior tests (Herdr), Stock macOS Bash, Repo invariants, and timing aggregate all pass (serial 1 first ended CANCELLED at its 30-minute cap on a loaded runner, wheretests/fm-watch-triage.test.shalone took 25 minutes against a 4.4-minute hint; the job rerun passed in 20 minutes). The only red check is the by-design no-mistakes attestation described above. Local commands and results are under "Tests run".Conflict resolutions (43 files)
AGENTS.mdconfig/pi-noderow; took upstream's rewrittenbacklog-backend,backend, andcalmrows (Calm is now Pi plus Claude Code mod).docs/configuration.mdFM_BRIDGE_*and upstream'sFM_BACKLOG_ROW_TIMEOUT_SECS.bin/fm-spawn.sh__CLAUDEPERMFLAG__, and originless-worktree launches. Took upstream's file and re-applied every fork change onto it:--base(parser, validation,freshen_spawn_worktree_base,base=in meta and relaunch), free-tier lane wiring (__FREELANE__,free_lane_for_model,free_lane_preflight),config/pi-node(__PINODE__), the Herdr recovery lock budget, and the chrome-bridge owner watcher. Upstream already hadRAW_LAUNCH, so the fork's copy was dropped.bin/fm-teardown.sh--legacy-record, and endpoint-close refusals; the fork added--release-shared-record, the same-home shared-worktree preflight, the browser-bridge sweep, and the attribution note. Took upstream's file and ported the fork onto it: the same-home literal-path preflight still runs first and prints the release flag;--release-shared-recordrides upstream's reassigned-slot skip path (teardown_owns_worktree) instead of a second set of guards; the browser-bridge family is excluded from the generic reap and swept only after upstream's endpoint close is confirmed dead/missing; attribution notes still land on every close. The fork'sworktree_owned_by_taskheuristic is superseded by upstream's claim-based ownership.bin/fm-brief.sh## Captain's intent/## Firstmate specTask subsections (they replace the fork'sIntent:line for ship and scout; the secondmate charter keeps it); kept fork's--base,--ui,--env-file, evidence rules 8-10, andfm_shell_quote; kept upstream'sASK_USER_BLOCK,RULE1, and the expanded daemon rule 7.bin/fm-dod-lib.shbin/fm-pr-merge.sh--expect-baseplus upstream's--attended-overrideand--allow-red; kept fork's required-checks and base-recording header text; kept upstream'srequire_released_captain_holdand the fork'srecord_pr_baseafter it; dropped the duplicatefm-wake-libsource. Follow-up commit:pr_base=is written ahead of thepr=block because upstream's identity parser requires that block to be the file's tail, andfm-wake-libis sourced after the state-directory check, as upstream does.bin/fm-watch.shpaused:wait outranks a working run-step and a live pane (PR #12), withfm_watch_threshold(PR #27); kept upstream'suntil <UTC>condition-aware recheck text and the captain-hold bound branch. Dropped the legacy.paused-recheckedwrite nothing reads any more.bin/fm-classify-lib.shstatus_paused_until; kept fork'sfm_watch_thresholdresolution note and the evidence-only finding classifier.bin/backends/herdr.shstale-agentprocess-state classifier (pane process-infoplus the real process table, nested-shell aware). Removed the fork's now-unusedfm_backend_herdr_pane_agent_gonewrapper; the fork's foreground-group-leader anchoring infm_backend_herdr_pane_foreground_shell_pidstays for the idle-shell proof.bin/fm-afk-return.shlocaldeclarations.bin/fm-backlog-transition-lib.shmode=closeandmode_countlocals.bin/fm-config-inherit-lib.shlaunch-env-allowlist claude-permission-modeplus fork'swatch-thresholds.bin/fm-fleet-snapshot.shhold_set/local_noteplus the fork's attribution-note filter onbody_excerpt.bin/fm-promote.shpromote_delivery_contract; inside it,fm_dod_blockreceives the fork's meta path argument.bin/fm-test-run.shbin/fm-bearings-board.sh,bin/fm-bearings-snapshot.shname_marker,optional_subject,optional_filed,as_gate, hold buckets plus fork'soptional_bool("idea"),idea_item, pool/usage validators and thekind != ideagate filter..github/workflows/ci.yml/bin/bash3.2.docs/architecture.mduntilsentences; upstream's spawn paragraph with a--baseclause; upstream's dod-lib paragraph plus fork's attribution sentence; upstream's slot-ownership paragraph plus the fork's release-flag sentence.docs/configuration.md(env),docs/documentation-audiences.json,docs/scripts.mddocs/fm-test-portable-shards.mddocs/herdr-backend.mddocs/turnend-guard.mddocs/verification/runtime-backends.md.agents/skills/afk/SKILL.md.agents/skills/bearings/SKILL.md,assets/board-template.html.agents/skills/secondmate-provisioning/SKILL.mdfm_config_inherit_items(the fork's inline list is now owned by that lib).tests/lib.shtests/assets/board-render-harness.mjstests/fm-backend-herdr.test.shtests/fm-bearings-board.test.sh,tests/fm-bearings-board-render.test.sh,tests/fm-bearings-snapshot.test.shmap(.title).tests/fm-bootstrap-network-parallel.test.shtests/fm-brief.test.shtests/fm-control-herdr-smoke.test.shtests/fm-pr-merge.test.shgh pr merge --match-head-commitline and always-read head rollup.tests/fm-secondmate-reconcile.test.shtests/fm-spawn-pool-base-freshen.test.sh--basecases.tests/fm-watch-triage.test.shtest_exited_declared_pause_is_bounded_but_live_gate_surfacescall replaced by the fork'stest_declared_wait_is_bounded_for_exited_and_live_agents, and upstream'stest_paused_authoritative_working_preserves_wedge_timernot added, because both assert the live-pane override the fork's PR #12 removed. Upstream's other new triage tests are included.Three conflicts could not keep both behaviours literally and are the deliberate fork-over-upstream calls (all in the watcher's stale triage, fork PR #12, captain-merged):
bin/fm-watch.sh'spause_state_class(a declared wait outranks a live pane and a working run-step), and the two upstream triage cases that asserted the opposite (tests/fm-watch-triage.test.sh, see the table). Two fork helpers were dropped as superseded rather than duplicated: the fork'sfm_backend_herdr_pane_agent_gone(upstream'sstale-agentprocess-state classifier covers it, nested-shell aware) and the fork'sworktree_owned_by_taskteardown heuristic (upstream's slot-owner claims cover it).Upstream features that arrived
.claude/mods/firstmate-calm,config/calmshared with Pi;docs/configuration.md"Calm preference").config/claude-permission-mode(bypass orauto;docs/configuration.md"Claude permission mode") andconfig/launch-env-allowlist("Worker launch environment"), both inherited into secondmate homes.bin/fm-mail.sh,check: mail <uid>wakes;docs/configuration.md"Mail plane").--legacy-record, endpoint-close refusals (fix(bin): refuse teardown when a task's endpoint close fails kunchenguid/firstmate#4510), slot-owner claims written by spawn and honoured by teardown, and cross-home pool-slot collision checks.--attended-override,--allow-red <check>, away-posture merge grants and holds, red-check supersession rules, and merge-poll preservation across volume remounts (fix(bin): preserve PR merge polls across volume remounts kunchenguid/firstmate#4656).--effort ultranative validation; originless-worktree launches;FM_TASK_IDmarker refusing the behaviour suite in a primary checkout.## Captain's intent/## Firstmate specTask subsections and the no-mistakes--intentcontract (bin/fm-dod-lib.sh), validation-round pause teaching (fix(bin): teach validation-round pauses in generated briefs kunchenguid/firstmate#2752), background-and-poll drive-call guidance.paused: ... until <UTC>condition-aware rechecks, four-hour recheck default, quiet mode (/quiet), away-posture record (state/.afk-contract) and mandate clauses.--check-coverageunhinted-share guard, sharedfm_live_gatefor live guards, 30-minute serial cap, bash 3.2 public-followup regression.Fork features verified present
| Required checks (bin/fm-pr-merge.sh reader) |
bin/fm-pr-merge.sh:1246|| Required checks (doc) |
docs/configuration.md:137|| Env mirrors (bin/fm-env-sync.sh) |
bin/fm-env-sync.sh:2|| Env mirrors (doc) |
docs/configuration.md:526|| Away-mode evidence-only self-handling (classifier) |
bin/fm-classify-lib.sh:327|| Away-mode evidence-only self-handling (daemon) |
bin/fm-supervise-daemon.sh:394|| Away-mode evidence-only self-handling (skill) |
.agents/skills/afk/SKILL.md:87|| Brief acceptance evidence rules 8-9 |
bin/fm-brief.sh:729|| Context budget nudge (turn-end guard) |
bin/fm-turnend-guard.sh:215|| Context budget nudge (doc) |
docs/configuration.md:301|| Task base branches (spawn --base) |
bin/fm-spawn.sh:646|| Task base branches (base recorded in meta) |
bin/fm-spawn.sh:4328|| Task base branches (brief --base) |
bin/fm-brief.sh:230|| Task base branches (pr-merge --expect-base) |
bin/fm-pr-merge.sh:193|| Integration base branches (doc) |
docs/configuration.md:127|| Watcher thresholds (resolver) |
bin/fm-classify-lib.sh:168|| Watcher thresholds (watcher use) |
bin/fm-watch.sh:296|| Watcher thresholds (doc) |
docs/configuration.md:248|| Brief lean-out (cheap pre-flight under no-mistakes) |
bin/fm-dod-lib.sh:325|| Teardown shared-worktree refusal and --release-shared-record |
bin/fm-teardown.sh:338|| Teardown release rides the reassigned-slot skip |
bin/fm-teardown.sh:2383|| Herdr presentation recovery lock budget |
bin/fm-spawn.sh:1101|| Herdr treehouse-subshell foreground anchoring |
bin/backends/herdr.sh:1400|| Free-tier lane wiring (spawn) |
bin/fm-spawn.sh:1669|| Free-tier lane preflight (spawn) |
bin/fm-spawn.sh:1701|| Free-tier repository allowlist (doc) |
docs/configuration.md:568|| Free-tier guard |
bin/fm-free-tier-guard.sh:3|| Free-lane runner |
bin/fm-free-lane-run.sh:3|| Quota history sampler |
bin/fm-quota-sampler.sh:7|| Quota history sampler (doc) |
docs/configuration.md:624|| ChatGPT quota reader |
bin/fm-gpt-quota.sh:2|| Browser bridge sweep (teardown Fix 4) |
bin/fm-teardown.sh:3650|| Browser bridge sweep (spawn owner watcher) |
bin/fm-spawn.sh:4633|| Browser bridge sweep (session start summary) |
bin/fm-session-start.sh:700|| Browser bridge sweep (doc) |
docs/configuration.md:632|| Pi node interpreter (spawn) |
bin/fm-spawn.sh:2143|| Pi node interpreter (doc) |
docs/configuration.md:470|| no-mistakes-prod-only mode |
AGENTS.md:317|| Model attribution note on backlog close |
bin/fm-teardown.sh:1508|| Model scorecard |
bin/fm-model-scorecard.sh:2|| Declared pause outranks run-step and live pane |
bin/fm-watch.sh:1273|| Untrusted embedded instructions guard |
bin/fm-brief.sh:426|| Artificial-load rule in briefs |
bin/fm-brief.sh:421|| Crewmate orientation (Toolkit section) |
bin/fm-brief.sh:383|| Stow continuation prompt |
.agents/skills/stow/SKILL.md:274|| Bearings ideas capture |
bin/fm-bearings-board.sh:288|| Bearings usage panel |
bin/fm-bearings-board.sh:272|| Bearings pools gauge |
bin/fm-bearings-board.sh:243|| Away-mode daemon accepted by turn-end guard |
bin/fm-turnend-guard.sh:39|| Hardened Herdr recovery (shell-only panes read dead) |
bin/backends/herdr.sh:2376|Tests run (local, macOS arm64, Bash 5 unless stated; all on the final branch head unless stated)
Everything
.github/workflows/ci.ymlruns was run here. The dedicated Herdr lane drives real Herdr sessions and this task was briefed without--herdr-lab, so it was not run on this host; it was run instead inside a throwaway Ubuntu 24.04 container (Docker, Herdr 0.7.4 and Treehouse v2.0.1 installed by the repo's own pinned installers, its ownherdr server), which cannot reach the captain's live Herdr.CI=true bin/fm-lint.shbin/*.sh bin/backends/*.sh tests/*.shwith extended analysis; actionlint 1.7.12, 3 workflow files validbin/fm-test-run.sh --check-coverageFM_TEST_COVERAGE ok total=229 parallel=24 parallel_unhinted=0 serial=189 serial_shards=6 serial_unhinted=0 herdr=16bin/fm-test-run.sh --lane portable-parallel-1 --exclude-family real-herdr-gated --json p1.jsonbin/fm-test-run.sh --lane portable-parallel-2 --exclude-family real-herdr-gated --json p2.jsonbin/fm-test-run.sh --lane portable-serial --exclude-family real-herdr-gated --json ser.json/bin/bash -noverbin/fm-lint.sh --list-files;/bin/bash tests/fm-fleet-snapshot-view.test.sh;/bin/bash tests/fm-bearings-snapshot.test.sh;FM_TEST_ONLY=test_first_register_succeeds_with_empty_lock_list_under_bash32 /bin/bash tests/fm-public-followup.test.shCLAUDE.mdis the canonical pointer,.claude/skills->../.agents/skills, no personal fleet paths trackedbin/fm-doc-audience-check.shok surfaces=104 local_links=437Serial-lane first-pass failures and what each was:
tests/fm-bearings-board-render.test.shname41df23d6); rerun: 32 ok, exit 0tests/fm-live-gate.test.sh41df23d6): opens withfm_live_gate default-on; rerun 10 ok, and the guard itself now runs for real here against pi 0.84.3, 3 oktests/fm-quota-pools.test.sh,tests/fm-usage-history.test.sh41df23d6); reruns 7 ok and 10 oktests/fm-watch-thresholds.test.sh628e88e7); rerun 7 oktests/fm-watch-triage.test.sh628e88e7,be61f793); full suite rerun 116 ok, exit 0tests/fm-public-followup.test.shtests/fm-remote-secondmate-lifecycle-e2e.test.sh.claude/settings.local.json, into the fake remote root, which then reads "dirty"ALL TESTS PASSED); passes on a pristineupstream/maincheckout here too. Note: upstream's own main CI at9f8ad95ahad this script red in serial shard 2 (run 35139828921), so watch it on this PR's CItests/fm-backlog-atomicity.test.shtimeout, absent on macOStests/fm-remote-doctor.test.sh,tests/fm-remote-herdr-guard.test.sh,tests/fm-secondmate-liveness.test.shupstream/maincheckout on this Mac (platformjq, a realherdranswering the status probe)First CI run of this PR (run 35174052233) found two red jobs, both reproduced and fixed in
757dcbe0:tests/fm-backend-herdr-presentation-e2e.test.sh: "focus wave 1 lost stable contiguous ordering" - a torn-down worker's└ order-aprojected workspace was still listedorder-b's teardown printingREFUSED: another Treehouse slot allocation or return is in progress(upstream's new project slot lock, contended by its concurrent sibling teardown) yet exiting 0: the fork'steardown_tasktest helper (PR #25) readrc=$?after anifwhose condition failed, which is theifstatement's own status, 0. The refused teardown was reported as success, so the test's retry-on-contention never ran and the workspace survived into the next wave. On the fork alone no lock existed to contend.... || rc=$?). Container rerun of the wholereal-herdr-gatedfamily: presentation e2e green (all focus waves), every other Herdr script green excepttests/fm-afk-inject-herdr-e2e.test.sh, which fails identically on a pristineupstream/maincheckout in the same container and passes on CI, so it is a container artefact.tests/fm-control-relaunch.test.sh: "relaunch did not reach its pre-publication endpoint check" (a 2-second budget), twicedocker run --cpus=0.5: merged tree red, pristine upstream green. Timestampedset -xtraces of therl30relaunch showed the same trace-line counts on both trees but ~50x slower per line on the merged tree (spawn reached its first pane-cwd read at +15.2 s versus +0.3 s).psinside the container showed the cause: one fork bridge-owner watcher per earlier spawn (bin/fm-chrome-bridge-sweep.sh --watch-owner, started even though the tmux double answersfakepanefor the pane pid), each alive up to five minutes and walking the whole process table in a shell loop every five seconds at ~2% CPU each, starving the case under test late in the suite.grep. Throttled container rerun:tests/fm-control-relaunch.test.shexit 0. Local:tests/fm-chrome-bridge-sweep.test.sh9 ok,tests/fm-teardown.test.sh89 ok.Suites touched by the port fixes, rerun on the final head:
tests/fm-pr-merge.test.sh109 ok,tests/fm-captain-hold-lifecycle.test.shexit 0,tests/fm-backend-herdr.test.sh216 ok,tests/fm-test-run.test.sh39 ok,tests/fm-turnend-guard.test.sh92 ok,tests/fm-brief.test.sh(in parallel lane 1, green).Hand checks (step 4 of the brief):
bin/fm-brief.sh --helpexit 0 and lists--base,--ui,--env-file;bin/fm-spawn.sh --helpexit 0 and documents--base;bin/fm-session-start.sh --helpexit 0;bin/fm-context-budget.sh --helpexit 0 and--percentruns (exit 0);bin/fm-env-sync.sh --helpexit 0.bin/fm-pr-merge.shhas no--helpflag on either side (it never did;--helpis refused as an invalid request); its header documents required checks atbin/fm-pr-merge.sh:124and--expect-baseatbin/fm-pr-merge.sh:112.Red-then-green evidence for the two behaviour ports:
tests/fm-pr-merge.test.shcaseexpect-base-matchwas red (merge authority could not be persisted, becausepr_base=sat after thepr=block upstream's parser now requires last) and is green after3600435f;tests/fm-captain-hold-lifecycle.test.shcase "the captain could not answer before cleanup replay" was red (tasks-axi updaterejected the fork's--note) and is green after3600435f.Follow-ups (not done here)
docs/fm-test-portable-shards.md).tests/fm-remote-secondmate-lifecycle-e2e.test.shwas red on upstream's own main at the merged commit; if it is red on this PR's CI, that is inherited, not introduced.Built by: claude/opus at high
🤖 Generated with Claude Code