Skip to content

feat: absorb upstream supervision and fleet updates - #21

Merged
Mauryanx merged 320 commits into
mainfrom
fm/upstream-absorb-oct
Oct 9, 2026
Merged

Mauryanx merged 320 commits into
mainfrom
fm/upstream-absorb-oct

Conversation

@Mauryanx

@Mauryanx Mauryanx commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

Land this as a merge commit, not a squash. It records upstream kunchenguid/firstmate 19fcbbde as a real parent (merge 5094dad5), so the next absorb starts from the right merge base; #9 was squashed, which is why this one had to re-resolve from 2026-09-17.

The full absorb write-up (overlaps and the evidence for each, fork features carried, test results, declined upstream findings) is in the first PR comment.

Behavior changes for the captain

  • Changed defaults: the supervision host is on by default for a Claude Code primary (opt out with config/supervision-host-off); AI co-author trailers are stripped from fleet-launched commits (config/keep-ai-trailers opts out); away mode acts on the captain's away words; /quiet is a statement when attended supervision runs; the merge guard refuses unreported required checks and retries a still-computing mergeability; AGENTS.md is much shorter, with detail moved into new trigger-loaded skills.
  • New: Devin worker runtime; idempotent captain-inbox capture, replies and receipts; Lavish feedback routed to the owning worker; per-rule dispatch confidence floors; quota-axi multi-account snapshots; mid-session second-mate relaunch.
  • New opt-ins: fleet ledger, pipeline spend at cleanup, per-project capacity, branch prefix and base branch, Claude/Pi account pins, worker tool exclusions, wait-without-turns, Gerrit, startup growth check.
  • Removed: nothing user-facing; the fork's CI pin of Pi 0.85.1 is dropped because upstream fixed the parity tests for Pi 1.x.
  • Every fork feature is kept: Codex account axis, voice-turn answering and live-call scoping (voice turns now also stay with main in away mode, on Pi and on the supervision host), courier iMessage pickup (feat(bin): add courier spool pickup for iMessage conversations #17-fix: harden iMessage pickup timing and binding output #19), fast phone/text turn surfacing, pooled-worktree teardown, GitHub App checks reader, brain desk, journal export timer (feat: add daily brain journal export timer #20).

Needed in this home before restart

  • Upgrade tasks-axi to at least 0.2.6 (hermes has 0.2.5; merges need it), quota-axi to at least 0.1.51 (has 0.1.36), and lavish-axi to at least 0.1.80 (has 0.1.64).
  • Decide the supervision host for this Claude Code primary: leave it on, or create config/supervision-host-off.

Intent

The captain, 2026-10-09: "the main repo has advanced a lot more. I'm thinking we should do an upstream absorb." Firstmate recommended starting it now as its own lane and landing it before the hardening rollout (so the rollout runs on the code we keep), at a quiet point; the captain said "yes start the absorb lane now". Upstream kunchenguid/firstmate has about 232 commits since our last absorb (#9, 2026-09-19, which merged upstream kunchenguid/firstmate into this fork and carried the fork's Codex account axis through typed dispatch, but landed as a squash, so upstream's commits are not ancestors of the fork); our fork carries its own features since then. Nothing of ours may be lost; prefer upstream's design wherever it already covers something we built. Smallest durable result, the way Kun builds Firstmate.

The captain, later: "I don't want to lose any of the functionality we've built."

The captain, later: "If upstream has a better way of doing it, definitely absorb upstream's. I agree we should prefer upstream's version."

What Changed

  • Add attended and away supervision hosts, durable outcome reporting, and updated AFK/quiet and Pi/Claude Calm behavior.
  • Extend dispatch and delivery with Devin workers, Claude/Pi account pins, configurable ship branches, Gerrit changes, and an opt-in fleet activity ledger.
  • Improve wake and inbox delivery, remote secondmate recovery, merge checks, and teardown safety; update operational skills, documentation, CI, and regression coverage.

Risk Assessment

🚨 High: The change can duplicate handled instructions, strand live voice calls, and remove project-memory functionality explicitly required to survive the absorb.

Testing

Corrected baseline fixture-path and dependency issues, then drove targeted runtime suites and manual CLI checks that passed the prior regressions. The remote inheritance serialization fixture failed repeatedly without establishing a live scenario result; test cleanup was fixed and verified, CLI and generated-brief evidence retained, and disposable files removed. Native harness, remote-host, forge, audio-delivery, and gate/CI checks remain untested; no UI image was captured because this phase forbids launching a real harness.

  • Live validation: ⚠️ inconclusive - 26 of 34 scenarios driven live against the product
Scenario Result Live Evidence
Acknowledge a normally announced note, then replay or announce it; both interfaces report durable acknowledgement without promising another pickup ✅ pass live round4-regressions.log:8; round4-manual-cli.log:82-109; predecessor reproduction in round4-ack-before-fix.log
Repair an unannounced saved note without duplicating it, and refuse empty bodies or unsafe request IDs ✅ pass live round4-regressions.log:6-7,16
Capture and accept committed conversation text, then publish ordered replies; crash retries preserve identity and wrong-session callers are refused ✅ pass live round4-product-tests.log:88-93; subprocess cases in tests/fm-inbox-conversation-cases.py
Publish a conversation reply only to its policy-bound text destination; incompatible or withdrawn destinations create no new reply ✅ pass live round4-product-tests.log:94-96
Drain mixed voice and ordinary wakes; saved conversation turns remain pending and appear before ordinary rows ✅ pass live round4-product-tests.log:3-9
Drain with stdout closed; return failure and retain the outcome for one successful subsequent presentation ✅ pass live round4-closed-output-cli.log
Submit a branch report; reject foreign actors, ended turns, and out-of-scope tasks while storing authorized and silent outcomes correctly ✅ pass live round4-regressions.log:26-27
Deliver an away status wake through the local supervision engine contract; the host handles it without waking main ✅ pass live round4-regressions.log:71; completed supervision-host run at line 98
Queue voice-only, mixed, or newly arriving conversation turns while away; hand them to main without starting an engine turn ✅ pass live round4-regressions.log:69
Drain an away queue before either offer boundary; keep an empty queue quiet and hand a corrupted queue to main ✅ pass live round4-regressions.log:70
Return during a supervision turn or leave an outcome unacknowledged; preserve visible outcomes for main and recover incomplete or failed local engine turns safely ✅ pass live round4-regressions.log:40-45,73-74,80-82
Write dialog through the mirror CLI and resume its feed; preserve text, exclude operational input, enforce private storage, and refuse malformed records ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-host-mirror-final.log: operational-input exclusion, owner-only storage, committed-cursor resumption, and malformed-sequence refusal
Run a registered local process-event source; capture its result once and retain it across drains until explicit handling ✅ pass live round4-regressions.log:106-108,112
Break a process-event registration, repair it, then reconcile; announce one failure episode and confirm the repaired launch ✅ pass live round4-regressions.log:162,164-165
End a detached listener's owning session; reap its descendant tree within the documented bound while leaving a healthy home's listener running ✅ pass live round4-regressions.log:223-227,234
Dispatch beyond project capacity or race for its last place; defer without allocating state and admit at most one competing spawn ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-capacity-compatible.log: capacity deferral before allocation, concurrent last-place admission, and unreadable-state refusal
Select a Codex account root; carry its canonical path into launch output and metadata, and refuse unusable roots before publication ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-codex-account-final.log: CODEX_HOME metadata, canonical aliases, malformed-root refusals, and batch forwarding
Apply a Claude or Pi worker-account pin; construct the correct launch environment and preserve the secondmate home's own pin ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-isolated-home-tests.log: Claude/Pi launch environments and secondmate-owned pin preservation
Resolve an invalid dispatch configuration; refuse it before making a network request ✅ pass live round4-product-tests.log:38-39
Resolve registered delivery posture and generate each ship brief; enforce required modes, reject incompatible Gerrit posture, and keep shell metacharacters literal ✅ pass live round4-manual-cli.log:1-35; round4-targeted-cli-contracts.log, SOURCE round4-brief-final.log; three generated brief artifacts
Enter and archive away or quiet posture; preserve the captain's words verbatim and render return information from durable records ✅ pass live round4-manual-cli.log:38-78; round4-targeted-cli-contracts.log, SOURCE round4-delivery-tests.log lines 1-57
Enable the fleet ledger for a local task lifecycle; record ordered events once, preserve status through ledger failure, and leave no ledger artifacts when disabled ✅ pass live round4-product-tests.log:57,60,63-66
Commit through installed Git hooks and override channels; strip known bot co-authors, retain human attribution, and preserve project-hook refusals ✅ pass live round4-product-tests.log:69-83
Invoke spawn, send, or teardown from the gate worktree with its environment marker absent; refuse all three without changing isolated state ✅ pass live round4-gate-refusal-cli.log
Pass invalid arguments to PR registration or teardown; reject them before side effects ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-publication-checks.log: invalid entrypoints have zero side effects
Read the memory diagnostic and vary its thresholds; report bounded utilization, return OK above physical utilization, and alarm at zero thresholds ✅ pass live round4-memory-guard-cli.log
Run the portable remote inheritance serialization check; reach the blocked write and prove a later config push wins ⏸️ untested no The prior payload did not establish a live result: the deterministic SSH fixture repeatedly timed out before the blocked write, and its concurrent convergence assertion never ran. Synthetic SSH is not…
Provision, launch, steer, and retire a secondmate over real SSH while preserving inherited configuration ⏸️ untested no Built isolated Git homes and exercised the synthetic SSH boundary. Although sshd is available, this phase explicitly prohibits launching a real secondmate or fleet worker. A dedicated remote host and…
Resolve dispatch against real model and per-account quota services, then authenticate the selected native account ⏸️ untested no Local protocol and launch checks ran with isolated account files and service fixtures. The live quota guard requires FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E=1, which this phase forbids forcing, and real acco…
Handle supervision wakes and terminal submission through real vendor engines, hooks, and native signal surfaces ⏸️ untested no The real local host and mirror commands ran against disposable fixtures. Native guards explicitly skipped for FM_SUPERVISION_HOST_LIVE_E2E, FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E, FM_HOST_MIRROR_LIVE_E…
Use the changed Pi and Claude calm surfaces; retain queued input and display operational notes in their intended visual positions ⏸️ untested no Portable extension checks cannot establish rendering in the vendor runtime. Even the credential-free Pi route starts a real agent harness, which this phase prohibits. Supply an authorized disposable n…
Publish and register work against a real forge, including Gerrit's current patch-set tree and merge result ⏸️ untested no Real registration entrypoints ran against local Git origins and forge fixtures, which do not establish network publication. No dedicated forge endpoint or credentials were supplied; gerrit-axi is also…
Answer a real voice call and deliver ordered speech or a policy-bound message to its caller ⏸️ untested no The durable capture, acceptance, publication, and recovery protocol was driven through real CLI processes. No isolated live call endpoint, audio delivery connection, or messaging credentials were avai…
Run the changed gate and CI consumers through validation, publication, and required remote checks ⏸️ untested no This assigned phase may not initialize, rerun, control, push, or execute other pipeline phases. Parsing configuration would not prove consumer behavior. The outer executor must drive those phases and…
Evidence: Inbox, supervision-host, and process-event runtime transcripts

Source: Inbox, supervision-host, and process-event runtime transcripts

FM_TEST_BEGIN 2026-10-09T16:58:41Z tests/fm-inbox.test.sh family=unclassified expected_gate_skip=none
ok - plain note, list, and wake stay on the historical human path
ok - plain note keeps exit 1 for a saved-but-unannounced failure
ok - the same request id returns the original note as a distinguishable replay
ok - a crash between recording the request id and publishing the note reuses the original id
ok - saved-but-unannounced notes are repairable without creating a second note
ok - repair and replay do not wake firstmate for an already-acknowledged note
ok - normally announced notes report durable acknowledgement on replay and announce
ok - receipts JSON is bounded by fixed bounds and discloses what it omitted
ok - notes that predate the announcement marker are unknown, not re-announced
ok - the reply channel is durable and its cursor is a strict order
ok - the reply cursor never goes backwards when the sequence counter is lost
ok - a reply without a valid sequence is reported as malformed
ok - a non-UTF-8 note does not break the receipts view
ok - readiness says unknown (or not-receivable) instead of inferring liveness from a lock
ok - empty bodies and unsafe request ids are refused
ok - a note body that opens with a double dash is queued as text
ok - drain --ack still moves the note to handled
FM_TEST_END 2026-10-09T16:59:05Z tests/fm-inbox.test.sh exit=0 duration_ms=23949 gate_skip=false
FM_TEST_BEGIN 2026-10-09T16:59:05Z tests/fm-supervision-host.test.sh family=afk expected_gate_skip=none
ok - host+hook: a successor closed before exit_to_main does not suppress the branch-outcome rewake
ok - host+hook: a successor that closes during a held engine turn does not suppress the branch-outcome rewake
ok - host+hook: a refused hand-back becomes a delivered failure notice
ok - host+hook: a refused hand-back on an announced handling marker becomes a delivered failure notice
ok - host: parked child-exit sampling uses ordinary half-second sleeps
ok - report surface: only the branch actor's current turn may report, and only on the tasks its wake names
ok - report surface: visible late outcomes queue a relay, while silent outcomes remain stored without a wake or note
ok - dispatch entry: the host reads branch eligibility, the offer rule, and the wake prompt from the Pi branch's own owner
ok - drain: BRANCH OUTCOMES runs on a Claude home by default and on another primary with the file, never with off, and never on Pi
ok - drain: captain outcomes come first, and routine overflow collapses into a count one drain clears
ok - drain: repeated captain outcomes collapse per task, and the byte cap presents only the run its acknowledgement covers
ok - drain: a long away window costs one short drain, captain outcomes collapsed per task and routine overflow counted, and nothing from it is shown again
ok - drain: the BRANCH OUTCOMES budgets count bytes, cutting multibyte summaries by whole characters in any locale
ok - drain: branch outcomes stay unread when a projection of the store fails
ok - drain: branch outcomes stay unread and the drain fails when jq is missing
ok - drain: branch outcomes stay unread when the drain cannot print them
ok - drain: a legacy backlog is presented with each outcome's age and a check-first instruction, never adopted
ok - drain: a keyed decision survives acknowledgement through a newer outcome for its task
ok - drain: an outcome carried across a switch off Pi comes back with its age, never adopted
ok - drain: an outcome nothing has shown is presented until acknowledged, and a repeated acknowledgement changes nothing
ok - drain: an outcome the host drain presented but main never acknowledged stays unprocessed across a switch to Pi
ok - drain: an outcome the host drain presented but main never acknowledged survives an outcome index repair
ok - host: an attended wake the branch may take is handled on the engine, and its routine outcome never wakes main
ok - host: an attended captain outcome wakes main once and stays in its drain until main acknowledges it
ok - host: a captain outcome recorded after the captain left waits for the return, then reaches main's drain
ok - host: a quiet record without its daemon is a present captain, so outcomes and decisions reach main
ok - host: an attended decision close stays main's exactly as the plain arm delivers it
ok - host: an off written while the host is parked sends the next attended close to main, naming the opt-out
ok - host: a main-only pass-through leaves the successor watcher running and the close undelivered for main
ok - host: an attended close whose main session cannot be identified reaches main and runs no engine turn
ok - host: a decision close accepted away whose turn starts attended still reaches main unchanged
ok - host: an attended close whose task turns main-only before its turn still reaches main unchanged
ok - host+hook: an attended main-only pass-through rewakes main and keeps its successor watcher
ok - host+hook: a captain outcome beside a quiet record rewakes the present captain with no away note
ok - host+hook: a Claude home without config/supervision-host runs the host at the default engine, and an off file restores the plain arm
ok - host+hook: a close that turns main-only at its turn rewakes main and keeps its successor watcher
ok - host+hook: the successor a close that turns main-only at its turn leaves for main survives the hook's process group teardown
ok - host+hook: failed at-turn downtime write notifies main despite a healthy successor
ok - host+hook: a successor close that lands during main's turn is delivered at the next turn end
ok - host+hook: the successor a pass-through leaves for main survives the hook's process group teardown
ok - host+hook: the next park takes over the cycle a main-only pass-through left, so one arm owns it
ok - host+hook: a park stopped mid take-over leaves the take-over to the next park
ok - host+hook: a successor that cannot be recorded is stopped, and main's next turn end arms a fresh cycle
ok - host: a primary with no verified dialog mirror keeps every attended close on main, and its away posture still runs
ok - host: each wake carries the captain's dialog since the last wake, without operational input
ok - host: the dialog mirror, its feed, and the wake file are owner-only
ok - host: dialog a turn never completed with its report (a park boundary, a stopped turn, no report) reaches the next turn
ok - host: an attended wake whose mirror is missing, cannot be read (at the feed or at the prompt), or holds a malformed entry reaches main before any engine turn, and the cursor stays put
ok - host: away voice-only, mixed, and newly arriving conversation turns reach main without an engine turn
ok - host: already-drained away wakes stay quiet at both offer boundaries, while corrupted queues reach main
ok - host: an away wake is handled on the engine through the branch contract and never reaches main
ok - host: an engine turn that records no outcome hands its durable wake to main
ok - host: a captain return during an engine turn hands that turn's outcomes to main
ok - host: silent outcomes are excluded from both captain-return handoff paths
ok - host: an early visible outcome survives more than 1,000 same-turn receipts
ok - host: an outcome lookup failure forces a visible main handoff
ok - host: an outcome recorded after the return reaches main even when its host dies at the turn's end
ok - host: a killed predecessor's engine is reaped and the next host removes its turn files
ok - host: a turn that reports but leaves its granted rows queued hands the wake to main
ok - host: a captain return during a failed engine turn still hands that turn's outcomes to main
ok - host: an engine turn whose result is incomplete hands its wake to main
ok - host: two engine errors latch the session, main keeps every away wake in the cooldown, a failed probe doubles it up to its cap, and a report recovers it silently
ok - host: an attended close in a latched session reaches main as the arm printed it and leaves the latch as it was, and a home that opted out with off never reads it
ok - host: attended, two engine errors latch the session, main keeps every close

... [5539 bytes truncated] ...

owner
not-autohandled: cross-home-src (left for the handler; still unacknowledged)
ok - cross-home stale recovery removes abandoned output from the old state directory
not-autohandled: race-src (left for the handler; still unacknowledged)
not-autohandled: race-src (left for the handler; still unacknowledged)
ok - concurrent stale-claim replacement starts exactly one runner
ok - an ambiguous leaderless group is preserved without replacement
ok - a truly dead generation with no surviving group is still safely reclaimed
ok - reconcile reclaims a dead generation whose state-root identity no longer matches
ok - retire releases a dead generation's claim instead of refusing forever
ok - a live generation is never reclaimed, drifted state root or not
ok - a reused pid never makes its surviving process group reclaimable
ok - start reclaims a reused-pid claim whose leftovers can still be tidied
ok - a launch that cannot confirm is announced once per failure episode
ok - a 64-char source id keeps the launch-failed key within the watcher's marker bound
ok - reconcile reports a launch it could not confirm instead of counting it as a start
ok - a launch that finished before confirmation looked is still reported as started
ok - a zero-padded launch confirm window is honored as base 10
ok - an unusable launch confirm window is refused by name instead of blamed on the sources
ok - a dead generation whose leftovers cannot be tidied never keeps owning its source
ok - claim replacement cannot produce a torn ownership snapshot
ok - retirement and start share one serialized lifecycle boundary
ok - detected PID reuse is refused before signalling
ok - transient identity failure preserves the live source for retry
ok - bounded home sweep preflights then retires every locally owned source
ok - home sweep leaves foreign-home claims and runners untouched
ok - home sweep refuses safely until runner identity is readable
ok - healthy runtime behavior remains registration-only
ok - registration rejects unrepresentable newline arguments
ok - nonzero exit with no output stays armed and silent
ok - oversized output is bounded rather than published whole or dropped
ok - live output stays bounded and retirement reaps the whole source group
ok - stop mismatch evidence preserves the proved-stop boundary
ok - stop unreadable evidence preserves the proved-stop boundary
ok - stop unreadable-pgid evidence preserves the proved-stop boundary
ok - stop nonleader evidence preserves the proved-stop boundary
ok - invalid output bounds fail closed
ok - the adapter derives physical identity without newline path corruption
ok - source-only homes trigger the general supervision guard
ok - the adapter classifies published poll output safely
ok - poll derives host and port from the artifact session, not ambient or configured routing
ok - quiet retries use the board session regardless of configuration changes
ok - missing or unreadable session routing preserves replies and never guesses another server
ok - the adapter owns which Lavish results end a source, and payload text cannot forge one
ok - the adapter owns which Lavish results are silent, and fails closed on everything else
ok - read distinguishes a live captain message from a session-ending message
ok - read names the message count with the same label as the message section
ok - read presents every annotation and a distinct session-ending message
ok - read never certifies rows missing declared fields as complete
ok - read keeps every annotation when the session-ending message is absent
ok - read surfaces a typed comment on an annotated element
ok - read still surfaces a typed comment that matches the element text
ok - read still presents a pure annotation with no comment
ok - read does not present choice context as a comment
ok - read still presents a pure message with no selector
ok - read distinguishes a feedback capture from an ended-with-nothing close
ok - an adapter with no silence verdict keeps announcing every result
ok - the published interfaces state the loss limitation and claim no lossless delivery
registered: floor-src (lavish)
ok - an orphaned source command obeys the launch floor during its grace window
ok - a replacement registration starts with one fresh launch floor
ok - a superseded sleeping runner cannot recreate stale pacing state
ok - post-commit pacing cleanup cannot veto registration publication
ok - a pre-reboot monotonic stamp is treated as expired
ok - an expired owner lease stops a self-relaunching source generation
ok - a recreated state path does not preserve the old runner
registered: guard-fail-src (lavish)
ok - a runner fails closed when its owner guard cannot initialize
registered: attached-src (lavish)
ok - a foreground start refreshes its lease while its caller remains attached
ok - wall-clock corrections do not alter owner lease age
registered: detached-attached-src (lavish)
ok - an attached-start keeper stops refreshing after its parent exits
ok - a detected ambiguous reused-PID group is not signalled
ok - a detached listener starts reparented, with a live descendant tree under it
ok - a listener whose owning session is gone stops itself and its whole process group
ok - reaping the listener stops the process churn under it
ok - an identical listener in a home whose session is still there is untouched
ok - retiring a source reaps its reparented listener and every descendant under it
ok - a stop the guard cannot prove is retried until the expired runner is reaped
ok - retirement escalates after TERM leaves a surviving child (absent leader)
ok - retirement escalates after TERM leaves a surviving child (zombie leader)
ok - an expired runner's guard escalates past a signal-proof child
guard phase: fresh read 6.376-6.408s, expiry 8s, two full intervals could not finish before 18.376s (deadline 17s)
guard bound: lease=7s check=6s reaped 13.4s after the last owner activity, documented bound 15s
ok - an orphaned runner is reaped within the lease plus ONE check interval
decimal interval: 08 halves to 4s and its listener started
decimal interval: 010 halves to 5s and its listener started
ok - a zero-prefixed decimal interval starts its listener and halves as decimal
ok - one unreadable read does not end a live runner
ordinary stop: start status=143 retirement=935ms sampled windows=2177/2201ms
ok - a runner exits on the ordinary stop signal instead of outliving it
ok - a group whose leader died to something else is still refused, not signalled
ok - arm reports ready only after the listener is running
ok - arm waits out a delayed listener start before reporting ready
ok - arm fails when the listener cannot claim, and leaves the registration retire refused
ok - arm waits out the confirm window before reporting that the listener is not running
ok - re-arm over a live earlier listener reports it still serving the board
ok - older compatible Lavish versions retain the poll-with-reply behavior
ok - a refused synchronous reply fails arm before source registration
ok - a refused synchronous reply leaves the same arm retryable
ok - an arm refused for ownership, round, or endpoint never posts its reply
ok - direct poll confirms a new-version reply before polling
ok - an unknown Lavish version fails arm and poll closed, keeping the staged reply
ok - arm posts and confirms the staged reply before reporting listener readiness
ok - re-arm launches the new generation once a draining earlier claim is released
ok - arm does not launch beside a stale claim whose process group is alive

all procevent tests passed
FM_TEST_END 2026-10-09T17:41:56Z tests/fm-procevent.test.sh exit=0 duration_ms=785804 gate_skip=false
FM_TEST_SUMMARY total=3 failed=0 skipped_gate=0 duration_ms=2594982
FM_TEST_SUMMARY_FAMILY family=afk count=1 duration_ms=1784856 failed=0
FM_TEST_SUMMARY_FAMILY family=standalone count=1 duration_ms=785804 failed=0
FM_TEST_SUMMARY_FAMILY family=unclassified count=1 duration_ms=23949 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-supervision-host.test.sh duration_ms=1784856
FM_TEST_SLOWEST rank=2 script=tests/fm-procevent.test.sh duration_ms=785804
FM_TEST_SLOWEST rank=3 script=tests/fm-inbox.test.sh duration_ms=23949
Evidence: Initial targeted runtime checks; account and capacity setup failures superseded by corrected runs

Source: Initial targeted runtime checks; account and capacity setup failures superseded by corrected runs

FM_TEST_BEGIN 2026-10-09T16:58:43Z tests/fm-wake-drain-voice-first.test.sh family=watcher-wake-lock expected_gate_skip=none
FM_TEST_BEGIN 2026-10-09T16:58:43Z tests/fm-wake-drain-voice-hold.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - an acknowledgement holds a voice row whose request is still saved and retires the ordinary row beside it
ok - a held voice row is presented again by the next drain
ok - the hold ends once the request is accepted and the ordinary acknowledgement retires the row
FM_TEST_END 2026-10-09T16:58:52Z tests/fm-wake-drain-voice-hold.test.sh exit=0 duration_ms=8608 gate_skip=false
ok - a voice note queued after an ordinary wake is still presented first under VOICE
ok - a voice note queued first stays first and the rows behind it keep their order
ok - a typed inbox note is an ordinary wake and prints no VOICE heading
FM_TEST_END 2026-10-09T16:58:56Z tests/fm-wake-drain-voice-first.test.sh exit=0 duration_ms=12510 gate_skip=false
FM_TEST_BEGIN 2026-10-09T16:58:56Z tests/fm-dispatch-resolve.test.sh family=standalone expected_gate_skip=none
ok - absent key is off: one stderr line, exit 0, no network call
ok - TYPESAFE_API_KEY= in .env activates the tool; environment and config overrides work
ok - clear: one rule Choice request, key on the fd header only, spendPriority argmax over every candidate
ok - never-send list withholds the request on a match or a bad list, and never prints the value
ok - rules snapshots and shell quoting preserve the profile protocol
ok - no-rule fallback, Agy, Gemini, and documented configurations resolve
ok - the Codex account axis survives typed resolution: identity, per-account quota, and carry-through
ok - per-account Codex reads cover only reachable homes and survive a stalled account
ok - reachable per-account reads share one aggregate bound and keep per-account evidence
ok - ambiguous: confidence below the fixed floor hands the decision back
ok - per-rule confidence floors fall to the most probable runner-up that clears its own floor
ok - only the brief's task sections and scout tag reach the model, with a whole-brief fallback
ok - escalate: a rule declared approval: captain never yields a profile
ok - rule floor: known shortfall falls through while unavailable evidence escalates
ok - declared provider and profile floor evidence are applied in code
ok - nonnumeric spendPriority evidence is never ranked
ok - partial and missing quota evidence remain eligible but unranked
ok - provider-wide and exact quota rows combine into one limiting candidate
ok - default: no rule matched resolves among the default profiles
ok - tie: equal spendPriority never breaks by array order
ok - no rankable candidate: the tool escalates instead of guessing
ok - native Codex binds to codex-home before default, independently of Pi accounts and row order
ok - Pi native adapters bind to codex-home with existing fallbacks and schema 5 compatibility
ok - schema 6: each candidate binds to its account row; schema 5 is unchanged
ok - quota evidence comes from one quota-axi --json read, and its failure is an error outcome
ok - API, transport, and response failures are error outcomes with exit 0
ok - configuration errors exit 2 before any network call
# all fm-dispatch-resolve tests passed
FM_TEST_END 2026-10-09T16:59:46Z tests/fm-dispatch-resolve.test.sh exit=0 duration_ms=49476 gate_skip=false
FM_TEST_BEGIN 2026-10-09T16:59:46Z tests/fm-worker-account.test.sh family=backend-dispatch expected_gate_skip=none
ok - an absent pin leaves Claude and Pi launches exactly as they were
ok - a Claude pin selects its root and sheds the credentials that would outrank it
ok - a Claude pin refuses a signed-out root even when the invoking process has a usable login and API key
ok - an ordinary Claude pin selects the default login and drops an ambient root
ok - malformed, relative, CR-terminated, empty, extra-line, missing-root, and non-file pins refuse before launch
ok - a Pi pin selects its root and passes the declared provider
ok - a Pi pin refuses unqualified, missing, undeclared, signed-out, and raw launches
ok - extension providers and a Pi without auth check fall back to an exact model-listing match
ok - a Claude pin leaves codex and Pi launches unchanged
ok - a raw Claude launch command receives the home's pin
ok - a pinned home refuses a raw Claude command that overrides the account
ok - an unpinned home keeps a raw Claude account override
not ok - a local secondmate spawn under the launching home's pin should succeed: error: secondmate home cannot be inside the firstmate repo: ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/.test-phase-tmp/fm-worker-account.4ByaQ0/secondmate/secondmate-home: expected exit 0, got 1
FM_TEST_END 2026-10-09T17:01:16Z tests/fm-worker-account.test.sh exit=1 duration_ms=90360 gate_skip=false
FM_TEST_BEGIN 2026-10-09T17:01:16Z tests/fm-fleet-ledger.test.sh family=unclassified expected_gate_skip=none
ok - flag on: dispatch, polled status lines, the local merge after its task's pending lines, and cleanup are recorded in order
ok - flag on: a PR merge is recorded once, after the task's pending status lines
ok - flag on: registering a PR records task.pr_ready with its full URL after the task's pending status lines, and the merge-time re-record adds nothing
ok - flag on: a worker's status command records its line at once, and the watcher backstop does not record it again
ok - flag on, state override outside the home: the worker's status command records its line at once
ok - flag on, relative config override: a worker running elsewhere still records its line at once
ok - append failing: the worker's status command exits nonzero and records nothing
ok - ledger failing: the worker's status line still lands exactly, quietly, and the backstop records it later
ok - flag off: the worker's status command is a plain append and leaves no ledger file, offset, or lock
ok - flag off: the whole lifecycle leaves no ledger file, offset, or lock
FM_TEST_END 2026-10-09T17:02:16Z tests/fm-fleet-ledger.test.sh exit=0 duration_ms=59539 gate_skip=false
FM_TEST_BEGIN 2026-10-09T17:02:16Z tests/fm-git-strip-ai-trailers.test.sh family=backend-dispatch expected_gate_skip=none
ok - a Cursor --trailer commit object has no AI co-author and keeps the captain identity
ok - a human Co-authored-by trailer survives next to a stripped Cursor trailer
ok - a human co-author at a vendor domain survives; only the exact bot address is stripped
ok - a hook manager resolving the pane hooks dir fails instead of displacing the strip
ok - a relaunch reinstall replaces the read-only strip dir and still strips
ok - install chains the previous commit-msg hook after stripping
ok - a relative project core.hooksPath resolves against the worktree and still runs
ok - an inherited GIT_CONFIG hooksPath does not become the chained previous hooks
ok - a project hook that appears after install still runs for the rest of the task
ok - a pane GIT_CONFIG hooksPath still chains the repository git is actually in
ok - an empty project core.hooksPath runs no repository hook and still strips the trailer
ok - an unresolvable project core.hooksPath still refuses the commit
ok - a valueless project core.hooksPath still refuses the commit
ok - the repository's pre-push runs and can refuse under every hooksPath override channel
ok - a git -c hooksPath override still strips the trailer and chains the project's hooks
ok - commit-msg file mode strips the trailer and keeps the subject
# all fm-git-strip-ai-trailers tests passed
FM_TEST_END 2026-10-09T17:02:28Z tests/fm-git-strip-ai-trailers.test.sh exit=0 duration_ms=12029 gate_skip=false
FM_TEST_BEGIN 2026-10-09T17:02:28Z tests/fm-inbox-conversation.test.sh family=unclassified expected_gate_skip=none
PASS: 9 inputs accounted for; 8 single dispatch claims, 1 stale bound input explicitly rejected; 6 replies retained
PASS: capture/accept/publication/playback crash windows, duplicate races and wrong-session refusals
PASS: the newest capture names the live call; a call nobody answered is superseded, a nameless one supersedes nothing
PASS: a redial captured mid-pass is never superseded; the pass restarts and retires the dropped call
PASS: live publication is owner-authored, digest-bound, and ordered; portions surface as waiting speech
PASS: the session holding the lock answers what its predecessor left saved; a lock-less caller is refused
PASS: an ElevenLabs-only home refuses an iMessage reply and records nothing new for a voice conversation
PASS: a v2 policy carries a text destination; each conversation publishes only to the one it is bound to
PASS: a destination withdrawn from the policy can no longer be published to
ok - conversation contract: durable capture, session routing, reply and playback recovery
FM_TEST_END 2026-10-09T17:04:48Z tests/fm-inbox-conversation.test.sh exit=0 duration_ms=140136 gate_skip=false
FM_TEST_BEGIN 2026-10-09T17:04:48Z tests/fm-project-capacity.test.sh family=backend-dispatch expected_gate_skip=none
not ok - an undeclared project refused a spawn: warning: ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/.test-phase-tmp/fm-project-capacity.8rS6Th/undeclared/home/data/task-c/launch-brief.md records no ship branch; defaulting to legacy branch fm/task-c
error: task task-c cannot be dispatched because its backlog data directory is inaccessible: ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/.test-phase-tmp/fm-project-capacity.8rS6Th/undeclared/home/data (automatic backlog transitions require tasks-axi 0.2.6 or newer with the required update and mv features): expected exit 0, got 1
FM_TEST_END 2026-10-09T17:04:53Z tests/fm-project-capacity.test.sh exit=1 duration_ms=4734 gate_skip=false
FM_TEST_SUMMARY total=8 failed=2 skipped_gate=0 duration_ms=370472
FM_TEST_SUMMARY_FAMILY family=backend-dispatch count=3 duration_ms=107123 failed=2
FM_TEST_SUMMARY_FAMILY family=standalone count=1 duration_ms=49476 failed=0
FM_TEST_SUMMARY_FAMILY family=unclassified count=2 duration_ms=199675 failed=0
FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=2 duration_ms=21118 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-inbox-conversation.test.sh duration_ms=140136
FM_TEST_SLOWEST rank=2 script=tests/fm-worker-account.test.sh duration_ms=90360
FM_TEST_SLOWEST rank=3 script=tests/fm-fleet-ledger.test.sh duration_ms=59539
FM_TEST_SLOWEST rank=4 script=tests/fm-dispatch-resolve.test.sh duration_ms=49476
FM_TEST_SLOWEST rank=5 script=tests/fm-wake-drain-voice-first.test.sh duration_ms=12510
FM_TEST_SLOWEST rank=6 script=tests/fm-git-strip-ai-trailers.test.sh duration_ms=12029
FM_TEST_SLOWEST rank=7 script=tests/fm-wake-drain-voice-hold.test.sh duration_ms=8608
FM_TEST_SLOWEST rank=8 script=tests/fm-project-capacity.test.sh duration_ms=4734
Evidence: Closed stdout returns failure and preserves the unread outcome

Source: Closed stdout returns failure and preserves the unread outcome

$ ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/bin/fm-branch-outcome.sh append --task demo --verdict routine --summary retained after closed stdout
1
[exit 0]

$ bash -c bash "$0" >&- 2>/dev/null ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/bin/fm-wake-drain.sh
[exit 1]

$ ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/bin/fm-branch-outcome.sh unread
{"seq":1,"epoch":1791565437,"task":"demo","wake":"","verdict":"routine","summary":"retained after closed stdout","silent":false,"statusEndpoint":0,"statusIdent":"-"}
[exit 0]

$ ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/bin/fm-wake-drain.sh
BRANCH OUTCOMES, ROUTINE (handled by the supervision session since your last drain; for your awareness, nothing to acknowledge):
[seq 1] demo: retained after closed stdout
[exit 0]

$ ~/.no-mistakes/worktrees/8352e211db7e/01M4G9DCM2DEDN0RDNFCWWW0YJ/bin/fm-branch-outcome.sh unread
[exit 0]

PASS: undeliverable output returns failure and preserves unread state; the next successful presentation consumes the routine outcome once.
Evidence: Gate-worktree lifecycle refusal with the environment marker absent

Source: Gate-worktree lifecycle refusal with the environment marker absent

$ fm-spawn.sh isolated demo --mode local-only --yolo off
error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/8352e211db7e.git)
[exit 3]

$ fm-send.sh isolated Do not send.
error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/8352e211db7e.git)
[exit 3]

$ fm-teardown.sh isolated
error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/8352e211db7e.git)
[exit 3]

PASS: the worktree backstop refuses all three lifecycle entry points with the environment marker absent and leaves isolated state empty.
Evidence: Memory diagnostic output and threshold exit statuses

Source: Memory diagnostic output and threshold exit statuses

$ fm-jev-mem-guard.sh --json
{
  "name": "fm-jev-mem-guard",
  "checked_at": "2026-10-09T17:06:14Z",
  "status": "OK",
  "recommendation": "Memory and swap utilization are within thresholds; the caller should continue unchanged.",
  "reason": null,
  "summary": {
    "mem_total_gb": 15.24,
    "mem_available_gb": 9.08,
    "mem_used_pct": 40.4,
    "swap_total_gb": 10.0,
    "swap_used_gb": 5.86,
    "swap_used_pct": 58.6
  },
  "top_processes": [
    {
      "pid": 206213,
      "comm": "bun",
      "rss_mb": 398.4
    },
    {
      "pid": 2007396,
      "comm": "claude",
      "rss_mb": 375.0
    },
    {
      "pid": 1010035,
      "comm": "claude",
      "rss_mb": 344.6
    },
    {
      "pid": 322050,
      "comm": "claude",
      "rss_mb": 340.9
    },
    {
      "pid": 1263745,
      "comm": "hermes",
      "rss_mb": 290.4
    },
    {
      "pid": 2809743,
      "comm": "claude",
      "rss_mb": 288.5
    },
    {
      "pid": 2436517,
      "comm": "claude",
      "rss_mb": 280.5
    },
    {
      "pid": 1629,
      "comm": "hermes",
      "rss_mb": 268.2
    },
    {
      "pid": 2508682,
      "comm": "claude",
      "rss_mb": 259.9
    },
    {
      "pid": 1263786,
      "comm": "hermes",
      "rss_mb": 240.8
    }
  ]
}
[exit 0]

$ fm-jev-mem-guard.sh --check --warn-mem-pct 1000 --crit-mem-pct 1000 --warn-swap-pct 1000 --crit-swap-pct 1000
fm-jev-mem-guard — 2026-10-09T17:06:14Z
  • RAM: 40.7% used (9.04 GB available / 15.24 GB total)
  • Swap: 58.6% used (5.86 GB used / 10.0 GB total)
  • Status: OK
  • Recommendation: Memory and swap utilization are within thresholds; the caller should continue unchanged.

  Top 10 RSS Processes:
    - PID 206213 (bun): 398.4 MB
    - PID 2007396 (claude): 375.0 MB
    - PID 1010035 (claude): 344.6 MB
    - PID 322050 (claude): 340.9 MB
    - PID 1263745 (hermes): 290.4 MB
    - PID 2809743 (claude): 288.5 MB
    - PID 2436517 (claude): 280.5 MB
    - PID 1629 (hermes): 268.2 MB
    - PID 2508682 (claude): 259.9 MB
    - PID 1263786 (hermes): 240.8 MB
[exit 0]

$ fm-jev-mem-guard.sh --check --warn-mem-pct 0 --crit-mem-pct 0 --warn-swap-pct 0 --crit-swap-pct 0
fm-jev-mem-guard — 2026-10-09T17:06:14Z
  • RAM: 40.9% used (9.01 GB available / 15.24 GB total)
  • Swap: 58.6% used (5.86 GB used / 10.0 GB total)
  • Status: CRITICAL
  • Recommendation: Memory or swap utilization is at or above a critical threshold; this host condition can explain worker silence while it holds.

  Top 10 RSS Processes:
    - PID 206213 (bun): 398.4 MB
    - PID 2007396 (claude): 375.0 MB
    - PID 1010035 (claude): 344.6 MB
    - PID 322050 (claude): 340.9 MB
    - PID 1263745 (hermes): 290.4 MB
    - PID 2809743 (claude): 288.9 MB
    - PID 2436517 (claude): 280.5 MB
    - PID 1629 (hermes): 268.2 MB
    - PID 2508682 (claude): 259.9 MB
    - PID 1263786 (hermes): 240.8 MB
[exit 1]

PASS: the real diagnostic reports bounded percentages; thresholds beyond physical utilization stay OK, while zero thresholds alarm on this readable host.

Pipeline

Updates from git push no-mistakes

... (18 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Test** - 2 warnings

🔧 Fix applied.
2 warnings still open:

  • ⚠️ tests/fm-remote-secondmate-lifecycle-e2e.test.sh:988 - The remote inheritance serialization fixture repeatedly fails before reaching its deliberately blocked inheritance write. Correcting fixture placement and gate-refusal setup did not resolve it; both the full fixture and focused reruns failed, including the final run after cleanup repairs. The concurrent config-push convergence assertion therefore never executes. The captured spawn output contains only a watcher-down reminder, leaving the cause unresolved between product behavior and fixture infrastructure. Diagnose the blocked spawn and establish a reliable completed serialization check. Evidence: round4-remote-inherit-final.log. Test-only cleanup was fixed and verified separately.
  • ⚠️ live validation verdict: inconclusive (26 of 34 scenarios were driven live against the product); untested: Run the portable remote inheritance serialization check; reach the blocked write and prove a later config push wins, Provision, launch, steer, and retire a secondmate over real SSH while preserving inherited configuration, Resolve dispatch against real model and per-account quota services, then authenticate the selected native account, Handle supervision wakes and terminal submission through real vendor engines, hooks, and native signal surfaces, Use the changed Pi and Claude calm surfaces; retain queued input and display operational notes in their intended visual positions, Publish and register work against a real forge, including Gerrit's current patch-set tree and merge result, Answer a real voice call and deliver ordered speech or a policy-bound message to its caller, Run the changed gate and CI consumers through validation, publication, and required remote checks
  • Live validation: ⚠️ inconclusive - 26 of 34 scenarios driven live against the product
Scenario Result Live Evidence
Acknowledge a normally announced note, then replay or announce it; both interfaces report durable acknowledgement without promising another pickup ✅ pass live round4-regressions.log:8; round4-manual-cli.log:82-109; predecessor reproduction in round4-ack-before-fix.log
Repair an unannounced saved note without duplicating it, and refuse empty bodies or unsafe request IDs ✅ pass live round4-regressions.log:6-7,16
Capture and accept committed conversation text, then publish ordered replies; crash retries preserve identity and wrong-session callers are refused ✅ pass live round4-product-tests.log:88-93; subprocess cases in tests/fm-inbox-conversation-cases.py
Publish a conversation reply only to its policy-bound text destination; incompatible or withdrawn destinations create no new reply ✅ pass live round4-product-tests.log:94-96
Drain mixed voice and ordinary wakes; saved conversation turns remain pending and appear before ordinary rows ✅ pass live round4-product-tests.log:3-9
Drain with stdout closed; return failure and retain the outcome for one successful subsequent presentation ✅ pass live round4-closed-output-cli.log
Submit a branch report; reject foreign actors, ended turns, and out-of-scope tasks while storing authorized and silent outcomes correctly ✅ pass live round4-regressions.log:26-27
Deliver an away status wake through the local supervision engine contract; the host handles it without waking main ✅ pass live round4-regressions.log:71; completed supervision-host run at line 98
Queue voice-only, mixed, or newly arriving conversation turns while away; hand them to main without starting an engine turn ✅ pass live round4-regressions.log:69
Drain an away queue before either offer boundary; keep an empty queue quiet and hand a corrupted queue to main ✅ pass live round4-regressions.log:70
Return during a supervision turn or leave an outcome unacknowledged; preserve visible outcomes for main and recover incomplete or failed local engine turns safely ✅ pass live round4-regressions.log:40-45,73-74,80-82
Write dialog through the mirror CLI and resume its feed; preserve text, exclude operational input, enforce private storage, and refuse malformed records ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-host-mirror-final.log: operational-input exclusion, owner-only storage, committed-cursor resumption, and malformed-sequence refusal
Run a registered local process-event source; capture its result once and retain it across drains until explicit handling ✅ pass live round4-regressions.log:106-108,112
Break a process-event registration, repair it, then reconcile; announce one failure episode and confirm the repaired launch ✅ pass live round4-regressions.log:162,164-165
End a detached listener's owning session; reap its descendant tree within the documented bound while leaving a healthy home's listener running ✅ pass live round4-regressions.log:223-227,234
Dispatch beyond project capacity or race for its last place; defer without allocating state and admit at most one competing spawn ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-capacity-compatible.log: capacity deferral before allocation, concurrent last-place admission, and unreadable-state refusal
Select a Codex account root; carry its canonical path into launch output and metadata, and refuse unusable roots before publication ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-codex-account-final.log: CODEX_HOME metadata, canonical aliases, malformed-root refusals, and batch forwarding
Apply a Claude or Pi worker-account pin; construct the correct launch environment and preserve the secondmate home's own pin ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-isolated-home-tests.log: Claude/Pi launch environments and secondmate-owned pin preservation
Resolve an invalid dispatch configuration; refuse it before making a network request ✅ pass live round4-product-tests.log:38-39
Resolve registered delivery posture and generate each ship brief; enforce required modes, reject incompatible Gerrit posture, and keep shell metacharacters literal ✅ pass live round4-manual-cli.log:1-35; round4-targeted-cli-contracts.log, SOURCE round4-brief-final.log; three generated brief artifacts
Enter and archive away or quiet posture; preserve the captain's words verbatim and render return information from durable records ✅ pass live round4-manual-cli.log:38-78; round4-targeted-cli-contracts.log, SOURCE round4-delivery-tests.log lines 1-57
Enable the fleet ledger for a local task lifecycle; record ordered events once, preserve status through ledger failure, and leave no ledger artifacts when disabled ✅ pass live round4-product-tests.log:57,60,63-66
Commit through installed Git hooks and override channels; strip known bot co-authors, retain human attribution, and preserve project-hook refusals ✅ pass live round4-product-tests.log:69-83
Invoke spawn, send, or teardown from the gate worktree with its environment marker absent; refuse all three without changing isolated state ✅ pass live round4-gate-refusal-cli.log
Pass invalid arguments to PR registration or teardown; reject them before side effects ✅ pass live round4-targeted-cli-contracts.log, SOURCE round4-publication-checks.log: invalid entrypoints have zero side effects
Read the memory diagnostic and vary its thresholds; report bounded utilization, return OK above physical utilization, and alarm at zero thresholds ✅ pass live round4-memory-guard-cli.log
Run the portable remote inheritance serialization check; reach the blocked write and prove a later config push wins ⏸️ untested no The prior payload did not establish a live result: the deterministic SSH fixture repeatedly timed out before the blocked write, and its concurrent convergence assertion never ran. Synthetic SSH is not…
Provision, launch, steer, and retire a secondmate over real SSH while preserving inherited configuration ⏸️ untested no Built isolated Git homes and exercised the synthetic SSH boundary. Although sshd is available, this phase explicitly prohibits launching a real secondmate or fleet worker. A dedicated remote host and…
Resolve dispatch against real model and per-account quota services, then authenticate the selected native account ⏸️ untested no Local protocol and launch checks ran with isolated account files and service fixtures. The live quota guard requires FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E=1, which this phase forbids forcing, and real acco…
Handle supervision wakes and terminal submission through real vendor engines, hooks, and native signal surfaces ⏸️ untested no The real local host and mirror commands ran against disposable fixtures. Native guards explicitly skipped for FM_SUPERVISION_HOST_LIVE_E2E, FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E, FM_HOST_MIRROR_LIVE_E…
Use the changed Pi and Claude calm surfaces; retain queued input and display operational notes in their intended visual positions ⏸️ untested no Portable extension checks cannot establish rendering in the vendor runtime. Even the credential-free Pi route starts a real agent harness, which this phase prohibits. Supply an authorized disposable n…
Publish and register work against a real forge, including Gerrit's current patch-set tree and merge result ⏸️ untested no Real registration entrypoints ran against local Git origins and forge fixtures, which do not establish network publication. No dedicated forge endpoint or credentials were supplied; gerrit-axi is also…
Answer a real voice call and deliver ordered speech or a policy-bound message to its caller ⏸️ untested no The durable capture, acceptance, publication, and recovery protocol was driven through real CLI processes. No isolated live call endpoint, audio delivery connection, or messaging credentials were avai…
Run the changed gate and CI consumers through validation, publication, and required remote checks ⏸️ untested no This assigned phase may not initialize, rerun, control, push, or execute other pipeline phases. Parsing configuration would not prove consumer behavior. The outer executor must drive those phases and…
  • TMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh --jobs 1 tests/fm-inbox.test.sh tests/fm-supervision-host.test.sh tests/fm-procevent.test.sh
  • TMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh tests/fm-worker-account.test.sh tests/fm-fleet-ledger.test.sh tests/fm-git-strip-ai-trailers.test.sh tests/fm-wake-drain-voice-hold.test.sh tests/fm-wake-drain-voice-first.test.sh tests/fm-inbox-conversation.test.sh tests/fm-project-capacity.test.sh
  • TMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh --jobs 1 tests/fm-afk-contract.test.sh tests/fm-afk-return.test.sh tests/fm-brief.test.sh tests/fm-gate-refuse.test.sh
  • Materialized a disposable copy of tracked source with sibling fixture homes; reran worker-account and remote lifecycle tests through bin/fm-test-run.sh --jobs 1.
  • npm install --prefix .test-phase-tmp/round4-deps --no-audit --no-fund tasks-axi@0.2.6 for the isolated capacity fixture.
  • PATH="$PWD/.test-phase-tmp/round4-deps/node_modules/.bin:$PATH" TMPDIR="$PWD/.test-phase-tmp/round4-fixtures" bin/fm-test-run.sh --jobs 1 .test-phase-tmp/round4-source/tests/fm-project-capacity.test.sh
  • bin/fm-test-run.sh .test-phase-tmp/round4-codex-account.test.sh using selected existing account cases and unique disposable task IDs.
  • bin/fm-test-run.sh .test-phase-tmp/round4-brief-behavior.test.sh using selected public generator cases and a relative shell-injection marker valid under the fixture Git ref contract.
  • FM_TEST_ONLY=<selector> bin/fm-test-run.sh --jobs 1 tests/fm-pr-check-security.test.sh for test_invalid_entrypoints_have_zero_side_effects, test_gerrit_ready_gate_reads_the_published_tree, test_unpushed_named_head_refuses_registration, and test_valid_recording_and_merge_derivation.
  • bin/fm-test-run.sh tests/fm-host-mirror.test.sh
  • python3 .test-phase-tmp/round4-manual.py: isolated project-mode, generated briefs, verbatim posture records, and acknowledged inbox replay/announce checks.
  • python3 .test-phase-tmp/round4-closed-output.py: real drain with closed stdout, unread-state inspection, and subsequent successful presentation.
  • Ran real fm-spawn.sh, fm-send.sh, and fm-teardown.sh against an empty isolated home with the gate environment marker absent; verified refusal and unchanged state.
  • Ran real fm-jev-mem-guard.sh --json and --check with all warning/critical thresholds set to 1000, then 0.
  • Replayed the acknowledgement failure with the predecessor's real inbox executable, then compared the target's durable acknowledgement output.
  • Ran opt-in supervision-host, attended-host, mirror, quota-dispatch, Devin, and Herdr guards without forcing their control variables; recorded their skips. Also ran tests/fm-afk-inject-e2e.test.sh with its portable terminal fixture.
  • Ran the full remote lifecycle fixture in the corrected disposable source, then repeatedly ran TMPDIR="$PWD/state/round4-fixtures" bin/fm-test-run.sh tests/fm-remote-inherit-focused.test.sh; retained failure diagnostics.
  • bin/fm-test-run.sh .test-phase-tmp/round4-cleanup-check.test.sh: verified cleanup stops a signal-resistant owned writer and removes protected fixture directories; the final inheritance run also completed cleanup without errors.
  • Removed disposable source copies, fixture roots, local dependencies, and temporary drivers; verified only the intentional test cleanup change remains in the worktree.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 17, 2026 15:58
* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs kunchenguid#4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs kunchenguid#4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
…uid#4873) (kunchenguid#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes kunchenguid#3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The kunchenguid#4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees
…#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
tiago-peixoto and others added 28 commits October 6, 2026 19:27
…is retired (kunchenguid#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it
…#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes kunchenguid#6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership
…id#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed
….2 (kunchenguid#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.
* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes kunchenguid#6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run
kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.
… section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743
…henguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.
…te (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503
…nchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.
* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call
Absorbs the 234 upstream kunchenguid/firstmate commits since the last absorb
(upstream 9bc051f, squashed into the fork as 7a03435). That squash dropped
the upstream parent, so conflicts were resolved against 9bc051f as the
effective base; this commit records upstream/main as a real parent so the next
absorb starts from the correct merge base.

Fork features carried onto upstream's restructured code: the Codex account
axis (dispatch resolver, spawn, relaunch, quota watch), voice-turn answering
and live-call scoping, the iMessage conversation destination, fast surfacing
of queued phone/texted turns to a running watcher, teardown of records whose
pooled worktree was returned or reassigned, the GitHub App checks reader in
the merge guard, Bearings, and PR state, and the brain-room/shared-interface
pieces. Where upstream built the same thing, upstream's version is kept and
only the fork's missing behavior is ported onto it.
…m absorb

Main's deliberate removal of the shared-interface lock hooks wins: the
steal-lock path returns to upstream's exact form, and the shared-interface
test call this absorb had re-homed is dropped with the test itself.
…pstream's guidance

The absorb left two test.instructions keys in .no-mistakes.yaml: the fork's
what-the-product-is runbook and upstream's live-lab rules. Go YAML refuses
the duplicate, so no-mistakes could not load the repo config at all. Both
runbooks now live in the one block, upstream's under its own heading.
@Mauryanx

Mauryanx commented Oct 9, 2026

Copy link
Copy Markdown
Owner Author

Land this as a merge commit, not a squash.
The merge commit 5094dad5 has upstream/main (kunchenguid/firstmate 19fcbbde) as a real second parent.
The branch also merges current main (#17-#20) through merge commits. The last absorb (#9) was squashed, which dropped upstream's parent and left the fork's merge base stuck at 2026-09-10; a squash here would repeat that and the next absorb would re-conflict on everything this one resolved.

What this absorbs

234 upstream commits, from 9bc051ff (2026-09-17, the point #9 actually absorbed) through 19fcbbde (2026-10-08).
Conflicts were resolved against 9bc051ff rather than the stale 2026-09-10 base, which cut the conflicted files from 136 to 28.

What changes in Firstmate's behavior

Changed defaults

  • Supervision host is on by default for a Claude Code primary. Routine wakes are handled by a background supervision session instead of the main conversation. Opt out with a config/supervision-host-off file. Other primaries stay opt-in.
  • AI co-author trailers are stripped from fleet-launched commits by a per-task commit hook. config/keep-ai-trailers opts back out.
  • Away mode acts on the captain's away words. Under /afk the supervision branch (Pi) or host (Claude/Cursor) reads the away words as the whole mandate and may act on them through the guarded scripts; on Pi the main session is parked until return.
  • /quiet is a statement, not a mode, when attended supervision is already running; quiet records are treated as a present captain, never hold-for-return.
  • The PR merge guard is stricter: it refuses when a check the base branch requires has not reported at all (an attended --allow-missing <check> waiver mirrors --allow-red), and it retries a bounded number of times while GitHub reports mergeability as UNKNOWN.
  • Status events carry their emission time, and a fleet status line is recorded the moment it is written.
  • AGENTS.md is much shorter: situational detail moved into new loaded-on-trigger skills (operational-home-layout, session-start-recovery, validation-supervision, ship-landing, scout-completion, away-quiet-supervision, agent-skill-trigger-index).

New

  • Devin CLI as a worker/scout runtime.
  • Captain inbox: idempotent capture by request id, durable replies to a note, receipts, and a readiness view.
  • Lavish feedback routes directly to the owning worker.
  • Typed dispatch resolution sends only the brief's task sections and supports per-rule min_confidence floors with a runner-up fallback.
  • quota-axi schema 6 (multiple accounts per provider) is read row-per-account.
  • A dead second mate is relaunched mid-session by the watcher, not only at session start.

New opt-ins (no effect until configured)

  • Fleet activity ledger (config/fleet-ledger), per-task no-mistakes pipeline spend at cleanup, per-project worker capacity, per-project ship-branch prefix, start-from-a-named-base-branch, Claude/Pi worker account pins (config/claude-account, config/pi-account), worker tool exclusions (config/crew-exclude-tools), wait-without-turns (config/wait-no-turns), home-local brief include, Gerrit publishing, daily startup growth check, Herdr resume lock wait.

Removed

  • Nothing user-facing. The fork's CI pin of the Pi package to 0.85.1 is dropped because upstream updated the Calm parity tests for Pi 1.x (see the overlap table).

Needed in this home before restart

  • Upgrade three tools - upstream raised their minimums above what hermes has installed, and bootstrap will report them as incompatible:
    • tasks-axi 0.2.5 -> at least 0.2.6 (needed for automatic backlog moves at dispatch and cleanup, and for merges, which check whether the task is held for the captain)
    • quota-axi 0.1.36 -> at least 0.1.51 (needed for quota-aware dispatch, including the Codex account axis)
    • lavish-axi 0.1.64 -> at least 0.1.80 (needed for Lavish boards)
  • Decide the supervision host for this Claude Code primary: it now runs by default. Leave it on, or create config/supervision-host-off to keep today's single-conversation supervision.
  • No skill or config edits are required for any fork feature; the Codex dispatch profiles, GitHub App checks config, conversation transport, and brain-room files keep their existing formats.

Where both sides built the same thing

The rule applied: take upstream's version only when it does everything the fork's does (proven by running the fork's own tests against upstream's code); otherwise keep the fork's, or port the fork's missing piece onto upstream's version.

Area Fork Upstream Chosen Evidence
Teardown of a task whose pooled worktree moved on #13: return receipt (treehouse_returned=) + ownership-first ordering; handles a slot already returned but not yet reclaimed, and a reassigned one kunchenguid#6213: skips the record-exclusivity scan only when the slot's owner claim names another task Fork's, with upstream's extra early-exit kept (harmless with fork ordering) Fork's tests/fm-teardown-endpoint-safety.test.sh run against pure upstream code fails at returned: retained record cannot retire safely (upstream refuses to retire the stale record). On this branch the whole suite passes, including upstream's kunchenguid#6213 case.
no-mistakes passed-with-override outcome Reads as done and keeps an "(approved override)" marker visible Reads as done like a clean pass (marker dropped); adds passed-with-skips Upstream's mapping + fork's marker ported Fork's crew-state test on pure upstream fails: the override stays visible beside the PR outcome. Passes on this branch, along with upstream's skips cases.
no-mistakes daemon-down probe v1.91.0 answers "daemon not running" with exit 0; fork reads that text as down Restructured probe (one cached call, up/unanswered/down) but treats any exit 0 as up Upstream's probe + fork's text check ported Fork's case on pure upstream fails: a zero-exit daemon-not-running answer is still down. Passes on this branch.
Pi multicall sleep in the Herdr descendant tests gnusleep fallback fm_agent_standin long-running stand-in Upstream's (covers the same failure more generally) tests/fm-backend-herdr.test.sh passes on this branch with upstream's version.
Slow watcher-checkpoint fixtures in the wake-queue suite Widened checkpoints from 1s to 4s Refactored into shared stall-watch legs (kunchenguid#6887) Upstream's tests/fm-wake-queue.test.sh passes on this branch.
CI Pi version Pinned 0.85.1 because Pi 1.0 broke two Calm parity tests Parity tests updated for Pi 1.0/1.0.1 (kunchenguid#6338, kunchenguid#6530) Upstream's (unpinned) tests/fm-pi-branch-extension.test.sh (one of the two pinned-for tests) passes against Pi 1.0.4 here; tests/fm-pi-primary-types.test.sh passes against Pi 1.1.0.
Test-agent runbook in .no-mistakes.yaml "What the product is" runbook under test.instructions Live-lab and fleet-safety rules under the same key Both, in one block (upstream's under its own heading) The duplicate key made no-mistakes refuse to load the config; tests/fm-nm-test-contract.test.sh passes on the combined block.
Validation verdict wording in AGENTS.md "Use the helper's state line as the verdict" Same rule, now in validation-supervision Upstream's Wording-only.

Close cousins that are not the same feature, so both were kept: upstream's Claude/Pi worker account pin vs the fork's per-task Codex account axis; upstream's quota-axi schema-6 account rows vs the fork's per-CODEX_HOME reads (the resolver now reads each Codex account's snapshot through upstream's row join); upstream's inbox reply vs the fork's voice-turn answering; upstream's required-check reads vs the fork's GitHub App checks reader.

Fork features carried onto upstream's code

  • Codex account axis: typed dispatch resolver (per-account reads now also cover the runner-up rule that upstream's new min_confidence fallback can pick), spawn --codex-home alongside upstream's new flags, relaunch carry-forward ahead of upstream's account-pin and tool-exclusion pre-stop checks, quota watch --codex-home on upstream's new retry-streak logic (a signed-out account stays a distinct, non-retried error).
  • Voice turns: answer-voice-turn, live-call scoping, VOICE-first drain, the conversation transport and its iMessage destination; their AGENTS.md lines moved into upstream's new skills (operational-home-layout, session-start-recovery, agent-skill-trigger-index).
    One integration fix: upstream's away posture lets the background supervision branch/host claim notification rows, which would have handed a phone or texted turn to an actor that cannot answer it. Voice-turn rows now stay with the main session in both postures (.pi/extensions/lib/fm-branch-dispatch.ts, shared by Pi and the supervision host), with a regression case in tests/fm-pi-branch-extension.test.sh. Validation review found the same gap in the supervision host's own away path on a Claude primary; the pipeline fixed it by routing the host's away acceptance through that same shared offer check, while keeping upstream's quiet no-op for an empty queue (tests/fm-supervision-host.test.sh).
  • Fast phone/texted turn surfacing: the watcher's ring-tolerant wait now also covers upstream's new pane-capture wait, so a ring cannot cut a capture short.
  • GitHub App checks reader: still used by the merge guard, Bearings, and PR state, beside upstream's new required-check reads.
  • Brain desk (feat: add opt-in brain-room and shared voice interfaces #15): unchanged. The feat: add opt-in brain-room and shared voice interfaces #15 shared-interface hooks were deliberately removed on main by the courier pickup (feat(bin): add courier spool pickup for iMessage conversations #17); this branch takes that removal, so the steal-lock path is upstream's exact form.
  • Courier iMessage pickup (feat(bin): add courier spool pickup for iMessage conversations #17) and its reboot launcher (fix: keep courier iMessage pickup running across reboots #18): merged in from main (commit 844a26e3); tests/fm-courier-pickup.test.sh passes on this branch.

Open fork work

  • The courier pickup branch landed on main while this was in progress, so it is merged here rather than needing a rebase; its two conflicts (bin/fm-wake-lib.sh, tests/fm-wake-queue.test.sh) were resolved in main's favor.

Testing

Run on hermes against this branch.

  • bin/fm-lint.sh: ShellCheck 0.11.0 and actionlint clean. bin/fm-doc-audience-check.sh: ok (121 surfaces, 770 local links). bin/fm-test-run.sh --check-coverage: ok (256 scripts).
  • Full portable suite, all three CI lanes, real-Herdr family excluded: parallel lane 1 13/13 passed; parallel lane 2 11 scripts, serial lane 215 scripts.
  • Every failure was re-run. Each either passed when re-run with task temp dirs under a private base, or fails the same way on pure upstream (19fcbbde) on this host. Host causes: about 500 stale group-writable /tmp/fm-<task> dirs from old runs, which upstream's new private-temp check refuses; a per-user systemd that adopts orphans before PID 1; no Chrome or ruby; a shared 4 GB /tmp that filled during the run; and old global tool versions.
  • tests/fm-watch-triage.test.sh hit the 30-minute per-script bound; the same hang is already known in main's CI serial lane.
  • Fork features on the final tree: courier pickup, conversation transport, voice-first and voice-hold drain, inbox, Pi branch dispatch (with the new voice-turn case), dispatch resolver, quota watch, crew state, merge guard, GitHub App checks reader, PR state, Bearings snapshot, both teardown suites, brain desk, voice relay, and spawn dispatch profiles (97 cases, every Codex-account case) all pass. Relaunch passes (every Codex-account carry-forward and refusal case) once its two-second checkpoint wait is off a loaded machine.
  • Pi compatibility after dropping the CI pin: tests/fm-pi-primary-types.test.sh passes against Pi 1.1.0, and tests/fm-pi-branch-extension.test.sh passes against Pi 1.0.4.

Validation notes

  • Review findings in pre-existing upstream code that this absorb leaves unchanged were declined here rather than patched, to keep this copy aligned with upstream; they are worth an upstream report or a separate fork fix: an inbox capture retry race that can resurrect a handled instruction (bin/fm-inbox.sh), Gerrit re-registration accepting stale publication evidence (bin/fm-pr-check.sh/bin/fm-dod-lib.sh), a concurrent send losing its retry ring under config/wait-no-turns (bin/fm-task-inbox-lib.sh), a quiet-mode record read as away after switching to Pi (.pi/extensions/lib/fm-branch-dispatch.ts), an unbounded option loop in bin/fm-live-lab.sh, and two upstream tests (tests/fm-backend-herdr.test.sh, tests/fm-supervision-host-attended-live-e2e.test.sh).
  • Upstream deliberately stopped workers from adding knowledge to a project's AGENTS.md on their own (factual corrections only). The fork never built anything there, so this absorb keeps upstream's policy; it can be restored later as a small fork patch if wanted.
  • The Test step passed with a recorded exception: 26 of 34 live scenarios ran and passed; the 8 untested ones need real external services (a real SSH second mate, vendor engines, model quota and account auth, a real forge including Gerrit, a real voice call, the Pi/Claude Calm surfaces, remote CI) and will be exercised by the post-merge restart and normal use.
  • The pipeline's own fixes: the away-host voice handoff and empty-queue no-op (bin/fm-supervision-host.sh), inbox replay now reporting an already-acknowledged note as acknowledged (bin/fm-inbox.sh, a small divergence from upstream), and fixture cleanups in four test suites.

@Mauryanx
Mauryanx merged commit 2b1bc8e into main Oct 9, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.