Repository navigation
feat: absorb upstream supervision and fleet updates - #21
Conversation
* docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation
…guid#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895) * fix(bin): keep supervisor status closes from waking the same home A drain that already folded OPEN DECISIONS has presented those bytes even when the watcher has no matching seen marker. Treat that fold, and the presentation cursor, as known so the bookkeeping close stays quiet while later worker lines still signal. * no-mistakes(review): Keep folded worker failures waking past supervisor closes * no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes * no-mistakes(review): Stop folded worker resolved lines from counting as already read * no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr * no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973) * fix(bin): treat a live no-mistakes run as current after rebase A running run on the task's branch is authoritative regardless of head. Matching only the local head made a rebased in-flight run look failed. * no-mistakes(review): restrict coarse live-any-head to foreign-branch answers * no-mistakes(review): reject gate-parked runs from the executing predicate * no-mistakes(review): hoist gate-marker patterns into single run-lib owner * no-mistakes(review): require live daemon for head-free run binding * no-mistakes(review): require answered daemon-down before unbinding live runs * no-mistakes(review): extend daemon guard to anchored continuation routes * no-mistakes(review): delete live-any-head; restore dead-daemon verdict * no-mistakes(review): keep parked gates parked; name dead daemon everywhere * no-mistakes(review): set dead-daemon verdict instead of emitting early * no-mistakes(review): align selected route with legacy dead-daemon handling * no-mistakes(review): drop unproven-record binds; narrow coarse gate reading * no-mistakes(review): narrow header, drop vestigial guard, retarget tests * no-mistakes(review): revert coarse gate override; require answered-down probe * no-mistakes(review): cache one daemon probe; stop duplicating run id * no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows * no-mistakes(review): delete coarse dead-daemon extension and gate note * no-mistakes(review): delete remaining coarse dead-daemon block and stale docs * no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994) * fix(bin): stage the launch command in a private file and type a short source line A long launch line typed while the fresh pane shell is still busy waits in the terminal's canonical line buffer, which drops input past about 1,024 bytes on macOS, so the pane was left at an unfinished command with no agent running. fm-spawn now writes the assembled command to the task's own temp root under umask 077 and types only a short line that sources it. Refs kunchenguid#4559 * fix(bin): keep the per-task temp root private before staging the launch command The root lives at a predictable path under /tmp and now holds the whole launch command. Create it with mode 0700, refuse one that already exists as anything but a directory owned by this user that nobody else can write, and tighten an owned one, so no other local user can plant or swap the staged file. Refs kunchenguid#4559 * fix(bin): enforce private staged launch file mode * test(spawn): cover long staged Claude launches * no-mistakes(review): Namespace launch files and prove truncation staging * no-mistakes(review): Use immutable per-spawn launch filenames * no-mistakes(document): Document staged launch delivery safeguards * no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass --------- Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions * no-mistakes(review): Drop substring matching from test.instructions contract * no-mistakes(review): Assert commands.test key absence in YAML * Drop unit-first sentence and instructions contract test Captain-scoped follow-up on the Herdr-lab test.instructions ship: keep the lab safety runbook only, and leave the no-mistakes contract test focused on commands.test absence.
…uid#4873) (kunchenguid#5001) * docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873) Replace the pixels-of-today's-UI rule with a quarantined, version-pinned surface-adapter exception recorded as standing debt. Cap the always-loaded contract at 9,000 words and require prune-or-trigger before a crossing change lands. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> * docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873) Replace the compressed paraphrase with the issue's accepted wording: a named quarantined version-pinned adapter, expected to break, recorded as standing debt that never hardens into a shared contract. Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes kunchenguid#3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The kunchenguid#4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead of carrying a second copy of the rule in awk. The copies had already drifted: the shell side stripped tags from lines with no colon, which the awk rule left whole, so a colonless line could be mistaken for one already recorded. One definition, checked against the awk rule it replaces over the edge cases and a 4000-line fuzz. tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In hang mode the poll set DEADLINE to the real now plus a one-second budget, and when the second ticked before the first forge call the loop broke without ever calling gh: forge/calls was never written and the assertion failed reading a missing file. Freezing the clock in both modes removes the dependence on wall time; the bounded call is still cut by the real timeout, so the observation the test asserts still starts. Emission time stays optional on new status records, and legacy or malformed lines keep an unknown age. * no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion * no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs * test: fold emission-time snapshot coverage into the fixture case Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer touches workflows. Keep every emission-time assertion by folding it into test_fixture_snapshot_json. * no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch * no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape * no-mistakes(review): strip undelimited at-tags; correct brief stamp header * no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness * no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles * no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb * no-mistakes(review): read note and key past colon-bearing stamps * test(status): keep inactive reconcile assertions stamp-tolerant These two oracles were made stamp-tolerant while resolving one of the branch's merges from main. The rebase drops merge commits, so that adaptation was lost and both assertions went back to matching an exact substring that a stamped line no longer contains: the tag lands before the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...". Strip a well-formed tag before matching, as the branch's other oracles do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap * no-mistakes(document): correct stale unstamped status-line spellings in docs * no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments * no-mistakes(ci): rename subshell-local epoch in delivery-race stub The serialization test overrides fm_pending_reply_mark_delivered inside a (..) subshell. Its `epoch` local collided with the same name in status_line_at_epoch/status_stamp_line, which this branch added and this suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and failed Lint 2. The stub already prefixes its other locals with `pending_` for the same reason; `epoch` was the leftover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes * no-mistakes(review): Fix Lavish classifications and fail-closed host loading * no-mistakes(review): Restore Lavish host state across retries and launches * no-mistakes(review): Preserve destination Lavish host when configuration is absent * no-mistakes(document): Document Lavish status and host guarantees
…#5076) * feat(afk): make the captain's away words the whole mandate Retire the clause fields, verb list, never-set scan, refused records, and the per-task merge-grant list from the away-posture record. The record is now version 2: the captain's words verbatim plus expected return, spend cap, and reach line; a version 1 record still validates, reads, and archives so a live away window is never broken by the upgrade. The supervision branch reads the words at the tail of every wake and acts on them by its own judgment through the guarded scripts under standing authority, never by analogy, holding for the return on doubt, and opens each such outcome summary with "per your away instructions:" so the return brief can render the words beside the session's account. While the record exists any green merge runs under away authority (ledger tag "away"); red merges, --allow-red, asynchronous and queued merges, and local-only landing stay refused. The branch may file a backlog item the words explicitly call for before dispatching it under the spend cap. Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and fm-pr-merge.sh as commands: version 2 written, version 1 read, retired flags and subcommands refused by name, green merges landing under the record, red and waived-red refused, the record lock still closing the authority-read window, and the Pi away tail carrying the words. * no-mistakes(review): carry the away read-back to the session verbatim * no-mistakes(review): match the exact away-action marker in the return brief * no-mistakes(review): refuse a words block truncated by a damaged line * no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049) * fix(bin): render the remote charter's steering-inbox path host-local A freshly provisioned remote secondmate read a parent-home absolute steering-inbox path in its charter - a location that exists on no route - and spent its first turn discovering the gap and filing a blocked decision for what was a render defect. The seed's remote-copy rewrite now maps the inbox to the route's host-local parent-route inbox, exactly as it already maps the reply-log path, so every mention - bare path, listing, and handled/ acknowledgement - lands host-local. Both rewrites also become plain assignments, because a quoted substitution nested inside a double-quoted printf argument leaks literal quotes into the replacement text on stock macOS bash. The lifecycle suite pins the corrected render both directions against the real seed, provisioning, and delivery route, sharing one fixture value between the render truth and the delivery truth. Closes kunchenguid#5012 * no-mistakes(document): document remote charter's host-local steering inbox
) * feat(procevent): route worker-owned Lavish rounds * no-mistakes(review): drop duplicate artifact field from task-owned registration * no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic * no-mistakes(review): keep worker board owned until terminal round acknowledged * no-mistakes(review): refuse every retirement of an open worker-owned round * no-mistakes(review): use real lavish reply flag, isolate reply generations * no-mistakes(review): drop .posted marker for best-effort reply posting * no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures * no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms * no-mistakes(review): re-arm only to acknowledge an open round * no-mistakes(review): conclude only a still-open terminal round * no-mistakes(review): record the acknowledgement before retiring the board * no-mistakes(review): retain the registration across a conclude, qualify terminal docs * no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107) * fix(bin): reserve contribution observation budget * no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103) * feat(bin): add idempotent inbox orders, receipts, replies, and readiness Let a caller supply a request id when publishing a captain inbox note so a retry returns the original note instead of creating a second one, including across the crash window between save and wake announcement. Separate saved from announced so a failed wake is repairable without enqueueing again. Add bounded receipts JSON with omission disclosure, a durable primary reply against a note id, and a read-only readiness projection that can say unknown instead of inferring liveness from a lock file. * no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict * fix(bin): resolve ready from lock-holder ancestry; drop lock status --json Remove the extra JSON surface from fm-lock.sh so its human status still always exits zero. Have the readiness projection classify the inspected home from the lock-holder pid via fm-harness.sh ancestry, with an explicit FM_SUPERVISION_MODEL still winning and an unknown model when there is no holder. Prove the yes path when that ancestry names a known harness. * no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor * no-mistakes(document): Note read-only lock inspection in scripts inventory * no-mistakes(lint): Pass missing id argument to malformed-reply test printf --------- Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118) * fix(composer): stop a harness footer row from reading as a composer holding text A harness draws its own furniture below the composer - a user statusLine, a permission-mode hint - and the cursorless "bottom-most shape wins" rule looks exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text everywhere else, so a statusLine opening with `→` was selected as a bare composer, swallowed the hint row beneath it as wrapped input, and answered `pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first doorbell and every retry were skipped and the worker never saw the steer. Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on Herdr 0.8.0 had genuinely empty composers and every one of them was refused. A separator pair that closed over a bare agent-glyph row is a proven composer container, so the contiguous non-blank rows below its closing rule are that composer's footer and are no longer composer candidates. The demotion is bounded by all three of its own preconditions: a blank row ends the zone, a pair that closed over no glyph row demotes nothing, and a shape with no separator pair at all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same composer, including a stray SGR mouse report left by a click in the pane, still reads `pending`. Pinned by two portable regressions and by a new cursorless arm on the live composer-matrix guard, which re-reads each harness's already-proven-idle pane the way every non-tmux backend reads it and fails naming the harness and version when that read is `pending`. * no-mistakes(review): make composer footer-zone demotion shape-independent * no-mistakes(review): make footer-zone demotion refuse-only and drop rescan * no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100 --------- Co-authored-by: Koen Muller <koen@catapult.nl>
…5115) Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114) * fix(bin): stop misreading a no-run branch as an unreadable runs table Defect: when `no-mistakes axi status`'s overview is truncated (a task's own branch has zero rows among the shown ones), fm_nm_select_run's Python fallback derived the repo identity for its direct SQLite query from a `repo: <path>` line it expected in the overview text. The real CLI never emits that line, truncated or not (see the genuine capture at tests/captures/no-mistakes-v1.70.1/overview.toon, which has only `count:`/`runs[...]:`), so the lookup always failed and reported "unreadable runs table" for a task that simply has no run on its branch. On a fleet with many concurrent runs, every idle-branch task hits the truncated-overview path routinely, so this fired every few minutes and drowned genuine unreadable/blocked verdicts in noise. Fix: derive the repo identity from the task worktree path instead, which is exactly the value `no-mistakes` records as a repo's `working_path` (confirmed against the existing capped-overview test fixtures, which already register repos by worktree path). A worktree path that is not absolute cannot be matched and still reads as unreadable rather than being guessed at. Also raise the reader's SQLite busy timeout from 1s to 30s so ordinary lock contention on a busy fleet cannot masquerade as an unreadable database. Safety: every other verdict byte-for-byte unchanged - the repo lookup still requires exactly one matching row (a genuinely corrupt or mismatched repos table still reports unreadable, per the existing `repo` failure-mode test), the branch query and row validation are untouched, and a zero-row result for the branch still flows through the same recursive re-parse that already turns an empty `runs[0]{...}` table into `absent`. Added a regression test (test_capped_overview_without_repo_line_and_no_runs_reports_absent) that reproduces the real overview shape - capped, zero rows for the task's branch, no `repo: ` line - and asserts the crew state falls through to the pane/busy verdict instead of reporting unknown or "unreadable". Full fm-crew-state.test.sh suite passes unchanged otherwise. * fix: recovered same-branch inventory awk misreads empty result as unreadable fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:` overview and re-runs it through the same awk selection pass. When that rebuilt inventory has zero rows for the branch, the row-matching loop never executes, so its counters (`seen`) stay at awk's uninitialized empty string while `expected` and `shown` are plain strings parsed from the header text. Comparing an uninitialized value against a non-numeric string uses string comparison, so "" != "0" is true, and the END block takes the "unreadable runs table" branch instead of falling through to the correct "absent" verdict for a branch with genuinely zero runs. Coerce the affected END comparisons with `+0` so they are always numeric, matching seen/expected/shown/total regardless of whether awk classified them as strings or numeric strings. A truncated or genuinely malformed inventory still differs numerically and still reports unreadable. * no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup * no-mistakes(review): match recorded repo path first, tolerate duplicate spellings * no-mistakes(review): revert repo lookup to exact working_path match * no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141) * fix(bin): require a non-draft pull request before a PR-based done report A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge. The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done. bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before. The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged. Closes kunchenguid#4757 * fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more than one account: every provider row carries an accountKey and one provider id may appear on several rows. fm_quota_json_valid accepted only schema 5 with unique provider ids, so fm-dispatch-resolve.sh, fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live snapshot and quota-informed dispatch was dead against the current tool. - bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with accountKey required on every row and uniqueness on provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ is the one join every consumer uses: schema 5 binds by provider alone, schema 6 binds to the row keyed by the candidate's Pi lane, else the provider's default row, else no row (unmeasured, never blocked, never by position or summed across accounts). - bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey column, and joins through the shared function. - bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through the shared function; an expanded provider with no row for the candidate's account is reported as such. - tests: schema 6 fixtures shaped like the real snapshot, each paired with a schema 5 case on the same path; every new case fails on the previous scripts and passes now. - docs: the two sentences naming the row join describe the schema 6 key. * no-mistakes(review): Fix native Codex quota and expanded provider watches * no-mistakes(review): Align native Codex account matching across dispatch paths * no-mistakes(document): Align quota documentation with account-aware snapshots * no-mistakes(document): Align quota dispatch documentation with account matching * fix(bin): keep CI lint and the quota watch test portable - bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts that source this library, so full-mode ShellCheck reported SC2034 on the assignment; mark it alongside the existing SC2016 disable. - tests/fm-procevent-quota.test.sh: the schema 6 provider-watch assertions used rg, which CI runners do not install, so the case failed with 'rg: command not found' rather than on behavior; use grep like the rest of the file. * no-mistakes(document): Documented schema-version account-row compatibility
…is retired (kunchenguid#6733) * fix(control): drop busy_gen when an incarnation is retired A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable. * no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch * no-mistakes(review): correct lock comment to name the deadlock * no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it
…#6484) * test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe Closes kunchenguid#6456 * no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes --------- Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
* fix(calm): share the standalone Pi Calm working-ship widget slot Firstmate Calm and the user-global standalone Pi Calm both install an animated working-ship widget during agent runs. Each claimed its own Pi widget key, so a session loading both (the main Firstmate home) rendered two boats. Pi replaces widgets under one key, so claiming the shared "calm-working-ship" slot keeps dual-install sessions to a single boat while a Firstmate-only session is unchanged. Pins the shared slot contract in the working-ship module test so the key cannot silently diverge again. * test(calm): pin the shared working-ship widget key in CI, document dual-install The key-parity assertion inside the Pi fixture only runs where the @earendil-works/pi-coding-agent package is installed, so CI never exercised it. Add a source-level twin that needs nothing but the tracked file, and note in docs/calm.md that the boat shares the standalone Pi Calm working-row widget slot. * no-mistakes(review): Add executable dual-install widget replacement coverage * no-mistakes(review): Guard shared widget cleanup with disposal ownership * no-mistakes(document): Document shared Calm working-ship slot behavior * test(calm): read the standalone Calm slot from its own module The dual-install check registered both boats itself under the shared slot, so it could only prove that Pi replaces a widget under one key: it would still pass if the standalone Pi Calm extension installed its boat under a different key, which is the two-boat regression the check exists to prevent. Read the standalone extension's own working-ship module when it is installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm - and drive the check with the key that module exports, so a rename on either side registers two widgets and fails naming both keys. A pinned shared-slot contract still covers a machine without the extension, and the run reports which side it used instead of passing silently over an absent extension. Verified: the touched Pi Calm suite passes and reads the installed standalone extension; with a copy of it whose key is renamed to calm-working-ship-v2 the suite fails naming the drift. * no-mistakes(review): Gate stock-row restoration by shared-widget ownership * no-mistakes(review): Removed redundant widget-key source assertions * no-mistakes(document): Document shared Calm working-ship widget ownership
…id#6649) * fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then- give-up lock acquire as the new-task-create path, but the two paths are not equivalent on contention: a create has no prior state to strand and can safely fall back to a flat layout, while a resume is recovering a specific existing identity that a concurrent recovery may legitimately be holding the lock for. Giving up there does not degrade gracefully, it hard-fails the resume outright. The suite's own concurrent cross-home recoveries test already asserts both concurrent recoveries succeed with a genuine reclaim, and the file's header comment already (inaccurately) claimed lock contention falls back to the ordinary flat layout for both paths alike, so the intended contract was always that recoveries serialize and both succeed, not that either one refuses under a short bound. Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses this file's own established fm_lock_acquire_wait idiom (already used for its other fleet-shared locks) instead of the bounded loop, and use it only at the exact-resume call site. The new-task-create call site is unchanged and keeps its bounded-then-flat-fallback behavior, which is already covered by its own passing test. Dead-owner PID-liveness reclaim inside fm_lock_try_acquire still bounds the wait against a holder that crashed mid-hold. Adds a deterministic regression test that holds the shared session lock from an unrelated process for well past the old bound, then asserts the resume succeeds with a genuine reclaim and took close to the full hold duration, so a fix that merely widens the bound rather than genuinely waiting is still caught. The existing concurrent cross-home recovery test exercises this under real timing but does not reliably outlast a fixed bound on its own. Corrects the header comment's claim that create and resume share one bounded-then-flat-fallback behavior on lock contention; they no longer do. * no-mistakes(document): Document Herdr recovery waiting for presentation lock * no-mistakes(document): Update stale hard-refusal claim in verification log * no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass) * no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain * fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait Keep refuse-by-default on presentation-order lock contention for Herdr exact resume. Callers that need concurrent recoveries to serialize must pass --herdr-resume-lock-wait; unbounded blocking on a third-party session lock is never the default. Update docs and the real-Herdr e2e suite so the default path asserts the refusal and the opt-in path asserts the wait. * no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch * docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch The relaunch path reuses the recorded endpoint and never takes the presentation-order lock, so the flag is inert there. Drop it from the --relaunch usage line and state where the flag applies. * no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper * no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed
….2 (kunchenguid#6762) * fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 Stock macOS bash 3.2 holds a HUP or TERM until a running command substitution's child exits, and the watcher read every pane through $(fm_backend_capture ...). A blocked backend read therefore held the watcher's stop for as long as the read lasted, and a stopped watcher left the hung read orphaned. tests/fm-watch-triage.test.sh test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2 for this reason while passing on bash 5. Pane captures now go through watcher_capture, which runs the read as a waited background process group recorded like a check's, so the stop is honored at once and watcher_cleanup stops a read still in flight along with its per-call output file. * no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test * no-mistakes(document): Document watcher TERM handling for blocked checks and captures * no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures * fix(bin): verify the capture group and scope the stop claim to pane reads watcher_capture now confirms its background read leads its own process group, as run_check_capture already does, so watcher_cleanup never relies on a group that set -m failed to create. The comment and continuity doc now say only fm_backend_capture pane reads go through watcher_capture; agent-state and composer-state reads still run inside command substitutions.
* feat(spawn): add per-home worker tool exclusions Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed. It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents. Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names. Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it. Malformed entries are refused before provisioning, and before a relaunch stops a running worker. Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker. Closes kunchenguid#6744 * no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior * no-mistakes(document): Clarify worker tool exclusion documentation * no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run
kunchenguid#5343) * refactor(bin): share the local Firstmate home walk from the wake library Teardown's walk over the root home and its registered local secondmate homes moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to fm_firstmate_root_home, so a second consumer can count task records across this machine's homes without a copy. Teardown keeps its exact refusal wording through a thin wrapper. * feat(bin): defer spawns beyond a project's declared machine capacity A project whose machine-local resource only serves a few workers at once had no way to tell Firstmate so: every queued item was launched, and the surplus workers spent full-context turns retrying the resource. config/project-capacity in the root home now declares how many workers each named project admits at once on this machine. bin/fm-spawn.sh counts the ship and scout records on the same project origin across the root and its local secondmate homes, skipping ones whose ready PR is recorded, while holding the shared project lock through publication. A spawn with every place held exits 75 before any brief render, endpoint, worktree, record, or backlog move, so the item stays queued; batches report it as deferred. Undeclared projects keep today's uncapped dispatch, and an unreadable declaration refuses rather than guessing the limit. Refs kunchenguid#4237 * no-mistakes(review): Document that capacity matches the clone directory name * no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment * no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting * no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment * no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh * fix(bin): declare capacity for a project name that begins with # A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment. * no-mistakes(document): Rewrap project-capacity library header comment * no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh * fix(bin): release the Herdr session lock when reclaim finishes A concurrent resume in another home waits five seconds for that lock. Reclaim is the last presentation change on the recovery path, so holding the lock through the launch tail made the waiter time out. The contributions arm check also freezes its one-second clock, the same way the budget tests do, because an unfrozen clock can tick past before the first forge read. * no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim * no-mistakes(review): Skip the spawning task's own record in capacity count * no-mistakes(review): Restore release test comment above its test * docs: scope PR-ready re-evaluation to a declared project capacity A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.
… section (kunchenguid#6785) fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE only when session_ended is true and CAPTAIN MESSAGE otherwise, but the count line always said session_ending_message_count. Several composer messages on a still-open board were therefore counted as session-ending. The count line now follows the same session_ended switch: session_ending_message_count once the session ended, captain_message_count otherwise. Message rows stay out of the annotation count, per triage. Fixes kunchenguid#6743
…henguid#6780) The Herdr presentation lock namespace was the fixed machine-global /tmp/firstmate-herdr-presentation, so on a host where two OS users run Firstmate on Herdr the first account to create it owned it and every teardown from the other account was refused with no way to clear it. Suffix the namespace with the account uid. The owner-uid and mode-700 checks are unchanged, so a foreign-owned or wrong-mode name at this account's path is still refused and never adopted, chowned, or removed. Fixes kunchenguid#4716.
…te (kunchenguid#6809) The OpenCode session plugin's shouldArm kept its own copy of the need test that only looked for in-flight task records, while the turn-end guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh, which also counts registered process-event sources and trusted custom checks. With an empty fleet but any registered source or check, the guard blocked every turn end while the plugin declined to arm - a loop the guard's own repair line could not resolve because it names the plugin as the fix. The plugin now delegates the decision to the shared predicate through bash, keeping the local away-record decline and the x-mode.env arm override. OpenCode plugin test fixtures now carry the real predicate their arming path sources, and the arm suite gains six cases asserting the plugin's decision against the shared verdict over the same synthetic state directories. Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792) * fix(bin): resolve a pending reply only from its own task's status line Remote reply ingestion handed every corr= token in a mate's payload to fm_pending_reply_try_resolve together with that mate's own status log, so one mate echoing another mate's token resolved the other request. Honor a status-file override only when it is the record's own parent_status, and match the corr= token as a whole word. Fixes kunchenguid#6538 * no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784) agent-skill-trigger-index claims to be the complete agent-only trigger index but omitted operational-home-layout, session-start-recovery, validation-supervision, ship-landing, scout-completion, and away-quiet-supervision. Add each with its own description's trigger, placed beside the related entries. The decision-hold-lifecycle redirect stub stays out, per triage. Fixes kunchenguid#6503
…nchenguid#6814) * test: share a rename-safe agent stand-in across liveness suites On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to run when invoked through a symlink named after another utility. The Herdr descendant process-walk tests built their agent-named process as a `pi` symlink to the host `sleep`, so the process exited at once, its parent shell was gone before the walk ran, and both cases read `unknown unreadable` and failed on that host. The suite stops at its first failure, so every later case went unrun. The Herdr control smoke test's `claude` symlink has the same construction. The tmux liveness suite already solved this with a host-compiled spinner and a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the Herdr control smoke test now use it. When no stand-in can survive a foreign name, a case skips with the reason instead of failing. tests/fm-test-fixtures.test.sh gains a portable regression with a fake multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose. * no-mistakes(document): Correct Herdr verification fixture reference * ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787) * fix(bin): keep the supervision host's pass-through successor out of the hook's process group The successor a main-only pass-through leaves for main shared the Stop hook's process group, so the harness tearing that group down after the exit-2 rewake stopped it. The stop published downtime and the next park's first cycle announced an empty check: rearm-resurface, which woke main again in a loop. Start that successor in a process group of its own, as the hook's own handling successor already is. * no-mistakes(review): Give the at-turn successor left for main its own group * no-mistakes(document): Document own-group successor for turn-start hand-back too * no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823) * test: use idle composer readiness for Claude tmux guards * no-mistakes(test): Fix attended supervision test expectations and isolate worker state * no-mistakes(document): Correct live guard coverage and readiness documentation * no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass * test: preserve attended successor close assertions * no-mistakes(test): Fix attended live test watcher takeover expectations * no-mistakes(document): Correct stale attended guard documentation * Revert "no-mistakes(document): Correct stale attended guard documentation" This reverts commit 8e59d89. * Revert "no-mistakes(test): Fix attended live test watcher takeover expectations" This reverts commit c0b8510.
* test: order stale watcher-lock fixture races explicitly * no-mistakes(review): Removed duplicate stale-steal reap call
Absorbs the 234 upstream kunchenguid/firstmate commits since the last absorb (upstream 9bc051f, squashed into the fork as 7a03435). That squash dropped the upstream parent, so conflicts were resolved against 9bc051f as the effective base; this commit records upstream/main as a real parent so the next absorb starts from the correct merge base. Fork features carried onto upstream's restructured code: the Codex account axis (dispatch resolver, spawn, relaunch, quota watch), voice-turn answering and live-call scoping, the iMessage conversation destination, fast surfacing of queued phone/texted turns to a running watcher, teardown of records whose pooled worktree was returned or reassigned, the GitHub App checks reader in the merge guard, Bearings, and PR state, and the brain-room/shared-interface pieces. Where upstream built the same thing, upstream's version is kept and only the fork's missing behavior is ported onto it.
…m absorb Main's deliberate removal of the shared-interface lock hooks wins: the steal-lock path returns to upstream's exact form, and the shared-interface test call this absorb had re-homed is dropped with the test itself.
…pstream's guidance The absorb left two test.instructions keys in .no-mistakes.yaml: the fork's what-the-product-is runbook and upstream's live-lab rules. Go YAML refuses the duplicate, so no-mistakes could not load the repo config at all. Both runbooks now live in the one block, upstream's under its own heading.
…e-repair verification
|
Land this as a merge commit, not a squash. What this absorbs234 upstream commits, from What changes in Firstmate's behaviorChanged defaults
New
New opt-ins (no effect until configured)
Removed
Needed in this home before restart
Where both sides built the same thingThe rule applied: take upstream's version only when it does everything the fork's does (proven by running the fork's own tests against upstream's code); otherwise keep the fork's, or port the fork's missing piece onto upstream's version.
Close cousins that are not the same feature, so both were kept: upstream's Claude/Pi worker account pin vs the fork's per-task Codex account axis; upstream's quota-axi schema-6 account rows vs the fork's per- Fork features carried onto upstream's code
Open fork work
TestingRun on hermes against this branch.
Validation notes
|
Land this as a merge commit, not a squash. It records upstream kunchenguid/firstmate
19fcbbdeas a real parent (merge5094dad5), so the next absorb starts from the right merge base; #9 was squashed, which is why this one had to re-resolve from 2026-09-17.The full absorb write-up (overlaps and the evidence for each, fork features carried, test results, declined upstream findings) is in the first PR comment.
Behavior changes for the captain
config/supervision-host-off); AI co-author trailers are stripped from fleet-launched commits (config/keep-ai-trailersopts out); away mode acts on the captain's away words;/quietis a statement when attended supervision runs; the merge guard refuses unreported required checks and retries a still-computing mergeability; AGENTS.md is much shorter, with detail moved into new trigger-loaded skills.Needed in this home before restart
tasks-axito at least 0.2.6 (hermes has 0.2.5; merges need it),quota-axito at least 0.1.51 (has 0.1.36), andlavish-axito at least 0.1.80 (has 0.1.64).config/supervision-host-off.Intent
The captain, 2026-10-09: "the main repo has advanced a lot more. I'm thinking we should do an upstream absorb." Firstmate recommended starting it now as its own lane and landing it before the hardening rollout (so the rollout runs on the code we keep), at a quiet point; the captain said "yes start the absorb lane now". Upstream kunchenguid/firstmate has about 232 commits since our last absorb (#9, 2026-09-19, which merged upstream kunchenguid/firstmate into this fork and carried the fork's Codex account axis through typed dispatch, but landed as a squash, so upstream's commits are not ancestors of the fork); our fork carries its own features since then. Nothing of ours may be lost; prefer upstream's design wherever it already covers something we built. Smallest durable result, the way Kun builds Firstmate.
The captain, later: "I don't want to lose any of the functionality we've built."
The captain, later: "If upstream has a better way of doing it, definitely absorb upstream's. I agree we should prefer upstream's version."
What Changed
Risk Assessment
🚨 High: The change can duplicate handled instructions, strand live voice calls, and remove project-memory functionality explicitly required to survive the absorb.
Testing
Corrected baseline fixture-path and dependency issues, then drove targeted runtime suites and manual CLI checks that passed the prior regressions. The remote inheritance serialization fixture failed repeatedly without establishing a live scenario result; test cleanup was fixed and verified, CLI and generated-brief evidence retained, and disposable files removed. Native harness, remote-host, forge, audio-delivery, and gate/CI checks remain untested; no UI image was captured because this phase forbids launching a real harness.
Evidence: Inbox, supervision-host, and process-event runtime transcripts
Source: Inbox, supervision-host, and process-event runtime transcripts
Evidence: Initial targeted runtime checks; account and capacity setup failures superseded by corrected runs
Source: Initial targeted runtime checks; account and capacity setup failures superseded by corrected runs
Evidence: Closed stdout returns failure and preserves the unread outcome
Source: Closed stdout returns failure and preserves the unread outcome
Evidence: Gate-worktree lifecycle refusal with the environment marker absent
Source: Gate-worktree lifecycle refusal with the environment marker absent
Evidence: Memory diagnostic output and threshold exit statuses
Source: Memory diagnostic output and threshold exit statuses
Pipeline
Updates from git push no-mistakes
... (18 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 Fix applied.
2 warnings still open:
tests/fm-remote-secondmate-lifecycle-e2e.test.sh:988- The remote inheritance serialization fixture repeatedly fails before reaching its deliberately blocked inheritance write. Correcting fixture placement and gate-refusal setup did not resolve it; both the full fixture and focused reruns failed, including the final run after cleanup repairs. The concurrent config-push convergence assertion therefore never executes. The captured spawn output contains only a watcher-down reminder, leaving the cause unresolved between product behavior and fixture infrastructure. Diagnose the blocked spawn and establish a reliable completed serialization check. Evidence: round4-remote-inherit-final.log. Test-only cleanup was fixed and verified separately.TMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh --jobs 1 tests/fm-inbox.test.sh tests/fm-supervision-host.test.sh tests/fm-procevent.test.shTMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh tests/fm-worker-account.test.sh tests/fm-fleet-ledger.test.sh tests/fm-git-strip-ai-trailers.test.sh tests/fm-wake-drain-voice-hold.test.sh tests/fm-wake-drain-voice-first.test.sh tests/fm-inbox-conversation.test.sh tests/fm-project-capacity.test.shTMPDIR="$PWD/.test-phase-tmp" bin/fm-test-run.sh --jobs 1 tests/fm-afk-contract.test.sh tests/fm-afk-return.test.sh tests/fm-brief.test.sh tests/fm-gate-refuse.test.shMaterialized a disposable copy of tracked source with sibling fixture homes; reran worker-account and remote lifecycle tests throughbin/fm-test-run.sh --jobs 1.npm install --prefix .test-phase-tmp/round4-deps --no-audit --no-fund tasks-axi@0.2.6for the isolated capacity fixture.PATH="$PWD/.test-phase-tmp/round4-deps/node_modules/.bin:$PATH" TMPDIR="$PWD/.test-phase-tmp/round4-fixtures" bin/fm-test-run.sh --jobs 1 .test-phase-tmp/round4-source/tests/fm-project-capacity.test.shbin/fm-test-run.sh .test-phase-tmp/round4-codex-account.test.shusing selected existing account cases and unique disposable task IDs.bin/fm-test-run.sh .test-phase-tmp/round4-brief-behavior.test.shusing selected public generator cases and a relative shell-injection marker valid under the fixture Git ref contract.FM_TEST_ONLY=<selector> bin/fm-test-run.sh --jobs 1 tests/fm-pr-check-security.test.shfortest_invalid_entrypoints_have_zero_side_effects,test_gerrit_ready_gate_reads_the_published_tree,test_unpushed_named_head_refuses_registration, andtest_valid_recording_and_merge_derivation.bin/fm-test-run.sh tests/fm-host-mirror.test.shpython3 .test-phase-tmp/round4-manual.py: isolated project-mode, generated briefs, verbatim posture records, and acknowledged inbox replay/announce checks.python3 .test-phase-tmp/round4-closed-output.py: real drain with closed stdout, unread-state inspection, and subsequent successful presentation.Ran realfm-spawn.sh,fm-send.sh, andfm-teardown.shagainst an empty isolated home with the gate environment marker absent; verified refusal and unchanged state.Ran realfm-jev-mem-guard.sh --jsonand--checkwith all warning/critical thresholds set to 1000, then 0.Replayed the acknowledgement failure with the predecessor's real inbox executable, then compared the target's durable acknowledgement output.Ran opt-in supervision-host, attended-host, mirror, quota-dispatch, Devin, and Herdr guards without forcing their control variables; recorded their skips. Also rantests/fm-afk-inject-e2e.test.shwith its portable terminal fixture.Ran the full remote lifecycle fixture in the corrected disposable source, then repeatedly ranTMPDIR="$PWD/state/round4-fixtures" bin/fm-test-run.sh tests/fm-remote-inherit-focused.test.sh; retained failure diagnostics.bin/fm-test-run.sh .test-phase-tmp/round4-cleanup-check.test.sh: verified cleanup stops a signal-resistant owned writer and removes protected fixture directories; the final inheritance run also completed cleanup without errors.Removed disposable source copies, fixture roots, local dependencies, and temporary drivers; verified only the intentional test cleanup change remains in the worktree.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.