feat: sync upstream supervision and harness capabilities - #26
Merged
Merged
Conversation
* fix(pi): deliver supervision outcomes off Pi's render thread The supervision branch runs inside the captain's own Pi process, and Pi runs extensions, their tools, and their event handlers on the single JavaScript thread that also draws the TUI and reads the keyboard. Every delivered outcome ran roughly five bash script invocations plus several `ps` calls through spawnSync on that thread, so the TUI could not repaint or echo a keystroke for the whole chain - the subsecond freeze the captain saw every time a routine or captain-facing outcome arrived. Convert the delivery path's subprocess calls to an awaited spawn behind a serializing queue. lib/fm-async-exec.ts is the single owner of the awaited-spawn replacement and returns the same capture shape and failure verdicts spawnSync returned. Awaiting yields the thread, so what the single thread used to guarantee for free is now an explicit queue: every delivery, acknowledgement, and turn-boundary reconciliation runs as one unit of it, preserving the durable append before anything visible, one delivery at a time in sequence order, the read cursor advanced before the next reader sees a row, and one ownership activation per generation. Cancellation is preserved by the generation and lock-ownership rechecks the awaits are placed around. Two reads stay synchronous because Pi's own API is synchronous there, not as an optimization: its bash spawn hook is typed as a plain function, and the watcher reads offer.accepted the moment its dispatch event returns, so a session that does not own the fleet lock must still refuse a wake without waiting. Both walk the lock's process ancestry in full every time, never cached, because reparenting and pid reuse can invalidate a remembered chain and that answer decides ownership rather than hinting at it. The store scripts and their durability contracts are unchanged. Measured through the real fm_branch_report tool and real bin/ scripts with a 1 ms interval timer, the largest block of the JS thread falls from 273 to 2.0 ms for a routine outcome, 286 to 2.0 ms for a captain outcome, and 134 to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle floor. In a real Pi 0.82.0 TUI the worst keystroke echo while two outcomes arrive falls from 676.9 ms to 36.8 ms, against a 22.6 ms extension-free floor. Regressions: a delivery must leave the event loop running (zero timer ticks before this change, in 250 ms), interleaved reports stay ordered and exactly once, a session replaced mid-delivery neither loses nor duplicates an outcome, and a failing store script surfaces without losing or doubling one. The real-TUI half is an opt-in live guard that types into an isolated Pi pane while outcomes are delivered and fails if echo leaves the class of the same machine's own floor. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(review): Revalidate ownership and bound asynchronous subprocess output * no-mistakes(ci): Fixed CI defects: routine outcomes now persist a sequence-keyed delivery receipt before awaiting cursor advancement, preventing duplicate delivery after mark-read failure. Corrected the session-replacement test to exercise an actual asynchronous ps ancestry lookup. Targeted behavioral tests, strict Pi typecheck, ShellCheck, and diff checks pass. The full extension test remains locally blocked by an unrelated stock-render assertion under the installed Pi runtime. The no-mistakes attestation failure is external pipeline state (test was previously skipped), not a source defect * fix(pi): keep the declined routine receipt out and skip the renderer case below its Pi floor Four follow-ups on the same branch, plus one revert. Revert the routine-delivery receipt a CI auto-fix round added. It introduced a new persisted `fm-branch-routine-delivery` entry, written into the captain's transcript for every routine note, to deduplicate a note whose cursor write failed. That is a change to the delivery contract, which this task is not authorized to make: the approved work is the asynchronous conversion with the existing durability contract preserved. The ownership re-read and output bounding from the review round are kept - both are genuine asynchronous correctness, not contract changes - as is that round's use of a real parent pid so the replacement regression traverses an actual ps subprocess. Record the routine gap instead of closing it. A routine note is a plain message with no sequence-keyed record, so a mark-read failure after delivery makes the next reconciliation send it once more; a captain row cannot duplicate that way because its visible entry is found by store sequence. That asymmetry predates moving delivery off the render thread. It is now stated at the call site and in the delivery-contract docs, tracked as fm-pi-routine-delivery-idempotency-followup-r1, and pinned by a regression that proves the routine note is re-delivered exactly once more and never again, the captain entry stays single, and the store keeps both rows. Give the stock-renderer case a Pi version floor. It compares the extension's renderers against Pi's stock rendering, so its verdict only means anything against the contract those renderers target: since 0.84.4 the stock renderer no longer supplies an implicit reset at multiline boundaries and the extension emits that reset itself, so an older installed Pi differs legitimately. It now names the installed version and the floor and skips, while a package whose version cannot be read at all still fails. Make the responsiveness regression's second signal a fraction rather than a millisecond budget. A loaded machine that deschedules the process inflates an absolute stall budget into a false failure, but it inflates the delivery's own wall time too, so requiring the worst stall to be a minority of that wall time holds under load. Synchronous delivery sits near 1.0 there whatever the load, and the tick-count signal still reads zero on it. Replace the test-family mapping for the Pi extension libraries with per-script targeting. Routing them to whole families - or leaving them unmapped, which widens through the reference scan to each referencing suite's entire family - selected dozens of suites with nothing to do with Pi and pulled an unrelated flake into the run. The changed-file selection drops from 112 scripts to 61. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(document): Clarify asynchronous execution documentation --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…nguid#3763) * fix(memory): honor explicit project maintenance guidance * no-mistakes(test): Blocked by pre-existing Bash and Muse fixture failures * no-mistakes(test): Remove accidentally tracked test attribution report * no-mistakes(ci): Restricted the marker to the exact first line, preventing fenced examples from suppressing governance, and corrected the documentation. Regression failed before the fix; all 18 helper tests, focused ShellCheck, documentation validation, and diff checks pass. CI and Require no-mistakes report action_required with zero jobs executed; those external checks remain unresolved
…nguid#3783) * fix(spawn): refuse repository primary from linked spawning homes Compare the resolved task git directory with the spawning repository's common git directory before refreshing a fresh copy or relaunching a task. This protects the primary even when the spawning project is a linked home. Keep pooled copies accepted and preserve recorded work on relaunch. Fixes kunchenguid#3741. Verification for the pipeline PR body: - Red on origin/main 1820316 with the new regression and unchanged production code: bin/fm-test-run.sh tests/fm-spawn-pool-base-freshen.test.sh exited 1 with "linked spawning home accepted primary as a disposable copy". - Green after the guard: the complete pool-base-freshen and control-relaunch suites passed through bin/fm-test-run.sh, covering primary and symlink refusal before fetch/reset, spawning-directory refusal, scout acceptance, and committed plus unfinished work preserved during linked-home relaunch. - The worktree-settle suite passed on pristine main and the final branch. An earlier loaded-host run exceeded its five-second assertion (6s); the final retry passed without changing code or the assertion. - Test fixture commits ran with GIT_CONFIG_COUNT=1, GIT_CONFIG_KEY_0=commit.gpgsign, GIT_CONFIG_VALUE_0=false. - bin/fm-lint.sh and /bin/bash -n for all three changed scripts passed. The upstream cwd-selection cause remains outside this change. * no-mistakes(document): Clarify spawn isolation ownership and relaunch preservation
…kunchenguid#3785) * fix: read a failed herdr CLI as unreachable, not a gone backend target The no-run fallback in bin/fm-crew-state.sh collapsed every failed pane capture into 'backend target gone', which downstream consumers treat as positive death evidence - so a herdr CLI that errors or stalls under load briefly scored dozens of live claims dead on a busy box. Only a successful herdr answer proving the pane absent (fm_backend_agent_state's 'missing', backed by pane get answering pane_not_found) may now read as gone; every other verdict reports 'backend unreachable' with the endpoint state, which is never positive death evidence. Adds a behavior test: an always-failing fake herdr reads unknown/unreachable, never gone. * test: pin the herdr suite's ambient home to a marker-free fixture FM_HOME defaults to the suite's own root when unset, and any secondmate- marked checkout (every treehouse crew home carries .fm-secondmate-home) flips the default workspace label to 2ndmate-*, so the ambiguous-label placement test found zero firstmate matches and fell into the create path instead of refusing (expected exit 3, got 1) - deterministically green in CI, deterministically red from a crew home. Export a marker-free ambient FM_HOME fixture; per-test FM_HOME prefixes still override it. * fix: classify herdr endpoint answers instead of every non-missing verdict Review decision (firstmate, 2026-09-05): a failed pane capture is not itself evidence of death, but neither is every non-missing classifier verdict a failed answer. missing (pane get answered pane_not_found) and dead (pane present, agent_not_found husk) keep gone-class text so a stale-claim sweep may still reclaim them; an alive answer falls through to the normal busy/state flow instead of being discarded when only the heavy 200-line scrollback read failed; only when the cheap pane get / agent get calls themselves fail to answer does the line read 'backend unreachable'. Adds the two missing cases: alive with a failed scrollback read stays live, and a husk pane still reads gone. * no-mistakes(review): route tmux through agent-state classifier; drop test stall * no-mistakes(review): narrow inaccurate tmux socket and alive-arm fallback comments * no-mistakes(document): document classifier-backed endpoint verdicts in crew-state contract
) * fix: distinguish subshell wake-lock owners on stock Bash Restore distinct process ownership for issue kunchenguid#3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff. The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass. * no-mistakes(document): Correct lock grace-period documentation * no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged
…backends (kunchenguid#3782) * fix(bin): close legacy records on the Beads backend honestly Two pre-Beads reads blocked honest closure of leftover records: 1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids only against the live backend and the pre-collapse derived identity, so a home whose holds fm-hold-migration rehomed under fm- ids failed with an empty-name absence message (the resolve failure was swallowed by the command substitution feeding verify_hold_durable). Resolution now falls back, on the Beads backend only, to the legacy id under the configured beads prefix and to the row whose notes carry the exact marker line 'migrated from data/backlog.md id <legacy id>'; every refusal names the id it could not resolve, and the markdown path is unchanged. 2. fm-teardown.sh refused any record without spawn_gen forever. A record that predates the field can now be torn down with an explicit --legacy-record flag once the recovery-grade endpoint classifier confirms the recorded endpoint dead or agent-less; the accepted incarnation is stamped into the record right before its close marker binds to it and named in the teardown line. Refusals leave the record byte-identical, the unlanded-work refusal is not relaxed, and a corrupt (multi-valued) spawn_gen is never accepted. The companion repair this branch carries (follow-up commit) is the backend-gated --file and markdown-file requirement in the mutate path and lifecycle gates: fm_backlog_mutate passed --file and required the markdown backlog file regardless of the resolved backend, and the transition gate plus row probe required that file before any backend work, so a home on a non-markdown backend could neither gate, probe, nor close its rows. Behavior tests: self-contained beads fixtures over a scratch bd graph (self-skipping on markdown-only tasks-axi installs), legacy meta fixtures for every teardown gate, and the relocated markdown backlog coverage stays green. * no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records * fix(backlog): address the configured backend for lifecycle writes Completes the fm-backlog-transition-lib repair the first commit's message claims: on this base fm_backlog_mutate passed --file and required the markdown backlog file regardless of the resolved backend, and fm_backlog_transition_applies plus fm_backlog_row_probe required that file before any backend work, so a home on a non-markdown backend could neither gate, probe, nor close its backlog rows. All three now gate the markdown file on the resolved tasks-axi backend: markdown keeps exactly its explicit <data>/backlog.md behavior, non-markdown homes address the backend their own configuration selects with no markdown file requirement. fm_backlog_row_show and fm_backlog_row_list already gated correctly and are unchanged. docs/configuration.md owns the contract line. Also extends the same backend gate to fm-captain-hold.sh's own mutation wrapper - hold/add/update/answer/done append the markdown --file only when the resolved backend is markdown, so a captain call on a Beads home reaches the Beads store end to end - and applies the review round's two direct remedies there: the [beads] graph path resolves against the backlog root when relative (never the process CWD), and a failed bd graph read reports bd's own trimmed stderr reason in the refusal. Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads completion case proving the transition gate applies, the row probe reads, and done runs without any markdown file or --file override; the relocated markdown backlog test stays green. * no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution * test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib The legacy-records change made fm-backlog-transition-lib.sh resolve the configured backend via fm_tasks_axi_backend before the markdown-only skip. The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the markdown check fell through and teardown hit the incompatible-backend error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as markdown and define the floor, restoring the intended no-backlog skip. * fix(teardown): roll the legacy stamp back when the close marker fails A legacy-record teardown stamps its accepted incarnation into the record right before the close marker binds to it; when that marker write then fails, the stamp survived, so a retried teardown sailed past the dead-or-agent-less endpoint gate the stamp now proved unnecessary. The failed marker write now truncates the record back to its exact pre-stamp bytes (verified by size), restoring the byte-identical-refusal invariant; when the rollback itself fails the operator is told to re-run with --legacy-record after reconciling the endpoint. Also completes the recorded review decision's coverage wording: the beads stub test now drives the answer close end to end (update and done through the gated wrapper), asserting no markdown file override reaches either verb. * fix(review): harden the legacy stamp rollback and resolve derived migrated ids The legacy-record stamp rollback now uses perl (already in the teardown curated PATH; truncate is not, and is absent on stock macOS), routes every failure branch inside the stamp block through the same size-verified rollback so the byte-identical-refusal invariant holds on those paths too, and gains behavior coverage: an unrecordable close (an invalid pr= link) fails the teardown, leaves the record byte-identical, keeps the backlog row in flight, and a flag-less retry still refuses. Migrated-hold resolution now probes the derived pre-collapse identity (<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration recorded the DERIVED id in every migrated row's marker note - in both the prefix and the migration-note forms, with the ambiguity refusal naming every identity tried, plus behavior coverage for a bare decision key resolved through its derived identity's marker. Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE paragraphs with the backend-gated contract, drops an unreachable FORCE validity guard the parser rewrite left behind, and switches the new stub fixture to the portable sed -i.bak idiom. * no-mistakes(review): Name the configured backend in teardown's backlog reminder * no-mistakes(review): Scan migration markers before the prefix guess * no-mistakes(review): Document marker-first resolution and cover the prefix branch * no-mistakes(document): Record prefix-attestation audit and marker-line forms * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean
* fix(lint): drop source following on the local changed-file gate The local lint step was inlining library closures through --external-sources and peaking above 8 GB on a single root. Keep full analysis in CI, on main, and without a merge-base; exclude the four cross-file codes from the local pass so those findings still land in CI. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Run local ShellCheck per root and document measurements * no-mistakes(review): Correct local source-following telemetry * no-mistakes(document): Clarify context-sensitive lint documentation --------- Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(pi): route needs-decision wakes and mixed batches wholly to main Skip the supervision branch for every needs-decision status append, the same way a check-kind wake already skips it. A coalesced signal/stale trigger batch containing any needs-decision row is delivered wholly to main, not split between the branch and a later main wake - the whole batch, including any co-present routine rows for a different task, travels together. Heartbeat and unread-status scans stay independent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * test(pi): cover distinct-file mixed batches and heartbeat independence Add a regression using two distinct files (not the same status file twice) in one coalesced trigger so a some-vs-every regression on the file-list cross-reference cannot hide behind a degenerate same-key case, and a heartbeat/needs-decision co-presence test proving a needs-decision row neither vetoes nor rides along with an otherwise eligible heartbeat scan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE * no-mistakes(document): Clarify needs-decision and heartbeat routing * no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect * no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect * no-mistakes(review): Route stale open decisions directly to main * no-mistakes(review): Honor configured verbs in stale decision routing * no-mistakes(review): Route second-mate escalations and configured decisions to main * no-mistakes(review): Ignore trailing whitespace after captain holds * no-mistakes(review): Cache stale decision classification per status file * no-mistakes(review): Document unread decision precedence for later task wakes * no-mistakes(review): Cache unchanged stale decisions across scope scans * no-mistakes(review): Resolve decision aliases and reject symlinked statuses * no-mistakes(review): Route surfaced captain-held signals directly to main * no-mistakes(document): Document decision-owned main routing * no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034) * no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…d#3802) * feat(spawn): add an opt-in worker environment allowlist Honor a home-local launch-env-allowlist at the shared worker command boundary and inherit it into secondmate homes. Preserve the existing launch behavior when the file is absent. Keep the operational environment and explicit launch assignments, and account for filtered Muse credentials. Refs kunchenguid#3742 Verification: - Red on origin/main 1820316: the new enabled-allowlist regression observed synthetic-unrelated in the worker; the absent-file control passed. - Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness, and fm-trace-context-spawn; all three passed without skips. - Synthetic emitted-command probes ran through sh, stock Bash, and zsh. - Canonical lint, documentation audience checks, and stock Bash syntax checks passed. * no-mistakes(review): Reject inaccessible launch environment configuration * no-mistakes(review): Preserve inherited allowlists on source inspection errors * no-mistakes(document): Clarify worker environment grants and inheritance documentation * no-mistakes(lint): Fix inheritance test ShellCheck source boundary
…kunchenguid#3575) * feat(bin): add rovo as a verified crewmate/scout worker harness Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr adapter contract: detection with marker-precedence ordering, one-shot positional launch with --startup-receipt readiness polling instead of composer scraping, model/effort flags, a screen-scrape busy fallback scoped like grok's, and crew/scout-only lifecycle control that refuses secondmate launches. Ships with a portable regression suite, a live PTY guard against the real binary, a per-harness reference doc, and a dated verification record covering the silent OAuth refresh, the interrupt-ack divergence from the originating scout report, and the still-open composer-ghost and tmux/herdr pane-liveness gaps. * no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt * no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape * no-mistakes(document): add rovo to stale worker-harness enumerations in docs * docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence Placement, launch-then-send, and busy/idle rendering are now verified live in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since the cross-session launcher-identity guard refuses this task's own ambient Herdr identity for a full fm-spawn.sh run. fm_backend_agent_state reported dead for a live, responding rovo pane at every point checked, because herdr's own agent-integration registry has no rovo entry (herdr integration status), so herdr agent get returns agent_not_found regardless of whether rovo is actually running. This is recorded as a Herdr-side integration gap rather than a firstmate bug, left unpatched to avoid a false-positive alive verdict for other idle shells. Updates docs/verification/rovo.md's backend-liveness section and its two cross-references (docs/verification/runtime-backends.md, docs/configuration.md) accordingly. * fix(bin): close rovo's failed-spawn leak and busy-scrape false idle Greptile P1s on PR kunchenguid#3575: a failed rovo readiness/submission/delivery gate exited without tearing down the just-created endpoint, leaving the launched --yolo rovo process running as an orphaned agent outside task control. Separately, the busy classifier's rendered-tail fallback returned definitive idle whenever the "Rovo is thinking" marker scrolled out of the last 12 nonblank lines of a long turn, which could make supervision wrongly conclude a still-working worker had gone idle. fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and terminal remain owned by the separate ORCA_ABORT_CLEANUP trap. fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex" instead of "idle rovo-regex" when the marker is absent, matching how muse and cursor already express "can't tell" for their own fallbacks. The positive busy match is unchanged. Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure tests now assert the endpoint is torn down (and the success test asserts it is not), and a new test drives the busy marker out of the tail window to confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both changed files. * test(rovo): align spawn fixture with the launch-brief validation contract Upstream main now requires a brief's ## Captain's intent and ## Firstmate spec subsections (or a nonempty legacy # Task body) before spawn, and rewrites ship+no-mistakes briefs into launch-brief.md. Update the rovo harness fixture and pointer assertions to match, mirroring the kimi harness fixture. * no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence * fix: prevent stale supervision wake loops (kunchenguid#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): grant rovo the per-task home paths its standard crewmate flow needs rovo confines every file-tool operation to its worktree by default, and its bash tool independently refuses the same external paths regardless of any grant (confirmed live), so a rovo worker could not read its own brief or steering messages or write its status/report - all of which live in the firstmate home outside the worktree - without hand-feeding it. Grant toolPermissions.allowedExternalPaths for exactly the task's brief directory, steering inbox, and status file at launch time via --config-override, merged with agent.efficiencyLevel into one JSON object since that flag is single-value and silently discards a second occurrence. Extends the live PTY guard to prove, against the real binary, that the grant lets rovo read an external brief and append to an external status file, and that the same flow is blocked without the grant. * no-mistakes(document): align rovo reference Effort row with merged single --config-override * no-mistakes(document): document rovo file-access grant in harness reference --------- Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…guid#3813) * fix(bin): read a crew's pipeline-death claim against the live run A crew's no-mistakes drive call blocks until the next gate or outcome, routinely far longer than its harness lets one command live, so the call gets killed or times out while the daemon runs the fix round on in the background. Crews read that as daemon death and block on it, and firstmate had nothing that contradicted them. Rule 7 of every generated brief now says a drive-call error or a harness command timeout is not a daemon error, requires `no-mistakes daemon status` plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves that report for a refused socket or a run record failed with a daemon error. The no-mistakes definition of done adds the harness command limit and the background-and-poll shape that fits inside it. fm-crew-state gains one classification case: a `blocked:` line blaming the daemon, a timeout, or unreachability, while the run is running or fixing AND the pipeline reports fresh activity, now reads as superseded because the run is alive. Recency comes from the client's own `quiet` marker on active_steps.last_activity rather than a threshold invented here, and positive evidence is required, so a run record that outlives a genuinely dead daemon keeps the plain reading. stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook: firstmate reads both statuses itself, steers a reattach, never restarts the shared daemon on a crew's claim, and escalates only a refused socket. Nothing here depends on an unshipped no-mistakes capability. * no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching * no-mistakes(review): Honor socket refusal across coarse status and crew guidance * no-mistakes(test): Replace flaky settle timing assertion with pane-read count * no-mistakes(document): Document daemon timeout recovery contract * no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks
…unchenguid#3797) * fix(brief): scope Firstmate workers to their launch contract * no-mistakes(document): Clarify supervisor scope and worker contract ownership * no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed * no-mistakes(review): make launch overlay sole owner of worker role contract * no-mistakes(review): narrow heading test dimension and fix publish error wording * no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin * no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header
…d#3823) * fix(bin): stop ringing steering doorbells into dead panes The steering-inbox doorbell was a plain sentence plus Enter typed into a worker's pane, and the watcher re-rang it on the assumption that a ring is free. In a pane whose agent has exited that line is a shell command, and the re-ring ladder kept typing it into a shell that can never acknowledge it. - Prefix the doorbell with the shell no-op `: ` so a bare shell executes nothing while a live worker still reads the same self-describing line. `#` is not used because interactive zsh does not treat it as a comment by default and the claude harness binds it to memory mode. - fm_task_inbox_ring skips the pane (return 3) when the backend positively classifies the agent as dead; missing, ambiguous, unreadable, and unverified endpoints still ring so a blind classifier never starves a live worker. - The watcher caps the ladder for a dead pane: one stale wake for recovery, no ring, no ladder walk, and the durable record stays for stuck-crewmate-recovery. fm-send and the remote steer leg report the skip. Tests cover the no-op in real shells, the dead/live/unclassifiable ring verdicts, and the single-surfacing watcher path. * no-mistakes(review): Quote doorbell paths against shell injection * no-mistakes(review): Reject terminal-control paths before ringing * no-mistakes(review): Document accepted partial doorbell delivery race * no-mistakes(review): Skip unavailable endpoints before busy-state handling * no-mistakes(test): Respect shell startup PATH in environment allowlist test * no-mistakes(test): Fix doorbell test fixtures for endpoint liveness * no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax * no-mistakes(document): Document dead and missing doorbell recovery * no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks
…tree poll (kunchenguid#3834) * fix(spawn): keep the worktree poll from adopting the repository primary After `treehouse get` is sent, the worktree-discovery poll reads the pane's foreground-process cwd. While treehouse is still fetching and checking a slot out, the foreground process is treehouse itself and it reports the repository's PRIMARY checkout as its cwd for several seconds. The poll accepted any path that merely differed from the spawning project, so from a linked spawning home - whose project is itself a worktree of that repository - it adopted the primary, and the isolation guard then refused a launch whose slot treehouse went on to create normally. Screen every candidate with the isolation guard's own conditions, extracted as spawn_worktree_isolated, so a read the guard would reject stays a transient the poll keeps waiting through. The two-consecutive-reads rule and the guard as final backstop are unchanged; a pane that never reaches an isolated worktree still fails at the existing 60s deadline, now naming the last path it reported. The already-settled timing assertion counted whole-spawn wall time against a 5s budget and failed on unmodified HEAD on slower machines; it now counts pane reads, which is what "one confirming read, not an extra cycle" actually means. * fix(spawn): say which path the worktree wait rejected, and why Screening every discovery-poll candidate means a host that never reaches an isolated worktree spends the whole 60s window before refusing. That wait is deliberate - separating a transient from a terminal misconfiguration needs machinery this path does not want - so the refusal explains itself instead: the isolation check records why a candidate failed, and the deadline names the last path seen together with that reason. Message and diagnostics only; the poll's control flow is unchanged. Two suites asserted the guard's wording on paths the poll now rejects rather than adopts, so their refusal arrives from the deadline instead: realign fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also asserting the stated reason, and the second the metadata absence it was missing) and the herdr projection e2e's forced non-worktree cwd. * no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test * no-mistakes(document): document spawn poll isolation screen in fm-spawn header * test(spawn): make the non-git isolation case non-git anywhere The refusal-reason assertion for a path outside any repository assumed TMPDIR is not inside a git repository. Where it is, git walks up from the temporary directory, finds that repository, and the spawn reports the subdirectory cause instead - so the case passed or failed on a property of the host rather than on the behaviour under test. Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES, which git documents as not chdir-ing up into a listed directory while looking for a repository. Git never excludes the directory being searched, so the ceiling is the parent of the path handed to the spawn. The assertions pin which cause fired rather than the sentence that explains it, leaving the operator wording free to improve. * no-mistakes(document): point spawn poll comment at the isolation screen's comparison * no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure
Co-authored-by: Talon Stark <talonstark@gmail.com>
…d as not failed (kunchenguid#3846) * fix(bin): read an orphaned green ci monitor as held-for-merge, not failed A no-mistakes run held for a captain merge decision keeps its ci step polling until merged or closed; when the shared daemon restarts under that poll, the run is recorded failed although every substantive step completed and GitHub reports the PR green. A monitor whose only remaining job is to observe a human decision must not convert the absence of that decision into a failure verdict. fm-crew-state.sh now reclassifies a terminal failed run as done (held-for-merge), surfacing the run's PR URL, when the steps table shows every step completed except exactly ci failed and the ci log's last recognized marker reads checks green. A genuinely red check, an unreadable ci log, or a second failed step keeps the failure. * no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed * no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification
…livery (kunchenguid#3852) * fix(bin): read preserved spawn state back before the interrupted exit claims it The deferred-signal exit path asserted the paired task record and In-flight backlog state were preserved without reading either back, exactly when a reader is least able to check (fm-yi4j evidence, 2026-09-05). The commit's exit status alone has been observed to agree with a row that did not actually move. The exit path now re-reads the record and the row under the same per-task lock as the commit, repairs a row the commit believed it moved, and phrases the error as exactly what was verified or attempted - verified preserved, repaired and verified, or an explicit preservation-could-not-be-verified with the reason and hand-closeout instruction. Two behavior tests drive a lying tasks-axi start through a real interrupted spawn and assert the printed claim and the real backlog state agree. * no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask * no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner * no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change * no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean
…3821) * docs: correct stale tmux/herdr backend maturity claims Herdr now has 21 test files, its own required CI job (tests-herdr) that installs a pinned build and hard-fails on "skip: herdr not found", while tmux has 3 test files and is only required as a dependency of the portable-serial e2e lane. zellij, orca, and cmux still have no CI lane at all. AGENTS.md and docs/herdr-backend.md still called Herdr merely "experimental" alongside those three, misleading every session and reader about actual coverage. Update AGENTS.md's config/backend entry, the opening lines of docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend section of docs/configuration.md, and the matching claims in docs/architecture.md, CONTRIBUTING.md, and README.md so they agree and distinguish tmux (default), herdr (own required CI lane, largest suite, Windows still spike-only), and zellij/orca/cmux (still experimental, no CI lane). No behavior, selection order, or dispatch logic changes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj * no-mistakes(review): docs: fix stale herdr label and CI-lane wording * no-mistakes(document): docs: align tmux adapter label in scripts.md * no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page * no-mistakes(review): docs: drop windows claim, align contributing backend wording * no-mistakes(review): docs: trim duplicated CI claim from herdr opening line * no-mistakes(review): docs: drop unguarded largest-test-suite superlative * no-mistakes(review): docs: restore tmux verified label and README experimental scope --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test verbatim and unconditionally after every fix round, so that walk multiplied by round count: measured at 32.7 minutes per validation versus 3.6 minutes intent-targeted. Delete the pin and restore the 3.6-minute posture. Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing .no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed. Record the rule in docs/configuration.md's "Gate defaults" section (the authoritative owner CONTRIBUTING.md already points at) and strengthen CONTRIBUTING.md's existing local-Test guidance to state it plainly: never configure commands.test to a deterministic test command, complete or partial. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC * no-mistakes(review): Centralize no-mistakes test policy and narrow guard --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): verify pool-slot ownership before returning a worktree slot Workers were killed when cleanup returned a Treehouse pool slot that a different, live task had already taken. Teardown now proves the slot is genuinely this task's before releasing it: it refuses when another task record claims the same live worktree path, or when the endpoint's working directory contradicts the recorded slot, and that refusal holds under --force. Slot allocation, metadata publication, ownership verification, and slot return are serialized across linked firstmate homes, and forced secondmate cleanup verifies descendant slot ownership before returning any child worktree. Regression coverage drives the scripts with two task records naming one slot path and asserts the live worker survives and its slot is not reset. * no-mistakes(review): Protect slots across cloned Firstmate homes * no-mistakes(test): Gate teardown locking on genuine Treehouse slots * no-mistakes(test): Clarify pooled descendant slot gating * no-mistakes(test): Synchronize watcher re-arm test on process exit * no-mistakes(test): Wait for watcher cleanup before timeout escalation * no-mistakes(document): Document pool-slot ownership safeguards * no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion * fix(bin): resolve relative origins from repository root * no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout
…ndmate, and primary (kunchenguid#3867) * feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary Add omp as a verified harness: anchored process-name detection with a Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry, the fm-spawn launch template with foreign-marker clearing, the tracked .omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and pre-launch model validation scoped to providers 'omp models --json' lists. Workers get a state-resident busy-state extension keyed on agent_end without willContinue (omp has no agent_settled). The primary gets two tracked .omp/extensions: a turn-end guard that answers omp's blocking session_stop hook by compelling one continuation per turn, with the pre-tool seatbelts and Run-tier session-start delivery, and a watcher extension ported from the Pi one with fm_watch_arm_omp. Control tables, composer busy footers, omp's status row as a bare-composer boundary, the extension supervision model with an omp-keyed ownership proof, the session-start diagnostic, and the supervision protocol snippet follow. Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout through spawn, busy state, steer, interrupt, exit, and teardown, and the isolated rpc primary lab through extension auto-discovery, digest delivery, lock identity, watcher arm, successor and wake delivery, and the compelled guard continuation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test: prove the omp guard continuation through a guard spy Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * fix(spawn): clear the gemini marker at the omp launch boundary Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): force the guard stage by freezing the watcher and clear lint findings Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): reap the live lab by path and record omp's rpc shutdown as a note Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces Replace the template-extraction check with a genuine --secondmate launch pinned to the fake tmux backend, assert the worker extension's handler set through the executable rather than its bytes, classify the two new omp surfaces in the documentation inventory, and record the Herdr worker evidence in the runtime-backends verification doc. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin * no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs * no-mistakes(review): omp: pin config-model validation with a test, trim overlay * no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family * no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes * no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux * no-mistakes(document): docs: add omp subagent-guard row, fix live test header * no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π||pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically) * no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#3843) * Fix Pi shell invocation on native Windows * no-mistakes(document): Document Pi Windows Bash transport * no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally * no-mistakes(review): Cover async Windows branch-outcome Bash invocation * no-mistakes(review): Preserve Cygwin checks and refresh Windows timing * no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows * validation-fixture * no-mistakes(document): Document Windows Bash helper invocation * no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression --------- Co-authored-by: test <test@example.invalid>
…unchenguid#3870) * fix(bin): recognise squash-merged rebased work as landed at teardown A pipeline rebase can leave the local worktree on pre-rebase commits while GitHub squash-merges the rebased head. The landed-work test then compared those stale commits against a squashed main and refused cleanup of work that had already landed. When the forge reports the recorded PR merged and its merge commit is on the default branch, treat a local branch that only repeats paths from the pipeline push as stale rather than unlanded. If the forge is unreachable, the same coverage check runs against a PR head whose content is already on default. Extra local paths still refuse. * fix(bin): drop unprovable squash-rebase landed-work coverage Path-set coverage treated a diverged local branch as landed whenever it touched the same files as the squash merge. That accepts the reviewer's failing sequence: same path, different content, work discarded. git cherry and merge-tree containment were already too strict on the real rebase-fold case. No remaining check is both safe and permissive enough to recognise a stale pre-rebase copy without also accepting unlanded edits, so that case still refuses. Keep the proofs that hold: a merged PR head that contains local work, or a clean content-in-default tree match. Tests now refuse same-path different content and extra unlanded commits, and still allow a local branch that followed the pipeline rebase. * no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers * no-mistakes(review): silence squash-merge stdout corrupting test PR head * no-mistakes(review): make unlanded follow-up commit sole cause of refusal * no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests * no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified --------- Co-authored-by: Morten Gad <mogad@itm8.com>
…nguid#3860) * fix(bin): keep supervision armed for registered custom checks A custom check bound by bin/fm-check-register.sh only ever runs inside the watcher's check sweep, but fm_supervision_status counted in-flight tasks, the relay poll shim, and process-event sources as supervision need, and not registered checks. Tearing down the last task therefore stopped every home-level check silently until the next spawn. Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as supervision need. The relay shim keeps its own trust path and task PR polls carry no such binding and are torn down with their task, so neither arms a home by accident. Presence of the binding is the whole test: the sweep validates the bytes at execution time and wakes firstmate when it rejects one, which is the outcome an idle home needs. Closes kunchenguid#3856 * no-mistakes(review): name registered checks in turn-end block banner and doc invariant * no-mistakes(review): narrow PR poll predicate test to what it proves * no-mistakes(document): point Grok re-arm step at supervision-need owner
…nguid#3883) * fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes Every spawn and teardown inside a remote-seeded secondmate home refused, because the project lock's anchor could not be resolved there. fm_firstmate_root_home walks a home's parent bindings upward to find the anchor the lock lives in, and treated a remote parent binding as an error. A remote-seeded home's parent is on another machine, so that walk can never succeed from there - and neither can the home's own local descendants, whose chain terminates at the same record. Both fail closed on every Treehouse-backed spawn and every pool-slot teardown. A remote parent now terminates the walk at the home holding it, which is the correct anchor: a lock taken on this filesystem is neither held nor observable across that boundary, and that home is already the top of the local tree teardown's collect_local_firstmate_states enumerates, since that walk skips remote registry entries for the same reason. Mutual exclusion is unchanged - every home reachable through local parent links still derives one identical lock file per project, and an unreadable binding, an unsupported route, an unreachable local parent, a cycle, and an over-deep chain all still refuse. Origin-less local-only projects keep resolving through their worktree top. Regression coverage pins the anchor for the main-home layout, a local secondmate, a remote-seeded home, and its local child; drives teardown end-to-end in a remote-seeded home; keeps the cross-home slot-ownership refusal across that boundary; and proves two homes still serialize on the one shared lock file. * no-mistakes(document): Clarify machine-local Treehouse lock ownership
…nchenguid#3872) * fix(bearings): repair the board's listening, card hygiene, and reconcile path Three defects made the fleet board go quiet and then lie about what still needs the captain. Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0 even when it refuses to reopen a session the captain ended from the browser, reporting `status: user-ended` with the same session id, so the build's exit-status check accepted a dead session, printed `already-armed`, and left the board reading "not listening". The build now proves the session is live from a fresh authoritative listing immediately before arming - not from the establish call's status alone, which is already stale by then - reopens once when it finds the session ended, and refuses rather than arming when it stays ended. A reopen also replaces the pre-reopen source generation before reporting success, so a runner on its way out cannot be mistaken for a listener, and a board whose source is registered but unowned gets a replacement started before the build returns. Let a dead generation's ownership actually move. Reclaiming a claim ran its capture-reservation cleanup first, and that cleanup re-verifies the recorded state-root identity, so a claim naming a pid and a process group that were both provably gone could not be cleared: reconcile reported a start while nothing attached, and retire refused with "cannot release source ownership". Reservation records are keyed by claim token and every replacement claims a fresh one, so they are hygiene, not an ownership invariant. Reclamation now additionally requires the owning process group to be absent independently, which keeps a reused pid whose poll child still runs from ever reading as a gone generation. A live owner and a crashed leader whose owned group survives are still never reclaimed. Stop carding decisions whose subject already landed. The build drops a decision card whose work item or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call, naming each drop on stderr. A task whose state cannot be established is kept, because a call wrongly hidden is worse than a card wrongly shown. Add the reconcile choice, and make it structurally incapable of closing a call. Every decision card carries a standard `reconcile` option, injected by the build rather than left to the composer. The board now emits the picked option and any freeform note as separate structured fields instead of fusing them, so a reconcile selection is not expressible as an answer value at all - the defect that let `reconcile - <note>` reach the intake as an ordinary answer. The adapter routes selections from that structured field, creation of a reconcile request is bound to a verified board source rather than the shared keyed-answer intake, and the intake still refuses the reserved value on every channel. Each authorization is bound to the captain-hold generation that produced the card, so an obsolete card cannot close a later call, and both terminal outcomes require a pending request: `reconcile close` records the evidence under its own `reconciled` mode so it never reads as the captain's words, and `reconcile note` leaves the call open. Anything unprovable - an unversioned row, a missing generation, an unreadable state - refuses rather than acting. Regression coverage fails without each fix, and pins every leak path: a bare reconcile, a standalone close or note with no pending request, an any-channel reconcile, an annotated selection from a freeform card, and a generation-skewed authorization. An opt-in guard re-proves the lavish-axi shapes and the reopen against the installed tool. * fix(bin): quote the done comparison in the reconcile intake shellcheck SC1010 reads the bare word as the loop keyword. The failed run never reached its lint step, so this shipped in the recovered content. * no-mistakes(review): Publish reconciled parent resolution before request retirement * no-mistakes(review): Clarify committed cleanup and reconcile reservation scope * no-mistakes(review): Preserve remote cards and legacy answer compatibility * no-mistakes(test): Separate live claim release from stale reclamation * no-mistakes(test): Allow terminal self-retirement during active capture * no-mistakes(document): Document Bearings repair contracts * no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks
* fix(bin): skip pooled-worktree freshness fetch when no origin is configured An origin-less local-only project has nothing remote to be stale against, so fm-spawn's freshen_spawn_worktree_base refused to launch crews for it. Detect a missing origin remote and skip the fetch freshness gate entirely; an existing-but-unreachable origin keeps refusing as before. * no-mistakes(review): Preserve pool safety for absent and unusable origins * no-mistakes(review): Refuse empty origin configurations during pooled spawn * no-mistakes(review): Detect empty origin sections across config includes * no-mistakes(review): Honor globbed includes when detecting origin configuration * no-mistakes(review): Document conservative conditional include handling * no-mistakes(review): Use Git-resolved config files for origin detection * no-mistakes(review): Document included empty-origin detection boundary * no-mistakes(document): Document originless pooled spawn behavior
…enguid#3889) * feat(tests): run live harness guards by default where the harness is installed The 24 live-harness guards each opened with their own env check, so on the machine that has every harness - the one the product and its validation actually run on - all of them skipped and passed. Fourteen had never been run by the pipeline at all. tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard that spends no model tokens runs wherever its tools are installed, a guard that submits prompts stays opt-in, an absent tool is a named capability skip, and a guard's own variable or FM_LIVE forces it on (turning an absent tool into a failure) or off. Every live guard now opens with it, which also carries the test-suite gate-refusal bypass into the guards that never sourced the shared helpers and were therefore refused whenever a gate agent ran them. bin/fm-test-run.sh records what a skip means: the family's expected class is live-capability rather than a bare env opt-in, and each gate skip's reason is logged and written to the timing artifact, so a lane can say which tool this host could not exercise. Only the token-free guards flip to default-on: composer-matrix, the harness liveness drift guard, and the Herdr version floor. cursor-primary submits three prompts, so it stays opt-in. Running the drift guard unasked immediately found a real defect it existed to catch: it resolved the harness through a generic `command -v cursor`, which on a machine that also has the Cursor editor finds the editor launcher rather than cursor-agent. That binary exits at once, leaving a bare shell in the pane and a liveness-drift failure no classifier change could fix. It now asks fm_cursor_resolve_binary first, the same verified owner fm-spawn uses. CI installs the public Pi package in the portable serial lane and fails on its skip token, so the Pi extension tests stop passing silently against a package that is not there. No secret is added. Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and classifies 8 installed harnesses alive; the Herdr version-floor guard runs by default and checks 4 real releases; every live guard refuses together under FM_LIVE=0. * fix(tests): keep the composer-matrix guard opt-in Running it unasked is red on a healthy machine for reasons no code change here removes: a harness that has not trusted this checkout sits on its own trust dialog, which the guard treats as an unreadable composer and correctly fails. The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces identically on main and is filed as separate work. So this token-free guard stays opt-in with the reason stated in its header, and the coding guidelines record the narrow exception: a guard whose verdict depends on host state that installing its tools does not establish may stay opt-in, because one that is permanently red is one the fleet learns to ignore. The other two token-free guards keep running by default. * no-mistakes(review): Wire bearings guard and remove composer exception policy * no-mistakes(review): Run Pi responsiveness guard by default * no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep * no-mistakes(review): Sanitize live gate test environments * no-mistakes(document): Document default-on live guard behavior * no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks * no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass
… is set (kunchenguid#3891) * fix(bin): refuse the behavior suite in the repository primary checkout A task worker's isolated worktree placement is verified exactly once, when its task starts, and nothing re-checks it afterwards. A worker that later changes directory into the repository's primary checkout runs its Git commands, and this branch-switching suite, against the one checkout every linked worktree resolves against and every landing merges into. A run that dies mid-suite can leave that checkout on a stray branch. bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task worker and the runner resolves to the primary checkout, every executing mode exits non-zero before selecting a suite, with one line naming the primary path and pointing at the assigned task worktree. The predicate is the one bin/fm-spawn.sh already uses for launch placement: the working tree's own git dir is the repository's common git dir, which separates the primary from every linked worktree even when their top levels differ. A run with no FM_TASK_ID set is unchanged, and so are the inspection modes, which execute nothing. When git resolves neither directory - a non-repository fixture, a detached copy - nothing proves this is the primary, so the run proceeds. bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID into the pane shell on the same pre-launch channel as GOTMPDIR, and the name joins the sanitized launch environment allowlist so an isolated launch keeps it. * no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT * no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
) * fix(bin): bound a stale alarm with the backlog hold, not only the status line A legitimate wait has two records and the stale alarm reads only one. `status_is_paused_or_captain_held` takes a status line, so it sees a wait the worker declared. It cannot see the wait firstmate records when it hands work to the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and leaves the status log alone, so a delivered task keeps `done: PR ...` as its last line for the whole time the captain is deciding. Both stale branches were blind to it, and each churned a new pane hash back into its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale branch, while a held task whose last line is `working:` reaches `surface_nonterminal_stale` and fails its declared-wait test. Consult that second record where the watcher is about to alarm, through `bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and bound the alarm on the shared `.paused-resurfaced-<key>` marker and `PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first sight still alarms, the window's end alarms once more, and a held crew that goes genuinely silent still escalates through the wedge timer. Only an established open captain call bounds anything: an unreadable backlog, an absent or incompatible tasks-axi, a row this home does not carry, and every task with no hold keep alarming exactly as before. The backlog hold is deliberately not recorded as a declared pause, because the loop-top reconciliation and `pause_state_class` both read the status line and would clear a flag that line does not support. Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status line itself can express. * fix(bin): identify the captain call a stale alarm is bounded by Three gaps in the bound added by the previous commit, all in how the throttle is scoped and where the backlog is consulted. The scope carried only the status-log signature. A task can be held, answered with `--release`, and re-held as a genuinely different captain call without any status append, so the second call inherited the first one's marker and its first sight was absorbed - the one thing this bound must never do. The task id is not the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the call's own lifecycle - its hold-set stamp and the number of recorded answers - on an exit 0 and only then, leaving the silent predicate every existing caller reads unchanged. The throttle scope now carries that identity. The terminal path recorded the throttle before publishing the durable wake. A failed append exits the watcher with nothing queued, and the next sighting then read that fresh marker and absorbed the retry, turning a delayed alarm into a lost one. Recording moves behind the append, as the non-terminal path already had it, and the comment claiming the marker could not outlive its wake is gone because it was false. The backlog was consulted only on a new terminal pane hash. A captain call can open after a hash was absorbed as provably working, changing neither the pane nor the status log, so nothing re-read the backlog and the wedge timer kept firing possible-wedge alarms through a legitimate wait. That timer now consults the call at its own alarm boundary and takes the same bounded cadence - and only at that boundary, so an ordinary repeat poll under the bound stays the local-only read it was. Regression coverage for each, all driving churn through one watcher process rather than relaunching per pane change: relaunch cost dominated the earlier shape, and an absorbing watcher stays in its poll loop across churn in production anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer with no open captain call still escalates as a possible wedge. * fix(review): Compose stale throttles with captain-call lifecycle identity * fix(review): Preserve bounded same-hash captain-call resurfacing * revert(bin): narrow the captain-hold stale bound to its observed defect Lifts the lifecycle-identity and cadence-ownership work back out, leaving the change at the shape that matches the defect actually observed: the stale alarm did not consult the backlog captain hold, on either stale branch. Reviewing the wider version surfaced a series of adjacent gaps in the watcher's alarm state machine - a call opening after the first alarm, marker invalidation at the hold lifecycle boundary, and which deadline a terminal timer represents. They are real, but fixing them turns a small extension into a state-machine change to the alarm path, which is a different review on a subsystem that is being actively reworked. They are named as known limitations rather than carried here, and none of them is load-bearing for what remains: the bound does strictly less than the reverted version, leaves the wedge path escalating on STALE_ESCALATE_SECS exactly as before, and introduces no silence that the existing terminal-alarm path did not already have. Kept from the reverted work is the record-after-append ordering, because that is a defect in the code being shipped rather than an adjacent one: recording the cadence marker before publishing the durable wake let a failed append lose an alarm outright instead of delaying it. History is preserved: the earlier commits stay on the branch and this removal sits on top of them. * fix(review): Document secondmate captain-hold scope boundary * fix(document): Document captain-hold stale alarm scope * fix(bin): bind the stale throttle to the captain call, not the status log The throttle this change introduces was scoped to the task's status-log signature. Answering a call with `--release` and holding the task again creates a genuinely different captain call without necessarily appending to that log, so the second call inherited the first one's marker and its first sight was absorbed. That is the one alarm this bound must never swallow. A delivery announced twice is noise; a decision waiting on the captain that is never surfaced is invisible, because nobody asks for what they do not know to ask for. Measured rather than assumed, on the same fixture - a delivered task held for the captain, released, and re-held with no status append, driven through bin/fm-watch.sh: base c499f84 call-1 first=ALARM call-1 churn=ALARM new call first sight=ALARM before this fix call-1 first=ALARM call-1 churn=absorbed new call first sight=absorbed after call-1 first=ALARM call-1 churn=absorbed new call first sight=ALARM Base never suppresses the new call, so the suppression came from this change and closing it completes the fix rather than widening it. `bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle - its hold-set stamp and count of recorded answers - on an exit 0 and only then, so the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope carries that identity beside the status signature. The sibling case was measured too and is NOT included: on the status-declared path, where the last line is `captain-held:`, base already absorbs a re-held call's first sight. That behaviour predates this change and stays documented as a known limitation rather than repaired here. * fix(document): Document captain-call throttle lifecycle scope * fix(ci): isolate the Herdr restart fixtures from a claimed worktree The Herdr behaviour test intermittently reused a local worktree still claimed by an earlier fixture after a restart. The restart scenarios now use an isolated Treehouse project. The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as well.
…unchenguid#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs kunchenguid#4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes kunchenguid#3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained.
…chenguid#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one.
…chenguid#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes kunchenguid#3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge upstream Firstmate through
b430bf50d9aea5d1a810cab1bd293670bbca64beinto the fork, bringing Claude Code Calm and the other upstream changes into the running fork while preserving its local capabilities.The fork parent is
f50a7997c5b98f460615792223db0c8ccdd200e6.The fetched upstream tip is 109 commits ahead of that parent; it includes Calm commit
bdcacb9fand the subsequently landed declared-wait fixb430bf50.This is a true two-parent merge; no rebase, squash, default-branch rewrite, or force-push is used.
Preserve the two-parent merge when landing this PR so the next fork sync retains upstream ancestry.
Fork commit classification
Classification is against the exact upstream tip above, using the commit patches, the remaining fork-versus-upstream diff, and current executable owners.
No fork-only capability is dropped.
4c8d97bdcmux failed-list guardlist-panesguard infm_backend_cmux_surface_exists; upstream still pipes an unchecked result to jq.e31fd313review-diff branch resolutionfm/$ID.8d304c6cHerdr capability probe259be8e0Claude session identity105ff038hosted Codex lock ownercodex:<thread-id>ownership and conservative reclamation in the lock owners and their consumers; integrate upstream namespace-PID-1 ancestry with the numeric owner validator.ae77b98dlegacy Herdr binding repairfm-work-landed-lib.shbecause upstream independently added a display selector at the old filename.045e5382tasks-axi compatibility gatesfm_tasks_axi_compatiblein the two handoff suites; upstream checks only binary presence there.d0f820d7Bitwarden migration tooling30705e93Bitwarden rotation boundary1af5ea04Automic Vault worker auth3294f656Vault development-only boundaryc3b6535bproject lifecycle registrye353d9bdsole Vault settings160c812bVault settings invariant docs20240d73Vault verification docs71522695Vault filtering regression663a5f68Herdr/Vault verification consolidation99839a98Bearings--contractnameandfiledfields and cached remote-snapshot enum values to it.cdd5462aPi 0.85.1 renderer baseline36fd955b) instead of the fork's equivalentcreateAllToolDefinitionsbaseline. Preserve the independent watcher fixture cleanup.0a14719cOrchestra live-board refreshf50a7997lifecycle across Bearings projectionsExact conflict inventory
.agents/skills/harness-adapters/references/harness/claude.md.agents/skills/secondmate-provisioning/SKILL.md.github/workflows/ci.ymlbin/fm-bearings-board.shbin/fm-bearings-snapshot.shbin/fm-config-inherit-lib.shbin/fm-fleet-snapshot.shbin/fm-landed-lib.shbin/fm-work-landed-lib.sh, updating both callers and fixtures.bin/fm-spawn.shbin/fm-test-isolation-proof.shbin/fm-test-run.shdocs/calm-mode-feasibility.mddocs/configuration.mddocs/documentation-audiences.jsondocs/herdr-backend.mddocs/scripts.mddocs/verification/runtime-backends.mdtests/fm-calm-pi-extension.test.shtests/fm-secondmate-harness.test.shtests/fm-spawn-dispatch-profile.test.shtests/fm-task-delivery.test.shtests/fm-watch-arm.test.shAGENTS.mdauto-merged without manual edits: it retains upstream sections and only fork additions with live owners (Vault, Orchestra, and project lifecycle).The required memory helper reported it unchanged.
The prior
firstmate-source-reconciliation/report.mdwas absent; the available reconciliation-plan, multi-brain, independent-continuation, and earlier Calm-sync brief were read.Claude Code Calm after landing
docs/configuration.md, “Calm preference (config/calm)”, documents the shared home-local toggle and links todocs/calm.md#claude-codefor activation.Start or restart Claude Code with
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1, let the trusted project's.claude/skills/firstmate-calmlink load the experimental mod, then use/calmto toggle the preference.The mod requires that variable to equal
1; module loading through thetengu_plugin_hooks_modulesrollout flag alone does not activate it.Without the explicit opt-in the module is a complete no-op, even if
config/calmis already on.The merge does not enable Calm or change the running home's configuration.
Validation
Local validation uses
bin/fm-test-run.shas documented inCONTRIBUTING.md, with isolated temporary homes andFM_LIVE=0for portable checks.CI=true bin/fm-lint.sh; all three workflows valid.bin/fm-session-start.sh --help, andbin/fm-spawn.sh --help.surfaces=102 local_links=422) and test coverage partition (total=218 parallel=24 serial=177 herdr=17, five serial shards).b430bf50d9aea5d1a810cab1bd293670bbca64be; this is an installed-runtime limitation, not a merge-only failure. Upstream records compatible evidence with Claude Code 2.1.272/2.1.273. No installed runtime or running-home setting was changed.No-mistakes review identified and fixed the namespace-PID-1 lock-validator integration, missing cached remote-snapshot enum values, and an imported source-text-only assertion; the existing generated worker-brief behavior coverage remains.
No-mistakes Test also fixed a missing changed-test mapping for the browser-rendering asset.
Its Claude Code 2.1.273 environment passed strict plugin validation and plugin tests; this differs from the local 2.1.261 guard limitation above.
Live validation
Firstmate approved proceeding at the Test gate with the following scenarios untested live, recorded verbatim:
Fixture and plugin-engine checks passed; these do not replace the live scenarios listed above.
The real Calm TUI check requires tmux, authenticated Claude access, and
FM_CLAUDE_CALM_LIVE_E2E=1in a suitable validation environment.Intent
Captain 2026-09-15 (verbatim): "dispatch the fork sync". Context the captain gave: "/calm mode just arrived for claude code in latest firstmate main branch - it depends on claude mods which is experimental and needs opt-in via env var CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1". The captain wants the fork (origin = https://github.com/rega10/firstmate, the running Firstmate) brought up to date with upstream (https://github.com/kunchenguid/firstmate, remote upstream already configured in the primary checkout) so that upstream's Claude Code Calm mode (upstream commit bdcacb9 "feat(calm): add flag-gated Claude Code Calm mode (kunchenguid#4565)") and everything else upstream has landed becomes available here. At intake (2026-09-15, upstream fetched): origin/main f50a799 is 108 commits behind upstream/main and 21 commits ahead of it. The fork's 21 own commits are captain-authorized work and must survive: Automic Vault authentication for Claude workers and its credential-scrub settings, Bitwarden migration tooling, the bearings lifecycle-posture registry and projections, the Orchestra live-board refresh, the Herdr legacy endpoint-binding repair, the hosted Codex lock ownership rework, and the smaller fixes.
The referenced prior fork reconciliation evidence establishes these retained behaviors: hosted Codex session locks recognize opaque codex: owners and reclaim conservatively; cmux surface checks refuse failed pane listings; Herdr capability probing handles large schema output without broken pipes; review-diff resolves the task worktree's actual HEAD branch; Claude workers clear inherited parent session identity so their transcripts are saved; legacy Herdr binding repair remains guarded by safe endpoint identity and the same landed-work reachability rules used by teardown; task handoff tests require compatible tasks-axi rather than mere binary presence. Earlier syncs also preserved metadata-publication failure handling, confirmed Herdr closure before state retirement, and stock macOS Bash compatibility while adopting upstream Calm and Pi multi-brain supervision. Prior independent-continuation evidence calls for preserving accepted fork behavior and already-validated history, with ordinary complete validation of the resulting fork. These are the substantive preservation requirements behind the referenced reconciliation-plan, upstream-multibrain-sync, source-reconciliation, independent-continuation, and earlier Calm-sync records; the source-reconciliation report is no longer present in the task evidence.
What Changed
Risk Assessment
Testing
The configured baseline initially failed before running tests because the changed browser harness lacked a mapping. The actual changed selector and live Bearings contract and installed-Claude plugin checks passed. Fixture-based remote fallback, simulated PID namespace behavior, fake-harness worker delivery, and the Calm real-TUI scenario were not established against live product surfaces.
bin/fm-test-run.sh --list --changed --exclude-family real-herdr-gatedcompleted against the actual worktree and selectedtests/fm-bearings-board-render.test.sh; see changed-test selection artifact…bin/fm-bearings-snapshot.sh --contractemittedstructured-home-cacheandcached; see Bearings contract artifact.Evidence: Bearings cached-vocabulary CLI contract
Source: Bearings cached-vocabulary CLI contract
Evidence: Changed-test selection after mapping fix
Source: Changed-test selection after mapping fix
Evidence: Installed Claude engine Calm validation
Source: Installed Claude engine Calm validation
Evidence: PID 1 and conservative session-lock behavior
Source: PID 1 and conservative session-lock behavior
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
Step was skipped.
🔧 **Review** - 3 issues found → auto-fixed ✅
bin/fm-session-lock-lib.sh:141- The new ancestry path can return namespace PID 1, but the retained lock validator rejects every numeric owner <= 1. A verified harness running as PID 1 therefore writes an unreadable lock and cannot recognize itself as owner. Permit PID 1 while retaining the harness-shape liveness check.bin/fm-bearings-contract-lib.sh:67- The Bearings contract omits values emitted for valid cached remote-secondmate snapshots. When live collection fails but a valid cache exists, output contains provenancestructured-home-cacheand freshnesscached, which contract-validating consumers must reject. Add both values to their closed enums.tests/fm-task-delivery.test.sh:1097- This test only greps the implementation prompt inAGENTS.mdand compares source-line order. It does not execute a worker-facing interface or prove role precedence, violating the test-quality rule. Remove it or replace it with an observable launch or worker behavior assertion.🔧 Fix applied.
✅ Re-checked - no issues remain.
tests/fm-calm-claude-mod-live-e2e.test.sh:20- The central Calm real-TUI scenario lacks live visual evidence because tmux is unavailable. Provide tmux on PATH, an authenticated Claude session, and FM_CLAUDE_CALM_LIVE_E2E=1, then rerun tests/fm-calm-claude-mod-live-e2e.test.sh. Workspace restrictions prevented installing tmux during this phase.bin/fm-test-run.sh --list --changed --exclude-family real-herdr-gatedcompleted against the actual worktree and selectedtests/fm-bearings-board-render.test.sh; see changed-test selection artifact…bin/fm-bearings-snapshot.sh --contractemittedstructured-home-cacheandcached; see Bearings contract artifact.bin/fm-test-run.sh --changed --exclude-family real-herdr-gatedbin/fm-test-run.sh --changed --exclude-family real-herdr-gated- reproduced exit 2 before tests due to unmappedtests/assets/board-render-harness.mjstests/fm-test-run.test.shbin/fm-test-run.sh --list --changed --exclude-family real-herdr-gatedbin/fm-bearings-snapshot.sh --contract | jq ...bash tests/fm-bearings-snapshot.test.shbash tests/fm-session-lock-ancestry.test.shbash tests/fm-task-delivery.test.shbash tests/fm-calm-claude-mod.test.shbash tests/fm-calm-claude-mod-plugin.test.shwith installed Claude Code 2.1.273✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.