fix: sync fork with upstream to restore portable CI behavior - #38
Merged
Merged
Conversation
…6064) * fix(bin): read a live quiet record as a present captain at the host and watcher A quiet record left without its daemon (a quiet start that never ran or was interrupted) was read as away by the supervision host, so it parked a present captain's main and held captain outcomes for a return that never comes, and the watcher and daemon silenced captain-held rechecks on record presence. The host's posture checks, the watcher's and daemon's captain-held silencing, and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated branch authority, the owners' away wake note, and the Codex checkpoint bound) now ask the record owner's away-or-quiet reading, so only an away record is away. A live away record keeps today's behavior. * no-mistakes(document): Correct quiet-record documentation and supervision guidance * no-mistakes(document): Clarify quiet-record posture and captain-held rechecks * no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043) * fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line The away return brief said nothing had failed after the supervision host latched on engine errors during the window, and printed a GAP: watcher downtime line whenever a wake was merely being handled or queued at return. The failures section now reads the host ledger and latch record and names the latch time, the window's engine-error count, and whether the session is still paused or recovered. An open recovery episode is reported as information, and as a gap only when a queued episode outlived the return grace or the marker cannot be read. * no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count * no-mistakes(review): Report paused latch without ledger trip row; bound errors * no-mistakes(review): Never report a failed probe's latch row as trip time * no-mistakes(review): Only a retained trip row marks a pre-window latch * no-mistakes(document): Clarify return-brief latch and watcher-gap documentation * no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed * no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass * no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass * no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass * no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: " Claude Code labels every mod transcript line with the plugin name, so the notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update the live guard to assert the fm: label, and document the one-time replay for sessions resumed across the rename. * no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037) * feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder * fix(bin): exact lab windows, per-lab task ids, self-safe teardown * fix(bin): target lab windows by id, stop lab descendants, add readiness tests * fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh * fix(bin): start the lab tmux server without user config * no-mistakes(review): Scope lab teardown to its store, root, and task ids * no-mistakes(review): Record selected user stores at up for check and down * no-mistakes(document): Clarify live lab documentation and remove stale narratives * no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH * no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged * no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass * no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh * no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet * no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times * no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass * no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified * no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103) * fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle. * no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110) * fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN Fixes kunchenguid#6020 bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for a short while after a push or a base-branch change while it recomputes mergeability, so a green, conflict-free pull request was refused as if it could not be merged. github_verify_mergeable now returns a distinct status when mergeable is the only failing condition and reads UNKNOWN. The caller retries up to 5 times, 3 seconds apart (overridable in tests), re-reading and re-checking every live condition on each attempt. Once the bound is spent it reports mergeability as still being computed rather than unmergeable, with the same nonzero exit as before. Every other refusal (closed, draft, conflicting, red or missing checks, away authority, queue protection) is unchanged and never retried. * no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112) * fix(bin): converge every open owner onto a known terminal contribution settle_final only cleared a stale error on retry, so an owner whose saved row still said open kept projecting a merged or closed pull request as open after another owner's row had already recorded the terminal observation. Copy the known terminal observation to every owner whose saved row is not itself terminal, keeping that owner's own pending and notified state, and clear its error. * no-mistakes(review): Carry terminal checked_at when converging existing owner rows * no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124) * feat: run the supervision host by default on a Claude primary An absent config/supervision-host on a Claude primary now reads as on with the default engine, and a file holding `off` opts any home out. Cursor, OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled there too. Every reader asks fm_supervision_host_enabled instead of testing the file, and non-bash readers query it through the lib's `enabled` entry. A primary's `off` is not inherited by secondmates: each home keeps its own supervision posture. * test: pin the watcher-path posture in fixtures that assume no supervision host Fixtures that drive the watcher arm or assert a non-host drain now write an explicit off file, and fixtures that copy the Stop auto-arm or the supervision instructions carry the engine lib they now source. The two drain suites also stop reading the code root's config. * fix: name the opt-out when an off home passes an attended wake to main A host parked when the home writes off now logs that the home does not run the supervision host, rather than claiming it has no engine. * no-mistakes(document): Clarify Claude supervision defaults and historical evidence * no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125) * fix(bin): create the state dir on a fresh primary before the session-start scope check fm_primary_scope_matches required an already-existing state directory, so bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could create it. Split out fm_primary_root_matches so the run wrapper can confirm primary-home identity first, create the gitignored state dir when it is missing, and only then run the unchanged scope check. * no-mistakes(document): Document session-start state dir creation on fresh clones * no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126) * fix(bin): measure pending-reply grace from turn completion, not delivery Fixes kunchenguid#6057 The pending-reply guard demanded a repost ("REPOST REQUIRED: previous marked request had no correlated parent report") while the second mate's correlated reply was already on its way. fm_pending_reply_send_recovery measured its grace window from delivery instead of from the request turn's completion, so any turn longer than the grace fired the demand the moment the turn ended, before the reply could have landed. The missed-report escalation had the same gap: it fired the instant the recovery turn's completion was observed, with no grace at all. Both now measure grace from the relevant turn's completion (request turn for the recovery repost, recovery turn for the escalation), and both take one fresh, uncached read of the parent status file immediately before firing, accepting a correlated line regardless of its verb. Transport-failure escalations stay immediate, and the one-repost limit is unchanged. * no-mistakes(review): Document grace window as measured from turn completion * no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction * no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001) * fix: provider-table lookup never writes a broken-pipe error to stderr Fixes kunchenguid#5956 fm_quota_single_provider_for_harness returned from its while read loop as soon as it found a match, closing the pipe while fm_quota_single_provider_table's printf could still be writing. Where SIGPIPE is ignored, as on GitHub Actions runners, bash then prints "printf: write error: Broken pipe" on the resolver's stderr, which intermittently broke the one-diagnostic-line assertions in tests/fm-dispatch-resolve.test.sh. Read the whole table before answering, the way fm_control_harness_supported already does, so the writer always finishes. Return values and output are unchanged. Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE ignored on a single pinned core under CPU contention: 30 of 30 runs failed before the fix, 0 of 30 after. Note: reproducing requires setting the trap inside the tested shell because nice(1) resets an inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh passes and bin/fm-lint.sh is clean. * no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162) * fix: survive Pi 0.99 rendering and Git 2.55 local-clone races Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed. * no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones * no-mistakes(document): Clarify Calm export visibility and tool rendering * no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available * no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion * no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs * no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior * no-mistakes(document): Clarify Pi Calm export and queued-turn documentation * no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169) * Prevent premature Lavish board handoffs * Prove Lavish arm lacks reply acknowledgement * Confirm Lavish replies before arming worker boards * no-mistakes(review): Post Lavish reply only after locked arm eligibility checks * no-mistakes(review): Fail Lavish reply closed on unknown version * no-mistakes(document): Correct Lavish reply documentation and remove stale guidance * no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154) * feat: inherit the supervision-host opt-out from the primary Move the supervision host's off opt-out out of config/supervision-host into its own presence flag, config/supervision-host-off, and add that flag to the primary-authoritative inherited config set. A primary that opts out now opts every secondmate home out at spawn and convergence, and clearing it converges them back. config/supervision-host stays the home-local engine choice. Shape: config/supervision-host mixed two things, a fleet posture (off) and a per-home engine and model. Only the posture should follow the primary, so it becomes a separate presence flag that rides the existing inherited-config mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new machinery, while the engine line stays local. The parse stays in its one owner, fm_supervision_host_enabled. There is no migration or compatibility handling for a home that still holds off in config/supervision-host. Primary off, mate on: inherited material is primary-authoritative by design, so a mate cannot keep the host while the primary is opted out, and a mate's own opt-out is removed at the next convergence while the primary has none. Running the host on a mate is the primary's choice for the fleet; no override mechanism is added. Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a real seeded secondmate, --supervision-host off): - up: every readiness check ok, including "host: none running, as expected" and a live mate session; the spawned mate home held the inherited config/supervision-host-off and the gate read primary OFF, mate OFF. - primary removed its opt-out, then bin/fm-config-push.sh reported "supervision-host-off: pushed - mirrored primary absence" and a config reread sent; the gate read primary ON, mate ON, and the live mate handled the reread. - primary opted out again and pushed: "supervision-host-off: pushed", mate gate OFF. - down stopped every lab process and left no lab process running. Out of scope, follow-up: default-on for the other harnesses, away-daemon retirement, rollout. * no-mistakes(document): Document inherited supervision-host opt-out ownership * no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed * no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up * no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179) * fix(tests): cut the fixed sleeps in supervision-host cycles The serial CI lane keeps brushing its 30-minute cap because fm-supervision-host.test.sh spends ~903s of the job, and per the run-36635306527 case profile the top nine cases are all multi-cycle ones (3-10 park/close/turn cycles each): every close waits out the host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan cycle, and every engine turn waits out the fixed sleep 1 descendant snapshot. That is ~3s of pure sleep per cycle before any real work. The host poll now accepts positive decimal seconds through a new seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a positive decimal defaulting to one second - the smallest seam at each wait's single owner. The suite drives them at 0.2 alongside the existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real poll loops still run. The park-boundary case moves onto the injected test clock instead of a real 3s wait, per-case cleanup polls the host pid rather than sleeping a full second, and the proof-by-absence windows (flood re-escalation, successor re-announce, watcher persistence, recovery staying off main) shrink from 2-3s to 1s, which still spans two watcher polls at the test cadence. Every assertion, process lifecycle, and reaping path is unchanged; production defaults stay at one second. Isolated case timings on a contended host, base vs branch: attended-latch 54.3->34.6s, undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence 47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s, registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s, latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and shellcheck clean. * no-mistakes(review): Wait for scan lock release before duplicate check * no-mistakes(document): Correct supervision snapshot cadence documentation * fix(tests): keep production poll cadence, probe exits at 0.1s The fractional poll cadences multiplied the cost of each loop body: full process-table scans in the engine turn and process refreshes in await_close ran five times more often, which swamped the thin CI runner and nearly doubled every multi-cycle case (serial 5 was cancelled at its 30-minute limit on run 36635306527's successor). Restore the production cadence and notice arm/engine exits with a cheap kill -0 probe at a tenth of a second between the one-second bodies instead: strictly less dead time than baseline with no added CPU. Also hold each injected-clock park bound well past its case's wall-clock checks so a host that ignored the test clock fails instead of silently passing at a real-time boundary, and restore the shortened proof windows (watcher liveness, recovery-off-main absence, first-cycle stream) to their baseline depth. * no-mistakes(document): Clarify supervision engine snapshot documentation
Takes the 15 upstream commits eb219c8..c35b9a6, including e2668de's portable CI repairs for Pi rendering and remote provisioning and 0a2cdf9's default-on supervision host for Claude primaries. Conflicts: - .agents/skills/afk/SKILL.md: kept the fork's trigger-only description. - .agents/skills/process-event-sources/SKILL.md: kept the fork's one-line trigger description; upstream's body changes merged cleanly. - tests/fm-brief.test.sh: kept the fork's harness-banner isolation test and took upstream's rewritten Lavish floor comment.
… packing estimates
…stent, no stale facts remain
…cused.test.sh scratch file"}
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
THIS PR MUST BE LANDED WITH A MERGE COMMIT - do not squash, do not rebase, do not force-push.
It carries a true merge of upstream kunchenguid/firstmate; a squash or rebase would drop the shared ancestry the last two syncs (PR 33 and PR 35) kept, and the next sync would conflict against history git can no longer see.
Intent
"you can upgrade firstmate now right? our fork is mostly up to date with upstream main?" and, on Slack, "You can upgrade firstmate as the upstream repo was folded in right?"
PR 35 true-merged upstream
eb219c80on 2026-09-29, and upstream main has moved on since.Upstream's
e2668de0("fix: restore portable CI behavior across Pi rendering and remote provisioning", kunchenguid#6162) repairs the portable CI lanes after the CI image moved to Pi 0.99 and Git 2.55; without it this fork's main fails Behavior portable serial 2, 4 and 8 (the Calm export-DOM boundary, the Pi outcomes-rendering stock-behavior check, and the two remote-seed provisioning tests), as a rerun of main's CI (run 36518848136) showed.This sync takes that fix, and the rest of upstream main, so those lanes turn green.
What
This brings the fork up to date with the upstream firstmate project again.
All 15 upstream changes made since the last sync are merged in, and every fork-specific feature is kept.
The practical effects: the fork's CI goes green again on the Pi and remote-setup tests that broke when GitHub's build machines got newer Pi and Git versions; on a Claude-run home, the background supervision helper now runs by default instead of only when switched on; and several supervision, merge, and Lavish review reliability fixes arrive.
Only three files needed hand resolution, all small.
CI's longest test group then ran past its 30-minute limit, so upstream's own rebalancing of the CI test groups was taken as well.
Merge shape
507229daMerge upstream kunchenguid/firstmate c35b9a6 into the fork (true merge,git merge --no-ff upstream/main).Parents (
git show -s --format=%P 507229da):d860a1bc8cba1ce4ee5ab255bb87dc8883ab63d1(fork main, the PR 35 merge) andc35b9a69be55c3c1147035701ad167405378dbb7(upstream/main tip).eb219c80(the PR 35 sync point).9c353d82,e8e5aad5,1124b274(author kunchenguid) are upstream PR ci: rebalance portable test groups and enforce a packing budget kunchenguid/firstmate#6192's three branch commits, cherry-picked; together they are the content of upstreameb77f02b"ci: rebalance portable test groups and enforce a packing budget". The one conflict,bin/fm-test-run.sh's serial weight-hint table, keeps upstream's remeasured values plus the fork's 15 fork-only test hints; the resulting runner equals upstream's tip plus exactly the fork's existing divergence.82fb99c6"no-mistakes(document): Docs already updated in-branch; verified consistent, no stale facts remain" is misleading: it actually added a stray 1849-linetests/fm-test-run-focused.test.sh, a scratch near-copy oftests/fm-test-run.test.shleft by the repair round.1ece78a0removes that stray file again (its subject is a JSON blob the lint step wrote). Net change after the merge: onlybin/fm-test-run.sh,docs/fm-test-portable-shards.md,tests/fm-session-start.test.sh,tests/fm-test-run.test.sh.1ece78a049503225c9359332e53cd4b881be9e2d, all 19 CI checks green, including "PR must be raised via no-mistakes" and Behavior portable serial 2, 4, 5 and 8.eb77f02bby ancestry only. Its content is on this branch through the cherry-picks above, so the next sync's merge of it should be content-equivalent; upstream is otherwise level.Upstream commits taken (15,
eb219c80..c35b9a69)2d833ff1fix: treat quiet records as attended across supervision (fix: treat quiet records as attended across supervision kunchenguid/firstmate#6064)00679ae3fix: report supervision host latches accurately in away return briefs (fix: report supervision host latches accurately in away return briefs kunchenguid/firstmate#6043)40e981dafix: shorten Claude Code Calm supervision note label (fix: shorten Claude Code Calm supervision note label kunchenguid/firstmate#6086)b5fdf74dfeat(bin): add a disposable live supervision lab builder (feat(bin): add a disposable live supervision lab builder kunchenguid/firstmate#6037)46d58d64fix: keep watcher arms and reply listeners alive through slow cycles (fix: keep watcher arms and reply listeners alive through slow cycles kunchenguid/firstmate#6103)c5f48e4cfix(bin): retry fm-pr-merge a bounded number of times when GitHub mergeable is UNKNOWN (fix(bin): retry fm-pr-merge a bounded number of times when GitHub mergeable is UNKNOWN kunchenguid/firstmate#6110)260c4f08fix(bin): converge every open owner onto a known terminal contribution (fix(bin): converge every open owner onto a known terminal contribution kunchenguid/firstmate#6112)0a2cdf95feat: enable supervision host by default for Claude primaries (feat: enable supervision host by default for Claude primaries kunchenguid/firstmate#6124)5bddfc44fix(bin): create the state dir on a fresh primary before the session-start scope check (fix(bin): create the state dir on a fresh primary before the session-start scope check kunchenguid/firstmate#6125)1f2c9548fix(bin): measure pending-reply grace from turn completion, not delivery (fix(bin): measure pending-reply grace from turn completion, not delivery kunchenguid/firstmate#6126)a774c448fix(bin): stop provider-table lookup from writing broken-pipe errors to stderr (fix(bin): stop provider-table lookup from writing broken-pipe errors to stderr kunchenguid/firstmate#6001)e2668de0fix: restore portable CI behavior across Pi rendering and remote provisioning (fix: restore portable CI behavior across Pi rendering and remote provisioning kunchenguid/firstmate#6162)b3d41332fix: confirm Lavish board replies before worker handoff (fix: confirm Lavish board replies before worker handoff kunchenguid/firstmate#6169)12e90e14fix: inherit supervision host opt-out across secondmates (fix: inherit supervision host opt-out across secondmates kunchenguid/firstmate#6154)c35b9a69fix: reduce supervision exit latency and stabilize host tests (fix: reduce supervision exit latency and stabilize host tests kunchenguid/firstmate#6179)Conflict files and resolutions (3)
.agents/skills/afk/SKILL.md.agents/skills/process-event-sources/SKILL.mdtests/fm-brief.test.shEvery other file auto-merged; the fork's divergences in the auto-merged scripts (
fm-config-inherit-lib.shexit-3 handling,fm-pr-merge.shREST fallback, the Slack surfaces) are intact beside upstream's changes.Supervision host now on by default for Claude primaries (0a2cdf9, 12e90e1)
What changes for this home: the main home runs a Claude primary and has no
config/supervision-host, so after this lands and the home updates, the supervision host starts running by default, on the Claude engine at its default model (sonnet).Attended, it takes the routine wakes the Pi branch would take on a headless
claude -p --safe-modesession, so routine outcomes stop reaching the main conversation; check-kind wakes (merge polls, credential failures, process-event wakes such as the Slack captain channel) still reach main while the captain is present.Away,
/afkno longer launches the away daemon: the host is the away session and takes every wake, including check-kind ones.It spends Claude quota headlessly for each handled wake.
The opt-out is the presence flag
config/supervision-host-off(12e90e1 moved it out ofconfig/supervision-host), and a primary's opt-out is inherited by secondmates; this home has no registered secondmates today.Fork settings that interact:
.claude/settings.jsoncarries both upstream'sfm-host-mirror.sh hook claudeand the fork'sfm-slack-mirror.sh stopStop hooks. The host mirror hook now records by default. The engine runs with--safe-mode, which loads none of the home's hooks, so the fork's Slack mirror never fires for an engine turn and cannot double-post.config/crew-harness,config/crew-dispatch.json, andconfig/slack-captaindo not interact with the host gate.Upstream commits that change firstmate's operating contract
0a2cdf95) supervision host on by default for Claude primaries; fix: inherit supervision host opt-out across secondmates kunchenguid/firstmate#6154 (12e90e14) opt-out becomesconfig/supervision-host-off, inherited by secondmates.2d833ff1) a quiet record reads as an attended captain across host, watcher, and daemon; only an away record is away.b3d41332) Lavish worker boards confirm the reply before handoff; the scout brief's Lavish offer now follows the legacy board-compatibility floor.1f2c9548) pending-reply grace is measured from turn completion, not delivery.c5f48e4c)fm-pr-mergeretries a bounded number of times while GitHub reports mergeable UNKNOWN.260c4f08) every open owner converges onto a known terminal contribution.46d58d64) watcher arms follow an attached cycle; live-holder stall bound semantics clarified.Local verification versus CI
Local runs were limited to the conflict-resolved files and the four tests named in the intent; CI runs the full lint and behaviour suites.
FM_PI_PACKAGE_DIR.tests/fm-pi-branch-extension.test.shunder Pi 0.99.1: fork maind860a1bcfails with CI's exact message ("Pi outcomes rendering consumers must preserve stock behavior"); this merge passes (51 ok).tests/fm-calm-pi-extension.test.shunder Pi 0.99.1, run without thetest_queued_operational_escape_e2ecase (see Findings not fixed): fork main fails with CI's exact message ("rendered export DOM violated the Calm conversation boundary"); this merge passes (14 ok, exit 0).tests/fm-remote-secondmate-trace-context.test.shandtests/fm-spawn-compact-adviser-disable-remote.test.sh: both pass on this merge. Their CI failure is Git 2.55's local-clone race, which upstream's--no-localclone fixes and which Git 2.53 cannot reproduce, so this PR's CI (Behavior portable serial 8) is the evidence for them.tests/fm-brief.test.sh(conflict-resolved): passes.bin/fm-lint.shwith the pinned ShellCheck 0.11.0 and actionlint 1.7.12, both on the conflict-resolved test file and in the pipeline's default changed-file mode: clean.bin/fm-doc-audience-check.sh: ok.bin/fm-test-run.sh --check-coverageon the final head reports the largest serial shard at 1074843 ms (about 17.9 min) against the 1200000 ms packing budget; CI's serial 5 passed on the head.Findings not fixed
author.*andcommitter.*for every git dir under~/devs/, and those beat the per-worktree crewuser.*thatbin/fm-git-identity.sh apply-worktreewrites, so git resolves the captain as author. The guard correctly refused; this merge was committed with the crew identity passed throughGIT_AUTHOR_*/GIT_COMMITTER_*.apply-worktreeshould probably also setauthor.*andcommitter.*in the worktree config; that needs its own test, so it is left out of this PR.tests/fm-calm-pi-extension.test.sh'squeued_mixedcase ("did not list the captain's queued follow-up") fails on this host with both Pi 0.87.1 and Pi 0.99.1, identically on pristine upstreamc35b9a69, fork main, and this merge; CI passes it. Host-specific (terminal or timing), not caused by this merge.--safe-modeand loads no hooks. Worth deciding whether away-mode Slack replies should go throughbin/fm-slack-post.sh, or whether this home should opt out withconfig/supervision-host-off.eb77f02b(ci: rebalance portable test groups and enforce a packing budget kunchenguid/firstmate#6192) was taken by cherry-picking its three PR-branch commits rather than merged, so git still lists it as behind; the next sync's true merge of it should be content-equivalent.82fb99c6misdescribes what it added,1ece78a0is a raw JSON blob); both are pipeline-written and already pushed.tests/fm-ci-workflow.test.shlocally (no ruby on this host); CI covered it.Risk Assessment
Testing
On the target commit,
fm-test-run.sh --check-coveragereports serial_max_ms=1074843 against serial_budget_ms=1200000 (previously unbudgeted/unreported on base, with serial_unhinted dropping from 17 to 1), directly demonstrating the rebalance closes the near-30-minute shard-5 overrun the user intent describes; the dedicated packing-boundary regression test passed standalone. A larger batch of related serial-packing tests was still executing in the background when this report was generated and its result is not yet known.test_portable_serial_packing_budget_boundaryin tests/fm-test-run.test.sh — exit 0, 'ok - serial packing accepts the exact budget and refuses one millisecond above it'Evidence: check-coverage on target commit
Evidence: check-coverage on base fm-test-run.sh (pre-rebalance)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
🔧 **Test** - 1 issue found ✅
bash tests/fm-remote-secondmate-trace-context.test.shstandalone to c…bash tests/fm-spawn-compact-adviser-disable-remote.test.shstan…bash tests/fm-calm-pi-extension.test.sh (exit=0, all subtests ok)bash tests/fm-pi-branch-extension.test.sh (exit=0, all subtests ok)bash tests/fm-remote-secondmate-trace-context.test.sh (started, still running)bash tests/fm-spawn-compact-adviser-disable-remote.test.sh (queued behind previous test, not yet started)✅ No issues found.
test_portable_serial_packing_budget_boundaryin tests/fm-test-run.test.sh — exit 0, 'ok - serial packing accepts the exact budget and refuses one millisecond above it'bin/fm-test-run.sh --check-coverageon the target commit (1124b274)bin/fm-test-run.sh --check-coveragewith bin/fm-test-run.sh reverted to base commit d860a1bc for comparison, then restoredtest_portable_serial_packing_budget_boundaryfrom tests/fm-test-run.test.sh run standalone (exit 0)tests/fm-test-run.test.sh: test_portable_shard_union_and_coverage_guard, test_portable_serial_shards_partition_the_serial_lane, test_portable_serial_hint_coverage_is_reported_and_bounded, test_portable_serial_shard_lane_refusals — launched, still running in background at time of this report✅ **Document** - passed
✅ No issues found.
✅ No issues found.
🔧 Fix applied.
1 warning still open:
✅ **Push** - passed
✅ No issues found.
✅ No issues found.