Skip to content

fix: sync fork with upstream to restore portable CI behavior - #38

Merged
digbycampbell merged 21 commits into
mainfrom
fm/fm-upstream-sync-r3
Sep 30, 2026
Merged

digbycampbell merged 21 commits into
mainfrom
fm/fm-upstream-sync-r3

Conversation

@digbycampbell

@digbycampbell digbycampbell commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

THIS PR MUST BE LANDED WITH A MERGE COMMIT - do not squash, do not rebase, do not force-push.
It carries a true merge of upstream kunchenguid/firstmate; a squash or rebase would drop the shared ancestry the last two syncs (PR 33 and PR 35) kept, and the next sync would conflict against history git can no longer see.

Intent

"you can upgrade firstmate now right? our fork is mostly up to date with upstream main?" and, on Slack, "You can upgrade firstmate as the upstream repo was folded in right?"

PR 35 true-merged upstream eb219c80 on 2026-09-29, and upstream main has moved on since.
Upstream's e2668de0 ("fix: restore portable CI behavior across Pi rendering and remote provisioning", kunchenguid#6162) repairs the portable CI lanes after the CI image moved to Pi 0.99 and Git 2.55; without it this fork's main fails Behavior portable serial 2, 4 and 8 (the Calm export-DOM boundary, the Pi outcomes-rendering stock-behavior check, and the two remote-seed provisioning tests), as a rerun of main's CI (run 36518848136) showed.
This sync takes that fix, and the rest of upstream main, so those lanes turn green.

What

This brings the fork up to date with the upstream firstmate project again.
All 15 upstream changes made since the last sync are merged in, and every fork-specific feature is kept.
The practical effects: the fork's CI goes green again on the Pi and remote-setup tests that broke when GitHub's build machines got newer Pi and Git versions; on a Claude-run home, the background supervision helper now runs by default instead of only when switched on; and several supervision, merge, and Lavish review reliability fixes arrive.
Only three files needed hand resolution, all small.
CI's longest test group then ran past its 30-minute limit, so upstream's own rebalancing of the CI test groups was taken as well.

Merge shape

  • 507229da Merge upstream kunchenguid/firstmate c35b9a6 into the fork (true merge, git merge --no-ff upstream/main).
    Parents (git show -s --format=%P 507229da): d860a1bc8cba1ce4ee5ab255bb87dc8883ab63d1 (fork main, the PR 35 merge) and c35b9a69be55c3c1147035701ad167405378dbb7 (upstream/main tip).
  • Merge base eb219c80 (the PR 35 sync point).
  • Follow-up commits on top of the merge, from the CI repair round (serial 5 hit its 30-minute job limit, twice, because the merged upstream tests pushed it past main's 28m21s):
    • 9c353d82, e8e5aad5, 1124b274 (author kunchenguid) are upstream PR ci: rebalance portable test groups and enforce a packing budget kunchenguid/firstmate#6192's three branch commits, cherry-picked; together they are the content of upstream eb77f02b "ci: rebalance portable test groups and enforce a packing budget". The one conflict, bin/fm-test-run.sh's serial weight-hint table, keeps upstream's remeasured values plus the fork's 15 fork-only test hints; the resulting runner equals upstream's tip plus exactly the fork's existing divergence.
    • 82fb99c6 "no-mistakes(document): Docs already updated in-branch; verified consistent, no stale facts remain" is misleading: it actually added a stray 1849-line tests/fm-test-run-focused.test.sh, a scratch near-copy of tests/fm-test-run.test.sh left by the repair round.
    • 1ece78a0 removes that stray file again (its subject is a JSON blob the lint step wrote). Net change after the merge: only bin/fm-test-run.sh, docs/fm-test-portable-shards.md, tests/fm-session-start.test.sh, tests/fm-test-run.test.sh.
    • The pipeline pushed these subjects, so they cannot be reworded without rewriting pushed history.
  • Validated head: 1ece78a049503225c9359332e53cd4b881be9e2d, all 19 CI checks green, including "PR must be raised via no-mistakes" and Behavior portable serial 2, 4, 5 and 8.
  • Commits still behind upstream/main at push time: eb77f02b by ancestry only. Its content is on this branch through the cherry-picks above, so the next sync's merge of it should be content-equivalent; upstream is otherwise level.

Upstream commits taken (15, eb219c80..c35b9a69)

Conflict files and resolutions (3)

File Resolution
.agents/skills/afk/SKILL.md Kept the fork's trigger-only description; upstream's wording change there ("opted into it" -> "runs it") lives on in the skill body, which merged cleanly.
.agents/skills/process-event-sources/SKILL.md Kept the fork's one-line trigger description; upstream's body changes (kunchenguid#6169 Lavish reply confirmation) merged cleanly beside the fork's Slack, quota-topic, and GitHub-assignment adapter paragraphs.
tests/fm-brief.test.sh Kept the fork's harness-banner isolation test and took upstream's rewritten Lavish-floor comment above the scout test.

Every other file auto-merged; the fork's divergences in the auto-merged scripts (fm-config-inherit-lib.sh exit-3 handling, fm-pr-merge.sh REST fallback, the Slack surfaces) are intact beside upstream's changes.

Supervision host now on by default for Claude primaries (0a2cdf9, 12e90e1)

What changes for this home: the main home runs a Claude primary and has no config/supervision-host, so after this lands and the home updates, the supervision host starts running by default, on the Claude engine at its default model (sonnet).
Attended, it takes the routine wakes the Pi branch would take on a headless claude -p --safe-mode session, so routine outcomes stop reaching the main conversation; check-kind wakes (merge polls, credential failures, process-event wakes such as the Slack captain channel) still reach main while the captain is present.
Away, /afk no longer launches the away daemon: the host is the away session and takes every wake, including check-kind ones.
It spends Claude quota headlessly for each handled wake.
The opt-out is the presence flag config/supervision-host-off (12e90e1 moved it out of config/supervision-host), and a primary's opt-out is inherited by secondmates; this home has no registered secondmates today.

Fork settings that interact:

  • .claude/settings.json carries both upstream's fm-host-mirror.sh hook claude and the fork's fm-slack-mirror.sh stop Stop hooks. The host mirror hook now records by default. The engine runs with --safe-mode, which loads none of the home's hooks, so the fork's Slack mirror never fires for an engine turn and cannot double-post.
  • The Slack captain channel: attended, its wakes stay on main, so replies keep flowing through the Slack mirror as today. Away, the host engine takes those wakes; since the engine loads no hooks, an engine turn's reply is not mirrored into Slack (see Findings not fixed).
  • config/crew-harness, config/crew-dispatch.json, and config/slack-captain do not interact with the host gate.

Upstream commits that change firstmate's operating contract

Local verification versus CI

Local runs were limited to the conflict-resolved files and the four tests named in the intent; CI runs the full lint and behaviour suites.

  • This host has Pi 0.87.1 and Git 2.53, while CI installs Pi 0.99.1 and runs Git 2.55. To reproduce CI's Pi failures, Pi 0.99.1 was installed into a private scratch prefix and passed to the tests through FM_PI_PACKAGE_DIR.
  • tests/fm-pi-branch-extension.test.sh under Pi 0.99.1: fork main d860a1bc fails with CI's exact message ("Pi outcomes rendering consumers must preserve stock behavior"); this merge passes (51 ok).
  • tests/fm-calm-pi-extension.test.sh under Pi 0.99.1, run without the test_queued_operational_escape_e2e case (see Findings not fixed): fork main fails with CI's exact message ("rendered export DOM violated the Calm conversation boundary"); this merge passes (14 ok, exit 0).
  • tests/fm-remote-secondmate-trace-context.test.sh and tests/fm-spawn-compact-adviser-disable-remote.test.sh: both pass on this merge. Their CI failure is Git 2.55's local-clone race, which upstream's --no-local clone fixes and which Git 2.53 cannot reproduce, so this PR's CI (Behavior portable serial 8) is the evidence for them.
  • tests/fm-brief.test.sh (conflict-resolved): passes.
  • Lint: bin/fm-lint.sh with the pinned ShellCheck 0.11.0 and actionlint 1.7.12, both on the conflict-resolved test file and in the pipeline's default changed-file mode: clean. bin/fm-doc-audience-check.sh: ok.
  • Rebalance: bin/fm-test-run.sh --check-coverage on the final head reports the largest serial shard at 1074843 ms (about 17.9 min) against the 1200000 ms packing budget; CI's serial 5 passed on the head.
  • no-mistakes: the test step's live-validation verdict was inconclusive for the two remote-seed scenarios (they need Git 2.55); firstmate approved it, with CI as the evidence. The lint step's only finding was this host's ShellCheck 0.10.0 failing the 0.11.0 version check; it was approved after the pinned-version lint above passed.

Findings not fixed

  • Crewmate commits in firstmate task worktrees are refused by the identity guard on this machine. The captain's global git config sets author.* and committer.* for every git dir under ~/devs/, and those beat the per-worktree crew user.* that bin/fm-git-identity.sh apply-worktree writes, so git resolves the captain as author. The guard correctly refused; this merge was committed with the crew identity passed through GIT_AUTHOR_*/GIT_COMMITTER_*. apply-worktree should probably also set author.* and committer.* in the worktree config; that needs its own test, so it is left out of this PR.
  • tests/fm-calm-pi-extension.test.sh's queued_mixed case ("did not list the captain's queued follow-up") fails on this host with both Pi 0.87.1 and Pi 0.99.1, identically on pristine upstream c35b9a69, fork main, and this merge; CI passes it. Host-specific (terminal or timing), not caused by this merge.
  • With the supervision host now on by default, a Slack captain-channel wake handled by the host while the captain is away gets no reply mirrored into Slack, because the engine runs with --safe-mode and loads no hooks. Worth deciding whether away-mode Slack replies should go through bin/fm-slack-post.sh, or whether this home should opt out with config/supervision-host-off.
  • eb77f02b (ci: rebalance portable test groups and enforce a packing budget kunchenguid/firstmate#6192) was taken by cherry-picking its three PR-branch commits rather than merged, so git still lists it as behind; the next sync's true merge of it should be content-equivalent.
  • Two commit subjects on this branch are poor (82fb99c6 misdescribes what it added, 1ece78a0 is a raw JSON blob); both are pipeline-written and already pushed.
  • The CI repair round could not run tests/fm-ci-workflow.test.sh locally (no ruby on this host); CI covered it.
  • This host's ShellCheck is 0.10.0 against the 0.11.0 pin, which the pipeline's lint step reports as a finding on every run.

Risk Assessment

⚠️ Medium: The branch true-merges 14 upstream commits plus three previously-uncertified fixer commits; the fixer commits (9c353d8, e8e5aad, 1124b27) are tightly scoped to bin/fm-test-run.sh CI-timing constants, the coverage guard's new serial_max_ms/serial_budget_ms bound, its docs, and an executable boundary test — the packing-guard logic (bin-packing sum, boundary comparison, exact-vs-over-budget refusal message) checks out correctly and the new test exercises the real runner rather than grepping source text, and the doc edits are non-functional cleanup, but the overall diff size and unreviewed-fixer provenance warrant a human pass before delivery.

Testing

On the target commit, fm-test-run.sh --check-coverage reports serial_max_ms=1074843 against serial_budget_ms=1200000 (previously unbudgeted/unreported on base, with serial_unhinted dropping from 17 to 1), directly demonstrating the rebalance closes the near-30-minute shard-5 overrun the user intent describes; the dedicated packing-boundary regression test passed standalone. A larger batch of related serial-packing tests was still executing in the background when this report was generated and its result is not yet known.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Coverage guard reports the new serial packing budget and stays under it ✅ pass live FM_TEST_COVERAGE ok ... serial_max_ms=1074843 serial_budget_ms=1200000 — largest packed serial shard (~17.9min) is comfortably under the 20-minute target, addressing the reported shard-5 near-30-minut…
Rebalance measurably reduces unhinted (guessed-duration) scripts versus the pre-fix hint table ✅ pass live serial_unhinted dropped from 17 (base bin/fm-test-run.sh) to 1 (target commit) with the same suite composition
Packing accepts a fixture exactly at budget and refuses one millisecond above it (existing regression test) ✅ pass live test_portable_serial_packing_budget_boundary in tests/fm-test-run.test.sh — exit 0, 'ok - serial packing accepts the exact budget and refuses one millisecond above it'
Serial-shard partitioning and hint-coverage reporting for the full suite (broader regression battery) ⏸️ untested no The batch of tests (test_portable_shard_union_and_coverage_guard, test_portable_serial_shards_partition_the_serial_lane, test_portable_serial_hint_coverage_is_reported_and_bounded, test_portable_seria…
Evidence: check-coverage on target commit
FM_TEST_COVERAGE ok total=257 parallel=24 parallel_max_ms=665545 parallel_imbalance_ms=2 parallel_unhinted=0 serial=217 serial_shards=9 serial_unhinted=1 serial_max_ms=1074843 serial_budget_ms=1200000 herdr=16
Evidence: check-coverage on base fm-test-run.sh (pre-rebalance)
FM_TEST_COVERAGE ok total=257 parallel=24 parallel_max_ms=417163 parallel_imbalance_ms=2894 parallel_unhinted=0 serial=217 serial_shards=9 serial_unhinted=17 herdr=16
- Outcome: 🔧 1 issue found ✅ across 2 runs (14m8s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

✅ No issues found.

🔧 **Test** - 1 issue found ✅
  • ⚠️ live validation verdict: inconclusive (2 of 4 scenarios were driven live against the product); untested: Behavior portable serial 8a — remote seed provisions the trace-context route (fm-remote-secondmate-trace-context.test.sh), Behavior portable serial 8b — remote seed provisions the compact-adviser-disable route (fm-spawn-compact-adviser-disable-remote.test.sh)
  • Live validation: ⚠️ inconclusive - 2 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Behavior portable serial 2 — Calm rendered-export DOM respects the conversation boundary (fm-calm-pi-extension.test.sh) ✅ pass live bash tests/fm-calm-pi-extension.test.sh exited 0 with all listed subtests 'ok', including 'the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures...'
Behavior portable serial 4 — Pi outcomes rendering consumers preserve stock behavior (fm-pi-branch-extension.test.sh) ✅ pass live bash tests/fm-pi-branch-extension.test.sh exited 0 with all listed subtests 'ok'
Behavior portable serial 8a — remote seed provisions the trace-context route (fm-remote-secondmate-trace-context.test.sh) ⏸️ untested no Test was still executing in a background shell (started 16:28, PID 2914589) when this report was forced before completion; rerun bash tests/fm-remote-secondmate-trace-context.test.sh standalone to c…
Behavior portable serial 8b — remote seed provisions the compact-adviser-disable route (fm-spawn-compact-adviser-disable-remote.test.sh) ⏸️ untested no Queued after fm-remote-secondmate-trace-context in the same background loop and had not started yet when this report was forced; rerun bash tests/fm-spawn-compact-adviser-disable-remote.test.sh stan…
  • bash tests/fm-calm-pi-extension.test.sh (exit=0, all subtests ok)
  • bash tests/fm-pi-branch-extension.test.sh (exit=0, all subtests ok)
  • bash tests/fm-remote-secondmate-trace-context.test.sh (started, still running)
  • bash tests/fm-spawn-compact-adviser-disable-remote.test.sh (queued behind previous test, not yet started)

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Coverage guard reports the new serial packing budget and stays under it ✅ pass live FM_TEST_COVERAGE ok ... serial_max_ms=1074843 serial_budget_ms=1200000 — largest packed serial shard (~17.9min) is comfortably under the 20-minute target, addressing the reported shard-5 near-30-minut…
Rebalance measurably reduces unhinted (guessed-duration) scripts versus the pre-fix hint table ✅ pass live serial_unhinted dropped from 17 (base bin/fm-test-run.sh) to 1 (target commit) with the same suite composition
Packing accepts a fixture exactly at budget and refuses one millisecond above it (existing regression test) ✅ pass live test_portable_serial_packing_budget_boundary in tests/fm-test-run.test.sh — exit 0, 'ok - serial packing accepts the exact budget and refuses one millisecond above it'
Serial-shard partitioning and hint-coverage reporting for the full suite (broader regression battery) ⏸️ untested no The batch of tests (test_portable_shard_union_and_coverage_guard, test_portable_serial_shards_partition_the_serial_lane, test_portable_serial_hint_coverage_is_reported_and_bounded, test_portable_seria…
  • bin/fm-test-run.sh --check-coverage on the target commit (1124b274)
  • bin/fm-test-run.sh --check-coverage with bin/fm-test-run.sh reverted to base commit d860a1bc for comparison, then restored
  • test_portable_serial_packing_budget_boundary from tests/fm-test-run.test.sh run standalone (exit 0)
  • tests/fm-test-run.test.sh: test_portable_shard_union_and_coverage_guard, test_portable_serial_shards_partition_the_serial_lane, test_portable_serial_hint_coverage_is_reported_and_bounded, test_portable_serial_shard_lane_refusals — launched, still running in background at time of this report
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 1)

  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
1 warning still open:

  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

✅ No issues found.

kunchenguid and others added 16 commits September 28, 2026 17:19
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
Takes the 15 upstream commits eb219c8..c35b9a6, including e2668de's
portable CI repairs for Pi rendering and remote provisioning and 0a2cdf9's
default-on supervision host for Claude primaries.

Conflicts:
- .agents/skills/afk/SKILL.md: kept the fork's trigger-only description.
- .agents/skills/process-event-sources/SKILL.md: kept the fork's one-line
  trigger description; upstream's body changes merged cleanly.
- tests/fm-brief.test.sh: kept the fork's harness-banner isolation test and
  took upstream's rewritten Lavish floor comment.
@digbycampbell digbycampbell changed the title fix(bin): sync upstream fixes for portable CI, AFK return, and live-lab support fix: sync fork with upstream to restore portable CI behavior Sep 30, 2026
@digbycampbell
digbycampbell merged commit 977325f into main Sep 30, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants