Skip to content

merge: catch up with kunchenguid/firstmate through 4a9979a (7 upstream changes) - #44

Merged
HelloWorldSungin merged 14 commits into
mainfrom
fm/fm-upstream-catchup-round-1
Aug 10, 2026
Merged

HelloWorldSungin merged 14 commits into
mainfrom
fm/fm-upstream-catchup-round-1

Conversation

@HelloWorldSungin

@HelloWorldSungin HelloWorldSungin commented Aug 4, 2026 •

Copy link
Copy Markdown
Owner

Merges kunchenguid/firstmate@4a9979a (upstream main) into this fork, catching up the 7 upstream changes that landed on 2026-08-04.

This is a true merge, not a rebase, squash or cherry-pick. Fork main stays append-only so every firstmate home keeps fast-forwarding normally, which is the property the whole self-update path depends on. Merge parents: aed567f (fork main) and 4a9979a (upstream main); merge base was 3d9d12d.

All issue and PR numbers below carry their repository, because the two repos' number spaces overlap and a bare #N is ambiguous.

The finding that reframes this merge

The scout report and the task brief both described bin/fm-remote-doctor.sh as an add/add collision between "two different programs" - a 53-line fork implementation against an ~800-line upstream one - and treated choosing between them as a design decision.

That framing turns out to be wrong, and I verified it rather than assuming it. The fork's remote-doctor and remote-entrypoint are byte-identical to upstream's own first version of that work, kunchenguid/firstmate#1623:

git show 733a504:bin/fm-remote-doctor.sh | sha256sum   ->  4decf626...26bd4   (kunchenguid/firstmate#1623)
git show aed567f:bin/fm-remote-doctor.sh | sha256sum   ->  4decf626...26bd4   (fork main)
git show 733a504:bin/fm-remote-entrypoint.sh | sha256sum -> fcf7cd9a...933ea   (kunchenguid/firstmate#1623)
git show aed567f:bin/fm-remote-entrypoint.sh | sha256sum -> fcf7cd9a...933ea   (fork main)

The fork received that work through its own HelloWorldSungin/firstmate#22 (b45b811). Upstream has since evolved the same lineage through kunchenguid/firstmate#1639, kunchenguid/firstmate#1659, kunchenguid/firstmate#1660 and kunchenguid/firstmate#1691. So there were never two competing designs to choose between - there is one design, and upstream is five revisions further along it. Git reported add/add only because the same content reached the two branches by different routes.

Taking upstream for this family is therefore not an implementation swap. It is taking the newer revisions of code this fork already runs.

Applicability of the 7 upstream changes

Upstream change What it does Outcome here
kunchenguid/firstmate#1623 (733a504) preflight remote runtime tool paths Adds the read-only remote doctor and the entrypoint's PATH composition Already present. Byte-identical content already on fork main via HelloWorldSungin/firstmate#22. Every conflict it produced was the same content arriving twice.
kunchenguid/firstmate#1639 (e5e8a67) gate remote second mates on Herdr readiness Grows the doctor to a readiness owner with --fix; adds fm-remote-readiness-lib.sh, Herdr gating in spawn and bootstrap Taken. Conflicted against kunchenguid/firstmate#1623-era content in the doctor, home-seed and their docs and tests; resolved to upstream as the newer revision of the same lineage.
kunchenguid/firstmate#1659 (c8edff3) isolate remote secondmates in shared Herdr session Per-secondmate Herdr session isolation; adds remote_herdr_session= task metadata Taken. Its one-line AGENTS.md change conflicted with fork metadata additions; both intents kept (see below).
kunchenguid/firstmate#1660 (a83be60) route remote commands through an Aqua job worker Adds fm-remote-job-lib.sh and fm-remote-job-worker.sh; entrypoint routes through the worker, doctor stays reachable over plain SSH Taken. Largest change; the entrypoint conflict was the kunchenguid/firstmate#1623 direct-exec path against this worker path.
kunchenguid/firstmate#1691 (fc3684a) clarify remote doctor bootstrap path Doctor and entrypoint wording plus the DOCTOR_SHA256 pin value Taken clean.
kunchenguid/firstmate#1699 (1939785) bound remote SSH dead-peer detection Bounded dead-peer window in fm-on.sh Taken clean. No fork content in fm-on.sh.
kunchenguid/firstmate#1701 (4a9979a) report stale AXI tools during bootstrap Stale tasks-axi/quota-axi detection in bootstrap Taken clean. Auto-merged with fork bootstrap changes.

Doctor bootstrap authentication - verified, not assumed

bin/fm-remote-entrypoint.sh pins the doctor by DOCTOR_SHA256 so an altered doctor cannot bootstrap when git is unavailable. A hand-blended or partially resolved doctor would fail closed there. The doctor was therefore taken byte-exact and the pin was recomputed against the merged tree rather than assumed to have survived:

pin    : 7bb13d9fad8455978bf109d4681a3aa3cb170565c8a74be4ec7b520427db14c2
actual : 7bb13d9fad8455978bf109d4681a3aa3cb170565c8a74be4ec7b520427db14c2   (sha256sum bin/fm-remote-doctor.sh)

Per-file reasoning for every contract file touched

AGENTS.md - one conflict, in the <id>.meta inventory line. Both sides had extended it since the merge base and the intents are compatible, so both were kept: the fork's issue=, work_item= and pr_target= fields and its .pr-status, .gbrain and .usage-sessions record lines, plus upstream's remote_herdr_session= field from kunchenguid/firstmate#1659. Nothing was dropped from either side.

.agents/skills/bootstrap-diagnostics/SKILL.md and .agents/skills/secondmate-provisioning/SKILL.md - auto-merged, no conflict. Upstream's readiness-gate and stale-AXI diagnostic text landed alongside the fork's existing entries.

docs/scripts.md - conflict in the script table. Resolved to upstream, which replaces the kunchenguid/firstmate#1623 one-line doctor description with the readiness description and adds rows for fm-remote-job-lib.sh and fm-remote-job-worker.sh. The fm-remote-readiness-lib.sh row auto-merged separately. Fork-only rows elsewhere in the table (fm-recall.sh, fm-usage.mjs, fm-bearings-snapshot.sh and the rest) were untouched and verified still present.

docs/remote-secondmates.md - four conflicts, all describing the same subsystem at its two revisions; all resolved to upstream. The resulting file matches upstream exactly, because everything the fork had added to this doc was the kunchenguid/firstmate#1623 text that upstream itself has since rewritten.

bin/fm-remote-doctor.sh, bin/fm-remote-entrypoint.sh, bin/fm-remote-home-seed.sh - taken byte-exact from upstream. The fork's content in all three was kunchenguid/firstmate#1623, superseded in-lineage. The home-seed's restore_registry_and_brief helper exists on both sides and survives.

What survived from the fork side, and how it was verified

The specific failure mode being guarded against is a conflict resolution that quietly reverts fork work. Checks run against the merged tree:

  • cursor and agy crew adapters - present. bin/fm-launch-lib.sh retains the crew-only, herdr-only adapter contract and the divergence note already carried in that file; bin/fm-agy-trust-lib.sh present; bin/fm-spawn.sh retains 37 agy references. Upstream has none of this and did not touch it.
  • The fork's GBrain, usage-accounting, PR-status, work-item and recall features - bin/fm-recall.sh, bin/fm-usage.mjs, bin/fm-pr-status.sh, bin/fm-issue-lib.sh, bin/fm-gbrain-lib.sh, docs/gbrain-scoping.md, docs/usage-accounting.md all present.
  • The fork's GBrain hardening inside the remote family - bin/fm-remote-inherit.sh retains its shared-brain-plane schema validation, which refuses a credential pasted into config/gbrain.json at the receiving code root. This sits in the merged remote subsystem and was explicitly preserved rather than swept up in the family resolution.
  • Its regression test - the fork-authored case in tests/fm-remote-secondmate-lifecycle-e2e.test.sh asserting that refusal was kept, while that file's three conflicting hunks were resolved to upstream. This file is the one place where the resolution is deliberately mixed.

One coverage change worth naming

tests/fm-on.test.sh was taken from upstream, which drops five fork assertions that pinned kunchenguid/firstmate#1623 behaviour. Four have direct upstream counterparts, verified by name in upstream's version of the file: deduplicated child PATH composition, doctor-reports-the-entrypoint-PATH, required/optional tool reporting, and tracked-command authorization excluding checkout-local git. Upstream also adds tests/fm-remote-doctor.test.sh and tests/fm-remote-job.test.sh, roughly 1,155 lines of coverage this fork did not have.

The fifth, "the entrypoint gives an actionable missing-git diagnostic", has no exact upstream twin. It is not a silent regression: the diagnostic string it asserted still exists verbatim in the merged entrypoint, but under kunchenguid/firstmate#1660 a missing git on the doctor path now takes the hash-authenticated bootstrap branch instead, which upstream covers with "doctor bootstrap remains authenticated when git is unavailable". The old assertion is false by design after this merge, which is why it could not be carried forward.

Retired fork design, for the divergence ledger

Recorded here because the ledger the sync design proposes does not exist yet. The fork's own remote-second-mate machinery - a 53-line read-only fm-remote-doctor.sh plus PATH composition inside fm-remote-entrypoint.sh, landed via HelloWorldSungin/firstmate#22 - is retired in favour of upstream's current implementation. The reason is that it was never an independent design: it was upstream's kunchenguid/firstmate#1623 verbatim, and upstream has carried the same code forward through four further revisions. This entry should replace, not accompany, any ledger line describing the fork's remote-doctor as a deliberate divergence, since the byte-identity shows it never was one.

Risk profile

This subsystem is not in live use in this home: there are no registered second mates and no remote-second-mate runtime state, so no working path changes behaviour today. The first remote second mate provisioned from here will exercise upstream's implementation, with upstream's own test coverage behind it, rather than the kunchenguid/firstmate#1623 version that has equally never run here. A large implementation change on machinery nothing currently runs is a materially different proposition from one on a live path.

Deliberately excluded, as parked on unresolved decisions and verified not ancestors of this branch: fm/fm-afk-injection-wedge, fm/fm-crew-state-blind-during-fix-round, fm/fm-parked-decision-stale-noise, fm/fm-subagent-model-routing-guard, fm/fm-vault-drift-check.

Validation

  • bin/fm-lint.sh - clean, exit 0, ShellCheck 0.11.0 pinned.
  • bin/fm-doc-audience-check.sh - ok, 81 surfaces, 275 local links.
  • bin/fm-test-run.sh --all - baselined on both merge parents in isolated clones first, so a pre-existing failure stays separable from a merge-caused one. Results below.
Tree Result
Fork main aed567f (clean clone) total=134 failed=0 skipped_gate=16
Upstream main 4a9979a (clean clone) total=118 failed=5 skipped_gate=13
This merge branch total=136 failed=2 skipped_gate=16

The two failures on this branch are pre-existing upstream failures, not merge-caused. Both are the same scripts, failing the same way, on a pristine upstream checkout:

  • tests/fm-on.test.sh - composed child PATH did not match the portable contract. Upstream's own kunchenguid/firstmate#1660 behaviour resolves the three Nix locations through their final bin symlink to the physical directory, which its test's expectation does not reconstruct. On this host a home-manager nix profile makes the two disagree: expected .../bin:/home/sungin/.local/bin:/usr/local/bin:..., actual inserts /nix/store/v6lk8jk...-home-manager-path/bin.
  • tests/fm-remote-doctor.test.sh - --fix did not create the first needed harness wrapper. Also host-dependent, in upstream's new doctor test.

Upstream's other three baseline failures (tests/fm-calm-pi-extension.test.sh, tests/fm-test-run.test.sh, tests/fm-session-start.test.sh) do not occur on this branch, because the fork's versions of that code and those tests are what the merge kept.

Stated plainly, because it is a real change in status: fork main is green on this host today and this branch is not. Nothing here reverts fork behaviour to cause that - the two failures arrive with upstream's own new test files and reproduce on upstream's own tree - but the suite does stop being green, and both look like upstream test bugs on Nix-style hosts rather than product bugs. Both did pass on CI, whose runners have no nix profile, confirming the host-dependence rather than leaving it assumed. CI surfaced a different failure instead, which was diagnosed to root cause and fixed on this branch; see finding 2 below. CI is now fully green apart from the deliberate no-mistakes gate.

Per-family detail is in the run output; failing-script names above were obtained by re-running the affected families with full logs on both trees.

Two findings from validating this merge

Neither is a conflict resolved the wrong way, and nothing from either side was reverted. Finding 1 is an upstream defect that reproduces on a pristine upstream checkout and arrives here with this merge. Finding 2 belonged to this merge specifically, and is fixed on this branch.

1. Upstream's job worker leaks a permanent background process per test run

bin/fm-remote-job-worker.sh is a daemon: its main loop is while :; do ... sleep; done with no idle timeout and no self-exit. Four upstream-authored tests start a real one - tests/fm-on.test.sh, tests/fm-remote-reply.test.sh, tests/fm-remote-backlog-handoff.test.sh and tests/fm-remote-secondmate-trace-context.test.sh. Each has an EXIT trap that tries to reap it:

trap 'if [ -f "$TMP_ROOT/remote-jobs/worker.pid" ]; then kill "$(cat ...)" ...; fi; rm -rf -- "$TMP_ROOT"' EXIT

That reap is not reliable. worker_shutdown traps TERM, and when worker_publish_quarantine fails it deliberately re-arms the trap and returns 0 instead of exiting (bin/fm-remote-job-worker.sh:252-258). Racing that against the trap's own rm -rf "$TMP_ROOT", which removes the state root the shutdown needs, leaves a daemon that ignored its TERM and now polls a deleted directory forever.

Measured on this host during validation: 17 orphaned workers accumulated across a handful of suite runs, all reparented to init, 552 CPU-seconds between them, 14 of 15 pointing at state roots that no longer exist. The count grew with every run, so the leak is per-run and unbounded.

Attribution: 6 of them came from a clean upstream main clone, not from this merge branch, so it is an upstream defect. The control is decisive in the other direction too - fork main does not contain fm-remote-job-worker.sh at all and leaked nothing.

This matters more than a test-hygiene nit because the machine that runs this suite is the machine firstmate itself runs on. All 17 were reaped after the investigation; none remain.

2. A list-length assumption in an upstream test fixture, found here and fixed here

Diagnosed to root cause and fixed on this branch. Behavior portable serial 1 and the whole CI run are green; the only remaining red check is the deliberate no-mistakes gate.

Symptom. tests/fm-remote-secondmate-lifecycle-e2e.test.sh failed on this branch's CI with not ok - first inheritance transaction never reached its blocked write, reproducibly, while passing on fork main's CI and passing every local run on this 24-core host.

Root cause. The fixture blocks a remote inheritance write and waits for it. Its inherit-block transport mode signals exactly one path (:151):

inherit-block:fm-remote-inherit.sh:data/captain-shared.md)

Propagation order is every FM_INHERITABLE_CONFIG item under config/, then the one shared data file (bin/fm-config-inherit-lib.sh:93). So the marker can only appear once the push walks past every config/ entry. The wait, however, was a fixed 250 polls - five seconds - which silently assumed how many entries precede it:

Tree FM_INHERITABLE_CONFIG tail Has the fixture
Upstream main ... trace-context yes
Fork main ... trace-context gbrain.json project-board no
This merge ... trace-context gbrain.json project-board yes

The fork's GBrain and project-board work put two extra remote round trips ahead of data/captain-shared.md. Only the merge has both the longer list and the fixture, which is why neither parent shows it. A fast host still reached the blocked write inside five seconds; a loaded shared runner did not. Instrumentation on CI caught the push mid-flight on exactly the expected file:

fm-config-push.sh -> fm-remote-inherit-push.sh ios 9
   └─ fm-on.sh ios fm-remote-inherit.sh absent config/gbrain.json 0 e3b0c442...855 9

The instrumentation commit was reverted before the fix, so the merge commit's tree was never altered to obtain this.

This was an integration defect belonging to this merge, not to either parent. An earlier revision of this description read it as probably inherited from upstream; that was wrong and is corrected here rather than quietly dropped.

Fix, one commit on top of the merge, touching only the fixture. Both blocked-inheritance waits now share await_blocked_inherit_write, which keys on progress rather than elapsed time: every remote call bumps the transport's invocation counter and resets the patience window, so declaring another inheritable file can lengthen the wait but can never fail it. The remaining window expires only if the transfer stops making calls at all, which is a genuine wedge rather than a slow runner.

Raising the budget was deliberately rejected. It would have left the identical trap for the next inherited file, and the evidence that this happens is in the fixture itself: the sibling spawn wait had already hit this and papered over it at thirty seconds, carrying a comment acknowledging that earlier inherited files traverse the worker first. Both waits are now deterministic with respect to list length, and the sibling's latent version of the bug is closed in the same pass.

Confirmation. Full CI green on 2dd6b3e, every behaviour lane included. Worth noting that the test took 183s on that run against roughly 83s on the failing runs, so the runner really was slow: the progress-keyed wait absorbed a delay that any fixed budget in the old shape would have failed on.

Shipped behaviour is untouched. Inheriting gbrain.json and project-board is correct and intended; only the test's waiting strategy changed.

The expected red check

This PR is raised direct-PR, deliberately bypassing this repository's no-mistakes pipeline, because that pipeline's rebase step would linearise the merge and force every conflict to be re-resolved one commit at a time - destroying the merge structure that is the entire point of this branch. The same call was made, with authorization, on the precedent round that landed as HelloWorldSungin/firstmate#17.

The consequence is one red check, the fork's own "PR must be raised via no-mistakes" gate. That check is expected here and is not a failure to fix.

It is now the only red check: every behaviour lane, lint, the coverage guard, the macOS snapshot check and repo invariants are green on 2dd6b3e. The Behavior portable serial 1 failure that appeared on earlier revisions is fixed, not waived - see finding 2. If more than one red check appears, the extras are duplicate runs of the body-compliance gate, one per edit of this description.

Merge method matters: this must be merged as a true merge commit, never squashed. A squash discards the merge parentage, so the merge base never advances and the next sync round re-presents and re-conflicts everything already taken here.

Refresh against current fork main - 2026-08-10

This section supersedes the earlier local branch-result counts and head references above.
It does not change the applicability decisions for the seven upstream changes.

Fork main advanced by 14 commits while this PR was open, through HelloWorldSungin/firstmate#72 at d2c157d.
The branch was refreshed by merging that fork head into this branch as merge commit e3bfe14, with parents 2dd6b3e and d2c157d.
There was no rebase, squash, cherry-pick, force-push, or history rewrite.
The original upstream merge remains 811de06, with parents aed567f and kunchenguid/firstmate@4a9979a.

Refresh conflict and contract reasoning

The refresh produced one textual conflict, in bin/fm-bootstrap.sh's sourced libraries.
The upstream-sync side requires fm-remote-readiness-lib.sh for the remote second-mate readiness contract, while the newer fork side requires fm-timeout-lib.sh for bounded usage refreshes.
Those intents are compatible and independent, so both sources were retained.

AGENTS.md, .agents/skills/bootstrap-diagnostics/SKILL.md, .agents/skills/harness-adapters/SKILL.md, and .agents/skills/stow/SKILL.md merged without textual conflict.
The newer fork contract is preserved structurally as the second parent of e3bfe14, while the original seven-change contract decisions remain in the first parent.
The cursor and agy crew-only, Herdr-only adapter contract remains present in AGENTS.md, harness-adapters, and bin/fm-spawn.sh.

bin/fm-classify-lib.sh remains the single definition of a declared wait through status_is_paused_or_captain_held.
bin/fm-fleet-snapshot.sh calls that predicate and carries no second token list, preserving the contract established by HelloWorldSungin/firstmate#72.

All five parked branches remain excluded and were rechecked as non-ancestors of the refreshed head: fm/fm-afk-injection-wedge, fm/fm-crew-state-blind-during-fix-round, fm/fm-parked-decision-stale-noise, fm/fm-subagent-model-routing-guard, and fm/fm-vault-drift-check.

Refresh validation

The first full local run exposed the two Nix-host-only upstream test defects documented earlier in this body.
They are now fixed rather than carried as baseline failures:

  • tests/fm-on.test.sh now reconstructs the implementation's documented symlink-resolved Nix directories and accepts a verified harness already installed in the fixed system tail.
  • tests/fm-remote-doctor.test.sh now isolates its synthetic account from the runner's per-user Nix profile and accepts an ambient verified system harness as legitimately satisfying readiness instead of demanding an unnecessary wrapper.

Final local evidence on the tree introduced by 71f1f72 and retained byte-identically at pushed head e8021ba6299ca3fcc222b9cfad89788978b6ad06:

  • bin/fm-test-run.sh --all - total=146 failed=0 skipped_gate=18.
  • bin/fm-lint.sh - clean with pinned ShellCheck 0.11.0.
  • bin/fm-doc-audience-check.sh - ok surfaces=88 local_links=344.
  • git diff --check - clean.
  • git diff --exit-code 71f1f72 e8021ba - clean, proving the check-refresh commit did not alter the validated tree.

The only expected failing checks are the two reported instances of the fork's deliberate direct-PR compliance gate.
Every behavior, lint, documentation, coverage, and repository-invariant check must pass on the pushed head above.
This PR still requires a true merge commit and must not be squashed.

kunchenguid and others added 8 commits August 3, 2026 17:10
* feat(bin): widen the remote runtime PATH and add a remote doctor preflight

The fixed remote entrypoint hard-coded a four-directory PATH, so a remote
account whose tools live under nix or a per-user profile could not run basic
Firstmate work without a login shell. The entrypoint now composes its child
PATH from the code root's bin, the account's ~/.local/bin, the common
package-manager directories that actually exist on the host, and the portable
system tail, deduplicated and in a fixed order, still under env -i with the
same variable allowlist and no shell command string.

fm-remote-doctor.sh reports that exact PATH by inheriting it from its own
entrypoint launch rather than recomposing it, so the ordering keeps one owner.
It is read-only, reports where each required and optional tool resolved, and
exits non-zero naming every required tool that did not. Remote seeding runs it
as a preflight before anything is created on the host and restores the registry
when it fails.

* no-mistakes(review): Harden remote git authorization and missing-tool diagnostics

* no-mistakes(document): Document remote PATH doctor and safe shims

* no-mistakes(lint): Fix ShellCheck findings in remote path tests

* no-mistakes(lint): Suppress exported fixture's false-positive ShellCheck warning
* feat(bin): gate remote second mates on herdr readiness

A remote second mate now always runs on the Herdr backend, whose server
belongs to the host's GUI login session and therefore outlives the SSH
connections that supervise it. fm-spawn's remote route forces that backend
and the host-local control script refuses any other, so the requirement
cannot be dropped from either side.

fm-remote-doctor.sh becomes the single owner of what "ready" means. It keeps
its PATH and tool reporting from kunchenguid#1623 and adds the Herdr, Aqua LaunchAgent,
GUI-session, server-reachability, and entrypoint-symlink checks, tagging each
gap fixable: or human: with the exact operator step. --fix closes only the
automatable gaps - writing and loading the Aqua-scoped dev.firstmate.herdr
launch agent, starting the server where no launch agent applies, and
recreating the entrypoint symlink - then re-derives every check from the host,
so a human gap is never presented as fixed. It never creates a login session,
writes an auto-login password, or touches FileVault.

Remote seed, remote spawn, and the startup liveness relaunch all run the same
check, repair, re-check sequence through one shared library and fail closed
with the doctor's own gap text. Recovery inherits the gate because it respawns
through the same route.

Tests drive the real doctor against a controlled account fixture with a
private HOME, a state-backed launchctl, and a fake herdr, and prove the
dangerous actions are never attempted. The remote lifecycle suites gain a
stateful Herdr CLI fixture and answer the readiness gate at the SSH boundary,
so they never inspect or repair the runner's own account.

* no-mistakes(review): Validate launch-agent contract and confirm Herdr startup

* no-mistakes(review): Validate loaded launch-agent contract before readiness

* no-mistakes(review): Refuse legacy remote backends without altering routes

* no-mistakes(review): Clarify conditional remote readiness repair sequence

* no-mistakes(review): Repair remote readiness before liveness probing

* no-mistakes(review): Preserve unknown seeds and reject legacy liveness

* no-mistakes(document): docs: clarify remote Herdr backend ownership
…1659)

* Pin remote secondmates to fm-remote

* no-mistakes(review): Fail closed on legacy remote Herdr endpoints

* no-mistakes(review): Isolate fm-remote launch agent from interactive default

* no-mistakes(document): Document shared remote Herdr retirement safety
)

* feat: run remote commands through Aqua job worker

* no-mistakes(review): Enforce remote job deadlines and safe worker shutdown

* no-mistakes(review): Refresh stale workers and harden dependency-free supervision

* no-mistakes(review): Harden worker ownership recovery and shutdown quarantine

* no-mistakes(review): Fix doctor bootstrap, harness repair, and output draining

* no-mistakes(review): Probe doctor tools through authenticated worker bootstrap

* no-mistakes(review): Refresh stale workers before doctor tool probes

* no-mistakes(review): Recover stopped quarantines and extend job deadlines

* no-mistakes(review): Separate queue and execution timeout windows

* no-mistakes(review): Supervise Linux worker crashes and bind root identity

* no-mistakes(review): Resolve authorized Nix profile bin links

* no-mistakes(review): Clarify Nix path resolution documentation

* no-mistakes(review): Harden PATH safety and nvm selection

* no-mistakes(review): Honor nvm system defaults and refresh doctor digest

* no-mistakes(review): Keep workers ready during active jobs

* no-mistakes(review): Bound pre-execution validation by job timeout

* no-mistakes(document): Clarify remote worker documentation

* no-mistakes(lint): Fix remote worker ShellCheck diagnostics

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(remote): arm SSH dead-peer detection in fm-on.sh

A vanished remote host mid-poll (a reboot, a dropped link) left ssh
blocked indefinitely on a half-open TCP connection, because fm-on.sh's
ssh invocation had no ServerAliveInterval/ServerAliveCountMax. This
wedged the remote-reply ferry: fm-procevent.sh's runner blocked inside
the ssh child and never reached its own no-result -> claim-release ->
reconcile re-arm self-healing path, which otherwise already handles a
nonzero exit with empty output correctly. Recovery required a manual
retire and re-arm.

Arm ServerAliveInterval=15 and ServerAliveCountMax=3 by default
(bounded ~45s detection window), both overridable via
FM_SSH_ALIVE_INTERVAL and FM_SSH_ALIVE_COUNT_MAX. This is a transport-
level fix in fm-on.sh, so it covers every remote command routed
through it, not just the reply ferry. The remote sshd answers
keepalive probes independently of whatever the remote command is
doing, so a legitimately long-but-alive command (a 55s poll, a clone,
the doctor) is never falsely killed - only a truly vanished peer trips
it, turning that case into a bounded, detectable ssh failure (exit
255) instead of an indefinite hang.

Extends tests/fm-on.test.sh with a behavioral regression asserting a
bounded, positive ServerAliveInterval/ServerAliveCountMax on the real
ssh argv captured through the FM_SSH_BIN process seam, plus coverage
that both are env-overridable.

* no-mistakes(document): Document SSH dead-peer detection ownership
* feat(bootstrap): gate stale axi CLIs at the floors firstmate actually uses

Add gh-axi 0.1.29 floor so bare --squash PR merges stop failing quietly on
older builds. Raise tasks-axi to FM_TASKS_AXI_MIN=0.2.2 (multi-id mv) while
keeping feature probes. Keep quota-axi at 0.1.16 after verifying schema 3
and per-model availability already ship there; runway remains optional.

* no-mistakes(document): Clarify AXI compatibility documentation ownership
…hup-round-1

# Conflicts:
#	AGENTS.md
#	bin/fm-remote-doctor.sh
#	bin/fm-remote-entrypoint.sh
#	bin/fm-remote-home-seed.sh
#	docs/remote-secondmates.md
#	docs/scripts.md
#	tests/fm-on.test.sh
#	tests/fm-remote-secondmate-lifecycle-e2e.test.sh
@cursor

cursor Bot commented Aug 4, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

Sungin Kim added 3 commits August 4, 2026 23:54
Reverted in the next commit; the branch head tree returns to 811de06 exactly.
Emits process liveness, the background push output, inherit markers, generation
records, fake-ssh history and tool availability at the point the wait exhausts,
so the CI-only failure can be diagnosed by mechanism rather than by tree.
Restores tests/fm-remote-secondmate-lifecycle-e2e.test.sh exactly; the branch
head tree is now content-identical to the merge commit 811de06.
…sed time

The concurrency fixtures waited a fixed number of polls for a deliberately
blocked write on data/captain-shared.md. That silently assumed how many files
precede it: config/ items propagate before the one shared data file, so every
entry added to FM_INHERITABLE_CONFIG ate into a constant allowance.

Adding gbrain.json and project-board pushed the config-push fixture past its
five-second budget on a loaded runner while it still passed on a fast host,
which is why the failure appeared only in CI. The spawn fixture had already met
this and papered over it by raising its own budget to thirty seconds, leaving
the same trap for the next inheritable file.

Both waits now share await_blocked_inherit_write, which treats a new remote call
as the liveness signal and resets its patience window on every one. Declaring
another inheritable file can lengthen the wait but can no longer fail it. The
remaining window expires only when the transfer stops making calls at all, which
is a genuine wedge rather than a slow runner, so the fixture is deterministic
with respect to list length instead of merely roomier.
@HelloWorldSungin
HelloWorldSungin merged commit 6eb5f1a into main Aug 10, 2026
12 of 14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants