Skip to content

feat(bin): merge upstream firstmate and add stow cascade, orphan reap - #7

Merged
pruge merged 7 commits into
mainfrom
fm/firstmate-upstream-merge
Aug 8, 2026
Merged

pruge merged 7 commits into
mainfrom
fm/firstmate-upstream-merge

Conversation

@pruge

@pruge pruge commented Aug 8, 2026

Copy link
Copy Markdown
Owner

Intent

Merge the upstream firstmate repository (kunchenguid/firstmate, 5 commits ahead) into this fork (pruge/firstmate, which had accrued 6 fork-only commits by the time of execution, not the 3 originally estimated) and resolve the conflicts, so this home stops drifting from the tool it is built on. Requirements: use a real merge (git merge --no-ff), never rebase or force-push, because rebasing would force a force-push that strands every live worktree based on the old main - that is destroying unlanded work. All fork-only commits must survive intact and reachable. Four files were flagged as changed on both sides needing careful combination rather than taking one side wholesale: AGENTS.md, bin/fm-teardown.sh (the fork's .codegraph/ teardown exemption must survive alongside whatever upstream changed there), .agents/skills/harness-adapters/SKILL.md, and docs/architecture.md; docs/scripts.md also touched by both sides. In the event, git's three-way merge auto-resolved all five with no conflict markers because each side's edits landed on non-overlapping lines/sections - verified this manually by diffing each side against the merge-base and confirming both sides' unique content survived in the merged file for every one of the five. Must prove the merge did not break the fleet: confirmed bin/fm-session-start.sh produces its digest with no new diagnostics (verified byte-identical against a pre-merge baseline comparison in an isolated detached-HEAD worktree), bin/fm-spawn.sh is syntax-clean and covered by its passing dedicated test suite, bin/fm-lint.sh passes, and the full existing test suite passes overall - 125 of 133 test files pass outright, and the other 8 were individually reproduced as failing identically against the pre-merge fork head in an isolated worktree, proving they are pre-existing environment-specific failures (e.g. this machine's default bash 3.2.57 nounset-on-empty-array bug, a timing-sensitive real-Herdr concurrency test, an Orca metadata-write-failure test) and not regressions introduced by this merge; none of them touch the four flagged conflict files' own passing assertions, notably fm-teardown.test.sh's 'untracked CodeGraph index does not block teardown' passes both before and after. Out of scope / explicitly excluded by the requester: do not push to upstream (no write access, and that decision belongs to the captain, not this task); do not change /updatefirstmate or its skill (a separate task owns making it upstream-aware); do not make unrelated improvements - this is an integration, not a cleanup. This is a 34-file integration touching bootstrap, harness detection, and lint, so the independent review and test steps should run for real - do not skip the test step. The branch carries a real merge commit (f068e50) preserving both parent histories; if this pipeline's rebase step would flatten, drop, or conflict with that merge structure such that either side's history would be lost, stop and report rather than forcing it through - a clean rebase-and-continue is fine, but the merge structure and both histories must survive.

What Changed

  • Merged upstream firstmate (kunchenguid/firstmate, 5 commits ahead) into this fork via a real git merge --no-ff (f068e50), preserving both the 6 fork-only commits and the full upstream history; the three-way merge auto-resolved AGENTS.md, bin/fm-teardown.sh, .agents/skills/harness-adapters/SKILL.md, docs/architecture.md, and docs/scripts.md with no conflict markers, combining both sides' non-overlapping edits.
  • Brought in upstream capabilities: bin/fm-lint.sh now lints only the changed shard locally while CI still runs the full lint; bin/fm-remote-job-worker.sh and bin/fm-remote-job-lib.sh gained orphaned-worker reaping backed by new bin/fm-remote-job-reap-orphans.sh; new bin/fm-stow-cascade.sh cascades the internal /stow to every registered secondmate; and bin/fm-startup-network.sh plus new bin/fm-timing-lib.sh record per-step elapsed times for the deferred startup stage.
  • Updated test coverage (fm-bootstrap, fm-lint, fm-remote-job-orphan-reap, fm-remote-backlog-handoff, fm-remote-job, fm-remote-reply, fm-startup-network, fm-stow-cascade, fm-turnend-guard) and docs (architecture.md, configuration.md, remote-secondmates.md, scripts.md, subagent-guard.md, turnend-guard.md, verification/supervision.md), including a follow-up doc fix adding the missing fm-stow-cascade.sh entry to the scripts table.

Risk Assessment

✅ Low: This is a real --no-ff merge (verified two parents) preserving both histories — all 6 fork-only commits and all 5 upstream commits are reachable from the merge commit; no conflict markers exist anywhere in the tree; each of the five flagged files (AGENTS.md, bin/fm-teardown.sh, .agents/skills/harness-adapters/SKILL.md, docs/architecture.md, docs/scripts.md) shows a clean additive combination of both sides with no content lost, and the fork's .codegraph/ teardown exemption is confirmed still present; the remaining ~29 changed files are pure upstream additions the fork never touched (verified via diff against the merge-base), so they carry no merge-conflict-resolution risk; a spot security/correctness sweep of the largest new upstream files (fm-remote-job-reap-orphans.sh, fm-stow-cascade.sh, fm-lint.sh, fm-bootstrap.sh) found careful process-safety guards (self/ancestor/pgid checks, re-read-before-signal to avoid pid-recycle races, mktemp-based cleanup) and nothing alarming.

Testing

Manually verified via git plumbing that this is a genuine --no-ff merge preserving both parent histories intact, that all five explicitly flagged dual-side files merged cleanly with both sides' unique content present and zero leftover conflict markers, and that the fork's .codegraph/ teardown exemption survived alongside its dedicated regression test; fm-spawn.sh, fm-teardown.sh, fm-session-start.sh, and fm-lint.sh are all syntax-clean. I also started the project's own targeted fm-test-run.sh --changed suite (86 tests) in the background; the portion observed before this phase's time budget ended showed no failures, but the run did not reach its final summary, so I cannot personally confirm the author's claimed full-suite result — flagged as a missing-evidence warning.

  • Outcome: ⚠️ 1 info across 1 run (13m32s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 info
  • ℹ️ bin/fm-test-run.sh - The repo's own fm-test-run.sh --changed --base 9e889e2 selection maps this merge's 34 changed files to 86 related test files (a family-based targeted set, well short of the full 133-file suite). This run was started in the background but had not reached FM_TEST_SUMMARY when this phase's time budget ended; the portion observed (first ~150 assertions, still inside the first test file, a large pretool-check permission matrix) showed zero failures (not ok/FAIL count = 0). I could not personally reproduce the author's claimed full-suite result (125/133 passing, 8 pre-existing failures reproduced identically pre-merge). This is a gap in automated evidence for this run, not an observed defect.
  • git show -s --format='%P' f068e50 — confirmed the merge commit has two parents (9e889e2 fork tip, 833a9a2 upstream tip), a real git merge --no-ff, not a rebase
  • git merge-base 9e889e2 f068e50 — confirmed base is the pre-merge fork head, i.e. both parent histories are reachable and intact (no history loss)
  • git log --oneline --graph — visually confirmed both the 6 fork-only commits and the 5 upstream commits are present as distinct, reachable commits off the shared merge base 70aeba8
  • diffed each of AGENTS.md, bin/fm-teardown.sh, .agents/skills/harness-adapters/SKILL.md, docs/architecture.md, docs/scripts.md against merge-base vs. fork-tip and vs. upstream-tip, then confirmed every added line from both sides is present verbatim in the merged file at f068e50
  • grep for conflict markers across the five flagged files — confirmed zero leftover markers
  • inspected bin/fm-teardown.sh's .codegraph/ untracked-artifact exemption and confirmed tests/fm-teardown.test.sh still contains the dedicated regression test for it
  • bash --version — confirmed this machine runs GNU bash 3.2.57(1), matching the intent's cited pre-existing failure environment
  • bash -n on bin/fm-spawn.sh, bin/fm-teardown.sh, bin/fm-session-start.sh, bin/fm-lint.sh — all syntax-clean
  • started timeout 1800 ./bin/fm-test-run.sh --changed --base 9e889e2 in the background (86-file targeted family selection); observed ~150 passing assertions with zero failures before the phase's time budget required stopping it
  • git status --porcelain post-run — working tree clean, no stray test artifacts left behind
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

jayjongcheolpark and others added 7 commits August 7, 2026 14:52
…917)

* fix(hooks): keep tracked Claude entries inert under grok 1.0.0 hooks

Grok loads Claude-compatible settings, so the tracked `.claude/settings.json`
hook entries also fire under Grok. They were meant to be inert there, guarded
by `[ -z "${GROK_AGENT:-}" ] || exit 0`. That guard silently stopped working.

Verified from the live process environment of a wedged grok 1.0.0 Stop hook on
2026-08-07: a grok 1.0.0 HOOK process carries GROK_HOOK_EVENT, GROK_HOOK_NAME,
GROK_SESSION_ID, and GROK_WORKSPACE_ROOT, but no GROK_AGENT. The observed hook
process was labelled `GROK_HOOK_NAME=project/settings:stop[0].hooks[1]`, which
is the Claude-only auto-arm entry.

Consequence: Grok ran `bin/fm-claude-stop-autoarm.sh` synchronously. Grok has
no `asyncRewake`, so it waited on the foregrounded watcher for that entry's
declared 28800-second timeout and the Grok turn never ended - the operator saw
an infinite "Responding".

Widen the guard to `[ -z "${GROK_AGENT:-}${GROK_HOOK_EVENT:-}" ] || exit 0` on
the five entries that have a `.grok/hooks/` counterpart: both Stop entries, the
SessionStart entry, and the two PreToolUse Bash entries.

Two deliberate limits:

- The guard is NOT widened to GROK_SESSION_ID. Grok injects it into every child
  process, so it can survive into a Claude session that Grok launched and would
  silently disable Claude's own watcher continuity. GROK_HOOK_EVENT is
  per-hook-invocation and does not leak that way.
- `bin/fm-subagent-pretool-check.sh` stays unguarded on purpose. It is the one
  tracked entry with no `.grok/hooks/` counterpart, so guarding it would remove
  the guard from Grok entirely rather than deduplicate it. The new test asserts
  it stays unguarded so the exception cannot be closed silently, and
  docs/subagent-guard.md is honest that the coverage it leaves is partial.

`bin/fm-harness.sh` corrects a comment that presented GROK_AGENT as reliably
present; it is a fast path only, and the ancestry walk is what actually
guarantees grok identification.

tests/fm-turnend-guard.test.sh adds test_tracked_claude_entries_inert_under_grok,
which runs every tracked entry under a real grok 1.0.0 hook environment, a
legacy GROK_AGENT environment, and a native Claude environment.

* no-mistakes(document): docs: sync grok hook-marker guard facts to owners

* no-mistakes(review): docs: state grok guard criterion by event coverage
… stage (#1918)

The deferred network stage published one aggregate started/finished pair, so
a run that took a minute could not be attributed to a phase, a host, or a
clone without re-running it by hand under manual tracing.

Add bin/fm-timing-lib.sh as the single owner of elapsed-time records, and
bracket each network owner with one: the gh auth probe, the secondmate
liveness sweep, secondmate convergence, pending handoff delivery, and the
project clone refresh, plus one record per secondmate for the remote-touching
steps (id and host) and one per project clone. Each record carries a start
offset from one shared origin, so the artifact reads as a timeline.

The stage publishes them beside its report as state/.startup-network.timings,
for a timed-out or failed run too, where the partial record is the answer.
Only the on-demand `report` command prints them: `harvest` composes the
session-start digest, so its output, the wake cadence, and every other part
of a normal session start are unchanged.

Recording is inert unless a run asks for it, so nothing else that sources
these scripts pays for it. Details are identities only - a detail carrying
whitespace is refused rather than cleaned up, which is what keeps a command
line, an environment dump, or a captured error out of the file.

Split two per-item loop bodies into their own functions so each iteration can
be timed; every `continue` became a `return 0` with the same meaning, and the
sweeps still run directly, in the same order, returning the same results.
… (#1928)

* feat(stow): cascade the internal /stow to every registered secondmate

Invoked in a primary home, /stow now sweeps every registered secondmate
after the primary's own required pass, enforcing the same startup-memory
threshold in each home against that home's own allowance rather than a
fleet total.

bin/fm-stow-cascade.sh owns the mechanical inputs: it enumerates each
registered secondmate exactly once from data/secondmates.md, reports that
home's own budget accounting, and resolves how the sweep reaches it. A
live agent sweeps its own home so its uncaptured session knowledge is
captured too; a local home without one is curated in place; a remote home
without one is accounted read-only and deferred, because there is no
generic remote write path for a home's own memory files. Every host-
crossing step and each home's accounting runs under one hard bound, so a
slow or unreachable home reports an exception and the sweep continues.

Nothing changes until /stow is invoked: no new notification, digest
section, or background work. The public skills/stow skill is untouched.

* no-mistakes(review): fix(stow): extend cascade --help range to include full exit-code contract
29 fm-remote-job-worker.sh processes were found running at ppid 1, 1-2 days
old, each still polling and appending to a log inside a no-mistakes gate
worktree that had already been returned.

Three things combined to make that possible:

- The recorded worker.pid is the serving child, not the restart supervisor
  above it, so a teardown that stops that one pid only makes the supervisor
  respawn. The Linux start path also left the worker tree in the launching
  command's process group, so there was no group to signal instead.
- Neither the serving loop nor the supervisor ever rechecked whether its
  configured FM_ROOT still existed, so a worker launched from a worktree
  outlived that worktree indefinitely.
- The supervisor restarted a failing child with a fixed 0.1s delay and no
  bound, which is what grew the logs (~66MB/day measured).

The Linux start path now puts the worker tree in its own process group, and
fm_remote_job_stop_worker_tree signals that whole group - refusing any group
whose leader is not itself a worker, so a worker from an older build or from
launchd's own session is still stopped safely as a single process. The worker
stops itself once its code root stops being a Firstmate checkout, confirmed
across a grace window so an ordinary transient cannot stop a healthy worker.
The supervisor backs off and gives up rather than restarting forever.

bin/fm-remote-job-reap-orphans.sh is the belt-and-suspenders sweep for workers
already orphaned that way, wired into fm-teardown.sh. Its reap condition is
exactly "the code root named in the worker's own command line is gone", which
is why the account's healthy LaunchAgent worker and every live remote
secondmate worker are never candidates.

The two suites that leaked these in the first place now stop the worker tree
rather than the recorded pid alone.
* fix(bin): lint only the changed shard locally, full lint in CI

Two ships hitting fm-lint.sh at once could spike CPU to 190% and load
to 8.58 on a captain's Mac, even though each run finishes quickly.
fm-lint.sh now defaults to linting only the canonical-set files
changed since the merge-base with origin/main (including uncommitted
edits) on an ordinary local branch, using plain local git with no
network calls. It still lints the full canonical set in CI
(GITHUB_ACTIONS=true or CI=true), on the main branch, or whenever no
merge-base can be found, so CI coverage never depends on a local diff.
Explicit paths keep bypassing this selection entirely.

* no-mistakes: apply CI fixes
@pruge

pruge commented Aug 8, 2026

Copy link
Copy Markdown
Owner Author

빨간 검사 하나를 남긴 채 병합함 (캡틴 승인 2026-08-08)

실패: Behavior portable serial 4 -> not ok - the serving child ignored TERM
(tests/fm-remote-job-orphan-reap.test.sh)

왜 넘겼나

  • 이 테스트는 이 합치기가 원본에서 가져온 be32879 fix(remote-job): stop workers abandoned by a pruned code root (#1927) 가 함께 들여온 것이다. 이 포크가 깨뜨린 것도, 합치기가 잘못 붙인 것도 아니다.
  • 환경을 탄다: macOS 로컬에서는 통과, 리눅스 CI 에서는 재실행까지 두 번 모두 실패. 테스트가 fm_remote_job_start_linux_worker 경로를 탄다.
  • 이 홈은 원격 부관 경로가 0개다(data/secondmates.md 에 host: 라우트 없음). 실패하는 기능을 쓰지 않는다.
  • 나머지 12건은 전부 통과. 합치기 자체는 실제 충돌 0건으로 git 이 자동 병합했고, 양쪽 의도를 수동 확인했으며, 세션 시작 진단이 합치기 전과 동일하다.

🔴 나중에 원격 부관을 쓰게 되면 이 실패부터 확인할 것. 그때는 더 이상 무해하지 않다.

@pruge
pruge merged commit b34313d into main Aug 8, 2026
23 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants