Skip to content

fix(bin): sync upstream supervision, pending-reply, and herdr lock fixes - #23

Merged
RajeshRajendiran merged 12 commits into
mainfrom
fm/fm-upstream-sync-1008
Oct 8, 2026
Merged

RajeshRajendiran merged 12 commits into
mainfrom
fm/fm-upstream-sync-1008

Conversation

@RajeshRajendiran

@RajeshRajendiran RajeshRajendiran commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Intent

go sync

Context: the scheduled sync that merges upstream/main (kunchenguid/firstmate) into the fork's main (RajeshRajendiran/firstmate) stopped with a merge conflict. Fork main is at efb0243 (after merging the fork's pool-reset-on-cleanup PR 22). Upstream has 9 new commits: b062eb9 (kunchenguid#6818), 0ec1c5a (kunchenguid#6823), 3c58ec8 (kunchenguid#6787 supervision host successor watcher), 329ad4e (kunchenguid#6814), 2ce57d0 (kunchenguid#6784 index six missing agent-only skill triggers), 9b85a95 (kunchenguid#6792), 0c83b3f (kunchenguid#6809), 8fa2538 (kunchenguid#6780), 228b27d (kunchenguid#6785). A trial merge conflicts only in .agents/skills/agent-skill-trigger-index/SKILL.md.

What Changed

🤖 Generated with Claude Code

Risk Assessment

✅ Low: This is a merge of the 9 upstream commits listed in the intent. The only conflict was in .agents/skills/agent-skill-trigger-index/SKILL.md, and it was resolved as a clean union: all six upstream entries are present, the fork-only telegram-captain-channel entry is kept, and every indexed skill directory exists. The other files merged automatically, and the new upstream teardown tests use helpers that already exist and are invoked, with no duplicate function names.

Testing

I re-ran the opt-in live attended supervision-host E2E against the committed HEAD with a real Claude primary in an isolated lab, and it passed. The primary was woken for all four hand-offs. The event 1 turn end took over successor 3682351, the watcher it owned (3688913) delivered event 2's close, the stand-in remote listener stayed owned the whole time, and no captain prompts were submitted. I also checked the conflicted skill trigger index structurally: it has no conflict markers, no duplicate entries, and every indexed skill exists. That check was not a live run, so its scenario is recorded as untested. Earlier rounds' targeted suite logs are already in the evidence directory. The test's pre-fix negative control was skipped because no control ref was set.

  • Live validation: ✅ go - 4 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Live: an idle attended Claude primary under the supervision host is woken for a main-only event 1 by the Stop-hook rewake and drains it ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 2: rewake delivered, ledger outcome=rewake, primary ran fm-wake-drain --ack-through 4
Live: the event 1 turn end takes over the pass-through's successor watcher, and the watcher it owns delivers the event 2 close to the primary ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 3 and 3/4: took over successor 3682351, owned watcher 3688913 closed with origin=started reason=actionable-signal, rewake delivered
Live: the stand-in remote-reply listener stays owned and delivers event 3 ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 4: listener runner 3679325 owned at every check, 1 ingested and delivered
Live: a routine close that turns main-only at its turn (event 4 on demo2) is delivered with no engine turn and no captain prompt ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 5 and the final line 'captain prompts after setup: 0'
The merged skill trigger index keeps both upstream's six new triggers and the fork's telegram trigger, with no conflict leftovers or dangling skill names ⏸️ untested no The prior payload did not establish a live result: the index is agent-facing instruction text that was only checked structurally against the skills on disk (skill-trigger-index-merge-check.log), not d…
Negative control: on a pre-fix ref the host fails to wake the idle primary ⏸️ untested no It needs FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF set to a pre-fix git ref. It is not required for this sync, which only adopts upstream behavior. Set that variable to run it.
Evidence: Live attended supervision-host E2E transcript at HEAD 8939c96

Source: Live attended supervision-host E2E transcript at HEAD 8939c96

skip: control: set FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF to a pre-fix ref to run the negative control
# 12:22:49 positive step 1: primary idle (claude pid 3674551, 2.1.294 (Claude Code), model haiku); tracked Stop hook registered; config/supervision-host present; host pid 3678631 parked on watcher 3678861; listener runner 3679325; captain prompts so far: 1
# 12:22:49 positive transcript: ~/.claude/projects/-tmp-fm-sh-attended-live-zub0Qp-positive-fm/a8e8e2b1-acc0-4754-9193-879a4fcb48f6.jsonl
# 12:22:49 positive step 2: event 1 appended at 1791480169 (demo.status needs-decision)
# 12:22:51 positive step 2: host log: 1791480171	pass-through	attended	main-only
# 12:22:52 positive step 2: successor watcher pid 3682351 alive; ledger: epoch=1 owner_pid=3678342 outcome=arming updated_at=1791480167; marker: pending:downtime:3678861.1791480170.jY5qKb
# 12:22:57 positive step 2: rewake delivered at 1791480172 (Stop hook exited 2 with the banner); ledger: epoch=1 owner_pid=3678342 outcome=rewake updated_at=1791480172 session_pid=3674551 recovery_generation=3678861.1791480170.jY5qKb
# 12:22:57 positive step 2: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 3678861.1791480170.jY5qKb;
# 12:23:00 positive step 3: turn end re-armed: 1791480179	start	gen=host-3687203-1791480179; took over successor 3682351, now owning watcher 3688913
# 12:23:03 positive step 3: event 2 appended at 1791480183
# 12:23:12 positive step 3/4: owned watcher 3688913 closed: arm_pid=3687920 watcher_pid=3688913 origin=started started_at=1791480180 ended_at=1791480185 exit_code=0 signal=none reason=actionable-signal
# 12:23:12 positive step 3/4: its close was delivered: rewake at 1791480187; host log: 1791480186	pass-through	attended	main-only
# 12:23:17 positive step 4: event 3 appended to the stand-in remote log at 1791480197
# 12:23:28 positive step 4: listener mirrored it (1 ingested) and it was delivered at 1791480203; listener runner 3679325 -> 3679325, owned at every check
# 12:23:32 positive step 5: event 4, a routine working line on task demo2, appended at 1791480212
# 12:23:36 positive step 5: the host accepted it and confirmed its successor's handling handoff (marker announced:handling:3710488.1791480214.UoarfQ at 1791480216); a needs-decision on demo2 landed then, before the turn's start
# 12:23:37 positive step 5: host log: 1791480217	pass-through	attended	main-only; no engine turn
# 12:23:37 positive step 5: rewake delivered at 1791480217; ledger: epoch=4 owner_pid=3709261 outcome=rewake updated_at=1791480217 session_pid=3674551 recovery_generation=3710488.1791480214.UoarfQ
# 12:23:42 positive step 5: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 12 --recovery-generation 3710488.1791480214.UoarfQ;
# 12:23:44 positive step 5: turn end re-armed: 1791480224	start	gen=host-3719443-1791480223; watcher 3719946 live; listener runner 3679325 still owned
# 12:23:44 positive: captain prompts after setup: 0 (mirror holds only the setup prompt)
ok - attended live (2.1.294 (Claude Code)): an idle primary is woken for four hand-offs, a take-over of the pass-through's successor and a close that turned main-only at its turn included, with the listener owned throughout
Evidence: Skill trigger index merge check

Source: Skill trigger index merge check

conflict markers: 0
duplicate entries: 0
ok operational-home-layout
ok session-start-recovery
ok telegram-captain-channel
ok bootstrap-diagnostics
ok diagnostic-reasoning
ok ask-user-authority
ok validation-supervision
ok ship-landing
ok scout-completion
ok quota-array-dispatch
ok harness-adapters
ok firstmate-orca
ok project-management
ok stuck-crewmate-recovery
ok secondmate-provisioning
ok captain-hold-lifecycle
ok away-quiet-supervision
ok process-event-sources
ok fmx-respond
ok firstmate-codexapp
ok firstmate-coding-guidelines
- Outcome: 🔧 1 issue found → auto-fixed → no changes applied ✅ across 3 runs (52m39s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

🔧 **Test** - 1 issue found → auto-fixed → no changes applied ✅
  • ⚠️ The Test agent did not finish within its invocation budget. Reported: agent run tests timed out after 30m0s: agent last produced output 485ms ago (91 observed); agent reported: claude parse events: context deadline exceeded. This is a budget or provider-slowness cut, not a code failure. Re-running the same request costs another full budget, so no further attempt is made automatically. If this repository's targeted tests or evidence gathering routinely approach the default 30m0s, raise test_agent_timeout in global config. Respond with fix to spend another budget: a repair turn runs only for selected findings other than this budget cut, then validation re-runs. Or abort and retry after raising the budget.

🔧 No changes applied.
4 issues (1 error, 3 warnings) still open:

  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh - The live supervision-host attended E2E fails on event 2 because the arm following the successor watcher reports taken-over and never delivers its close. It fails identically on pure upstream b062eb9, so the sync didn't introduce it. It is either an upstream fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down kunchenguid/firstmate#6787 gap or specific to this host, and should be raised upstream.
  • ⚠️ tests/fm-procevent.test.sh:3863 - tests/fm-procevent.test.sh:3863 assumes an orphaned listener reparents to pid 1. On hosts with a systemd --user child subreaper (here pid 1059) it fails with 'was not reparented away from its session'. This predates the merge and isn't touched by it. The test should accept any subreaper that isn't the session.
  • ⚠️ tests/fm-procevent.test.sh:2319 - tests/fm-procevent.test.sh:2319 'a repaired source did not confirm' (reconciled started=0) failed once when run in parallel with other suites and passed when the suite ran alone. This looks like a timing flake that predates the merge.
  • 🚨 live validation verdict: no-go (2 of 7 scenarios were driven live against the product); failed: Live: an attended Claude primary under the supervision host receives the successor-armed close on event 2
  • Live validation: ❌ no-go - 2 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Merged skill trigger index keeps the fork's telegram-captain-channel entry and gains all six upstream kunchenguid#6784 entries ⏸️ untested no The prior payload did not establish a live result. This was checked only by diffing the file against both merge parents, and the file has no runtime consumer to drive live.
Fork pool-reset teardown (PR 22) plus upstream teardown tests still pass together after the merge ⏸️ untested no The prior payload did not establish a live result. It was exercised only through the repo test suite (fm-teardown.log exit=0) against fixtures, not the live product.
Upstream supervision-host successor watcher fix (kunchenguid#6787) works in the merged tree ⏸️ untested no The prior payload did not establish a live result. It was exercised only through the repo test suite (fm-supervision-host.log exit=0). The live attended check is a separate scenario.
Live: an attended Claude primary under the supervision host receives the successor-armed close on event 2 ❌ fail live supervision-host-attended-live-e2e.log and the -upstream-b062eb9.log fail identically (arm taken-over, close not delivered)
Pending reply resolves only from its own task's status line (kunchenguid#6792); watch-arm, OpenCode plugin predicate, test fixtures and Herdr backend lock namespace behave as upstream intends ⏸️ untested no The prior payload did not establish a live result. These were exercised only through the repo test suites (fm-pending-reply, fm-watch-arm, fm-pi-watch-extension, fm-test-fixtures, fm-backend-herdr exi…
Live: real herdr control exit/relaunch behavior in a throwaway lab session ✅ pass live herdr-control-smoke-live.log: 10 ok, no not ok
Lavish read labels its message count with the same label as its section (kunchenguid#6785) ⏸️ untested no The prior payload did not establish a live result. It was exercised only through tests/fm-procevent.test.sh, where the Lavish label assertions passed and unrelated environment and timing assertions fa…
  • git diff 12513f8^2 12513f8 -- .agents/skills/agent-skill-trigger-index/SKILL.md (conflict resolution vs upstream)
  • git diff efb0243 12513f8 -- .agents/skills/agent-skill-trigger-index/SKILL.md (conflict resolution vs fork main)
  • index integrity: every - skill`` entry in the trigger index has an existing .agents/skills/<name>/SKILL.md, no duplicates (21 entries)
  • bash tests/fm-supervision-host.test.sh
  • bash tests/fm-pending-reply.test.sh
  • bash tests/fm-watch-arm.test.sh
  • bash tests/fm-teardown.test.sh
  • bash tests/fm-test-fixtures.test.sh
  • bash tests/fm-backend-herdr.test.sh
  • bash tests/fm-pi-watch-extension.test.sh
  • bash tests/fm-procevent.test.sh (run twice)
  • live: tests/fm-control-herdr-smoke.test.sh against real herdr 0.9.3 (prior round, log reused)
  • live: tests/fm-supervision-host-attended-live-e2e.test.sh with real Claude on the merge and on upstream b062eb9 (prior round, logs reused)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 4 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Live: an idle attended Claude primary under the supervision host is woken for a main-only event 1 by the Stop-hook rewake and drains it ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 2: rewake delivered, ledger outcome=rewake, primary ran fm-wake-drain --ack-through 4
Live: the event 1 turn end takes over the pass-through's successor watcher, and the watcher it owns delivers the event 2 close to the primary ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 3 and 3/4: took over successor 3682351, owned watcher 3688913 closed with origin=started reason=actionable-signal, rewake delivered
Live: the stand-in remote-reply listener stays owned and delivers event 3 ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 4: listener runner 3679325 owned at every check, 1 ingested and delivered
Live: a routine close that turns main-only at its turn (event 4 on demo2) is delivered with no engine turn and no captain prompt ✅ pass live supervision-host-attended-live-e2e-head-8939c96.log step 5 and the final line 'captain prompts after setup: 0'
The merged skill trigger index keeps both upstream's six new triggers and the fork's telegram trigger, with no conflict leftovers or dangling skill names ⏸️ untested no The prior payload did not establish a live result: the index is agent-facing instruction text that was only checked structurally against the skills on disk (skill-trigger-index-merge-check.log), not d…
Negative control: on a pre-fix ref the host fails to wake the idle primary ⏸️ untested no It needs FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF set to a pre-fix git ref. It is not required for this sync, which only adopts upstream behavior. Set that variable to run it.
  • FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 tests/fm-supervision-host-attended-live-e2e.test.sh at committed HEAD 8939c96 (real Claude 2.1.294 primary, haiku, in a disposable lab home on a private tmux socket)
  • Checked the merged .agents/skills/agent-skill-trigger-index/SKILL.md: no conflict markers, no duplicate entries, and every indexed skill name resolves to an existing .agents/skills/<name>/SKILL.md
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

falkoro and others added 12 commits October 7, 2026 15:22
… section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743
…henguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.
…te (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503
…nchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.
… tests/fm-watcher-lock.test.sh ("TERM after watcher lock acquisition left the singleton lock behind"). This was an intermittent race that predates the PR (neither bin/fm-watch.sh nor the test is touched by it); it reproduced locally in 1 of 5 runs. Root cause: bin/fm-watch.sh took the .watch.lock singleton lock before installing `trap watcher_cleanup EXIT`, so a TERM landing in that gap killed the watcher without cleanup and leaked the lock. Rule: once the watcher holds the lock, every exit must run watcher_cleanup. fm_lock_try_acquire is the only place the watcher gets the lock, so one change covers every path. Fix: moved `trap watcher_cleanup EXIT` and `watcher_stop_signals` to just before the acquisition loop. This is safe on the early exits ("already running", stale heartbeat) because watcher_cleanup releases the lock only when it records WATCHER_PID and every other cleanup step does nothing when there is nothing to clean up. Verified: with a 0.5s sleep added in the gap, the test fails with the old ordering and passes with the new one; fm-watcher-lock.test.sh passed 8 runs in a row; fm-watch-arm and fm-watch-checkpoint tests pass; shellcheck -x is clean
@RajeshRajendiran
RajeshRajendiran merged commit 91d7092 into main Oct 8, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants