Repository navigation
fix(bin): sync upstream supervision, pending-reply, and herdr lock fixes - #23
Merged
Merged
Conversation
… section (kunchenguid#6785) fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE only when session_ended is true and CAPTAIN MESSAGE otherwise, but the count line always said session_ending_message_count. Several composer messages on a still-open board were therefore counted as session-ending. The count line now follows the same session_ended switch: session_ending_message_count once the session ended, captain_message_count otherwise. Message rows stay out of the annotation count, per triage. Fixes kunchenguid#6743
…henguid#6780) The Herdr presentation lock namespace was the fixed machine-global /tmp/firstmate-herdr-presentation, so on a host where two OS users run Firstmate on Herdr the first account to create it owned it and every teardown from the other account was refused with no way to clear it. Suffix the namespace with the account uid. The owner-uid and mode-700 checks are unchanged, so a foreign-owned or wrong-mode name at this account's path is still refused and never adopted, chowned, or removed. Fixes kunchenguid#4716.
…te (kunchenguid#6809) The OpenCode session plugin's shouldArm kept its own copy of the need test that only looked for in-flight task records, while the turn-end guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh, which also counts registered process-event sources and trusted custom checks. With an empty fleet but any registered source or check, the guard blocked every turn end while the plugin declined to arm - a loop the guard's own repair line could not resolve because it names the plugin as the fix. The plugin now delegates the decision to the shared predicate through bash, keeping the local away-record decline and the x-mode.env arm override. OpenCode plugin test fixtures now carry the real predicate their arming path sources, and the arm suite gains six cases asserting the plugin's decision against the shared verdict over the same synthetic state directories. Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792) * fix(bin): resolve a pending reply only from its own task's status line Remote reply ingestion handed every corr= token in a mate's payload to fm_pending_reply_try_resolve together with that mate's own status log, so one mate echoing another mate's token resolved the other request. Honor a status-file override only when it is the record's own parent_status, and match the corr= token as a whole word. Fixes kunchenguid#6538 * no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784) agent-skill-trigger-index claims to be the complete agent-only trigger index but omitted operational-home-layout, session-start-recovery, validation-supervision, ship-landing, scout-completion, and away-quiet-supervision. Add each with its own description's trigger, placed beside the related entries. The decision-hold-lifecycle redirect stub stays out, per triage. Fixes kunchenguid#6503
…nchenguid#6814) * test: share a rename-safe agent stand-in across liveness suites On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to run when invoked through a symlink named after another utility. The Herdr descendant process-walk tests built their agent-named process as a `pi` symlink to the host `sleep`, so the process exited at once, its parent shell was gone before the walk ran, and both cases read `unknown unreadable` and failed on that host. The suite stops at its first failure, so every later case went unrun. The Herdr control smoke test's `claude` symlink has the same construction. The tmux liveness suite already solved this with a host-compiled spinner and a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the Herdr control smoke test now use it. When no stand-in can survive a foreign name, a case skips with the reason instead of failing. tests/fm-test-fixtures.test.sh gains a portable regression with a fake multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose. * no-mistakes(document): Correct Herdr verification fixture reference * ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787) * fix(bin): keep the supervision host's pass-through successor out of the hook's process group The successor a main-only pass-through leaves for main shared the Stop hook's process group, so the harness tearing that group down after the exit-2 rewake stopped it. The stop published downtime and the next park's first cycle announced an empty check: rearm-resurface, which woke main again in a loop. Start that successor in a process group of its own, as the hook's own handling successor already is. * no-mistakes(review): Give the at-turn successor left for main its own group * no-mistakes(document): Document own-group successor for turn-start hand-back too * no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823) * test: use idle composer readiness for Claude tmux guards * no-mistakes(test): Fix attended supervision test expectations and isolate worker state * no-mistakes(document): Correct live guard coverage and readiness documentation * no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass * test: preserve attended successor close assertions * no-mistakes(test): Fix attended live test watcher takeover expectations * no-mistakes(document): Correct stale attended guard documentation * Revert "no-mistakes(document): Correct stale attended guard documentation" This reverts commit 8e59d89. * Revert "no-mistakes(test): Fix attended live test watcher takeover expectations" This reverts commit c0b8510.
… tests/fm-watcher-lock.test.sh ("TERM after watcher lock acquisition left the singleton lock behind"). This was an intermittent race that predates the PR (neither bin/fm-watch.sh nor the test is touched by it); it reproduced locally in 1 of 5 runs. Root cause: bin/fm-watch.sh took the .watch.lock singleton lock before installing `trap watcher_cleanup EXIT`, so a TERM landing in that gap killed the watcher without cleanup and leaked the lock. Rule: once the watcher holds the lock, every exit must run watcher_cleanup. fm_lock_try_acquire is the only place the watcher gets the lock, so one change covers every path. Fix: moved `trap watcher_cleanup EXIT` and `watcher_stop_signals` to just before the acquisition loop. This is safe on the early exits ("already running", stale heartbeat) because watcher_cleanup releases the lock only when it records WATCHER_PID and every other cleanup step does nothing when there is nothing to clean up. Verified: with a 0.5s sleep added in the gap, the test fails with the old ordering and passes with the new one; fm-watcher-lock.test.sh passed 8 runs in a row; fm-watch-arm and fm-watch-checkpoint tests pass; shellcheck -x is clean
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
go sync
Context: the scheduled sync that merges upstream/main (kunchenguid/firstmate) into the fork's main (RajeshRajendiran/firstmate) stopped with a merge conflict. Fork main is at efb0243 (after merging the fork's pool-reset-on-cleanup PR 22). Upstream has 9 new commits: b062eb9 (kunchenguid#6818), 0ec1c5a (kunchenguid#6823), 3c58ec8 (kunchenguid#6787 supervision host successor watcher), 329ad4e (kunchenguid#6814), 2ce57d0 (kunchenguid#6784 index six missing agent-only skill triggers), 9b85a95 (kunchenguid#6792), 0c83b3f (kunchenguid#6809), 8fa2538 (kunchenguid#6780), 228b27d (kunchenguid#6785). A trial merge conflicts only in .agents/skills/agent-skill-trigger-index/SKILL.md.
What Changed
.agents/skills/agent-skill-trigger-index/SKILL.md, is resolved by keeping the fork's entries and adding the newly indexed agent-only triggers (operational-home-layout,session-start-recovery,validation-supervision,ship-landing,scout-completion,away-quiet-supervision).bin/fm-supervision-host.shstarts the successor watcher in its own process group with stdin from/dev/null, so it keeps running after the Stop hook's group is torn down.bin/fm-pending-reply-lib.shresolves a pending reply only from the asked task's own status log and only on a whole-wordcorr=match.bin/backends/herdr.shadds the OS account's uid to the presentation lock namespace path.bin/fm-procevent-lavish.shreportscaptain_message_countinstead ofsession_ending_message_countwhen the session has not ended.fm-primary-watch-arm.jsplugin now uses the sharedfm_supervision_neededcheck to decide whether to arm.tests/lib.shfixtures that work with multicallsleep, tmux Claude readiness checks that don't depend on permission footers, and herdr smoke cleanup. A follow-up commit updates the attended supervision-host live E2E test to expect the successor take-over.🤖 Generated with Claude Code
Risk Assessment
✅ Low: This is a merge of the 9 upstream commits listed in the intent. The only conflict was in .agents/skills/agent-skill-trigger-index/SKILL.md, and it was resolved as a clean union: all six upstream entries are present, the fork-only telegram-captain-channel entry is kept, and every indexed skill directory exists. The other files merged automatically, and the new upstream teardown tests use helpers that already exist and are invoked, with no duplicate function names.
Testing
I re-ran the opt-in live attended supervision-host E2E against the committed HEAD with a real Claude primary in an isolated lab, and it passed. The primary was woken for all four hand-offs. The event 1 turn end took over successor 3682351, the watcher it owned (3688913) delivered event 2's close, the stand-in remote listener stayed owned the whole time, and no captain prompts were submitted. I also checked the conflicted skill trigger index structurally: it has no conflict markers, no duplicate entries, and every indexed skill exists. That check was not a live run, so its scenario is recorded as untested. Earlier rounds' targeted suite logs are already in the evidence directory. The test's pre-fix negative control was skipped because no control ref was set.
Evidence: Live attended supervision-host E2E transcript at HEAD 8939c96
Source: Live attended supervision-host E2E transcript at HEAD 8939c96
Evidence: Skill trigger index merge check
Source: Skill trigger index merge check
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
🔧 **Test** - 1 issue found → auto-fixed → no changes applied ✅
🔧 No changes applied.
4 issues (1 error, 3 warnings) still open:
tests/fm-supervision-host-attended-live-e2e.test.sh- The live supervision-host attended E2E fails on event 2 because the arm following the successor watcher reports taken-over and never delivers its close. It fails identically on pure upstream b062eb9, so the sync didn't introduce it. It is either an upstream fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down kunchenguid/firstmate#6787 gap or specific to this host, and should be raised upstream.tests/fm-procevent.test.sh:3863- tests/fm-procevent.test.sh:3863 assumes an orphaned listener reparents to pid 1. On hosts with a systemd --user child subreaper (here pid 1059) it fails with 'was not reparented away from its session'. This predates the merge and isn't touched by it. The test should accept any subreaper that isn't the session.tests/fm-procevent.test.sh:2319- tests/fm-procevent.test.sh:2319 'a repaired source did not confirm' (reconciled started=0) failed once when run in parallel with other suites and passed when the suite ran alone. This looks like a timing flake that predates the merge.git diff 12513f8^2 12513f8 -- .agents/skills/agent-skill-trigger-index/SKILL.md(conflict resolution vs upstream)git diff efb0243 12513f8 -- .agents/skills/agent-skill-trigger-index/SKILL.md(conflict resolution vs fork main)index integrity: every-skill`` entry in the trigger index has an existing .agents/skills/<name>/SKILL.md, no duplicates (21 entries)bash tests/fm-supervision-host.test.shbash tests/fm-pending-reply.test.shbash tests/fm-watch-arm.test.shbash tests/fm-teardown.test.shbash tests/fm-test-fixtures.test.shbash tests/fm-backend-herdr.test.shbash tests/fm-pi-watch-extension.test.shbash tests/fm-procevent.test.sh(run twice)live:tests/fm-control-herdr-smoke.test.shagainst real herdr 0.9.3 (prior round, log reused)live:tests/fm-supervision-host-attended-live-e2e.test.shwith real Claude on the merge and on upstream b062eb9 (prior round, logs reused)🔧 Fix applied.
✅ Re-checked - no issues remain.
FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 tests/fm-supervision-host-attended-live-e2e.test.shat committed HEAD 8939c96 (real Claude 2.1.294 primary, haiku, in a disposable lab home on a private tmux socket)Checked the merged .agents/skills/agent-skill-trigger-index/SKILL.md: no conflict markers, no duplicate entries, and every indexed skill name resolves to an existing .agents/skills/<name>/SKILL.md✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.