Repository navigation
test: make tmux Claude readiness checks independent of permission footers - #6823
Merged
kunchenguid merged 9 commits intoOct 8, 2026
Merged
Conversation
…late worker state
…signment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass
This was referenced Oct 8, 2026
Owner
|
Speaking as Kun's firstmate: this is merged. Thank you @0x7067 — really appreciate you taking the time on this. |
nathanjgaul-agile
added a commit
to nathanjgaul-agile/firstmate
that referenced
this pull request
Oct 9, 2026
* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (kunchenguid#5343) * refactor(bin): share the local Firstmate home walk from the wake library Teardown's walk over the root home and its registered local secondmate homes moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to fm_firstmate_root_home, so a second consumer can count task records across this machine's homes without a copy. Teardown keeps its exact refusal wording through a thin wrapper. * feat(bin): defer spawns beyond a project's declared machine capacity A project whose machine-local resource only serves a few workers at once had no way to tell Firstmate so: every queued item was launched, and the surplus workers spent full-context turns retrying the resource. config/project-capacity in the root home now declares how many workers each named project admits at once on this machine. bin/fm-spawn.sh counts the ship and scout records on the same project origin across the root and its local secondmate homes, skipping ones whose ready PR is recorded, while holding the shared project lock through publication. A spawn with every place held exits 75 before any brief render, endpoint, worktree, record, or backlog move, so the item stays queued; batches report it as deferred. Undeclared projects keep today's uncapped dispatch, and an unreadable declaration refuses rather than guessing the limit. Refs kunchenguid#4237 * no-mistakes(review): Document that capacity matches the clone directory name * no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment * no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting * no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment * no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh * fix(bin): declare capacity for a project name that begins with # A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment. * no-mistakes(document): Rewrap project-capacity library header comment * no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh * fix(bin): release the Herdr session lock when reclaim finishes A concurrent resume in another home waits five seconds for that lock. Reclaim is the last presentation change on the recovery path, so holding the lock through the launch tail made the waiter time out. The contributions arm check also freezes its one-second clock, the same way the budget tests do, because an unfrozen clock can tick past before the first forge read. * no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim * no-mistakes(review): Skip the spawning task's own record in capacity count * no-mistakes(review): Restore release test comment above its test * docs: scope PR-ready re-evaluation to a declared project capacity A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case. * fix(bin): name the Lavish read message count by the same label as its section (kunchenguid#6785) fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE only when session_ended is true and CAPTAIN MESSAGE otherwise, but the count line always said session_ending_message_count. Several composer messages on a still-open board were therefore counted as session-ending. The count line now follows the same session_ended switch: session_ending_message_count once the session ended, captain_message_count otherwise. Message rows stay out of the annotation count, per triage. Fixes kunchenguid#6743 * fix(herdr): make the presentation lock namespace per OS account (kunchenguid#6780) The Herdr presentation lock namespace was the fixed machine-global /tmp/firstmate-herdr-presentation, so on a host where two OS users run Firstmate on Herdr the first account to create it owned it and every teardown from the other account was refused with no way to clear it. Suffix the namespace with the account uid. The owner-uid and mode-700 checks are unchanged, so a foreign-owned or wrong-mode name at this account's path is still refused and never adopted, chowned, or removed. Fixes kunchenguid#4716. * Fix OpenCode arm plugin to decide with the shared supervision predicate (kunchenguid#6809) The OpenCode session plugin's shouldArm kept its own copy of the need test that only looked for in-flight task records, while the turn-end guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh, which also counts registered process-event sources and trusted custom checks. With an empty fleet but any registered source or check, the guard blocked every turn end while the plugin declined to arm - a loop the guard's own repair line could not resolve because it names the plugin as the fix. The plugin now delegates the decision to the shared predicate through bash, keeping the local away-record decline and the x-mode.env arm override. OpenCode plugin test fixtures now carry the real predicate their arming path sources, and the arm suite gains six cases asserting the plugin's decision against the shared verdict over the same synthetic state directories. Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain> * fix(bin): resolve a pending reply only from its own task's status line (kunchenguid#6792) * fix(bin): resolve a pending reply only from its own task's status line Remote reply ingestion handed every corr= token in a mate's payload to fm_pending_reply_try_resolve together with that mate's own status log, so one mate echoing another mate's token resolved the other request. Honor a status-file override only when it is the record's own parent_status, and match the corr= token as a whole word. Fixes kunchenguid#6538 * no-mistakes(document): docs: scope remote reply settlement to the asked mate * docs(skills): index the six missing agent-only skill triggers (kunchenguid#6784) agent-skill-trigger-index claims to be the complete agent-only trigger index but omitted operational-home-layout, session-start-recovery, validation-supervision, ship-landing, scout-completion, and away-quiet-supervision. Add each with its own description's trigger, placed beside the related entries. The decision-hold-lifecycle redirect stub stays out, per triage. Fixes kunchenguid#6503 * test: make agent process fixtures compatible with multicall sleep (kunchenguid#6814) * test: share a rename-safe agent stand-in across liveness suites On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to run when invoked through a symlink named after another utility. The Herdr descendant process-walk tests built their agent-named process as a `pi` symlink to the host `sleep`, so the process exited at once, its parent shell was gone before the walk ran, and both cases read `unknown unreadable` and failed on that host. The suite stops at its first failure, so every later case went unrun. The Herdr control smoke test's `claude` symlink has the same construction. The tmux liveness suite already solved this with a host-compiled spinner and a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the Herdr control smoke test now use it. When no stand-in can survive a foreign name, a case skips with the reason instead of failing. tests/fm-test-fixtures.test.sh gains a portable regression with a fake multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose. * no-mistakes(document): Correct Herdr verification fixture reference * ci: retrigger cancelled shard * fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (kunchenguid#6787) * fix(bin): keep the supervision host's pass-through successor out of the hook's process group The successor a main-only pass-through leaves for main shared the Stop hook's process group, so the harness tearing that group down after the exit-2 rewake stopped it. The stop published downtime and the next park's first cycle announced an empty check: rearm-resurface, which woke main again in a loop. Start that successor in a process group of its own, as the hook's own handling successor already is. * no-mistakes(review): Give the at-turn successor left for main its own group * no-mistakes(document): Document own-group successor for turn-start hand-back too * no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run * test: make tmux Claude readiness checks independent of permission footers (kunchenguid#6823) * test: use idle composer readiness for Claude tmux guards * no-mistakes(test): Fix attended supervision test expectations and isolate worker state * no-mistakes(document): Correct live guard coverage and readiness documentation * no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass * test: preserve attended successor close assertions * no-mistakes(test): Fix attended live test watcher takeover expectations * no-mistakes(document): Correct stale attended guard documentation * Revert "no-mistakes(document): Correct stale attended guard documentation" This reverts commit 8e59d89. * Revert "no-mistakes(test): Fix attended live test watcher takeover expectations" This reverts commit c0b8510. * test(herdr): accept safe exit refusals and clean lab once (kunchenguid#6818) * test: stabilize watcher lock race fixtures (kunchenguid#6887) * test: order stale watcher-lock fixture races explicitly * no-mistakes(review): Removed duplicate stale-steal reap call * Rerun CI after the merge --------- Co-authored-by: Tiago <tiagop@hey.com> Co-authored-by: yairtech <39274208+falkoro@users.noreply.github.com> Co-authored-by: Asser AboElkhair <asser.aboelkhair@gmail.com> Co-authored-by: dubiousenvelope <greg.ecklin@gmail.com> Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain> Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com> Co-authored-by: menidi <menidi@users.noreply.github.com> Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Yes: fix the two tmux live tests (host mirror, attended supervision host) that wait on Claude footer text, using the same native-idle plus empty-composer approach, and ship it as its own upstream PR to kunchenguid/firstmate through full validation.
Context:
tests/fm-host-mirror-live-e2e.test.sh(around line 130) andtests/fm-supervision-host-attended-live-e2e.test.sh(around line 260) decide Claude is ready only when the screen contains the literal footerbypass permissions on.Under a managed Claude policy that disables bypass mode, or with Firstmate's supported
--permission-mode autolaunch, the footer readsauto mode oninstead, so the first test silently proceeds after its 60s wait and the second fails.The Herdr live guard had the same fragility and its fix (now validating as its own PR from branch
fm/fm-claude-live-guard-idle-footer) replaced the footer match with the agent reporting idle plus the shared composer classifier reading the composer as empty.What Changed
Risk Assessment
✅ Low: The changes are bounded to live-test readiness and documentation, respect the recorded rendered-state decision, and introduce no substantiated regression.
Testing
Both live tests passed after correcting fixture setup and an inherited worker marker; the before-fix timeout was reproduced, explicit auto mode and adversarial readiness checks passed, and terminal evidence was rendered as HTML. Disposable labs and temporary drivers were cleaned up.
Evidence: Host mirror live pass
Source: Host mirror live pass
Evidence: Attended supervision complete live pass
Source: Attended supervision complete live pass
Evidence: Explicit auto-mode readiness and adversarial checks
Source: Explicit auto-mode readiness and adversarial checks
Evidence: Pending composer produces expected timeout failure
Source: Pending composer produces expected timeout failure
Evidence: Before-fix footer regression reproduced
Source: Before-fix footer regression reproduced
Evidence: Rendered real terminal captures
Source: Rendered real terminal captures
Evidence: Isolation, setup corrections, and driver details
Source: Isolation, setup corrections, and driver details
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
🔧 **Test** - 2 issues found → auto-fixed ✅
tests/fm-supervision-host-attended-live-e2e.test.sh:373- The attended live test reproducibly fails after the second notification: it requires reason=attached-delivered-wake, but the followed watcher is terminated during host takeover and the attached arm records reason=taken-over. Both current launch modes and a base-version comparison reproduce this failure. Readiness succeeds, so this is not attributed to the footer fix, but it prevents complete targeted validation. Reconcile the test's successor-lifecycle expectation with the runtime before shipping; no unrelated lifecycle changes were made.FM_HOST_MIRROR_LIVE_E2E=1 FM_HOST_MIRROR_LIVE_HARNESSES=claude bash tests/fm-host-mirror-live-e2e.test.shwith workspace-local fixtures; manually cleared the fixture's external-import dialog.Ran a disposable copy of the host-mirror test with--permission-mode autoand explicit handling of the fixture's import dialog.FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/.attended-live-boundary.test.sh; disposable copy preserved Claude-owned transcripts outside the workspace. Corrected an initial fixture-copy setup error and retried.Repeated the attended test with explicit--permission-mode auto; reproduced the same second-notification assertion failure.Executed the base-version attended test with its footer expectation adapted toauto mode on, auto-mode launch, and boundary-safe cleanup; reproduced the same later failure.Drove Claude 2.1.293 on a marked disposable home and private tmux socket with a 160×45 terminal: checked empty composer, pending draft, cleared draft, active inference, and missing pane.Captured live tmux panes and rendered them as PNG evidence; removed disposable sessions, fixtures, and temporary drivers, then verified a clean working tree.🔧 Fix applied.
✅ Re-checked - no issues remain.
FM_HOST_MIRROR_LIVE_E2E=1 FM_HOST_MIRROR_LIVE_HARNESSES=claude bash tests/fm-host-mirror-live-e2e.test.shagainst real Claude; also repeated unaided using TMPDIR=/tmp.env -u FM_REMOTE_JOB_ACTIVE TMPDIR=/tmp FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/fm-supervision-host-attended-live-e2e.test.shcompleted all four notifications.Ran the base-commit attended test against real Claude and reproduced its footer-wait timeout.Ran the disposable readiness probe with realclaude --permission-mode auto, testing empty, pending, cleared, and busy states.Ran the mirror test with a tmux input driver inserting a real unsent draft; verified expected timeout and exit status 1.Captured private tmux panes, rendered them into terminal-evidence.html, and verified cleanup and a clean worktree.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.