Skip to content

test: make tmux Claude readiness checks independent of permission footers - #6823

Merged
kunchenguid merged 9 commits into
kunchenguid:mainfrom
0x7067:fm/fm-tmux-claude-guard-bypass-footer
Oct 8, 2026
Merged

kunchenguid merged 9 commits into
kunchenguid:mainfrom
0x7067:fm/fm-tmux-claude-guard-bypass-footer

Conversation

@0x7067

@0x7067 0x7067 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Yes: fix the two tmux live tests (host mirror, attended supervision host) that wait on Claude footer text, using the same native-idle plus empty-composer approach, and ship it as its own upstream PR to kunchenguid/firstmate through full validation.

Context: tests/fm-host-mirror-live-e2e.test.sh (around line 130) and tests/fm-supervision-host-attended-live-e2e.test.sh (around line 260) decide Claude is ready only when the screen contains the literal footer bypass permissions on.
Under a managed Claude policy that disables bypass mode, or with Firstmate's supported --permission-mode auto launch, the footer reads auto mode on instead, so the first test silently proceeds after its 60s wait and the second fails.
The Herdr live guard had the same fragility and its fix (now validating as its own PR from branch fm/fm-claude-live-guard-idle-footer) replaced the footer match with the agent reporting idle plus the shared composer classifier reading the composer as empty.

What Changed

  • Replace permission-footer matching in both tmux live tests with Claude-scoped idle detection and the shared empty-composer classifier; disable host-mirror prompt suggestions and report readiness diagnostics on timeout.
  • Update attended supervision assertions to require watcher replacement after re-arm and an actionable close from the replacement watcher.
  • Record the host-mirror run under managed auto mode and link attended test coverage to its scenario header.

Risk Assessment

✅ Low: The changes are bounded to live-test readiness and documentation, respect the recorded rendered-state decision, and introduce no substantiated regression.

Testing

Both live tests passed after correcting fixture setup and an inherited worker marker; the before-fix timeout was reproduced, explicit auto mode and adversarial readiness checks passed, and terminal evidence was rendered as HTML. Disposable labs and temporary drivers were cleaned up.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Run the host mirror test under managed auto mode: it submits the setup prompt, mirrors both replies, and excludes the automatic follow-up from captain input. ✅ pass live mirror-live.log; terminal-evidence.html
Run the attended test with an auto-mode footer: readiness succeeds where the before-fix footer wait times out. ✅ pass live attended-live-clean-env.log; baseline-attended-live.log; terminal-evidence.html
Deliver successive attended notifications: the replacement monitor's close, remote reply, and decision arriving during handoff all wake Claude and receive acknowledgements. ✅ pass live attended-live-clean-env.log
Launch Claude with --permission-mode auto: an idle, empty composer satisfies readiness. ✅ pass live readiness-probe.log; terminal-evidence.html
Type an unsent draft: readiness rejects pending input and accepts the composer after clearing it. ✅ pass live readiness-probe.log; terminal-evidence.html
Submit a real Claude request: readiness rejects the active turn even with an empty composer. ✅ pass live readiness-probe.log; terminal-evidence.html
Leave a draft throughout the mirror startup wait: the test exits with a readiness diagnostic instead of silently proceeding. ✅ pass live mirror-pending-timeout.log; terminal-evidence.html
Evidence: Host mirror live pass

Source: Host mirror live pass

ok - claude 2.1.293 (Claude Code): a turn the harness started itself was not mirrored as the captain's words
ok - claude 2.1.293 (Claude Code): the tracked registrations mirrored the captain prompt and main reply
ok - host mirror live: 1 harness(es) proved their writers
Evidence: Attended supervision complete live pass

Source: Attended supervision complete live pass

skip: control: set FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF to a pre-fix ref to run the negative control
# 01:45:26 positive step 1: primary idle (claude pid 2239708, 2.1.293 (Claude Code), model haiku); tracked Stop hook registered; config/supervision-host present; host pid 2251492 parked on watcher 2251992; listener runner 2252586; captain prompts so far: 1
# 01:45:26 positive transcript: ~/.claude/projects/-tmp-fm-sh-attended-live-R96oY0-positive-fm/ae9cb559-e1c6-4ab8-b206-ebe91a0e227e.jsonl
# 01:45:26 positive step 2: event 1 appended at 1791423926 (demo.status needs-decision)
# 01:45:27 positive step 2: host log: 1791423927	pass-through	attended	main-only
# 01:45:27 positive step 2: successor watcher pid 2256183 alive; ledger: epoch=1 owner_pid=2251088 outcome=arming updated_at=1791423925; marker: announced:downtime:2251992.1791423926.L6dhoi
# 01:45:33 positive step 2: rewake delivered at 1791423928 (Stop hook exited 2 with the banner); ledger: epoch=1 owner_pid=2251088 outcome=rewake updated_at=1791423928 session_pid=2239708 recovery_generation=2251992.1791423926.L6dhoi
# 01:45:33 positive step 2: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 2251992.1791423926.L6dhoi;
# 01:45:35 positive step 3: turn end re-armed: 1791423935	start	gen=host-2263196-1791423934; replacement watcher 2264107 live
# 01:45:38 positive step 3: event 2 appended at 1791423938
# 01:45:45 positive step 3/4: replacement watcher 2264107 closed: arm_pid=2263445 watcher_pid=2264107 origin=started started_at=1791423935 ended_at=1791423939 exit_code=0 signal=none reason=actionable-signal
# 01:45:45 positive step 3/4: its close was delivered: rewake at 1791423941; host log: 1791423940	pass-through	attended	main-only
# 01:45:50 positive step 4: event 3 appended to the stand-in remote log at 1791423950
# 01:45:59 positive step 4: listener mirrored it (1 ingested) and it was delivered at 1791423954; listener runner 2252586 -> 2252586, owned at every check
# 01:46:03 positive step 5: event 4, a routine working line on task demo2, appended at 1791423963
# 01:46:06 positive step 5: the host accepted it and confirmed its successor's handling handoff (marker announced:handling:2311894.1791423965.mmCUAn at 1791423966); a needs-decision on demo2 landed then, before the turn's start
# 01:46:06 positive step 5: host log: 1791423966	pass-through	attended	main-only; no engine turn
# 01:46:07 positive step 5: rewake delivered at 1791423967; ledger: epoch=4 owner_pid=2309832 outcome=rewake updated_at=1791423967 session_pid=2239708 recovery_generation=2311894.1791423965.mmCUAn
# 01:46:11 positive step 5: primary turn ran: bin/fm-wake-drain.sh;bin/fm-wake-drain.sh --ack-through 12 --recovery-generation 2311894.1791423965.mmCUAn;
# 01:46:13 positive step 5: turn end re-armed: 1791423972	start	gen=host-2332395-1791423972; watcher 2333123 live; listener runner 2252586 still owned
# 01:46:13 positive: captain prompts after setup: 0 (mirror holds only the setup prompt)
ok - attended live (2.1.293 (Claude Code)): an idle primary is woken for four hand-offs, the replacement watcher's close and a close that turned main-only at its turn included, with the listener owned throughout
Evidence: Explicit auto-mode readiness and adversarial checks

Source: Explicit auto-mode readiness and adversarial checks

busy=idle composer=empty
busy=idle composer=pending
PASS: pending input blocks readiness
busy=idle composer=empty
busy=busy composer=empty
PASS: active Claude turn blocks readiness
Evidence: Pending composer produces expected timeout failure

Source: Pending composer produces expected timeout failure

not ok - claude 2.1.293 (Claude Code): never reached an idle empty composer (last readiness: busy=idle composer=pending; last screen: ▎ Why so serious? It's just a hotfix.
  1 more notice hidden
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ unsent captain draft
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle))
observed exit=1 (expected 1 for pending composer)
Evidence: Before-fix footer regression reproduced

Source: Before-fix footer regression reproduced

skip: control: set FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF to a pre-fix ref to run the negative control
not ok - positive: Claude never reached its composer
--- host log

--- cycle exits

--- queue

--- ledger: 
--- marker: 
--- screen
Evidence: Rendered real terminal captures

Source: Rendered real terminal captures

<!doctype html><meta charset="utf-8"><title>Claude tmux live validation</title><style>body{background:#171717;color:#eee;font:14px sans-serif;margin:24px}pre{font:12px/1.5 monospace;white-space:pre;overflow:auto;background:#242424;padding:16px}section{margin-bottom:32px}</style><h1>Real Claude terminal captures</h1><p>Captured from isolated tmux panes during this run. Monochrome rendering preserves terminal text and layout.</p><section><h2>Host mirror: managed auto mode</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 · Claude Team
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/159e9a668321/01M4CHW5YHGMARGX7629WDRE4A/.test-live-tmp/fm-host-mirror-live.iVRpV2/claude

▎ Message from Ravn_Muninn:
▎ May the Force be with you.
  1 more notice hidden

❯ Reply with exactly the word mirror-ok and nothing else.

  Thought for 1s (ctrl+o to expand)

● mirror-ok

✻ Cooked for 2s · done 1:41 AM

● Stop hook feedback

· Graphing whale migration… (1s · thinking)
  ⎿  Tip: Use ctrl+v to paste images from your clipboard

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · esc to interrupt · ← for agents                                                                                                            ◐ medium · /effort

























</pre></section><section><h2>Attended notifications: managed auto mode</h2><pre>
● Stop hook feedback

● Bash(bin/fm-wake-drain.sh)
  ⎿  1791423926 3       signal  demo.status     needs-decision: /tmp/fm-sh-attended-live.R96oY0/positive/fm/state/demo.status /tmp/fm-sh-attended-live.R96oY0/positive/fm/state/demo2.status
     1791423926 4       signal  demo2.status    signal: /tmp/fm-sh-attended-live.R96oY0/positive/fm/state/demo.status /tmp/fm-sh-attended-live.R96oY0/positive/fm/state/demo2.status
     WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 2251992.1791423926.L6dhoi
     … +5 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier

  Thought for 1s (ctrl+o to expand)

● Bash(bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 2251992.1791423926.L6dhoi)
  ⎿  (No output)
  ⎿  Allowed by auto mode classifier

● ACKED

✻ Cogitated for 6s · done 1:45 AM

● Stop hook feedback

● Bash(bin/fm-wake-drain.sh)
  ⎿  1791423939 6       signal  demo.status     needs-decision: /tmp/fm-sh-attended-live.R96oY0/positive/fm/state/demo.status
     WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 6 --recovery-generation 2264107.1791423939.LY36wu
     wake annotation: latest wake-EVENT observed at drain, not current state: demo.status: needs-decision [at=1791423938] [key=lab-e2]: pick region east or west
     … +5 lines (ctrl+o to expand)
  ⎿  Allowed by auto mode classifier

● Bash(bin/fm-wake-drain.sh --ack-through 6 --recovery-generation 2264107.1791423939.LY36wu)
  ⎿  (No output)
  ⎿  Allowed by auto mode classifier

● ACKED

✻ Worked for 5s · done 1:45 AM

● Stop hook feedback

  Bash(bin/fm-wake-drain.sh)
  ⎿  Running…

· Propagating… (2s · ↓ 20 tokens)
  ⎿  Tip: Use /voice to enable push-to-talk dictation

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · esc to interrupt · ← for agents

</pre></section><section><h2>Explicit auto mode: empty composer</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 with low effort · Claude Team
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/159e9a668321/01M4CHW5YHGMARGX7629WDRE4A

▎ Message from Ravn_Muninn:
▎ Elementary, my dear Watson. It was a null pointer.
  1 more notice hidden

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents                                                                                                                                  ○ low · /effort






































</pre></section><section><h2>Unsent draft: readiness refused</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 with low effort · Claude Team
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/159e9a668321/01M4CHW5YHGMARGX7629WDRE4A

▎ Message from Ravn_Muninn:
▎ Elementary, my dear Watson. It was a null pointer.
  1 more notice hidden

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ unsent captain draft
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle)                                                                                                                                                 ○ low · /effort






































</pre></section><section><h2>Active turn: readiness refused</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 with low effort · Claude Team
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/159e9a668321/01M4CHW5YHGMARGX7629WDRE4A

▎ Message from Ravn_Muninn:
▎ Elementary, my dear Watson. It was a null pointer.
  1 more notice hidden

❯ Reply with the integers from 1 to 100, one per line. Do not use tools.

✽ Philosophizing…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · esc to interrupt · ← for agents                                                                                                  Ctrl+Y to paste deleted text


































</pre></section><section><h2>Before fix: idle composer while footer wait continues</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 with low effort · Claude Team
 ▝▝   ▝▝   /tmp/fm-sh-attended-live.vJZ6Ak/positive/fm

▎ Message from Ravn_Muninn:
▎ Say &#x27;hello&#x27; to my little friend!
  1 more notice hidden

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents                                                                                                                                                      ○ low · /effort






































</pre></section><section><h2>Unsent draft held during mirror readiness wait</h2><pre> ▐▛███▛█   Claude Code v2.1.293
▝▜██████▀  Haiku 5.5 · Claude Team
 ▝▝   ▝▝   /tmp/fm-host-mirror-live.EoFxkY/claude

▎ Message from Ravn_Muninn:
▎ Why so serious? It&#x27;s just a hotfix.
  1 more notice hidden

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ unsent captain draft
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⏵⏵ auto mode on (shift+tab to cycle)                                                                                                                Bypass permissions mode was disabled by settings
</pre></section>
Evidence: Isolation, setup corrections, and driver details

Source: Isolation, setup corrections, and driver details

Real installed Claude Code 2.1.293, Haiku, normal managed login; private tmux panes were 200x50 or 220x50. Managed policy displayed auto mode even for the tests default bypass launch. Mirror first run required declining the lab external-import dialog; a subsequent default /tmp run passed unaided. Attended initial local TMPDIR fixture collided with copied untracked fixtures and was stopped; normal /tmp fixed setup. The next attended run inherited FM_REMOTE_JOB_ACTIVE=1, causing the disposable remote job to queue without a worker; removing that marker from the test environment produced a complete pass. No product source changes were made. Baseline attended script from 2ce57d0067d4e860e8969fa8fa8289771b855ec5 ran against real Claude and timed out on its footer wait; baseline-idle-screen.txt shows its empty composer and auto footer. readiness-probe.log drives real Claude with explicit --permission-mode auto and the real shared tmux helpers. pending-input-driver.sh forwards every tmux command to /usr/bin/tmux and types a real unsent draft on the first rendered auto footer; it does not substitute product output. All fixture servers/homes are removed by their drivers.
- Outcome: 🔧 2 issues found → auto-fixed ✅ across 2 runs (22m51s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

🔧 **Test** - 2 issues found → auto-fixed ✅
  • 🚨 tests/fm-supervision-host-attended-live-e2e.test.sh:373 - The attended live test reproducibly fails after the second notification: it requires reason=attached-delivered-wake, but the followed watcher is terminated during host takeover and the attached arm records reason=taken-over. Both current launch modes and a base-version comparison reproduce this failure. Readiness succeeds, so this is not attributed to the footer fix, but it prevents complete targeted validation. Reconcile the test's successor-lifecycle expectation with the runtime before shipping; no unrelated lifecycle changes were made.
  • 🚨 live validation verdict: no-go (6 of 6 scenarios were driven live against the product); failed: Run the attended test through successive notifications: the test verifies the successor's close and completes successfully
  • Live validation: ❌ no-go - 6 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Run the host-mirror test with Claude: captain prompt and reply are mirrored, while the hook-started turn is excluded from captain messages ✅ pass live host-mirror.log; host-mirror-auto.log
Start attended supervision with auto-mode Claude: readiness completes and the first notification wakes the primary to drain and acknowledge it ✅ pass live attended-auto-repeat.log, steps 1–2
Type a draft into an idle auto-mode composer: readiness rejects it and becomes ready again after the draft is cleared ✅ pass live readiness-with-busy.log; empty-capture.png; draft-capture.png
Submit a real Claude request: readiness rejects the active turn even when the composer is empty ✅ pass live readiness-with-busy.log; busy-capture.png
Close the Claude pane: readiness reports unknown state and refuses to proceed ✅ pass live readiness-with-busy.log
Run the attended test through successive notifications: the test verifies the successor's close and completes successfully ❌ fail live attended-retry.log; attended-auto-repeat.log; attended-baseline-control.log; control-host.log
  • FM_HOST_MIRROR_LIVE_E2E=1 FM_HOST_MIRROR_LIVE_HARNESSES=claude bash tests/fm-host-mirror-live-e2e.test.sh with workspace-local fixtures; manually cleared the fixture's external-import dialog.
  • Ran a disposable copy of the host-mirror test with --permission-mode auto and explicit handling of the fixture's import dialog.
  • FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/.attended-live-boundary.test.sh; disposable copy preserved Claude-owned transcripts outside the workspace. Corrected an initial fixture-copy setup error and retried.
  • Repeated the attended test with explicit --permission-mode auto; reproduced the same second-notification assertion failure.
  • Executed the base-version attended test with its footer expectation adapted to auto mode on, auto-mode launch, and boundary-safe cleanup; reproduced the same later failure.
  • Drove Claude 2.1.293 on a marked disposable home and private tmux socket with a 160×45 terminal: checked empty composer, pending draft, cleared draft, active inference, and missing pane.
  • Captured live tmux panes and rendered them as PNG evidence; removed disposable sessions, fixtures, and temporary drivers, then verified a clean working tree.

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Run the host mirror test under managed auto mode: it submits the setup prompt, mirrors both replies, and excludes the automatic follow-up from captain input. ✅ pass live mirror-live.log; terminal-evidence.html
Run the attended test with an auto-mode footer: readiness succeeds where the before-fix footer wait times out. ✅ pass live attended-live-clean-env.log; baseline-attended-live.log; terminal-evidence.html
Deliver successive attended notifications: the replacement monitor's close, remote reply, and decision arriving during handoff all wake Claude and receive acknowledgements. ✅ pass live attended-live-clean-env.log
Launch Claude with --permission-mode auto: an idle, empty composer satisfies readiness. ✅ pass live readiness-probe.log; terminal-evidence.html
Type an unsent draft: readiness rejects pending input and accepts the composer after clearing it. ✅ pass live readiness-probe.log; terminal-evidence.html
Submit a real Claude request: readiness rejects the active turn even with an empty composer. ✅ pass live readiness-probe.log; terminal-evidence.html
Leave a draft throughout the mirror startup wait: the test exits with a readiness diagnostic instead of silently proceeding. ✅ pass live mirror-pending-timeout.log; terminal-evidence.html
  • FM_HOST_MIRROR_LIVE_E2E=1 FM_HOST_MIRROR_LIVE_HARNESSES=claude bash tests/fm-host-mirror-live-e2e.test.sh against real Claude; also repeated unaided using TMPDIR=/tmp.
  • env -u FM_REMOTE_JOB_ACTIVE TMPDIR=/tmp FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 bash tests/fm-supervision-host-attended-live-e2e.test.sh completed all four notifications.
  • Ran the base-commit attended test against real Claude and reproduced its footer-wait timeout.
  • Ran the disposable readiness probe with real claude --permission-mode auto, testing empty, pending, cleared, and busy states.
  • Ran the mirror test with a tmux input driver inserting a real unsent draft; verified expected timeout and exit status 1.
  • Captured private tmux panes, rendered them into terminal-evidence.html, and verified cleanup and a clean worktree.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@0x7067 0x7067 changed the title test: make tmux live readiness independent of Claude permission footers test: make tmux Claude readiness checks independent of permission footers Oct 8, 2026
@kunchenguid
kunchenguid merged commit 0ec1c5a into kunchenguid:main Oct 8, 2026
19 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @0x7067 — really appreciate you taking the time on this.

nathanjgaul-agile added a commit to nathanjgaul-agile/firstmate that referenced this pull request Oct 9, 2026
* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743

* fix(herdr): make the presentation lock namespace per OS account (kunchenguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate

* docs(skills): index the six missing agent-only skill triggers (kunchenguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503

* test: make agent process fixtures compatible with multicall sleep (kunchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard

* fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run

* test: make tmux Claude readiness checks independent of permission footers (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.

* test(herdr): accept safe exit refusals and clean lab once (kunchenguid#6818)

* test: stabilize watcher lock race fixtures (kunchenguid#6887)

* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call

* Rerun CI after the merge

---------

Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: yairtech <39274208+falkoro@users.noreply.github.com>
Co-authored-by: Asser AboElkhair <asser.aboelkhair@gmail.com>
Co-authored-by: dubiousenvelope <greg.ecklin@gmail.com>
Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: menidi <menidi@users.noreply.github.com>
Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants