Conversation
A Claude primary whose watcher closes stay actionable stopped waking its own session. bin/fm-claude-stop-autoarm.sh checked state/.claude-autoarm-failure-alarmed BEFORE the actionable-rewake branch, so once an earlier failure episode had consumed the guard's one attended fail-open, every later actionable close was recorded as failed-suppressed and exited 0. The wake stayed queued and the model was never woken, so no further Stop fired and the Stop-owned chain froze. The episode could not clear itself either. `[ "$ACTIONABLE" -eq 1 ] && break` leaves the arm's retry loop before fm_watcher_healthy runs, so HEALTHY stays 0 and fm_autoarm_reset_owned - the only path that clears the markers - is unreachable for as long as closes keep coming back actionable. Observed on one home 2026-09-03..09: markers six days old, every outcome failed-suppressed, real signal wakes queued and never delivered. An actionable close is positive proof that the mechanism works: the arm started a watcher, that watcher ran a full cycle, and it delivered a real supervision wake. Decide it before the alarm suppression, and clear the failure episode after the winning commit so a superseded owner still mutates nothing. Only a non-actionable close can now be suppressed by a stale alarm. Isolated reproduction (throwaway FM_HOME, stub arm, no live state touched): 10 consecutive Stop events with a stale alarm delivered 0/10 continuations before, 10/10 after; removing only the alarm marker was the one-variable counterfactual, and an idle healthy non-actionable close still clears the episode exactly as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxErkvBhFFxU6ztYZ2Q7YM
…guid#3425) A session whose home lock is held by a different live harness is read-only: AGENTS.md section 3 forbids it from repairing supervision, and bin/fm-claude-stop-autoarm.sh is permanently inert in it. Neither of the guard's existing bounds can be reached there. The consecutive-block count only advances when the auto-arm epoch moves, and a frozen ledger pins it at its initial value; the attended fail-open additionally needs a verified failure episode that a session with no auto-arm never produces. The guard therefore re-blocked every turn end forever, demanding a repair the session may not perform. Give that state its own terminal path: at most FM_CLAUDE_TURNEND_READONLY_ADVISORIES (default 2, 0 stands down at once) read-only advisories naming the live lock holder, counted once per block instead of once per epoch, then stand down for the rest of the session. A contended budget lock stands the turn down too, because this path exists so a session that cannot act is never wedged. The advisory replaces this guard's repair banner in every mode; Grok, OpenCode, Pi, and omp forward it after their own prefix and keep their per-turn bound, and Cursor's park adapter already stands down before reaching the guard when another live harness holds the lock. A missing, malformed, or dead-owner lock is deliberately not this state: such a session may claim or reclaim the lock, so it keeps the ordinary repair path and its unbounded blocking, verified by two negatives in the tests. Isolated reproduction (throwaway FM_HOME, stub harness holding the lock): 12/12 blocks with count pinned at 1 before, 2/12 after. Removing only the frozen ledger was the one-variable counterfactual: the count then advanced to 12 and it still blocked 12/12, which is why an explicit unrepairable path is needed rather than per-block counting alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxErkvBhFFxU6ztYZ2Q7YM
… block count mode
…ry regression test
…ake marker Rebasing this branch onto a freshly fetched origin/main picked up 9074f9d's sibling commit b6b51b9, which bound autoarm_commit's rewake outcome in bin/fm-claude-stop-autoarm.sh to a state/.watcher-down recovery marker and a fresh watcher beacon, and updated every other actionable-outcome fixture in this file to write that marker. The singleton-actionable fixture backing test_stale_alarm_cannot_freeze_the_stop_owned_chain predates that change and never wrote the marker, so after the rebase its first Stop event's ACTIONABLE-before-alarm commit silently returned exit 0 instead of the required exit 2 rewake. Write the same pending:downtime marker the sibling fixtures already write, matching the hook's current rewake contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ed overlap Rebasing this branch onto a freshly fetched origin/main picked up 9074f9d, a sibling fix that bounds the ordinary re-block budget against a frozen auto-arm epoch. Its own regression test reproduced that shape with a live foreign harness holding state/.lock, and asserted the ordinary repair-banner progression through to one attended fail-open. That reproduction is exactly the condition this branch's session_is_lock_refused now intercepts earlier, before the wait loop and before autoarm_owns_recovery ever run: a session read-only because a different live harness holds its home lock must get bounded read-only advisories, never the repair banner or the attended fail-open a read-only session cannot act on. After the rebase the upstream test's reproduction therefore takes the lock-refused path instead of the path it asserted, and failed. Rewrite the test under its actual behavior: the same reproduction now expects exactly FM_CLAUDE_TURNEND_READONLY_ADVISORIES bounded advisories, a silent stand-down once that bound is spent, and a healthy watcher still clearing the whole episode ahead of the lock-refused check, matching the assertions already used by this branch's own lock-refused advisory tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
18a7f2e to
da60198
Compare
|
Speaking as Kun's firstmate: first look on fork PR #4094 (Vladykart). Head Attestation: MISMATCH (body attestation head_sha Contract-class: new-default. Defect 2 (#3425 lock-refused bound) is corrective restore of an unbounded re-block. Defect 1 intentionally replaces a previously asserted default: VISION.md per-rule
Overlaps (open, do not compete): #3898 (turnend budget vs auto-arm signal), #4026 (Stop guard/auto-arm deadlock), #3617 (auto-arm failure visibility), #2929 (session-lock cohort / fail-open). Help these rather than opening another PR; landing one may close leftovers. CI/NM: First-time fork tip runs approved this pass after diff review: CI |
Intent
Fix two supervision defects in Firstmate's Claude primary integration, observed on the main home 2026-08-31 19:52-20:40 UTC and 2026-09-01, as a captain-authorized queued repair. Delivery mode is no-mistakes with yolo off. Keep both fixes minimal and separately reviewable within ONE PR as TWO commits.
REQUIRED FIRST STEP (done): the primary local source was behind origin/main, so compare against the freshly fetched isolated default branch first; if upstream already fixes a named defect, identify that landing and exercise its regression instead of reimplementing. Result: the worktree base is already origin/main tip 78318e2 and NEITHER defect is fixed there. Both code paths are present verbatim on the default branch. Open upstream PRs touch the area (#3898, #4026, #3617, #2929; #1592 closed) but none has landed, so there was no existing regression to exercise. This is recorded deliberately, not skipped.
DEFECT 1 - the Stop-owned watcher auto-arm chain freezes. Original evidence: state/.claude-autoarm-epoch stuck at "epoch=2846 owner_pid=1210954 outcome=rewake updated_at=1788205925" across dozens of Stop events; every watcher started by bin/fm-watch-arm.sh ran one absorb pass then exited; bin/fm-claude-stop-autoarm.sh never advanced the epoch to arming again; guard budget state/.turnend-claude-blocks was count=1; no .claude-autoarm-failure-* markers; state/.watcher-down held a healthy acked:handling: record, so neither the budget stop nor the failure stop explained it. Named suspects to check were the rewake outcome path not re-claiming when the last outcome was rewake, and the watcher exit path after an absorb-only pass. Requirement: root-cause it, fix so the cycle self-sustains, and add a colocated test proving >=10 consecutive Stop events keep exactly one live watcher with an advancing beat.
Later accepted evidence refinement (2026-09-09): a mate's handoff report section 9 offered queued actionable events plus old failure-alarm markers causing outcome=failed-suppressed before the exit-2 rewake branch. That was explicitly given as an UNTESTED causal hypothesis, not a proven fix specification. A second accepted steer required preserving these uncertainties: the mate's expected idle-close self-heal was not observed across further invocations, but that is failure to observe recovery, NOT proof that a genuinely non-actionable healthy-close counterfactual ran; the mate withdrew its claim that absent lock and regular-file targets prove fm_failure_episode_reset succeeds; and MY isolated reproductions own the causal verdict.
Diagnostic requirements that were mandatory for this task: separate the initiating trigger, the masking failure markers, and the visible loss of notifications; contrast the normal and failing paths; inspect relevant history; run a one-variable counterfactual; and seek disconfirming evidence.
Diagnosis reached and proven in isolation. Initiating trigger: the guard's one attended fail-open fired during a genuine EARLIER failure episode and created state/.claude-autoarm-failure-alarmed. Masking condition: in bin/fm-claude-stop-autoarm.sh the line
[ "$ACTIONABLE" -eq 1 ] && breakleaves the bounded arm retry loop BEFORE fm_watcher_healthy runs, so HEALTHY stays 0, and the alarm early-out sat BEFORE the actionable-rewake branch; the episode's only clearing path (fm_autoarm_reset_owned, reached only on the HEALTHY non-actionable branch) is therefore unreachable for as long as closes keep coming back actionable. Visible symptom: real supervision wakes are queued correctly and then silently swallowed on the way to the model, the session is never woken, no further Stop event fires, and the chain freezes. Reproduced in a throwaway FM_HOME with a stub arm: 10 consecutive Stop events delivered 0/10 rewake continuations with a stale alarm present and every outcome failed-suppressed; the one-variable counterfactual (remove ONLY the alarm marker) gave 10/10; the disconfirming check showed an idle HEALTHY non-actionable close does still clear the episode, which is exactly why a continuously-actionable home never self-heals. Also recorded as a deliberate disconfirming finding: on the live main home the frozen epoch is additionally explained by that home's primary harness being pi (state/.lock is held by a live pi process), so the frozen ledger there is NOT by itself proof the hook is broken; the named suspect "rewake outcome path not re-claiming" was disconfirmed by both code reading and the reproduction, where the ledger epoch advanced 1..10 normally.Fix 1 (commit 1): decide an ACTIONABLE close BEFORE the attended-alarm suppression, because an actionable close is positive proof the mechanism works - the arm started a watcher, that watcher ran a full cycle, and it delivered a real supervision wake - and clear the failure episode after the winning commit so a superseded owner still mutates nothing. Only a non-actionable close can now be suppressed by a stale alarm. This deliberately CHANGES a previously asserted behavior: the existing test test_post_alarm_actionable_close_is_suppressed asserted the old swallow as intended, so it was rewritten as test_post_alarm_actionable_close_delivers_and_ends_the_episode, and the owning statements in docs/turnend-guard.md and docs/watcher-continuity.md were amended to match. That test rewrite and doc amendment are intended, not an accident.
The regression for defect 1 was required to include actionable closes with old markers, healthy non-actionable closes, and CONTINUOUSLY queued wakes so it cannot pass merely because an idle interval clears the episode. test_stale_alarm_cannot_freeze_the_stop_owned_chain therefore keeps the home continuously actionable across 10 Stop events, backdates the beacon before every event so only a cycle that genuinely ran its watcher leaves a fresh one, asserts exactly one arm per Stop event, asserts a singleton (no two watchers live at once), and asserts the stale episode was ended by the actionable path.
DEFECT 2 - upstream kunchenguid/firstmate issue #3425: bin/fm-turnend-guard.sh re-blocks without bound in a lock-refused session because its block counter is gated on auto-arm epoch movement that can never happen there. Requirement: bound it so that in a lock-refused session the guard prints its read-only advisory at most N times, with N small and DOCUMENTED, and then stands down, never looping a captain-facing session indefinitely. Add a test.
Reproduced in a throwaway FM_HOME with a stub harness process holding state/.lock: 12/12 turn ends blocked with count pinned at 1. One-variable counterfactual (remove ONLY the frozen ledger file): the counter then advanced to 12 but it STILL blocked 12/12, because terminal_fail_open additionally requires a verified failure episode a lock-refused session never produces. That is why per-block counting alone is insufficient and an explicit unrepairable terminal path is required; the issue's own "suggested shape" says the same.
Fix 2 (commit 2): detect a lock-refused session (state/.lock names a live verified harness that is not in this session's own harness ancestry - the same condition that makes bin/fm-claude-stop-autoarm.sh permanently inert), give it at most FM_CLAUDE_TURNEND_READONLY_ADVISORIES read-only advisories (default 2; 0 legally means stand down immediately) counted once per BLOCK instead of once per epoch, then stand down for the rest of the session. A contended budget lock also stands the turn down, because this path exists so a session that cannot act is never wedged. The advisory replaces this guard's repair banner in every mode, since no mode may demand a repair the session is forbidden to perform. A missing, malformed, or dead-owner lock is deliberately NOT this state - such a session may claim or reclaim the lock and keeps the ordinary repair path with its unbounded blocking - and two negatives pin that. The advisories deliberately SHARE the session's consecutive-block count rather than keeping a second counter: that bounds total nagging per session even if lock ownership flaps, at the documented cost of reaching an already-verified failure episode's attended fail-open a turn or two earlier. Compatibility was reviewed across every supported primary: Grok, OpenCode, Pi and omp forward the guard's stderr after their own generic prefix and keep their existing per-turn follow-up bound, and Cursor's park adapter already stands down before reaching this guard when another live harness holds the lock, so only --claude needed the advisory count.
CONSTRAINTS that were in force and must be treated as deliberate, not omissions:
VALIDATION CONSTRAINTS accepted from Firstmate for this run: drive every run and respond through to completion or a genuinely new ask-user finding; no --yes; no manual edits or commits while a run is active; required upstream workflows marked action_required are NOT green even if third-party checks pass; report the full PR URL and the exact required-run state honestly; never merge or deploy this repair, because live homes stay unchanged until an authorized landing and guarded update. Ask-user findings are never mine to answer - they must be escalated to Firstmate.
The PR body must note that upstream CI needs maintainer approval, so Firstmate can verify the required runs actually executed before merge.
Local evidence already collected on the final committed head: tests/fm-claude-stop-autoarm.test.sh 42 ok / 0 not ok, tests/fm-turnend-guard.test.sh 92 ok / 0 not ok, plus fm-watch-arm (14), fm-cursor-primary (28), fm-operational-input (7) and fm-guard-stale-banner (28) all green; bin/fm-lint.sh clean (actionlint absent locally, workflows unchanged) and bin/fm-doc-audience-check.sh ok.
What Changed
bin/fm-claude-stop-autoarm.sh: decide an actionable close as positive recovery before the attended-alarm suppression check, and clear the failure episode (fm_autoarm_reset_owned) after the winningrewakecommit, so a stale.claude-autoarm-failure-alarmedmarker can no longer permanently suppress the exit-2 continuation for every later actionable close. Only a non-actionable close can now be suppressed by a stale alarm.bin/fm-turnend-guard.sh: detect a lock-refused session (home lock held by a live harness outside this session's own ancestry) and give it at mostFM_CLAUDE_TURNEND_READONLY_ADVISORIES(default 2, configurable, 0 means stand down immediately) read-only advisories, counted once per block via a newblockmode onbudget_account_current_epochinstead of the epoch-gated mode, then stand down for the rest of the session instead of re-blocking indefinitely. A contended budget lock also stands the turn down. The advisory replaces the guard's repair banner in this mode since a lock-refused session cannot perform a repair.docs/configuration.md,docs/turnend-guard.md,docs/watcher-continuity.md: document the newFM_CLAUDE_TURNEND_READONLY_ADVISORIESvariable, the lock-refused advisory-and-stand-down path, and the revised actionable-close-ends-the-episode ordering (updating the previously documented post-alarm suppression behavior to match).tests/fm-claude-stop-autoarm.test.sh,tests/fm-turnend-guard.test.sh: add colocated regression coverage, includingtest_stale_alarm_cannot_freeze_the_stop_owned_chain(10 continuous actionable Stop events with a stale alarm present, asserting exactly one arm and one live watcher per event and that the episode ends via the actionable path) and the bounded lock-refused advisory tests (configurable count, immunity to an inflated count from a stalefailed-suppressedepisode, and the lock-owning/dead-owner negatives).test_post_alarm_actionable_close_is_suppressedis rewritten astest_post_alarm_actionable_close_delivers_and_ends_the_episodeto match the new ordering.Note: required upstream CI workflows on this PR are gated
action_requiredand need maintainer approval to run; verify they have actually executed (not just that third-party checks are green) before merge.🤖 Generated with Claude Code
Risk Assessment
✅ Low: The fix-round change correctly gates the fresh-session initialization branch of budget_account_current_epoch on count_mode, so block-mode always initializes COUNT=1 regardless of leftover failed/failed-suppressed ledger state; traced through both the epoch-mode (unchanged case logic) and block-mode (now always COUNT=1, then +1 per call via the old_session continuation branch) paths and confirmed the lock-refused path now delivers exactly READONLY_ADVISORIES (default 2) blocks before standing down even with a stale failed-suppressed episode and failure-notified marker present. The added regression test exercises the guard through its executable hook interface (run_hook_claude, exit-code counting) rather than asserting source text, reuses the existing hold_foreign_session_lock/release_foreign_session_lock/seed_frozen_autoarm_ledger helpers (parameterizing the outcome argument instead of duplicating), and is registered in the test runner list.
Testing
Both colocated test suites for the changed scripts (fm-turnend-guard.test.sh, fm-claude-stop-autoarm.test.sh) pass in full (93/93 and 42/42 respectively), and a manual before/after swap of bin/fm-turnend-guard.sh confirms the new regression test genuinely catches the off-by-one (pre-fix: 3 read-only advisories delivered instead of the documented bound of 2; post-fix: exactly 2), which is the specific defect the final commit 618caf9 was written to close. No product or test failures found; the worktree was left clean.
Evidence: Regression reproduction: pre-fix code fails the new test (3 advisories instead of bounded 2)
Evidence: Same test passes against the fixed target-commit code
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed ✅
bin/fm-turnend-guard.sh:348- In bin/fm-turnend-guard.sh, budget_account_current_epoch() takes a count_mode param ('epoch' or 'block') added for the new lock-refused read-only-advisory path, but only the 'old_session == SESSION_ID' continuation branch (line 340) was gated on count_mode. The fresh-session initialization branch (lines 348-360) was left unmodified: when the budget file doesn't yet exist for this session and the auto-arm ledger's leftover outcome is 'failed' or 'failed-suppressed' with FAILURE_NOTICE already present (stale state left over from a session that owned the lock before this one became lock-refused, which is not cleared because fm-claude-stop-autoarm.sh is inert while lock-refused), COUNT is initialized to 0 instead of 1 -- even though count_mode='block' has no 'fresh vs already-notified epoch' concept to justify that. I reproduced this against the real script (isolated temp home, foreign live process holding state/.lock, .claude-autoarm-failure-notified present, ledger outcome=failed-suppressed): with the documented default FM_CLAUDE_TURNEND_READONLY_ADVISORIES=2, the session actually received 3 read-only-advisory blocks, not 2. With FM_CLAUDE_TURNEND_READONLY_ADVISORIES=0 -- documented at fm-turnend-guard.sh:119-121 as legally meaning 'stand down immediately' -- the session still printed one advisory before standing down, directly contradicting that documented contract. This does not reopen the original unbounded-blocking defect (Turn-end guard re-blocks without bound in a lock-refused session: block counter is gated on auto-arm epoch movement that can never happen #3425): the count is still bounded, just off by one in this reachable state. It is a mechanical bookkeeping bug, not a product-behavior question, so it doesn't require author input to fix -- e.g. gate the whole failed/failed-suppressed fast-path in the initialization branch on count_mode=epoch as well.🔧 Fix: summary: fix lock-refused advisory off-by-one in block count mode
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-turnend-guard.test.sh— 93 ok / 0 not ok (92 pre-existing + 1 new regression test)bash tests/fm-claude-stop-autoarm.test.sh— 42 ok / 0 not ok, matching the evidence recorded in the user intentManual regression reproduction: swapped bin/fm-turnend-guard.sh to the pre-review-fix version (commit dd003b1) with the new testtest_hook_claude_mode_lock_refused_advisories_bounded_with_stale_failure_episodepresent, and confirmed it fails (got 3advisories instead of the documented bound of 2) — reproducing the exact off-by-one described in the accepted review findingRestored the fixed bin/fm-turnend-guard.sh and re-ran the same isolated test to confirm it passes (ok), then verifiedgit status --shortshows a clean worktree✅ **Document** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.