fix(bin): prevent idle composer injection wedges - #23
Merged
Merged
Conversation
added 3 commits
August 1, 2026 23:18
The away-mode sub-supervisor buffered escalations correctly and replayed
every one on return, but could not inject into the primary session for up
to eleven hours at a time - three overnight wedges of 7.8h, 11.0h, and
10.1h. Durability held; responsiveness did not.
The recorded hypothesis was that a long mid-turn Claude pane is simply
outside what the injection guard was designed for. The daemon log
disproves it: all three wedges are dominated by `state=pending`
deferrals (2792 on the incident day alone), while the busy branch that a
genuine mid-turn pane takes fired 11 times in the entire log history.
Root cause is a misread of an idle, injectable composer. Claude 2.x
renders its EMPTY composer row as `❯` + U+00A0 NO-BREAK SPACE, and bash's
`[:space:]` trims treat no non-ASCII blank as whitespace under any locale
the fleet runs, so the pad survived every trim as apparent typed content
and classified `pending`. The injector types only into an affirmatively
`empty` composer, so it deferred forever against a pane that was ready
the whole time. Reproduced live on the primary pane, which read
busy=idle composer=pending before this change and composer=empty after.
fm_composer_blank_normalize folds non-ASCII blank and zero-width
characters to an ASCII space before the verdict, in the shared owner so
every adapter gets it. The fold is safe in one direction only: every
character it folds is invisible, so any visible byte still reads
`pending` and the dead-shell refusal is untouched.
The leading prompt glyph is now removed as a literal prefix. Under an
exported C/POSIX locale `${content#?}` drops one BYTE, leaving the two
trailing bytes of a multibyte glyph behind as spurious content - a second,
independent deferring condition, verified to turn an idle placeholder
composer from `empty` into `pending`.
The injection guard itself is unchanged: it was fed a wrong verdict, not
reasoning wrongly. A genuinely mid-turn primary still defers on the busy
guard, bounded by one turn and self-clearing, with the max-defer wedge
alarm still covering a pathological one.
Buffering and replay-on-return are untouched, and nothing is delivered
that was not delivered before.
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Re-landed on the fork from the already-validated branch (previously opened as kunchenguid#1632, which this account cannot merge upstream). Code unchanged.
Intent
Diagnose and fix away mode buffering escalations for about eleven hours without delivering them.
INCIDENT AS REPORTED: On the first real overnight away-mode run (2026-07-28), escalations buffered correctly and every one replayed cleanly on return - nothing was lost - but the daemon could not inject into the primary session for 39,773 seconds, about eleven hours. The wedge alarm fired as designed and surfaced on the return catch-up. Consequence: away mode delivered durability but not responsiveness. Three decisions that should have been handled overnight sat until morning, which is the entire reason for running away mode rather than simply stopping.
CAUSE WAS NOT ESTABLISHED and the report deliberately did not guess. The recorded hypothesis, worth checking FIRST but explicitly NOT to be assumed: the injection guard requires an affirmatively empty composer AND an idle pane, and every verdict other than empty defers, so a persistently non-empty or unreadable primary composer would produce exactly this signature. The primary was a Claude session running a long agentic turn for much of that window. The first question to answer empirically was: does a mid-turn Claude pane ever read as injectable, and if not, is an eleven-hour turn simply outside what the guard was designed for?
REQUIRED APPROACH: (1) Establish the cause from recorded evidence before proposing a fix; if evidence is insufficient, reproduce it. (2) Fix so escalations reach the captain during a long primary turn, OR - if genuinely impossible with the current injection mechanism - say so plainly and propose what would be needed. An honest 'this needs a different delivery channel' was an explicitly valid outcome; a fix that appears to work but still defers silently was explicitly NOT acceptable. (3) Preserve durability: nothing was lost in the recorded incident and nothing may be lost after the change.
ACCEPTANCE CRITERIA: cause established with evidence rather than asserted; a captain-relevant escalation raised while the primary is mid-turn is delivered, or the reason it cannot be is documented with a concrete proposal; buffering and replay-on-return still lose nothing; covered by a test that reproduces the deferring condition; the rule that firstmate never arms a second watcher during away mode is preserved.
CONSTRAINTS: reproduce in a throwaway home, never against the primary checkout's live state/ records, because the live session supervising this work owns those away-mode records. Never pkill watcher or daemon processes, since that can kill sibling firstmate homes. Load and follow the firstmate-coding-guidelines skill before editing firstmate tracked material.
STANDING INSTRUCTION: until fixed, the captain is told plainly that away mode is safe but should not be relied on for overnight responsiveness. If the fix does not fully restore responsiveness, state exactly what it does and does not deliver rather than implying the limitation is gone.
WHAT I FOUND AND DECIDED (deliberate choices a reviewer reading only the diff would not know):
The recorded mid-turn hypothesis is DISPROVEN, not adopted. The daemon log shows all three wedges dominated by 'state=pending' deferrals (2,792 on the incident day), while the busy branch a genuine mid-turn pane takes fired only 11 times in the entire log history. The incident is also not a one-off: three wedges of 7.8h (07-23), 11.0h (07-28) and 10.1h (07-31).
ACTUAL CAUSE: Claude 2.x renders its EMPTY composer row as the agent glyph followed by U+00A0 NO-BREAK SPACE (bytes e2 9d af c2 a0). Bash's [:space:] trims treat no non-ASCII blank as whitespace under any locale the fleet runs (verified C, C.utf8, en_US.utf8, POSIX), so the pad survived every trim as apparent typed content and the row classified 'pending'. The injector types only into an affirmatively 'empty' composer, so it deferred against a pane that was idle and injectable the whole time. Reproduced live: the primary pane read busy=idle composer=pending before the change and composer=empty after.
A SECOND, INDEPENDENT deferring condition was found and fixed in the same path: the leading prompt glyph was removed by character count, and under an exported C/POSIX locale bash's ${content#?} drops one BYTE, leaving the two trailing bytes of a multibyte glyph behind as spurious content. Verified to flip an idle placeholder composer from empty to pending. It is now removed as a literal prefix. This is not scope creep; it is a second instance of the same wedge class on the same code path.
The fix lives in the ONE shared composer owner (bin/fm-composer-lib.sh) so every backend adapter (tmux, herdr, orca, cmux) gets it and the adapters cannot drift again, matching that file's existing consolidation rationale.
SAFETY DIRECTION IS DELIBERATE AND ONE-WAY: the blank fold only ever folds INVISIBLE characters (non-ASCII blanks and zero-width characters), so any visible byte still reads 'pending' and the existing dead-shell refusal ('unknown' for a bare shell prompt) is untouched. Under-folding merely defers, which the max-defer alarm surfaces; over-folding would inject over real human input, which is the failure this must never cause. Tests cover both directions explicitly.
THE INJECTION GUARD ITSELF IS DELIBERATELY UNCHANGED. It was fed a wrong verdict; it was not reasoning wrongly. Changing the guard would have weakened a safety boundary to work around a classifier bug.
MID-TURN LIMITATION IS DELIBERATELY RETAINED AND DOCUMENTED, not silently fixed away. A genuinely mid-turn primary still defers on the busy guard, bounded by one turn and self-clearing, with the max-defer wedge alarm still covering a pathological one. Empirically negligible (11 occurrences ever against 5,548 deferrals), so no separate delivery channel is needed. This is recorded in docs/verification/supervision.md as a maintainer-verification fact with exact versions, commands and output, per the coding guidelines' knowledge-placement tree. AGENTS.md is deliberately NOT touched: this is a script-level bug fix with no contract change, and mechanics belong in script headers.
EVIDENCE GATHERING RESPECTED THE CONSTRAINT: the live primary pane was only ever READ (read-only ANSI capture and classification). No live state/ record was written, no daemon or watcher was started, stopped or killed, and no second watcher was armed.
TWO PRE-EXISTING, UNRELATED TEST FAILURES were found and deliberately NOT fixed here, to keep this change reviewable: tests/fm-calm-pi-extension.test.sh ('/calm left collapsed thinking labels') and tests/fm-test-run.test.sh (exits 1 on a stray 'comm: not in sorted order' while every assertion passes, so the runner reports failed=2 with zero 'not ok' lines). Both reproduce on a clean baseline with these changes stashed. They are reported upward for separate tracking rather than folded in.
Regression tests were confirmed to FAIL on the pre-fix code and pass after, so they genuinely reproduce the deferring condition rather than merely asserting current behavior. Lint (bin/fm-lint.sh) and bin/fm-doc-audience-check.sh are clean.
What Changed
Risk Assessment
✅ Low: Captain, the change is well-bounded, fixes the proven classifier failure at the shared owner, preserves fail-closed injection and durability, and now documents the accepted mid-turn limitation with the existing out-of-band alert proposal.
Testing
The base-to-target counterfactual established the cause, focused automated tests validated classification and safety boundaries, and isolated end-to-end checks demonstrated exactly-once Herdr delivery, durable buffering/replay, busy-pane deferral with preservation, wedge alarming, and the one-watcher invariant. No visual artifact was applicable because this is a shell daemon and terminal-integration change; CLI transcripts and persisted-buffer behavior are the end-user evidence.
Evidence: Cause and counterfactual evidence
Recorded Claude empty composer row bytes: e2 9d af c2 a0 0d Baseline verdict: pending Fixed verdict: empty Visible draft remains: pending Bare shell prompt remains: unknown Baseline C-locale placeholder: pending Fixed C-locale placeholder: emptyEvidence: Real Herdr away-mode delivery
All real-Herdr scenarios passed: clean delivery after idle, exactly-once delivery after a swallowed Enter, normal escalation delivery, and buffer preservation with a wedge alarm when delivery remained unsafe.Evidence: Durable watcher-to-delivery lifecycle
A terminal event survived watcher restart, buffered once without duplication, injected once, and cleared only after successful delivery.Evidence: One-watcher invariant
Concurrent starts left exactly one watcher; singleton recovery and lifecycle checks completed successfully.Evidence: Daemon guard and buffer-preservation checks
Evidence: Composer regression checks
Evidence: Herdr adapter checks
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed ✅
docs/verification/supervision.md:285- The required approach says escalations must "reach the captain during a long primary turn" or the change must document why that is impossible and provide a concrete proposal. This hunk instead says delivery waits until the turn ends. The reachable sequence remains:escalate_flushcallsinject_msg, the busy guard returns failure, the durable buffer is retained, and every housekeeping retry follows the same path for the entire potentially unbounded turn. Captain approval is needed either to waive this criterion or to add a concrete out-of-band proposal, such as sending the digest through the existing backend-independent alert channel while retaining it for replay.🔧 Fix: Document mid-turn escalation delivery limitation
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
tests/fm-composer-lib.test.shtests/fm-backend-herdr.test.shtests/fm-daemon.test.shtests/fm-wake-daemon-lifecycle-e2e.test.shtests/fm-watcher-lock.test.shtests/fm-afk-inject-herdr-e2e.test.shagainst Herdr 0.7.5 in its isolated named labCompared base1e247571aa75e00b00c7a01c4830025ecd44dc61and target classifiers using the capturede2 9d af c2 a0row and the exportedLC_ALL=Cmultibyte-prompt caseVerifiedgit status --shortremained clean and no test watcher or Herdr lab process remained✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.